# DeepSeek MTP full cuda graph support?

**URL:** <https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548>\
**Category:** Speculative Decoding\
**Created:** [April 12, 2026, 9:50am UTC](https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548 "2026-04-12T09:50:37Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![dingfen](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/dingfen/32/117_2.png) [@dingfen](https://discuss.vllm.ai/u/dingfen)\
**Post date:** [April 12, 2026, 9:50am UTC](https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548/1 "2026-04-12T09:50:37Z")

</div>

Hey guys!  
I am working on vLLM 0.17.1, and I am running DeepSeek R1 with Spec Config like that:

```json
{
  "method": "deepseek_mtp",
  "num_speculative_tokens": 3,
  "draft_tensor_parallel_size": 16
}

```

And of cource for best performance, I use the FULL CUDA GRAPH when running.  
But I found that the draft model part (or MTP layer part) is ruuning on the PIECEWISE CUDA GRAPH thought the main model is the FULL CUDA GRAPH in the all-decoding phase.

And I wonder if there is any method or config can turn draft model to FULL CUDA GRAPH, or has vLLM support this feature yet?🤔

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [April 12, 2026, 9:50am UTC](https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548/2 "2026-04-12T09:50:51Z")

</div>

As of vLLM 0.17.1, full CUDA Graph support for the drafter (including DeepSeek MTP) is not available in mainline vLLM; the drafter runs in PIECEWISE CUDA Graph mode even if the target model uses FULL CUDA Graphs. However, there is active development to enable FULL CUDA Graph mode for the drafter, with a PR implementing this feature under review and tracked in [this issue](https://github.com/vllm-project/vllm/issues/33341) and [this PR](https://github.com/vllm-project/vllm/pull/34880). No config or flag in v0.17.1 enables FULL CUDA Graph for the drafter yet.

Once merged, this feature will allow the drafter (including MTP layers) to leverage FULL CUDA Graph mode, improving speculative decoding performance. For now, you must wait for the relevant PR to be merged and released in a future vLLM version. See discussion and technical details in [the tracking issue](https://github.com/vllm-project/vllm/issues/33341) and [the PR](https://github.com/vllm-project/vllm/pull/34880).

Would you like more detail on the technical implementation or how to track when this feature lands?

Sources:

- [https://github.com/vllm-project/vllm/issues/33341](https://github.com/vllm-project/vllm/issues/33341)
- [https://github.com/vllm-project/vllm/pull/34880](https://github.com/vllm-project/vllm/pull/34880)

Hint: Mention @RunLLM in the post for followups.

---

<div class="post-metadata">

**Author:** ![yangshanjun](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/yangshanjun/32/900_2.png) [@yangshanjun](https://discuss.vllm.ai/u/yangshanjun)\
**Post date:** [April 13, 2026, 8:22am UTC](https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548/3 "2026-04-13T08:22:17Z")

</div>

@RunLLM scheduler 和 gpu\_model\_runner 通过什么通信

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [April 13, 2026, 8:22am UTC](https://discuss.vllm.ai/t/deepseek-mtp-full-cuda-graph-support/2548/4 "2026-04-13T08:22:37Z")

</div>

scheduler（调度器）和gpu\_model\_runner（GPU模型执行器）之间通过进程间通信（IPC）机制进行数据交换，主要依赖于multiprocessing队列、共享内存（如NCCL、SHM等）和Python的Future对象。调度器将调度结果（如SchedulerOutput）通过队列或共享内存发送给gpu\_model\_runner，后者执行模型推理后再通过队列返回结果（如ModelRunnerOutput）给调度器。具体实现可参考[架构设计文档](https://docs.vllm.ai/en/latest/design/arch_overview/#gpu-worker-processes)和相关源码说明。

这种通信方式保证了调度与模型执行的解耦和高效并发，支持多GPU和分布式场景。详细流程和接口定义可见于[vllm/v1/engine/core.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py)和[vllm/v1/worker/gpu\_model\_runner.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)。

需要更详细的通信流程或源码解读吗？

Sources:

- [vllm/v1/engine/core.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/engine/core.py)
- [架构设计文档](https://docs.vllm.ai/en/latest/design/arch_overview/#gpu-worker-processes)
- [vllm/v1/worker/gpu\_model\_runner.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)

Hint: Mention @RunLLM in the post for followups.
