# Questions on piecewise torch compile design

**URL:** <https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589>\
**Category:** torch.compile integration\
**Created:** [May 7, 2025, 12:52pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589 "2025-05-07T12:52:39Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 7, 2025, 12:52pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/1 "2025-05-07T12:52:39Z")

</div>

Hi, vLLM Team, I’m a student and interested in v1’s feature especially on piecewise torch compile mechanism. Though have reviewed your codes in a few days and try to improve my understanding, I was puzzled on some details. For example, you have implemented custom outplace\_all\_reduce in one node with piecewise cuda graph, I’m confused about the address registered in custom\_all\_reduce.py v.s. cuda graph’s address management in backends.py , are they related or not? And during the prefill stage, it seems that cudagraph is applied only if uncached tokens\_num is less than certain batchsize, so which one is this shape, the max capture decoding batchsize or the general shape(as 8192 in chunked prefill)?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 7, 2025, 12:53pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/2 "2025-05-07T12:53:07Z")

</div>

The addresses registered in `custom_all_reduce.py` (for custom out-of-place allreduce) and the CUDA graph’s address management in `backends.py` are related but serve different purposes. In `custom_all_reduce.py`, buffer addresses are registered to enable IPC and efficient allreduce operations during CUDA graph capture and replay, ensuring that the same memory addresses are used for communication across processes. This is crucial because CUDA graphs require all memory addresses used during capture to be identical during replay, and `register_graph_buffers()` is called at the end of the capture context to record these addresses ([custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py), [API docs](https://docs.vllm.ai/en/latest/api/vllm/vllm.distributed.device_communicators.custom_all_reduce.html)).

In `backends.py`, CUDA graph address management ensures that the input buffers for each captured graph (for each batch size) are reused and consistent, which is necessary for correct CUDA graph replay ([backends.py](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/backends.py)). During the prefill stage, CUDA graph is applied only if the batch size matches one of the captured sizes (e.g., `[1, 2, 4, 8]` or a user-specified list). The “shape” here refers to the batch size for which the CUDA graph was captured, not the general shape (like 8192 in chunked prefill). You can override which batch sizes are captured using the `cudagraph_capture_sizes` config ([design doc](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#cudagraph-capture)).

Would you like a step-by-step breakdown of how these address registrations interact during graph capture and replay?

Sources:

- [custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py)
- [API docs: custom\_all\_reduce](https://docs.vllm.ai/en/latest/api/vllm/vllm.distributed.device_communicators.custom_all_reduce.html)
- [backends.py](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/backends.py)
- [Design doc: Cudagraph Capture](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#cudagraph-capture)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 7, 2025, 3:48pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/3 "2025-05-07T15:48:43Z")

</div>

Thanks for your warm explanation! As you mentioned, they are related and serve different purposes. However I’ve noticed this comment in [custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py#L261) and haven’t figured it out yet. Can you elaborate more on this? Also, the compilation process function `self.compiled_graph_for_general_shape` shouldn’t incur any IPC or just mimic the behaviour? Please correct me if I misunderstood.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 7, 2025, 3:49pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/4 "2025-05-07T15:49:05Z")

</div>

It seems likely that the comment in `custom_all_reduce.py` about not deduplicating addresses during graph buffer registration refers to intentionally registering all buffer addresses used during CUDA graph capture, even if some addresses are reused. This is to ensure correctness across ranks, since different processes might have different allocation patterns, and CUDA IPC handles have internal reference counting to manage any overhead ([custom\_all\_reduce.cuh](https://github.com/vllm-project/vllm/blob/main/csrc/custom_all_reduce.cuh)).

Regarding `self.compiled_graph_for_general_shape` in `backends.py`, my understanding is that this function compiles the computation graph for a symbolic (general) shape and does not perform any IPC or CUDA graph capture. It only mimics the computation for generalization and warmup purposes; actual IPC and CUDA graph address registration only occur during concrete shape graph capture and replay ([backends.py](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/backends.py)).

Would you like a more detailed step-by-step explanation of the registration and compilation process?

Sources:

- [custom\_all\_reduce.cuh](https://github.com/vllm-project/vllm/blob/main/csrc/custom_all_reduce.cuh)
- [backends.py](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/backends.py)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 7, 2025, 3:51pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/5 "2025-05-07T15:51:11Z")

</div>

Yeah, more detailed step-by-step explanation of the registration and compilation process is welcome.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 7, 2025, 3:51pm UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/6 "2025-05-07T15:51:47Z")

</div>

My understanding is that the registration and compilation process in vLLM V1 with piecewise CUDA graphs and custom allreduce works as follows:

1. **Graph Compilation (General Shape):**

2. **Concrete Shape Compilation and CUDA Graph Capture:**

3. **Replay:**

Let me know if you want code snippets or further breakdown of any specific step.

Sources:

- [backends.py](https://github.com/vllm-project/vllm/blob/main/vllm/compilation/backends.py)
- [custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py)
- [design doc](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#computation-graph-compilation)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 13, 2025, 5:45am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/7 "2025-05-13T05:45:16Z")

</div>

why piecewise compilation here needs outplace allreduce not the inplace one?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 13, 2025, 5:45am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/8 "2025-05-13T05:45:38Z")

</div>

Piecewise compilation in vLLM requires out-of-place allreduce because PyTorch custom ops used in compiled graphs do not support in-place mutation or returning a new tensor in the same op. This restriction means the allreduce operation must always be out-of-place, producing a new tensor as output rather than modifying the input tensor directly ([parallel\_state.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/parallel_state.py), [custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py)).

Would you like more detail on how this is implemented or why PyTorch imposes this limitation?

Sources:

- [parallel\_state.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/parallel_state.py)
- [custom\_all\_reduce.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/device_communicators/custom_all_reduce.py)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 13, 2025, 5:46am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/9 "2025-05-13T05:46:16Z")

</div>

why PyTorch imposes this limitation

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 13, 2025, 5:46am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/10 "2025-05-13T05:46:37Z")

</div>

PyTorch imposes this limitation because custom ops in `torch.compile` (Dynamo/Inductor) cannot safely support in-place mutation or returning a new tensor in the same op. This is due to the way PyTorch’s graph capture and optimization work: in-place mutations can break graph assumptions, aliasing analysis, and memory safety, making it difficult for the compiler to reason about tensor lifetimes and optimize execution ([parallel\_state.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/parallel_state.py), [PR 9061](https://github.com/vllm-project/vllm/pull/9061)).

As a result, vLLM’s custom allreduce is implemented as an out-of-place operation to ensure compatibility with PyTorch’s compilation and graph capture mechanisms.

Sources:

- [parallel\_state.py](https://github.com/vllm-project/vllm/blob/main/vllm/distributed/parallel_state.py)
- [PR 9061](https://github.com/vllm-project/vllm/pull/9061)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 13, 2025, 5:47am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/11 "2025-05-13T05:47:32Z")

</div>

how about the performance between custom\_allreduce and pynccl or torch.distributed package

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 13, 2025, 5:47am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/12 "2025-05-13T05:47:50Z")

</div>

Custom allreduce in vLLM generally outperforms both PyTorch’s `torch.distributed` and the `pynccl` wrapper, especially for small batch sizes and on NVLink/NVSwitch-connected GPUs. Benchmarks show up to 5–10% throughput and latency improvement over NCCL-based allreduce, and it can also reduce memory usage ([PR 2192](https://github.com/vllm-project/vllm/pull/2192), [PR 2760](https://github.com/vllm-project/vllm/pull/2760)). For large batch sizes or PCIe-only multi-GPU setups, the performance gain is smaller, and NCCL may be preferable in some cases.

Would you like more detailed benchmark data or configuration examples?

Sources:

- [PR 2192](https://github.com/vllm-project/vllm/pull/2192)
- [PR 2760](https://github.com/vllm-project/vllm/pull/2760)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 17, 2025, 2:46am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/13 "2025-05-17T02:46:28Z")

</div>

where is the place to update self.kv\_cache in v1 engine?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 17, 2025, 2:46am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/14 "2025-05-17T02:46:43Z")

</div>

In vLLM v1, `self.kv_cache` is updated within the attention layer’s `forward` method, where the KV cache is accessed and written to during each forward pass. The actual update occurs in the attention implementation, using the `self.kv_cache` attribute, which is managed per layer ([layer.py](https://github.com/vllm-project/vllm/blob/main/vllm/attention/layer.py)).

Would you like a step-by-step trace of how and where this update happens in the code?

Sources:

- [layer.py](https://github.com/vllm-project/vllm/blob/main/vllm/attention/layer.py)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 21, 2025, 4:57am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/15 "2025-05-21T04:57:02Z")

</div>

find it slower than v0, what do you think is the root cause?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 21, 2025, 4:57am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/16 "2025-05-21T04:57:22Z")

</div>

vLLM v1 can be slower than v0 in some scenarios, especially for single requests or when multi-step streaming is used, due to architectural changes like unified scheduling and chunked prefill always being enabled. Some users report higher first-token latency (TTFT) and slower throughput for certain workloads ([issue 11649](https://github.com/vllm-project/vllm/issues/11649), [issue 17540](https://github.com/vllm-project/vllm/issues/17540)). The root causes include less aggressive multi-step streaming, different batching policies, and overhead from new features.

Would you like a detailed breakdown of these causes and possible mitigations?

Sources:

- [issue 11649](https://github.com/vllm-project/vllm/issues/11649)
- [issue 17540](https://github.com/vllm-project/vllm/issues/17540)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 21, 2025, 10:18am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/17 "2025-05-21T10:18:16Z")

</div>

how to use full\_cuda\_graph and what is the difference?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 21, 2025, 10:18am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/18 "2025-05-21T10:18:37Z")

</div>

To use `full_cuda_graph` in vLLM, set `compilation_config={"full_cuda_graph": True}` when launching the server or creating the LLM instance. This enables capturing the entire forward pass—including attention—in a single CUDA graph, rather than splitting at attention layers. Currently, only FlashAttention 3 is compatible, and cascade attention must be disabled ([design doc](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#cudagraph-capture), [PR 16072](https://github.com/vllm-project/vllm/pull/16072)).

The main difference is that `full_cuda_graph` can improve performance for smaller models by reducing kernel launch overhead, but it may disable some attention optimizations and is less flexible than piecewise capture. Use it only when your attention backend supports it.

Would you like a code example or more technical details?

Sources:

- [design doc](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#cudagraph-capture)
- [PR 16072](https://github.com/vllm-project/vllm/pull/16072)

---

<div class="post-metadata">

**Author:** ![Juelianqvq](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/juelianqvq/32/280_2.png) [@Juelianqvq](https://discuss.vllm.ai/u/Juelianqvq)\
**Post date:** [May 21, 2025, 10:21am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/19 "2025-05-21T10:21:29Z")

</div>

> [@RunLLM](#):
>
> more technical details

more technical details

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [May 21, 2025, 10:21am UTC](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589/20 "2025-05-21T10:21:53Z")

</div>

Certainly! Here’s a step-by-step technical breakdown of `full_cuda_graph` in vLLM v1:

1. **Piecewise vs. Full CUDA Graph:**

2. **How to Enable:**

3. **Technical Implications:**

4. **Fallback:**

Let me know if you want a code snippet or further details on graph capture internals.

Sources:

- [design doc](https://docs.vllm.ai/en/latest/design/v1/torch_compile.html#cudagraph-capture)
- [PR 16072](https://github.com/vllm-project/vllm/pull/16072)

[Next page](https://discuss.vllm.ai/t/questions-on-piecewise-torch-compile-design/589.md?page=2)
