# V1 has lower end-to-end performance than V0(--num-scheduler-steps=8)

**URL:** <https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795>\
**Category:** V1 Feedback\
**Created:** [June 10, 2025, 12:43pm UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795 "2025-06-10T12:43:13Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![lianjie](https://avatars.discourse-cdn.com/v4/letter/l/94ad74/32.png) [@lianjie](https://discuss.vllm.ai/u/lianjie)\
**Post date:** [June 10, 2025, 12:43pm UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/1 "2025-06-10T12:43:14Z")

</div>

**Testing Environment:**

- vllm version: 0.8.5
- Operating System: Ubuntu 22.04
- GPU: Nvidia A100

**Performance Comparison:**

- For Llama/qwen2 series models, the end-to-end inference performance of V1 is 5% to 10% lower than that of V0.
- Version 0 configuration: `--num-scheduler-steps=8`

**Question:**

Is there any community that has relevant performance comparison data?

---

<div class="post-metadata">

**Author:** ![tjtanaa](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/tjtanaa/32/161_2.png) [@tjtanaa](https://discuss.vllm.ai/u/tjtanaa)\
**Post date:** [June 11, 2025, 8:25am UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/2 "2025-06-11T08:25:50Z")

</div>

@lianjie Which backend are you using? FlashInfer Attention on V0 and V1?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [June 11, 2025, 8:26am UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/3 "2025-06-11T08:26:14Z")

</div>

vLLM supports the FlashInfer attention backend on both V0 and V1 engines. You can select it by setting the environment variable VLLM\_ATTENTION\_BACKEND=FLASHINFER. However, some features and performance optimizations (like CUDA Graph support) are more mature in V1, and FlashInfer is not enabled by default due to ongoing performance tuning. You must install FlashInfer separately, as it is not bundled with vLLM by default. For details, see [docs](https://docs.vllm.ai/en/stable/getting_started/quickstart.html#on-attention-backends) and [discussion](https://github.com/vllm-project/vllm/pull/4353).

Would you like more detail on differences between V0 and V1 or how to configure FlashInfer?

Sources:

- [https://docs.vllm.ai/en/stable/getting\_started/quickstart.html#on-attention-backends](https://docs.vllm.ai/en/stable/getting_started/quickstart.html#on-attention-backends)
- [https://github.com/vllm-project/vllm/pull/4353](https://github.com/vllm-project/vllm/pull/4353)

---

<div class="post-metadata">

**Author:** ![benchislett](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/benchislett/32/380_2.png) [@benchislett](https://discuss.vllm.ai/u/benchislett)\
**Post date:** [June 11, 2025, 5:26pm UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/4 "2025-06-11T17:26:48Z")

</div>

I’ve noticed that on H100/H200, V0 with multi-step is still a few % faster than V1, but I’ve only observed this with a single concurrent request. When running more than one request, I’ve observed that the gap closes and V1 is slightly faster. I don’t think there will be many easy fixes to improve V1 performance to close this gap.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [June 11, 2025, 5:27pm UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/5 "2025-06-11T17:27:30Z")

</div>

Your observation is correct: with a single concurrent request, V0 with multi-step can be a few percent faster than V1, especially on H100/H200, but as concurrency increases, V1 matches or slightly surpasses V0. This is a known and expected behavior due to architectural differences—V1 is optimized for high concurrency and scheduling efficiency, while V0’s multi-step mode can minimize per-step overhead for single requests. There are no simple fixes to make V1 consistently faster for single-request, multi-step scenarios, as discussed in [this issue](https://github.com/vllm-project/vllm/issues/11649) and [related discussions](https://github.com/vllm-project/vllm/issues/6801).

Would you like a deeper technical breakdown of why this performance difference exists?

Sources:

- [https://github.com/vllm-project/vllm/issues/11649](https://github.com/vllm-project/vllm/issues/11649)
- [https://github.com/vllm-project/vllm/issues/6801](https://github.com/vllm-project/vllm/issues/6801)

---

<div class="post-metadata">

**Author:** ![lianjie](https://avatars.discourse-cdn.com/v4/letter/l/94ad74/32.png) [@lianjie](https://discuss.vllm.ai/u/lianjie)\
**Post date:** [June 12, 2025, 7:07am UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/6 "2025-06-12T07:07:21Z")

</div>

I used flash\_attn2 and tested the batch sizes of 32, 64 and 128，thank you for your reply.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [June 12, 2025, 7:07am UTC](https://discuss.vllm.ai/t/v1-has-lower-end-to-end-performance-than-v0-num-scheduler-steps-8/795/7 "2025-06-12T07:07:48Z")

</div>

With FlashAttention 2 (flash\_attn2), increasing batch size generally improves GPU utilization and throughput, especially for small models or short contexts. For batch sizes 32, 64, and 128, you should see higher throughput and better FLOPS utilization as batch size increases, but the gains may plateau due to memory or scheduling overheads. This is consistent with vLLM’s design and is discussed in [this issue](https://github.com/vllm-project/vllm/issues/44058e50) and [optimization docs](https://docs.vllm.ai/en/stable/performance/optimization.html).

Would you like more detail on how to further tune batch size or interpret your results?

Sources:

- [https://docs.vllm.ai/en/stable/performance/optimization.html](https://docs.vllm.ai/en/stable/performance/optimization.html)
