# When will draft tokens append to input\_ids

**URL:** <https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099>\
**Category:** General\
**Created:** [July 16, 2025, 7:21am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099 "2025-07-16T07:21:02Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![wZuck](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/wzuck/32/498_2.png) [@wZuck](https://discuss.vllm.ai/u/wZuck)\
**Post date:** [July 16, 2025, 7:21am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/1 "2025-07-16T07:21:02Z")

</div>

If enabled pipeline parallelism, spec\_token\_ids will not be appended to input\_batch.token\_ids\_cpu in ModelRunner.\_update\_states() due to not found req\_id in self.input\_batch.req\_id\_to\_index, i’m curious when will the draft tokens insert to token\_ids\_cpu.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [July 16, 2025, 7:21am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/2 "2025-07-16T07:21:27Z")

</div>

When pipeline parallelism (PP) is enabled, the scheduler sends the sampled tokens (including speculative/draft tokens) back to the model runner, because there is no direct communication between the first-stage and last-stage workers. In this case, the update to `input_batch.token_ids_cpu` for spec\_token\_ids is handled in the `_update_states` method, but only for requests present in `self.input_batch.req_id_to_index`. If a request is not found there (e.g., new or resumed requests), it is added to the persistent batch via `self.input_batch.add_request(req_state)`, which copies both prompt and output token IDs—including any draft/speculative tokens—into `token_ids_cpu` at that point. Thus, for PP, draft tokens are inserted into `token_ids_cpu` when the request is (re-)added to the batch, not during the main update loop for running requests. See the relevant logic in [gpu\_model\_runner.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py) and discussion in [PR #20291](https://github.com/vllm-project/vllm/pull/20291).

Would you like a step-by-step breakdown of the code paths involved?

Sources:

- [gpu\_model\_runner.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)
- [PR #20291](https://github.com/vllm-project/vllm/pull/20291)

---

<div class="post-metadata">

**Author:** ![wZuck](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/wzuck/32/498_2.png) [@wZuck](https://discuss.vllm.ai/u/wZuck)\
**Post date:** [July 16, 2025, 7:28am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/3 "2025-07-16T07:28:31Z")

</div>

if the are 2 batches in queue, after batch\_0 prefills and batch\_1 prefills, when batch\_0 decoding, req\_id\_to\_index only contains batch\_1 reqs, and the spec\_token\_ids will not be appended.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [July 16, 2025, 7:29am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/4 "2025-07-16T07:29:01Z")

</div>

It seems likely that if `req_id_to_index` only contains batch\_1 requests during batch\_0’s decoding, then speculative tokens (`spec_token_ids`) for batch\_0 will not be appended to `input_batch.token_ids_cpu` in `_update_states`. This is because the code only updates `token_ids_cpu` for requests present in `req_id_to_index`, and new/resumed requests are added to the batch after this check. As a result, if a request is not in the persistent batch at the time of decoding, its speculative tokens may not be inserted as expected. This behavior is discussed as a source of inconsistency in [vLLM issue #20247](https://github.com/vllm-project/vllm/issues/20247).

Would you like a detailed explanation of how this can lead to token count mismatches between the scheduler and model runner?

Sources:

- [gpu\_model\_runner.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_model_runner.py)
- [vLLM issue #20247](https://github.com/vllm-project/vllm/issues/20247)

---

<div class="post-metadata">

**Author:** ![wZuck](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/wzuck/32/498_2.png) [@wZuck](https://discuss.vllm.ai/u/wZuck)\
**Post date:** [July 16, 2025, 7:29am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/5 "2025-07-16T07:29:44Z")

</div>

when InputBatch.remove\_request() will be called?

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [July 16, 2025, 7:30am UTC](https://discuss.vllm.ai/t/when-will-draft-tokens-append-to-input-ids/1099/6 "2025-07-16T07:30:04Z")

</div>

`InputBatch.remove_request()` is called when a request is finished, aborted, or unscheduled (e.g., preempted or not scheduled in the current step). This method removes the request from the batch and must be followed by a call to `condense()` to compact the batch. See [gpu\_input\_batch.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_input_batch.py) and [API docs](https://docs.vllm.ai/en/stable/api/vllm/vllm.v1.worker.gpu_input_batch.html).

Would you like more detail on the specific scenarios or code paths where this occurs?

Sources:

- [gpu\_input\_batch.py](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_input_batch.py)
- [API docs](https://docs.vllm.ai/en/stable/api/vllm/vllm.v1.worker.gpu_input_batch.html)
