# Does LLM.generate() differentiate between Prefill and Decode phases? Where to insert timers?

**URL:** https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167
**Category:** General
**Created:** [July 24, 2025, 9:16am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167 "2025-07-24T09:16:49Z")
**Posts on this page:** 6
**Page:** 1

<div class="post-metadata">

### Author: ![zhy1658858023](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/zhy1658858023/32/521_2.png) [@zhy1658858023](https://discuss.vllm.ai/u/zhy1658858023)
#### Post date: [July 24, 2025, 9:16am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/1 "2025-07-24T09:16:49Z")

</div>

I am working with `LLM.generate()` from vLLM and would like to measure the time spent in the **Prefill** and **Decode** phases separately.

After reading the code, I understand that `generate()` internally calls `_run_engine()`, which loops over `self.llm_engine.step()` to perform inference.

My questions are:

1. **Does `llm.generate()` internally distinguish between the Prefill and Decode phases?**

2. **If this distinction exists, where is the best place in the codebase to insert timing measurements for these two phases?**

This information would be very helpful for latency profiling and benchmarking, especially in performance-sensitive applications.

Thank you!

---

<div class="post-metadata">

### Author: ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)
#### Post date: [July 24, 2025, 9:17am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/2 "2025-07-24T09:17:27Z")

</div>

vLLM does distinguish between Prefill and Decode phases, but the LLM.generate() API does not directly expose separate timing for these phases. The recommended way to measure prefill and decode times is to instrument the LLMEngine.step() method in vllm/engine/llm\_engine.py, as this is where the engine processes both phases. You can insert timing logic around the relevant code blocks for prefill and decode within this method to record their durations per request. This approach aligns with the internal profiling and metrics design described in the documentation and codebase, where intervals like prefill and decode are calculated based on event timestamps recorded in the engine core and processed in the frontend or output handler loop.

For more advanced or automated profiling, vLLM provides profiling scripts and examples (see profiling.py and related docs), but for custom, per-request timing, modifying LLMEngine.step() is the most direct and supported method according to the official documentation and code comments. There is no built-in API parameter in generate() to return prefill/decode times directly as of now. See [vllm.engine.llm\_engine LLMEngine.step()](https://docs.vllm.ai/en/latest/api/vllm/vllm.engine.llm_engine.html), [metrics design doc](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection), and [profiling example](https://docs.vllm.ai/en/latest/getting_started/examples/profiling.html#profiling) for further details.

Would you like a step-by-step guide or code snippet for adding this timing logic?

Sources:

- [Metrics Collection Design](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection)
- [Profiling Example](https://docs.vllm.ai/en/latest/getting_started/examples/profiling.html#profiling)
- [vllm.engine.llm\_engine LLMEngine.step()](https://docs.vllm.ai/en/latest/api/vllm/vllm.engine.llm_engine.html)

---

<div class="post-metadata">

### Author: ![zhy1658858023](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/zhy1658858023/32/521_2.png) [@zhy1658858023](https://discuss.vllm.ai/u/zhy1658858023)
#### Post date: [July 24, 2025, 9:20am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/3 "2025-07-24T09:20:52Z")

</div>

> [@RunLLM](#):
>
> Would you like a step-by-step guide or code snippet for adding this timing logic?

Yes, I need a step-by-step guide or code snippet for adding this timing logic.

---

<div class="post-metadata">

### Author: ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)
#### Post date: [July 24, 2025, 9:21am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/4 "2025-07-24T09:21:24Z")

</div>

Certainly! Here’s a step-by-step guide and code snippet for adding prefill and decode timing logic to vLLM by instrumenting LLMEngine.step() in vllm/engine/llm\_engine.py:

**Step 1: Locate LLMEngine.step()**

- Open vllm/engine/llm\_engine.py and find the step() method. This method schedules, executes, and processes both prefill and decode phases for each batch of requests ([source](https://docs.vllm.ai/en/latest/api/vllm/vllm.engine.llm_engine.html)).

**Step 2: Insert Timing Logic**

- Use Python’s time.monotonic() or time.perf\_counter() to record timestamps before and after the prefill and decode operations. The prefill phase typically occurs in the first step for each request, while subsequent steps are decode ([metrics design](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection)).

**Step 3: Example Code Snippet**  
Below is a simplified example of how you might add timing logic. Adjust the placement based on your model’s scheduling and batch structure.

```python
import time

class LLMEngine:
    # ... existing code ...

    def step(self):
        # Start of step
        step_start = time.monotonic()
        # Schedule and execute prefill
        if self.is_prefill_phase():
            prefill_start = time.monotonic()
            # ... prefill logic ...
            prefill_end = time.monotonic()
            prefill_time = prefill_end - prefill_start
            print(f"Prefill time: {prefill_time:.6f} seconds")
        # Schedule and execute decode
        if self.is_decode_phase():
            decode_start = time.monotonic()
            # ... decode logic ...
            decode_end = time.monotonic()
            decode_time = decode_end - decode_start
            print(f"Decode time: {decode_time:.6f} seconds")
        # ... rest of step logic ...

```

**Step 4: Per-Request Timing**

- For per-request timing, store timestamps in each request’s metadata (e.g., in the sequence group or request object) and print or log them after each phase ([metrics design](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection)).

**Step 5: Validate and Benchmark**

- Run your modified vLLM and observe the printed or logged prefill and decode times for each request or batch.

Let me know if you need a more detailed breakdown or help with a specific model or batch structure.

Sources:

- [Metrics Collection Design](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection)
- [Profiling Example](https://docs.vllm.ai/en/latest/getting_started/examples/profiling.html#profiling)
- [https://docs.vllm.ai/en/latest/api/vllm/vllm.engine.llm\_engine.html](https://docs.vllm.ai/en/latest/api/vllm/vllm.engine.llm_engine.html)

---

<div class="post-metadata">

### Author: ![zhy1658858023](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/zhy1658858023/32/521_2.png) [@zhy1658858023](https://discuss.vllm.ai/u/zhy1658858023)
#### Post date: [July 24, 2025, 9:28am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/5 "2025-07-24T09:28:22Z")

</div>

> [@RunLLM](#):
>
> `is_prefill_phase`，`is_decode_phase`

Is this a real variable defined by vllm? Also STEP is triggered multiple times to perform token decoding, how do you get the stage status of Prefill and Decode in the STEP function.

---

<div class="post-metadata">

### Author: ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)
#### Post date: [July 24, 2025, 9:28am UTC](https://discuss.vllm.ai/t/does-llm-generate-differentiate-between-prefill-and-decode-phases-where-to-insert-timers/1167/6 "2025-07-24T09:28:48Z")

</div>

It seems likely that is\_prefill\_phase and is\_decode\_phase are not real variables or methods defined by vLLM; they were used as placeholders in the previous example. In vLLM, the step() function is called repeatedly, and each SequenceGroup or request tracks its own phase (prefill or decode) internally. The phase can be determined by checking the status or state of each sequence/request, such as using seq\_group.is\_prefill() or similar properties ([see code reference](https://github.com/vllm-project/vllm/blob/main/vllm/engine/llm_engine.py)).

To distinguish between prefill and decode within step(), you should inspect the status of each sequence group being processed in that iteration. For example, you can check if a sequence group is in prefill by calling seq\_group.is\_prefill(), and otherwise treat it as decode. This allows you to time the prefill and decode phases accurately for each request as step() is triggered multiple times ([metrics design](https://docs.vllm.ai/en/latest/design/v1/metrics.html#metrics-collection)).

Would you like a code snippet showing how to use seq\_group.is\_prefill() for timing?

Sources:

- [see code reference](https://github.com/vllm-project/vllm/blob/main/vllm/engine/llm_engine.py)
