# Methods for Server-Side Parameter Optimization in VLLM Performance Benchmarking

**URL:** <https://discuss.vllm.ai/t/methods-for-server-side-parameter-optimization-in-vllm-performance-benchmarking/3012>\
**Category:** General\
**Created:** [September 29, 2026, 9:04am UTC](https://discuss.vllm.ai/t/methods-for-server-side-parameter-optimization-in-vllm-performance-benchmarking/3012 "2026-09-29T09:04:08Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![xinbaby-1](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/xinbaby-1/32/1774_2.png) [@xinbaby-1](https://discuss.vllm.ai/u/xinbaby-1)\
**Post date:** [September 29, 2026, 9:04am UTC](https://discuss.vllm.ai/t/methods-for-server-side-parameter-optimization-in-vllm-performance-benchmarking/3012/1 "2026-09-29T09:04:09Z")

</div>

**1 - Overview**

Since system performance is affected by the available GPU memory for inference, and the available memory for inference is in turn constrained by the server-side parameters of the VLLM framework, choosing appropriate server-side parameters during performance testing is a relatively important issue. This solution provides a method for computing VLLM server-side parameters for equal-length contexts as well as for short-input/long-output contexts, making it possible to calculate the VLLM server-side parameters that yield optimal system performance at a fixed context length.

**2 - Method for Obtaining Optimal VLLM Server-Side Parameters**

**2.1 - Data Required for Computing Optimal VLLM Server-Side Parameters**

To compute the optimal VLLM server-side parameters using this method, certain GPU memory usage data must be obtained from the server side, as shown in the table below:

| Required Data | Remarks |
| --- | --- |
| Total GPU memory | Total GPU memory |
| GPU memory utilization | GPU utilization |
| Model weight memory | GPU memory occupied by model weights |
| Non\_torch memory | Memory not allocated by Torch (NCCL, low-level driver buffers, etc.) |
| Peak Activation memory | Peak activation memory usage |
| Available KV cache memory | Available KV cache size |
| GPU KV cache size | Number of tokens that can be accommodated under the current available KV cache |

**2.2 - Instrumentation Methods and Locations for Each Data Item on the Server Side**

Since different versions of the VLLM server produce different log outputs, the code instrumentation locations and methods for the GPU memory usage data listed in the table above are provided here together, so that everyone can print them out in a unified manner. If deployed in a container, the `find` command can be used to locate the relevant files.

**2.2.1 - Printing Model weight memory**

File location: `vllm/v1/worker/gpu_model_runner.py`

Print function: `def load_model(self, load_dummy_weights: bool = False) -> None:`

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/e/ec645cf3ca06a6db43d4f161d389230dcf8e973e.png)

**2.2.2 - Printing Non\_torch memory**

File location: `vllm/v1/worker/gpu_worker.py`

Print function: `def determine_available_memory(self) -> int:`

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/d/d400b1d717f51b6e2d622d63ed6f7ff162fcaf92.png)

**2.2.3 - Printing Peak Activation memory**

File location: `vllm/v1/worker/gpu_worker.py`

Print function: `def determine_available_memory(self) -> int:`

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/f/f7272f005265c3b9201a7db4c3b01c1a7411dc88.png)

**2.2.4 - Printing Available KV cache memory**

File location: `vllm/v1/worker/gpu_worker.py`

Print function: `def determine_available_memory(self) -> int:`

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/e/e1672754357d38acec9f90ec98ec345420fe3ad2.png)

File location: `vllm/v1/core/kv_cache_utils.py`

Print function: `def _report_kv_cache_config(vllm_config: VllmConfig, kv_cache_config: KVCacheConfig) -> None:`

![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/9/9d813fb2022810210e3b9ddc603ad6e2aeaacafa.png)

**2.3 - Method for Computing the Optimal VLLM Server-Side Parameters**

To compute the optimal parameters using this method, a total of 4 service launches are required. The needed data is recorded and then used for calculation. The 4 launches are divided into two groups: one group is used to compute the slope of the Mem-sample-dummy-run curve (i.e., the conversion coefficient between Peak Activation memory and max-num-seqs), and the other group is used to compute the slope of the Mem-model-dummy-run curve (i.e., the conversion coefficient between Peak Activation memory and max-num-batched-tokens). The groups are described below:

For convenience of expression, let Total GPU memory = TM, GPU memory utilization = U, Model weight memory = M, Non\_torch memory = N, Peak Activation memory = P, Available KV cache memory = KV, and GPU KV cache size = KVS.

**Group 1:** (To avoid being affected by max-num-batched-tokens, Group 1 should make max-num-seqs as large as possible and max-num-batched-tokens as small as possible.)  
Set a value max-num-seqs = mns1, launch the service once, and record M1, N1, P1, KV1, and KVS1;  
Keeping the other parameters unchanged, set another value max-num-seqs = mns2, launch the service once, and record M2, N2, P2, KV2, and KVS2.

**Group 2:** (To avoid being affected by max-num-seqs, Group 2 should make max-num-seqs as small as possible and max-num-batched-tokens as large as possible.)  
Set a value max-num-batched-tokens = mnbt1, launch the service once, and record M3, N3, P3, KV3, and KVS3;  
Keeping the other parameters unchanged, set another value max-num-batched-tokens = mnbt2, launch the service once, and record M4, N4, P4, KV4, and KVS4.

Once the above parameters are obtained, the calculation can be performed:  
A = TM \* U - M1 - N1 (since M and N do not vary with the parameters, any one group of data may be chosen)  
per\_token\_kv\_cache = KV1 / KVS1 (likewise, any one group of data may be chosen)  
K1 = (P4 - P3)/(mnbt2 - mnbt1)  
K2 = (P2 - P1)/(mns2 - mns1)  
max-num-batched-tokens = A / [(Inputlen + Outputlen) \* per\_token\_kv\_cache \* K1/K2 + K1]  
max-num-seqs = max-num-batched-tokens \* K1/K2

To avoid rounding errors in intermediate calculations, the final computed optimal max-num-batched-tokens and max-num-seqs can be expressed as:  
max-num-batched-tokens = (TM \* U - M1 - N1) / { (Inputlen + Outputlen) \* (KV1 / KVS1) \* [(P4 - P3)_(mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)] + (P4 - P3)/(mnbt2 - mnbt1) }  
max-num-seqs = max-num-batched-tokens \* [(P4 - P3)_(mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)]

**3 - Theoretical Analysis of the Optimal VLLM Server-Side Parameters**

To achieve higher system performance (end-to-end throughput), it is necessary to reduce the E2E latency while simultaneously increasing the concurrency level to raise the total number of tokens. Therefore, we want to ensure that the decode phase as a whole is not split into batches, so as to reduce latency, and on this basis to obtain the maximum concurrency. For the VLLM system, the factor that determines whether the decode phase is batched is the KV Cache size or the max-num-seqs value. The limit at which the decode phase as a whole is not batched lies in making the max-num-seqs value as large as possible, while at the same time having the decode phase exactly fill up the KV Cache at that concurrency level. Therefore, our approach is to maximize the system’s KV Cache by setting the server-side parameters, and to compute the maximum concurrency that the system can sustain under such a large KV Cache (i.e., max-num-seqs).

For the VLLM framework, the KV Cache is calculated as follows:  
KV Cache = G\_total\_memory \* K\_utilization - (W\_weights + Peak Act Mem + O\_non-torch) (Equation 1)  
where W\_weights refers to the GPU memory occupied by the model weights, which is a fixed value for a system with a determined model and partitioning scheme; O\_non-torch refers to memory not allocated by Torch, which is a fixed value for a given system; and Peak\_Act\_Mem refers to the peak activation, which is a variable. From this equation, it can be seen that for a large-model system, the smaller the GPU memory occupied by peak memory, the larger the KV Cache.

**.1 - Relationship between Peak\_Act\_Mem and max-batch-tokens, max-num-seqs**

Through analysis of the VLLM code, it can be seen that the size of Peak\_Act\_Mem can be expressed by the following equation:  
Peak\_Act\_Mem = dummy-run\_mem + simulated\_sampling\_mem (Equation 2)  
simulated\_sampling\_mem = topk\_topp\_sampler + B (B represents the precision conversion to FP32, which is a fixed value) (Equation 3)  
Since dummy-run\_mem is linearly related to max-batch-tokens, and simulated\_sampling\_mem is essentially linearly related to max-num-seqs, when max-num-batched-tokens is fixed (i.e., dummy-run\_mem remains unchanged), Peak\_Act\_Mem and max-num-seqs are linearly related.

The relationship between Peak\_Act\_Mem and max-batch-tokens, max-num-seqs can be illustrated in Figure 1 and Figure 2 as shown below.

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/4/43334e234cd0dc4f579b0052cd36d2103cc1f1ec.png)

Figure 1: Schematic diagram of the relationship between Peak Act Mem and Max\_Num-Seqs

 ![image](https://canada1.discourse-cdn.com/flex036/uploads/vllm/original/2X/0/0d685097a20b95786e745624ec654f3e1e2cb2a5.png)

Figure 2: Schematic diagram of the relationship between Max\_Num\_Batch\_Tokens and Mem-Model dummy run

Notes:  
Mem-Model dummy run refers to the GPU memory occupied by forward-computation activations.  
Mem-sampler dummy run refers to the GPU memory occupied by sampling-computation activations.  
Peak Act Mem refers to the currently effective activation memory usage,  
where Peak Act Mem = Max(Mem-Model dummy run, Mem-sampler dummy run).

Mem-Model dummy run is independent of the Max\_Num\_Seq configuration.  
Mem-Model dummy run is directly proportional to the Max\_Num\_Batched\_Tokens configuration.

**3.2 - Derivation and Calculation of the Optimal Throughput Configuration**

For higher performance, the VLLM framework is constrained by the maximum KV Cache or the max-num-seqs value. When the number of output tokens reaches the maximum KV Cache or max-num-seqs, the decode phase triggers overall batching, which increases the E2E latency (mainly reflected in TPOT, where the differences among TPOT percentiles increase), leading to a decline in performance (E2E output throughput). The configuration of max-num-seqs is therefore critical, in order to avoid the scenario where the KV Cache is not fully utilized. Filling the KV Cache as much as possible allows it to accommodate more tokens or greater concurrency, thereby improving optimal performance. Hence, point C in Figure 1 is the optimal concurrency value; at point C, Peak Act Mem is minimal and the KV Cache is maximal.

From the equation: G\_total\_memory \* K\_utilization = W\_weights + Peak Act Mem + KV Cache + O\_non-torch

we obtain: KV Cache = G\_total\_memory \* K\_utilization - W\_weights - O\_non-torch - Peak Act Mem

Since G\_total\_memory \* K\_utilization - W\_weights - O\_non-torch are all fixed values, we may let G\_total\_memory \* K\_utilization - W\_weights - O\_non-torch = A

Therefore, KV Cache = A - Peak Act Mem

As can be seen from Figure 1, the maximum value of KV Cache is obtained when Max\_Num\_Seqs is configured at the maximum value, point C, i.e., the intersection of Mem-Model dummy run and Mem-sampler dummy run.

At this point, Peak Act Mem = Mem-Model dummy run = Mem-sampler dummy run

As can be seen from Figure 1: Mem-Model dummy run = M = B \* K1, where K1 is the slope, K1 = (M2 - M1)/(B2 - B1)

As can be seen from Figure 2: Mem-sampler dummy run = C \* K2, where K2 is the slope of Mem-sampler dummy run, K2 = (P4 - P3)/(C4 - C3)

At this point, we obtain Peak Act Mem = B \* K1 = C \* K2

Therefore, KV Cache = A - Peak Act Mem = A - B\*K1 = (Inputlen + outputlen) \* Per\_token\_kv\_cache \* C

where C = B \* K1/K2

We obtain A - B \* K1 = (Inputlen + Outputlen) \* Per\_token\_kv\_cache \* B \* K1/K2

In the above equation, A, K1, Inputlen, Outputlen, Per\_token\_kv\_cache (see Appendix 1), and K2 are all known values, while B is the unknown value.

Let H = (Inputlen + Outputlen) \* Per\_token\_kv\_cache \* K1/K2

Then A - B_K1 = H_B

Therefore, Max\_Num\_Batched\_Tokens = B = A/(H + K1)

Max\_Num\_Seqs = C = B \* K1/K2 = A \* K1/[(H + K1) \* K2]

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [September 29, 2026, 9:04am UTC](https://discuss.vllm.ai/t/methods-for-server-side-parameter-optimization-in-vllm-performance-benchmarking/3012/2 "2026-09-29T09:04:38Z")

</div>

To obtain the optimal vLLM server-side parameters for maximizing system performance at a fixed context length, you need to empirically measure several GPU memory usage metrics (model weight memory, non-torch memory, peak activation memory, available KV cache memory, and GPU KV cache size) by launching the vLLM server multiple times with different values for max-num-seqs and max-num-batched-tokens. The relevant logging code for each metric is located in vllm/v1/worker/gpu\_model\_runner.py and vllm/v1/worker/gpu\_worker.py, as shown in your attached code snippets and described in detail in the method section of your question. These log lines will output the required values for each run ([see code locations and logging details](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py), [example log output](https://github.com/vllm-project/vllm/issues/11550)).

Once you have the measurements, you can calculate the optimal max-num-batched-tokens and max-num-seqs using the formulas provided in your question, which are derived from the linear relationships between peak activation memory and these parameters. The formulas are:

- max-num-batched-tokens = (TM \* U - M1 - N1) / { (Inputlen + Outputlen) \* (KV1 / KVS1) \* [(P4 - P3) (mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)] + (P4 - P3)/(mnbt2 - mnbt1) }
- max-num-seqs = max-num-batched-tokens \* [(P4 - P3) (mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)]

This approach ensures the decode phase is not unnecessarily batched and the KV cache is maximally utilized, yielding the highest concurrency and throughput for your workload ([see theoretical analysis and formula derivation](https://github.com/vllm-project/vllm/issues/18582), [memory profiling explanation](https://docs.vllm.ai/en/latest/api/vllm/utils/mem_utils/#__span-0-4)).

Would you like a step-by-step worked example using sample log outputs?

Sources:

- [see code locations and logging details](https://github.com/vllm-project/vllm/blob/main/vllm/v1/worker/gpu_worker.py)
- [example log output](https://github.com/vllm-project/vllm/issues/11550)
- [see theoretical analysis and formula derivation](https://github.com/vllm-project/vllm/issues/18582)
- [memory profiling explanation](https://docs.vllm.ai/en/latest/api/vllm/utils/mem_utils/#__span-0-4)

Hint: Mention @Herald in the post for followups.
