Methods for Server-Side Parameter Optimization in VLLM Performance Benchmarking

1 - Overview

Since system performance is affected by the available GPU memory for inference, and the available memory for inference is in turn constrained by the server-side parameters of the VLLM framework, choosing appropriate server-side parameters during performance testing is a relatively important issue. This solution provides a method for computing VLLM server-side parameters for equal-length contexts as well as for short-input/long-output contexts, making it possible to calculate the VLLM server-side parameters that yield optimal system performance at a fixed context length.

2 - Method for Obtaining Optimal VLLM Server-Side Parameters

2.1 - Data Required for Computing Optimal VLLM Server-Side Parameters

To compute the optimal VLLM server-side parameters using this method, certain GPU memory usage data must be obtained from the server side, as shown in the table below:

Required Data Remarks
Total GPU memory Total GPU memory
GPU memory utilization GPU utilization
Model weight memory GPU memory occupied by model weights
Non_torch memory Memory not allocated by Torch (NCCL, low-level driver buffers, etc.)
Peak Activation memory Peak activation memory usage
Available KV cache memory Available KV cache size
GPU KV cache size Number of tokens that can be accommodated under the current available KV cache

2.2 - Instrumentation Methods and Locations for Each Data Item on the Server Side

Since different versions of the VLLM server produce different log outputs, the code instrumentation locations and methods for the GPU memory usage data listed in the table above are provided here together, so that everyone can print them out in a unified manner. If deployed in a container, the find command can be used to locate the relevant files.

2.2.1 - Printing Model weight memory

File location: vllm/v1/worker/gpu_model_runner.py

Print function: def load_model(self, load_dummy_weights: bool = False) -> None:

2.2.2 - Printing Non_torch memory

File location: vllm/v1/worker/gpu_worker.py

Print function: def determine_available_memory(self) -> int:

2.2.3 - Printing Peak Activation memory

File location: vllm/v1/worker/gpu_worker.py

Print function: def determine_available_memory(self) -> int:

2.2.4 - Printing Available KV cache memory

File location: vllm/v1/worker/gpu_worker.py

Print function: def determine_available_memory(self) -> int:

File location: vllm/v1/core/kv_cache_utils.py

Print function: def _report_kv_cache_config(vllm_config: VllmConfig, kv_cache_config: KVCacheConfig) -> None:

image

2.3 - Method for Computing the Optimal VLLM Server-Side Parameters

To compute the optimal parameters using this method, a total of 4 service launches are required. The needed data is recorded and then used for calculation. The 4 launches are divided into two groups: one group is used to compute the slope of the Mem-sample-dummy-run curve (i.e., the conversion coefficient between Peak Activation memory and max-num-seqs), and the other group is used to compute the slope of the Mem-model-dummy-run curve (i.e., the conversion coefficient between Peak Activation memory and max-num-batched-tokens). The groups are described below:

For convenience of expression, let Total GPU memory = TM, GPU memory utilization = U, Model weight memory = M, Non_torch memory = N, Peak Activation memory = P, Available KV cache memory = KV, and GPU KV cache size = KVS.

Group 1: (To avoid being affected by max-num-batched-tokens, Group 1 should make max-num-seqs as large as possible and max-num-batched-tokens as small as possible.)
Set a value max-num-seqs = mns1, launch the service once, and record M1, N1, P1, KV1, and KVS1;
Keeping the other parameters unchanged, set another value max-num-seqs = mns2, launch the service once, and record M2, N2, P2, KV2, and KVS2.

Group 2: (To avoid being affected by max-num-seqs, Group 2 should make max-num-seqs as small as possible and max-num-batched-tokens as large as possible.)
Set a value max-num-batched-tokens = mnbt1, launch the service once, and record M3, N3, P3, KV3, and KVS3;
Keeping the other parameters unchanged, set another value max-num-batched-tokens = mnbt2, launch the service once, and record M4, N4, P4, KV4, and KVS4.

Once the above parameters are obtained, the calculation can be performed:
A = TM * U - M1 - N1 (since M and N do not vary with the parameters, any one group of data may be chosen)
per_token_kv_cache = KV1 / KVS1 (likewise, any one group of data may be chosen)
K1 = (P4 - P3)/(mnbt2 - mnbt1)
K2 = (P2 - P1)/(mns2 - mns1)
max-num-batched-tokens = A / [(Inputlen + Outputlen) * per_token_kv_cache * K1/K2 + K1]
max-num-seqs = max-num-batched-tokens * K1/K2

To avoid rounding errors in intermediate calculations, the final computed optimal max-num-batched-tokens and max-num-seqs can be expressed as:
max-num-batched-tokens = (TM * U - M1 - N1) / { (Inputlen + Outputlen) * (KV1 / KVS1) * [(P4 - P3)(mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)] + (P4 - P3)/(mnbt2 - mnbt1) }
max-num-seqs = max-num-batched-tokens * [(P4 - P3)
(mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)]

3 - Theoretical Analysis of the Optimal VLLM Server-Side Parameters

To achieve higher system performance (end-to-end throughput), it is necessary to reduce the E2E latency while simultaneously increasing the concurrency level to raise the total number of tokens. Therefore, we want to ensure that the decode phase as a whole is not split into batches, so as to reduce latency, and on this basis to obtain the maximum concurrency. For the VLLM system, the factor that determines whether the decode phase is batched is the KV Cache size or the max-num-seqs value. The limit at which the decode phase as a whole is not batched lies in making the max-num-seqs value as large as possible, while at the same time having the decode phase exactly fill up the KV Cache at that concurrency level. Therefore, our approach is to maximize the system’s KV Cache by setting the server-side parameters, and to compute the maximum concurrency that the system can sustain under such a large KV Cache (i.e., max-num-seqs).

For the VLLM framework, the KV Cache is calculated as follows:
KV Cache = G_total_memory * K_utilization - (W_weights + Peak Act Mem + O_non-torch) (Equation 1)
where W_weights refers to the GPU memory occupied by the model weights, which is a fixed value for a system with a determined model and partitioning scheme; O_non-torch refers to memory not allocated by Torch, which is a fixed value for a given system; and Peak_Act_Mem refers to the peak activation, which is a variable. From this equation, it can be seen that for a large-model system, the smaller the GPU memory occupied by peak memory, the larger the KV Cache.

.1 - Relationship between Peak_Act_Mem and max-batch-tokens, max-num-seqs

Through analysis of the VLLM code, it can be seen that the size of Peak_Act_Mem can be expressed by the following equation:
Peak_Act_Mem = dummy-run_mem + simulated_sampling_mem (Equation 2)
simulated_sampling_mem = topk_topp_sampler + B (B represents the precision conversion to FP32, which is a fixed value) (Equation 3)
Since dummy-run_mem is linearly related to max-batch-tokens, and simulated_sampling_mem is essentially linearly related to max-num-seqs, when max-num-batched-tokens is fixed (i.e., dummy-run_mem remains unchanged), Peak_Act_Mem and max-num-seqs are linearly related.

The relationship between Peak_Act_Mem and max-batch-tokens, max-num-seqs can be illustrated in Figure 1 and Figure 2 as shown below.

Figure 1: Schematic diagram of the relationship between Peak Act Mem and Max_Num-Seqs

Figure 2: Schematic diagram of the relationship between Max_Num_Batch_Tokens and Mem-Model dummy run

Notes:
Mem-Model dummy run refers to the GPU memory occupied by forward-computation activations.
Mem-sampler dummy run refers to the GPU memory occupied by sampling-computation activations.
Peak Act Mem refers to the currently effective activation memory usage,
where Peak Act Mem = Max(Mem-Model dummy run, Mem-sampler dummy run).

Mem-Model dummy run is independent of the Max_Num_Seq configuration.
Mem-Model dummy run is directly proportional to the Max_Num_Batched_Tokens configuration.

3.2 - Derivation and Calculation of the Optimal Throughput Configuration

For higher performance, the VLLM framework is constrained by the maximum KV Cache or the max-num-seqs value. When the number of output tokens reaches the maximum KV Cache or max-num-seqs, the decode phase triggers overall batching, which increases the E2E latency (mainly reflected in TPOT, where the differences among TPOT percentiles increase), leading to a decline in performance (E2E output throughput). The configuration of max-num-seqs is therefore critical, in order to avoid the scenario where the KV Cache is not fully utilized. Filling the KV Cache as much as possible allows it to accommodate more tokens or greater concurrency, thereby improving optimal performance. Hence, point C in Figure 1 is the optimal concurrency value; at point C, Peak Act Mem is minimal and the KV Cache is maximal.

From the equation: G_total_memory * K_utilization = W_weights + Peak Act Mem + KV Cache + O_non-torch

we obtain: KV Cache = G_total_memory * K_utilization - W_weights - O_non-torch - Peak Act Mem

Since G_total_memory * K_utilization - W_weights - O_non-torch are all fixed values, we may let G_total_memory * K_utilization - W_weights - O_non-torch = A

Therefore, KV Cache = A - Peak Act Mem

As can be seen from Figure 1, the maximum value of KV Cache is obtained when Max_Num_Seqs is configured at the maximum value, point C, i.e., the intersection of Mem-Model dummy run and Mem-sampler dummy run.

At this point, Peak Act Mem = Mem-Model dummy run = Mem-sampler dummy run

As can be seen from Figure 1: Mem-Model dummy run = M = B * K1, where K1 is the slope, K1 = (M2 - M1)/(B2 - B1)

As can be seen from Figure 2: Mem-sampler dummy run = C * K2, where K2 is the slope of Mem-sampler dummy run, K2 = (P4 - P3)/(C4 - C3)

At this point, we obtain Peak Act Mem = B * K1 = C * K2

Therefore, KV Cache = A - Peak Act Mem = A - B*K1 = (Inputlen + outputlen) * Per_token_kv_cache * C

where C = B * K1/K2

We obtain A - B * K1 = (Inputlen + Outputlen) * Per_token_kv_cache * B * K1/K2

In the above equation, A, K1, Inputlen, Outputlen, Per_token_kv_cache (see Appendix 1), and K2 are all known values, while B is the unknown value.

Let H = (Inputlen + Outputlen) * Per_token_kv_cache * K1/K2

Then A - BK1 = HB

Therefore, Max_Num_Batched_Tokens = B = A/(H + K1)

Max_Num_Seqs = C = B * K1/K2 = A * K1/[(H + K1) * K2]

To obtain the optimal vLLM server-side parameters for maximizing system performance at a fixed context length, you need to empirically measure several GPU memory usage metrics (model weight memory, non-torch memory, peak activation memory, available KV cache memory, and GPU KV cache size) by launching the vLLM server multiple times with different values for max-num-seqs and max-num-batched-tokens. The relevant logging code for each metric is located in vllm/v1/worker/gpu_model_runner.py and vllm/v1/worker/gpu_worker.py, as shown in your attached code snippets and described in detail in the method section of your question. These log lines will output the required values for each run (see code locations and logging details, example log output).

Once you have the measurements, you can calculate the optimal max-num-batched-tokens and max-num-seqs using the formulas provided in your question, which are derived from the linear relationships between peak activation memory and these parameters. The formulas are:

  • max-num-batched-tokens = (TM * U - M1 - N1) / { (Inputlen + Outputlen) * (KV1 / KVS1) * [(P4 - P3) (mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)] + (P4 - P3)/(mnbt2 - mnbt1) }
  • max-num-seqs = max-num-batched-tokens * [(P4 - P3) (mns2 - mns1) / (P2 - P1) / (mnbt2 - mnbt1)]

This approach ensures the decode phase is not unnecessarily batched and the KV cache is maximally utilized, yielding the highest concurrency and throughput for your workload (see theoretical analysis and formula derivation, memory profiling explanation).

Would you like a step-by-step worked example using sample log outputs?

Sources:

Hint: Mention @Herald in the post for followups.