|
Welcome to vLLM Forums! :wave:
|
|
1
|
1940
|
March 24, 2025
|
|
About the General category
|
|
0
|
127
|
March 17, 2025
|
|
FLASHINFER is FASTER then FA2 on Ampere hardware
|
|
5
|
44
|
August 16, 2026
|
|
Support for Qwen3.6 (27B & 35B-A3B) Hybrid Gated DeltaNet Models on TPU v5e-8 with vLLM and SGLang
|
|
11
|
81
|
August 14, 2026
|
|
How fast can you run DeepSeek V4 Spark 0731 with CPU-GPU hybrid inference when VRAM is not enough?
|
|
1
|
59
|
August 13, 2026
|
|
(Possible bug) --enable-prompt-tokens-details not working?
|
|
2
|
427
|
August 13, 2026
|
|
How use profile file to optimize the vllm inference as a model?
|
|
1
|
28
|
August 13, 2026
|
|
vllm部署minimax模型
|
|
1
|
40
|
August 11, 2026
|
|
Proxima - vLLM plugin for StarKV (3-4x more kv cache blocks)
|
|
1
|
36
|
August 10, 2026
|
|
Float32 Qwen3.5 4B Load
|
|
2
|
58
|
August 5, 2026
|
|
Community plugin library
|
|
1
|
55
|
August 5, 2026
|
|
vLLM - 1 Click Deployment on Cloudron
|
|
1
|
45
|
July 30, 2026
|
|
As of July 2026, what are some general tips in improving the speed of guided decoding in vllm?
|
|
2
|
91
|
July 29, 2026
|
|
Vllm lmcache server - how to run
|
|
11
|
197
|
July 28, 2026
|
|
Why does increasing max-num-seqs cause a runtime CUDA OOM?
|
|
2
|
91
|
July 28, 2026
|
|
Vllm omni and krea2
|
|
19
|
184
|
July 27, 2026
|
|
Vllm omni and 2x 5090
|
|
9
|
121
|
July 26, 2026
|
|
Structural GPU waste in DeepSeek-R1 inference anyone else measuring this
|
|
1
|
29
|
July 23, 2026
|
|
Does vllm *need* a restart once in a while?
|
|
1
|
73
|
July 23, 2026
|
|
PR #49434: Fix Poolside reasoning token initialization
|
|
1
|
86
|
July 22, 2026
|
|
Network Backend Simulation for vLLM
|
|
0
|
75
|
July 21, 2026
|
|
Registering with Transformers Auto Classes for out-of-tree models
|
|
4
|
324
|
July 15, 2026
|
|
LLM-Scaler vs. vllm
|
|
1
|
196
|
July 11, 2026
|
|
How can I use vllm with cuda 12.9
|
|
3
|
406
|
July 10, 2026
|
|
Should request priority also influence KV cache retention?
|
|
1
|
80
|
July 9, 2026
|
|
Configuration for PD disaggregation
|
|
2
|
107
|
July 7, 2026
|
|
Running out of memory despite very low `--gpu-memory-utilization`
|
|
2
|
152
|
July 5, 2026
|
|
HIP failure: the operation cannot be performed in the present state
|
|
7
|
227
|
July 2, 2026
|
|
Is there a way to perform DeepSeek-V4-Flash inference using vLLM on CUDA 12.6 and 8×H800?
|
|
1
|
146
|
July 1, 2026
|
|
Should vLLM consider prefix caching when chunked prefill is enabled?
|
|
2
|
822
|
June 30, 2026
|