|
Welcome to vLLM Forums! :wave:
|
|
1
|
2053
|
March 24, 2025
|
|
About the General category
|
|
0
|
149
|
March 17, 2025
|
|
vLLM 0.20.1, Radeon AI 9700, 1 CPU core at 100%
|
|
6
|
554
|
September 13, 2026
|
|
Muse glimmer - how to disable thinking
|
|
1
|
35
|
September 11, 2026
|
|
Not able to run amd/Qwen3.8-27B-Quark-AWQ-INT4-W4A16 on rocm
|
|
1
|
51
|
September 9, 2026
|
|
[RFC & PR Stack] Application-Directed Prefix Checkpoints for Hybrid GDN/Mamba (Up to 7.6x Speedup on L40S)
|
|
5
|
56
|
September 9, 2026
|
|
vLLM-Ascend在银河麒麟桌面版上不开Eager就报错,开了性能暴跌,换服务器版就正常,这是系统差异导致的吗?
|
|
1
|
69
|
September 4, 2026
|
|
我用单卡跑通了deepseek-v4-flash-0731的dspark模式
|
|
1
|
201
|
August 22, 2026
|
|
Flash-0731 在 vLLM 0.27.1 + SM120 上 dspark 不可启动
|
|
1
|
66
|
August 21, 2026
|
|
Why does changing VLLM_CACHE_ROOT not reproduce cold-start latency?
|
|
1
|
76
|
August 21, 2026
|
|
Model path in compute_hash() prevents torch.compile cache reuse across identical model loads
|
|
1
|
46
|
August 21, 2026
|
|
Mooncake prefix information in storage
|
|
3
|
74
|
August 21, 2026
|
|
Vllm是否支持加载deepseek_ocr模型的lora适配器进行推理
|
|
3
|
173
|
August 21, 2026
|
|
Dlash2 無法使用 在 vllm 0.27.1
|
|
1
|
182
|
August 19, 2026
|
|
Expert Parallelism All-to-All Communication without NVLink and DeepEP
|
|
4
|
724
|
August 17, 2026
|
|
FLASHINFER is FASTER then FA2 on Ampere hardware
|
|
5
|
135
|
August 16, 2026
|
|
Support for Qwen3.6 (27B & 35B-A3B) Hybrid Gated DeltaNet Models on TPU v5e-8 with vLLM and SGLang
|
|
11
|
418
|
August 14, 2026
|
|
How fast can you run DeepSeek V4 Spark 0731 with CPU-GPU hybrid inference when VRAM is not enough?
|
|
1
|
226
|
August 13, 2026
|
|
(Possible bug) --enable-prompt-tokens-details not working?
|
|
2
|
599
|
August 13, 2026
|
|
How use profile file to optimize the vllm inference as a model?
|
|
1
|
62
|
August 13, 2026
|
|
vllm部署minimax模型
|
|
1
|
87
|
August 11, 2026
|
|
Proxima - vLLM plugin for StarKV (3-4x more kv cache blocks)
|
|
1
|
55
|
August 10, 2026
|
|
Float32 Qwen3.5 4B Load
|
|
2
|
79
|
August 5, 2026
|
|
Community plugin library
|
|
1
|
63
|
August 5, 2026
|
|
vLLM - 1 Click Deployment on Cloudron
|
|
1
|
53
|
July 30, 2026
|
|
As of July 2026, what are some general tips in improving the speed of guided decoding in vllm?
|
|
2
|
129
|
July 29, 2026
|
|
Vllm lmcache server - how to run
|
|
11
|
305
|
July 28, 2026
|
|
Why does increasing max-num-seqs cause a runtime CUDA OOM?
|
|
1
|
107
|
July 28, 2026
|
|
Vllm omni and krea2
|
|
19
|
249
|
July 27, 2026
|
|
Vllm omni and 2x 5090
|
|
9
|
181
|
July 26, 2026
|