|
Welcome to vLLM Forums! :wave:
|
|
1
|
1818
|
March 24, 2025
|
|
About the General category
|
|
0
|
116
|
March 17, 2025
|
|
Registering with Transformers Auto Classes for out-of-tree models
|
|
4
|
271
|
July 15, 2026
|
|
LLM-Scaler vs. vllm
|
|
1
|
72
|
July 11, 2026
|
|
How can I use vllm with cuda 12.9
|
|
3
|
147
|
July 10, 2026
|
|
Should request priority also influence KV cache retention?
|
|
1
|
39
|
July 9, 2026
|
|
Configuration for PD disaggregation
|
|
2
|
52
|
July 7, 2026
|
|
Running out of memory despite very low `--gpu-memory-utilization`
|
|
2
|
71
|
July 5, 2026
|
|
HIP failure: the operation cannot be performed in the present state
|
|
7
|
139
|
July 2, 2026
|
|
Is there a way to perform DeepSeek-V4-Flash inference using vLLM on CUDA 12.6 and 8×H800?
|
|
1
|
79
|
July 1, 2026
|
|
Should vLLM consider prefix caching when chunked prefill is enabled?
|
|
2
|
737
|
June 30, 2026
|
|
Trying to passtrough 2 7900 XTX to proxmox ubuntu VM
|
|
13
|
144
|
June 27, 2026
|
|
Vllm serve拉起服务,到最后报错了
|
|
1
|
64
|
June 27, 2026
|
|
Window10 wsl2下不能使用vLLM的睡眠模式吗
|
|
1
|
78
|
June 27, 2026
|
|
Speech To Text Guidance
|
|
1
|
76
|
June 26, 2026
|
|
What is the correct chat template when serving gemma4?
|
|
2
|
786
|
June 25, 2026
|
|
Seeking Help: DeepSeek-V4-Flash Fails to Deploy on 8×4090 (48GB VRAM per Card)
|
|
1
|
292
|
June 25, 2026
|
|
Sparse Embedding Support
|
|
2
|
86
|
June 24, 2026
|
|
No available shared memory broadcast block found in 60 seconds
|
|
1
|
379
|
June 22, 2026
|
|
Max_model_len vs GPU memory Usage
|
|
2
|
157
|
June 22, 2026
|
|
Ahead of time compilation of "CUDA Kernels"?
|
|
3
|
123
|
June 21, 2026
|
|
The current vLLM CPU backend is not working properly
|
|
9
|
347
|
June 20, 2026
|
|
What happened to 0.23.0 release?
|
|
1
|
322
|
June 15, 2026
|
|
vLLM的MTP的标准测试方法是什么?
|
|
2
|
116
|
June 13, 2026
|
|
Vllm rocm gemma4 and MTP
|
|
9
|
269
|
June 12, 2026
|
|
PR backlog on GitHub
|
|
2
|
71
|
June 12, 2026
|
|
Rust cannot catch foreign exceptions
|
|
1
|
53
|
June 11, 2026
|
|
Do we Have ChatBots Here
|
|
2
|
69
|
June 11, 2026
|
|
Vllm服务后台的前缀缓存率和单个请求的前缀缓存率
|
|
1
|
80
|
June 11, 2026
|
|
RuntimeError: CUDA driver initialization failed, you might not have a CUDA gpu
|
|
1
|
102
|
June 5, 2026
|