|
About the Features category
|
|
0
|
134
|
March 20, 2025
|
|
Decode KV writeback: can multi-turn chat preserve cache reuse with default templates?
|
|
1
|
41
|
September 10, 2026
|
|
vLLM TurboQuant Implementation
|
|
1
|
57
|
September 7, 2026
|
|
Customize prefix caching
|
|
1
|
104
|
August 9, 2026
|
|
Sparse attention (e.g., H2O)
|
|
2
|
272
|
July 1, 2026
|
|
Is there a hook/flag to capture activation statistics during inference for use with llm-compressor AWQ?
|
|
3
|
121
|
June 4, 2026
|
|
Dose vllm support Qwen3.5 pd disaggregation with Mooncake?
|
|
1
|
179
|
May 28, 2026
|
|
Why Does Decode Forward on PP Stage 0 Appear to Precede Prefill Forward on PP Stage 1 for the Same Request?
|
|
1
|
61
|
May 26, 2026
|
|
GPTQModel 能量化 GLM-5 FP16 到 INT8 吗
|
|
9
|
243
|
April 24, 2026
|
|
DeepSeek MTP full cuda graph support?
|
|
3
|
243
|
April 13, 2026
|
|
Qwen3.5-27B-FP8 Speculative Decoding
|
|
2
|
2574
|
April 11, 2026
|
|
thinking_token_budget silently ignored when passed via extra_args in vLLM 0.18.0
|
|
1
|
972
|
April 11, 2026
|
|
GLM 5 / Kimi k2.5 on 4 x RTX 6000 Pro
|
|
1
|
561
|
March 22, 2026
|
|
Compressed Multimodal embeddings inputs
|
|
1
|
95
|
March 18, 2026
|
|
NVFP4 Support In Attention
|
|
1
|
1538
|
March 16, 2026
|
|
Distributed Speculative Decoding using Ray
|
|
3
|
296
|
February 11, 2026
|
|
Deployment example for a qwen3 model with hybrid thinking
|
|
10
|
2819
|
February 4, 2026
|
|
How do I precompute multimodal embeddings?
|
|
5
|
728
|
February 2, 2026
|
|
Implementing hidden state probes
|
|
1
|
134
|
January 30, 2026
|
|
Standalone draft model spec decode support in v0.x and v1
|
|
3
|
373
|
January 20, 2026
|
|
How to get kv cache value from vllm
|
|
5
|
527
|
January 19, 2026
|
|
Is there a plan for EVS to support Qwen3VL in response to the issue of sparse video tokens?
|
|
1
|
187
|
January 13, 2026
|
|
Exposing KV cache for recomposition / reuse beyond prefix caching?
|
|
1
|
289
|
January 13, 2026
|
|
Why I feel cuda-kernel marlin run not fast?
|
|
5
|
456
|
January 9, 2026
|
|
Can reasoning_effort parameter not ne used in vllm implementation via python?
|
|
1
|
735
|
January 2, 2026
|
|
Understanding vllm kv cache
|
|
5
|
2711
|
December 1, 2025
|
|
Has anyone successfully run DBO in a single node multi card environment?
|
|
1
|
169
|
December 1, 2025
|
|
EPLB behavior in elastic scaling
|
|
21
|
537
|
November 28, 2025
|
|
Qwen2.5 VL开启flashinfer失败
|
|
5
|
511
|
November 24, 2025
|
|
如何提升在单机多卡部署时的吞吐量
|
|
10
|
1522
|
November 24, 2025
|