|
About the Features category
|
|
0
|
95
|
March 20, 2025
|
|
Sparse attention (e.g., H2O)
|
|
2
|
156
|
July 1, 2026
|
|
Is there a hook/flag to capture activation statistics during inference for use with llm-compressor AWQ?
|
|
3
|
71
|
June 4, 2026
|
|
Dose vllm support Qwen3.5 pd disaggregation with Mooncake?
|
|
1
|
105
|
May 28, 2026
|
|
Why Does Decode Forward on PP Stage 0 Appear to Precede Prefill Forward on PP Stage 1 for the Same Request?
|
|
1
|
42
|
May 26, 2026
|
|
GPTQModel 能量化 GLM-5 FP16 到 INT8 吗
|
|
9
|
181
|
April 24, 2026
|
|
DeepSeek MTP full cuda graph support?
|
|
3
|
167
|
April 13, 2026
|
|
Qwen3.5-27B-FP8 Speculative Decoding
|
|
2
|
2297
|
April 11, 2026
|
|
thinking_token_budget silently ignored when passed via extra_args in vLLM 0.18.0
|
|
1
|
653
|
April 11, 2026
|
|
GLM 5 / Kimi k2.5 on 4 x RTX 6000 Pro
|
|
1
|
501
|
March 22, 2026
|
|
Compressed Multimodal embeddings inputs
|
|
1
|
74
|
March 18, 2026
|
|
NVFP4 Support In Attention
|
|
1
|
1086
|
March 16, 2026
|
|
Distributed Speculative Decoding using Ray
|
|
3
|
229
|
February 11, 2026
|
|
Deployment example for a qwen3 model with hybrid thinking
|
|
10
|
2313
|
February 4, 2026
|
|
How do I precompute multimodal embeddings?
|
|
5
|
534
|
February 2, 2026
|
|
Implementing hidden state probes
|
|
1
|
94
|
January 30, 2026
|
|
Standalone draft model spec decode support in v0.x and v1
|
|
3
|
294
|
January 20, 2026
|
|
How to get kv cache value from vllm
|
|
5
|
401
|
January 19, 2026
|
|
Is there a plan for EVS to support Qwen3VL in response to the issue of sparse video tokens?
|
|
1
|
155
|
January 13, 2026
|
|
Exposing KV cache for recomposition / reuse beyond prefix caching?
|
|
1
|
217
|
January 13, 2026
|
|
Why I feel cuda-kernel marlin run not fast?
|
|
5
|
331
|
January 9, 2026
|
|
Can reasoning_effort parameter not ne used in vllm implementation via python?
|
|
1
|
509
|
January 2, 2026
|
|
Understanding vllm kv cache
|
|
5
|
2068
|
December 1, 2025
|
|
Has anyone successfully run DBO in a single node multi card environment?
|
|
1
|
139
|
December 1, 2025
|
|
EPLB behavior in elastic scaling
|
|
21
|
395
|
November 28, 2025
|
|
Qwen2.5 VL开启flashinfer失败
|
|
5
|
440
|
November 24, 2025
|
|
如何提升在单机多卡部署时的吞吐量
|
|
10
|
1081
|
November 24, 2025
|
|
Does LMFE_STRICT_JSON_FIELD_ORDER not work?
|
|
2
|
140
|
November 19, 2025
|
|
Clarification Needed on Testing Elastic EP and Its Library Installation Dependencies
|
|
1
|
47
|
November 19, 2025
|
|
Custom modality
|
|
3
|
90
|
November 14, 2025
|