|
Prompt_embeds usage in vllm openai completion api
|
|
4
|
276
|
June 17, 2025
|
|
W8a8两种量化方式有详细介绍吗
|
|
1
|
257
|
June 15, 2025
|
|
Seqence Parallelism Support - Source Code Location
|
|
0
|
53
|
June 10, 2025
|
|
Something weired about the reading procedure of q_vecs in page attention kernel
|
|
3
|
67
|
June 9, 2025
|
|
Computation time remain consistent across chunks in chunked-prefill despite linearly growing attention complexity?
|
|
1
|
147
|
June 2, 2025
|
|
Why does computation time remain consistent across chunks in chunked-prefill despite linearly growing attention complexity?
|
|
3
|
205
|
June 2, 2025
|
|
APC Slowdown with block-size=1
|
|
1
|
97
|
May 26, 2025
|
|
RunBot's math-to-text on NVIDIA NeMo Framework AutoModel
|
|
11
|
280
|
May 19, 2025
|
|
Issue with DynamicYaRN and Key-Value Cache Reuse in vLLM
|
|
1
|
182
|
May 18, 2025
|
|
VUA - library code for LLM inference engines for external storage of KV caches
|
|
1
|
157
|
May 13, 2025
|
|
Specifying special tokens
|
|
5
|
744
|
May 8, 2025
|
|
vLLM cannot connect to existing Ray cluster
|
|
16
|
1384
|
May 8, 2025
|
|
Support for (sparse) key value caching
|
|
16
|
943
|
May 3, 2025
|
|
How to use speculative decoding?
|
|
3
|
1203
|
May 1, 2025
|
|
Grammar CPU bound performance
|
|
9
|
686
|
April 29, 2025
|
|
Does vLLM support multiple model_executor?
|
|
1
|
451
|
April 28, 2025
|
|
Spec decode with eagle get very low Draft acceptance rate
|
|
1
|
477
|
April 25, 2025
|
|
LoRA Adapter enabling with vLLM is not working
|
|
4
|
653
|
April 21, 2025
|
|
Goodput Guided Speculative Decoding
|
|
2
|
301
|
April 19, 2025
|
|
Is structured output compatible with automatic prefix caching?
|
|
1
|
164
|
April 14, 2025
|
|
Tool calling using Offline Inference?
|
|
1
|
209
|
April 14, 2025
|
|
How to crop kv_caches?
|
|
0
|
84
|
April 13, 2025
|
|
Minimum requirements for Disaggregated Prefilling?
|
|
0
|
117
|
April 9, 2025
|
|
Can Lora adapters be loaded on different GPUs
|
|
1
|
141
|
April 7, 2025
|
|
Do we have regression tests for structured output? Especially speed regression?
|
|
0
|
74
|
March 31, 2025
|
|
Why remove bonus token of requset in draft model?
|
|
0
|
67
|
March 30, 2025
|
|
Why is the prefix cache hit rate constantly increasing
|
|
3
|
1521
|
March 27, 2025
|
|
Multiple tools with Mistral Large 2411
|
|
4
|
399
|
March 26, 2025
|
|
Pipeline Parallelism Support - Source Code Location
|
|
1
|
217
|
March 25, 2025
|
|
GGUF quantized models Inference support
|
|
0
|
394
|
March 25, 2025
|