|
使用vllm_ascend0.9.1提示Failed to import vllm_ascend_C:
|
|
2
|
276
|
November 20, 2025
|
|
[Questions]Is there a plan to support the rerank model and embedding model
|
|
3
|
1249
|
November 20, 2025
|
|
Vllm ascend能够查看模型的计算图吗
|
|
2
|
198
|
November 20, 2025
|
|
When will the next class be released?
|
|
2
|
121
|
November 20, 2025
|
|
Can Support Qwen3-VL or Qwen2.5 VL 72B on Vllm-ascend 0.11.0?
|
|
2
|
314
|
November 20, 2025
|
|
Vllm-ascend是否支持mineru2.5
|
|
2
|
203
|
November 20, 2025
|
|
RuntimeError: Int8 not supported on SM120. Use FP8 quantization instead, or run on older arch (SM < 100)
|
|
1
|
321
|
November 19, 2025
|
|
Need help compiling and running on Jetson Thor
|
|
4
|
1046
|
November 1, 2025
|
|
How can vllm ascend support qwen3-vl-235b?
|
|
2
|
377
|
October 16, 2025
|
|
我能在Ascend310B芯片上通过vllm-ascend插件部署Qwen2.5-vl吗?
|
|
3
|
249
|
October 15, 2025
|
|
Vllm-ascend是否支持async推理?
|
|
2
|
169
|
October 15, 2025
|
|
昇腾920b是否支持通义千问2.5-vl
|
|
2
|
192
|
October 2, 2025
|
|
RTX Pro 6000 Tensor Parallelism CUBLAS_STATUS_ALLOC_FAILED
|
|
3
|
600
|
September 13, 2025
|
|
Is there any plan to organize the cuda-only configuration
|
|
1
|
72
|
August 15, 2025
|
|
Unable to use vLLM 0.10.1-gptoss on GH200 (aarch64) — source for custom wheel not available?
|
|
3
|
618
|
August 15, 2025
|
|
vLLM Benchmarking: Why Is GPUDirect RDMA Not Outperforming Standard RDMA in a Pipeline-Parallel Setup?
|
|
1
|
770
|
August 14, 2025
|
|
Why does quickReduce not need to use system-scope release write operations to update flags?
|
|
0
|
43
|
August 13, 2025
|
|
Can vLLM built for old GPU (GT 630M) ? It may use CUDA 9.1.85
|
|
1
|
452
|
August 4, 2025
|
|
How to deploy vllm-ascend in AutoDL's 910B instance?
|
|
7
|
582
|
August 2, 2025
|
|
GPU Time Slicing
|
|
0
|
256
|
July 16, 2025
|
|
How to modify the cuda graph capture sizes via vllm plugin
|
|
1
|
497
|
July 1, 2025
|
|
Can’t use ampere features
|
|
1
|
291
|
June 10, 2025
|
|
KV Cache quantizing?
|
|
3
|
1409
|
June 2, 2025
|
|
Does vllm support inference or service startup of CPU small model?
|
|
3
|
302
|
May 30, 2025
|
|
Struggling with my dual GPU setup. And getting chat template errors
|
|
2
|
488
|
May 30, 2025
|
|
How to get torch-npu >= 2.5.1.dev20250308
|
|
3
|
530
|
May 28, 2025
|
|
Question about vllm-ascend performance on server with 8*910B3
|
|
5
|
940
|
May 28, 2025
|
|
Why is this not working? I corrected it but still
|
|
1
|
1035
|
May 8, 2025
|
|
Can anyone help me? Why is this not working? It used 😭
|
|
1
|
1300
|
May 8, 2025
|
|
Docker explosion this morning after it worked fine for a long while
|
|
6
|
627
|
May 6, 2025
|