|
About the NVIDIA GPU Support category
|
|
0
|
169
|
March 20, 2025
|
|
Title: OutOfResources: shared memory (98304 > 65536) on Turing (SM75) — global attention layers with head_dim=512
|
|
2
|
46
|
September 20, 2026
|
|
Is this normal for vllm serve output waiting 1 hour b´before startup
|
|
7
|
147
|
September 9, 2026
|
|
Support for Nvidia A2 Tensor Core GPU, stuck on shm_broadcast, unsure how to debug
|
|
2
|
206
|
August 6, 2026
|
|
Performance Tuning for L40s with Qwen36-35B
|
|
1
|
214
|
July 29, 2026
|
|
Making best use of varying GPU generations
|
|
4
|
1316
|
July 4, 2026
|
|
VGPU on podman "No CUDA GPUs are available"
|
|
0
|
84
|
May 23, 2026
|
|
vLLM hangs during worker initialization on Blackwell PCIe GPUs unless --disable-custom-all-reduce is used
|
|
1
|
1596
|
April 11, 2026
|
|
# SM120 (RTX PRO 6000) NVFP4 MoE Performance Report -- Qwen3.5-397B
|
|
1
|
1930
|
April 11, 2026
|
|
Vllm启动时,日志卡在nccl相关部分,不继续往下
|
|
16
|
2152
|
April 8, 2026
|
|
SM120 (RTX PRO 4000): 6.5x throughput gain and v0.18.1 regression findings
|
|
1
|
1493
|
April 3, 2026
|
|
MoE config on GH200
|
|
9
|
834
|
February 4, 2026
|
|
vLLM on RTX5090: Working GPU setup with torch 2.9.0 cu128
|
|
18
|
7811
|
January 13, 2026
|
|
Support for RTX 6000 Blackwell 96GB card
|
|
5
|
9083
|
January 5, 2026
|
|
How to apply FA4 on B200?
|
|
3
|
800
|
December 18, 2025
|
|
RTX PRO 6000 users seek help, LLAMA 4 NVFP4
|
|
1
|
389
|
November 25, 2025
|
|
RuntimeError: Int8 not supported on SM120. Use FP8 quantization instead, or run on older arch (SM < 100)
|
|
1
|
365
|
November 19, 2025
|
|
Need help compiling and running on Jetson Thor
|
|
4
|
1123
|
November 1, 2025
|
|
RTX Pro 6000 Tensor Parallelism CUBLAS_STATUS_ALLOC_FAILED
|
|
3
|
661
|
September 13, 2025
|
|
vLLM Benchmarking: Why Is GPUDirect RDMA Not Outperforming Standard RDMA in a Pipeline-Parallel Setup?
|
|
1
|
852
|
August 14, 2025
|
|
GPU Time Slicing
|
|
0
|
268
|
July 16, 2025
|
|
KV Cache quantizing?
|
|
3
|
1559
|
June 2, 2025
|
|
Struggling with my dual GPU setup. And getting chat template errors
|
|
2
|
524
|
May 30, 2025
|
|
Why is this not working? I corrected it but still
|
|
1
|
1073
|
May 8, 2025
|
|
Can anyone help me? Why is this not working? It used 😭
|
|
1
|
1324
|
May 8, 2025
|
|
Docker explosion this morning after it worked fine for a long while
|
|
6
|
693
|
May 6, 2025
|
|
32GB vs 48GB vRam
|
|
1
|
1708
|
May 3, 2025
|
|
Run on B200/5090 without building from source?
|
|
1
|
388
|
May 1, 2025
|
|
Jetson orin, CUDA error: no kernel image is available for execution on the device
|
|
0
|
542
|
March 29, 2025
|