# 求救各位大佬看看是什么问题。cuda12.9，pytorch2.8，vllm0.11.0

**URL:** <https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910>\
**Category:** General\
**Created:** [November 14, 2025, 7:45am UTC](https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910 "2025-11-14T07:45:12Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![Jon](https://avatars.discourse-cdn.com/v4/letter/j/3bc359/32.png) [@Jon](https://discuss.vllm.ai/u/Jon)\
**Post date:** [November 14, 2025, 7:45am UTC](https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910/1 "2025-11-14T07:45:12Z")

</div>

(llm311) nvidia@localhost:/data/lqy/qwen$ python -m vllm.entrypoints.openai.api\_server \

> ```auto
> --model /data/lqy/qwen/Qwen3-VL-8B-Instruct \
> --served-model-name Qwen3-VL-8B-Instruct \
> --host 0.0.0.0 \
> --port 9000 \
> --gpu-memory-utilization 0.7 \
> --skip-mm-profiling
> 
> ```

INFO 11-14 15:07:40 [**init**.py:216] Automatically detected platform cuda.  
(APIServer pid=2926114) INFO 11-14 15:07:41 [api\_server.py:1839] vLLM API server version 0.11.0  
(APIServer pid=2926114) INFO 11-14 15:07:41 [utils.py:233] non-default args: {‘host’: ‘0.0.0.0’, ‘port’: 9000, ‘model’: ‘/data/lqy/qwen/Qwen3-VL-8B-Instruct’, ‘served\_model\_name’: [‘Qwen3-VL-8B-Instruct’], ‘gpu\_memory\_utilization’: 0.7, ‘skip\_mm\_profiling’: True}  
(APIServer pid=2926114) INFO 11-14 15:07:41 [model.py:547] Resolved architecture: Qwen3VLForConditionalGeneration  
(APIServer pid=2926114) `torch_dtype` is deprecated! Use `dtype` instead!  
(APIServer pid=2926114) INFO 11-14 15:07:41 [model.py:1510] Using max model len 262144  
(APIServer pid=2926114) INFO 11-14 15:07:41 [scheduler.py:205] Chunked prefill is enabled with max\_num\_batched\_tokens=2048.  
INFO 11-14 15:07:46 [**init**.py:216] Automatically detected platform cuda.  
(EngineCore\_DP0 pid=2926406) INFO 11-14 15:07:48 [core.py:644] Waiting for init message from front-end.  
(EngineCore\_DP0 pid=2926406) INFO 11-14 15:07:48 [core.py:77] Initializing a V1 LLM engine (v0.11.0) with config: model=‘/data/lqy/qwen/Qwen3-VL-8B-Instruct’, speculative\_config=None, tokenizer=‘/data/lqy/qwen/Qwen3-VL-8B-Instruct’, skip\_tokenizer\_init=False, tokenizer\_mode=auto, revision=None, tokenizer\_revision=None, trust\_remote\_code=False, dtype=torch.bfloat16, max\_seq\_len=262144, download\_dir=None, load\_format=auto, tensor\_parallel\_size=1, pipeline\_parallel\_size=1, data\_parallel\_size=1, disable\_custom\_all\_reduce=False, quantization=None, enforce\_eager=False, kv\_cache\_dtype=auto, device\_config=cuda, structured\_outputs\_config=StructuredOutputsConfig(backend=‘auto’, disable\_fallback=False, disable\_any\_whitespace=False, disable\_additional\_properties=False, reasoning\_parser=‘’), observability\_config=ObservabilityConfig(show\_hidden\_metrics\_for\_version=None, otlp\_traces\_endpoint=None, collect\_detailed\_traces=None), seed=0, served\_model\_name=Qwen3-VL-8B-Instruct, enable\_prefix\_caching=True, chunked\_prefill\_enabled=True, pooler\_config=None, compilation\_config={“level”:3,“debug\_dump\_path”:“”,“cache\_dir”:“”,“backend”:“”,“custom\_ops”:,“splitting\_ops”:[“vllm.unified\_attention”,“vllm.unified\_attention\_with\_output”,“vllm.mamba\_mixer2”,“vllm.mamba\_mixer”,“vllm.short\_conv”,“vllm.linear\_attention”,“vllm.plamo2\_mamba\_mixer”,“vllm.gdn\_attention”,“vllm.sparse\_attn\_indexer”],“use\_inductor”:true,“compile\_sizes”:,“inductor\_compile\_config”:{“enable\_auto\_functionalized\_v2”:false},“inductor\_passes”:{},“cudagraph\_mode”:[2,1],“use\_cudagraph”:true,“cudagraph\_num\_of\_warmups”:1,“cudagraph\_capture\_sizes”:[512,504,496,488,480,472,464,456,448,440,432,424,416,408,400,392,384,376,368,360,352,344,336,328,320,312,304,296,288,280,272,264,256,248,240,232,224,216,208,200,192,184,176,168,160,152,144,136,128,120,112,104,96,88,80,72,64,56,48,40,32,24,16,8,4,2,1],“cudagraph\_copy\_inputs”:false,“full\_cuda\_graph”:false,“use\_inductor\_graph\_partition”:false,“pass\_config”:{},“max\_capture\_size”:512,“local\_cache\_dir”:null}  
(EngineCore\_DP0 pid=2926406) /data/conda/envs/llm311/lib/python3.11/site-packages/torch/cuda/ **init**.py:326: UserWarning:  
(EngineCore\_DP0 pid=2926406) NVIDIA Thor with CUDA capability sm\_110 is not compatible with the current PyTorch installation.  
(EngineCore\_DP0 pid=2926406) The current PyTorch install supports CUDA capabilities sm\_80 sm\_90 sm\_100 sm\_120.  
(EngineCore\_DP0 pid=2926406) If you want to use the NVIDIA Thor GPU with PyTorch, please check the instructions at [https://pytorch.org/get-started/locally/](https://pytorch.org/get-started/locally/)  
(EngineCore\_DP0 pid=2926406)  
(EngineCore\_DP0 pid=2926406) warnings.warn(  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] EngineCore failed to start.  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] Traceback (most recent call last):  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 699, in run\_engine\_core  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] engine\_core = EngineCoreProc(\*args, \*\*kwargs)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 498, in **init**  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] super(). **init** (vllm\_config, executor\_class, log\_stats,  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 83, in **init**  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] self.model\_executor = executor\_class(vllm\_config)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/executor\_base.py”, line 54, in **init**  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] self.\_init\_executor()  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/uniproc\_executor.py”, line 54, in \_init\_executor  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] self.collective\_rpc(“init\_device”)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/uniproc\_executor.py”, line 83, in collective\_rpc  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] return [run\_method(self.driver\_worker, method, args, kwargs)]  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/utils/ **init**.py”, line 3122, in run\_method  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] return func(\*args, \*\*kwargs)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/worker/worker\_base.py”, line 259, in init\_device  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] self.worker.init\_device() # type: ignore  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/worker/gpu\_worker.py”, line 161, in init\_device  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] current\_platform.set\_device(self.device)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/platforms/cuda.py”, line 83, in set\_device  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] \_ = torch.zeros(1, device=device)  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] torch.AcceleratorError: CUDA error: no kernel image is available for execution on the device  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] For debugging consider passing CUDA\_LAUNCH\_BLOCKING=1  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708] Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.  
(EngineCore\_DP0 pid=2926406) ERROR 11-14 15:07:50 [core.py:708]  
(EngineCore\_DP0 pid=2926406) Process EngineCore\_DP0:  
(EngineCore\_DP0 pid=2926406) Traceback (most recent call last):  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/multiprocessing/process.py”, line 314, in \_bootstrap  
(EngineCore\_DP0 pid=2926406) self.run()  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/multiprocessing/process.py”, line 108, in run  
(EngineCore\_DP0 pid=2926406) self.\_target(\*self.\_args, \*\*self.\_kwargs)  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 712, in run\_engine\_core  
(EngineCore\_DP0 pid=2926406) raise e  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 699, in run\_engine\_core  
(EngineCore\_DP0 pid=2926406) engine\_core = EngineCoreProc(\*args, \*\*kwargs)  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 498, in **init**  
(EngineCore\_DP0 pid=2926406) super(). **init** (vllm\_config, executor\_class, log\_stats,  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core.py”, line 83, in **init**  
(EngineCore\_DP0 pid=2926406) self.model\_executor = executor\_class(vllm\_config)  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/executor\_base.py”, line 54, in **init**  
(EngineCore\_DP0 pid=2926406) self.\_init\_executor()  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/uniproc\_executor.py”, line 54, in \_init\_executor  
(EngineCore\_DP0 pid=2926406) self.collective\_rpc(“init\_device”)  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/executor/uniproc\_executor.py”, line 83, in collective\_rpc  
(EngineCore\_DP0 pid=2926406) return [run\_method(self.driver\_worker, method, args, kwargs)]  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/utils/ **init**.py”, line 3122, in run\_method  
(EngineCore\_DP0 pid=2926406) return func(\*args, \*\*kwargs)  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/worker/worker\_base.py”, line 259, in init\_device  
(EngineCore\_DP0 pid=2926406) self.worker.init\_device() # type: ignore  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/worker/gpu\_worker.py”, line 161, in init\_device  
(EngineCore\_DP0 pid=2926406) current\_platform.set\_device(self.device)  
(EngineCore\_DP0 pid=2926406) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/platforms/cuda.py”, line 83, in set\_device  
(EngineCore\_DP0 pid=2926406) \_ = torch.zeros(1, device=device)  
(EngineCore\_DP0 pid=2926406) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(EngineCore\_DP0 pid=2926406) torch.AcceleratorError: CUDA error: no kernel image is available for execution on the device  
(EngineCore\_DP0 pid=2926406) CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.  
(EngineCore\_DP0 pid=2926406) For debugging consider passing CUDA\_LAUNCH\_BLOCKING=1  
(EngineCore\_DP0 pid=2926406) Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.  
(EngineCore\_DP0 pid=2926406)  
(APIServer pid=2926114) Traceback (most recent call last):  
(APIServer pid=2926114) File “”, line 198, in \_run\_module\_as\_main  
(APIServer pid=2926114) File “”, line 88, in \_run\_code  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/entrypoints/openai/api\_server.py”, line 1953, in  
(APIServer pid=2926114) uvloop.run(run\_server(args))  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/uvloop/ **init**.py”, line 92, in run  
(APIServer pid=2926114) return runner.run(wrapper())  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/asyncio/runners.py”, line 118, in run  
(APIServer pid=2926114) return self.\_loop.run\_until\_complete(task)  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “uvloop/loop.pyx”, line 1518, in uvloop.loop.Loop.run\_until\_complete  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/uvloop/ **init**.py”, line 48, in wrapper  
(APIServer pid=2926114) return await main  
(APIServer pid=2926114) ^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/entrypoints/openai/api\_server.py”, line 1884, in run\_server  
(APIServer pid=2926114) await run\_server\_worker(listen\_address, sock, args, \*\*uvicorn\_kwargs)  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/entrypoints/openai/api\_server.py”, line 1902, in run\_server\_worker  
(APIServer pid=2926114) async with build\_async\_engine\_client(  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/contextlib.py”, line 210, in **aenter**  
(APIServer pid=2926114) return await anext(self.gen)  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/entrypoints/openai/api\_server.py”, line 180, in build\_async\_engine\_client  
(APIServer pid=2926114) async with build\_async\_engine\_client\_from\_engine\_args(  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/contextlib.py”, line 210, in **aenter**  
(APIServer pid=2926114) return await anext(self.gen)  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/entrypoints/openai/api\_server.py”, line 225, in build\_async\_engine\_client\_from\_engine\_args  
(APIServer pid=2926114) async\_llm = AsyncLLM.from\_vllm\_config(  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/utils/ **init**.py”, line 1572, in inner  
(APIServer pid=2926114) return fn(\*args, \*\*kwargs)  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/async\_llm.py”, line 207, in from\_vllm\_config  
(APIServer pid=2926114) return cls(  
(APIServer pid=2926114) ^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/async\_llm.py”, line 134, in **init**  
(APIServer pid=2926114) self.engine\_core = EngineCoreClient.make\_async\_mp\_client(  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core\_client.py”, line 102, in make\_async\_mp\_client  
(APIServer pid=2926114) return AsyncMPClient(\*client\_args)  
(APIServer pid=2926114) ^^^^^^^^^^^^^^^^^^^^^^^^^^^  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core\_client.py”, line 769, in **init**  
(APIServer pid=2926114) super(). **init** (  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/core\_client.py”, line 448, in **init**  
(APIServer pid=2926114) with launch\_core\_engines(vllm\_config, executor\_class,  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/contextlib.py”, line 144, in **exit**  
(APIServer pid=2926114) next(self.gen)  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/utils.py”, line 732, in launch\_core\_engines  
(APIServer pid=2926114) wait\_for\_engine\_startup(  
(APIServer pid=2926114) File “/data/conda/envs/llm311/lib/python3.11/site-packages/vllm/v1/engine/utils.py”, line 785, in wait\_for\_engine\_startup  
(APIServer pid=2926114) raise RuntimeError("Engine core initialization failed. "  
(APIServer pid=2926114) RuntimeError: Engine core initialization failed. See root cause above. Failed core proc(s): {}  
(llm311) nvidia@localhost:/data/lqy/qwen$

---

<div class="post-metadata">

**Author:** ![jeejeelee](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/jeejeelee/32/23_2.png) [@jeejeelee](https://discuss.vllm.ai/u/jeejeelee)\
**Post date:** [November 14, 2025, 8:01am UTC](https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910/2 "2025-11-14T08:01:44Z")

</div>

可以参考这个issue: [[Bug]: NVIDIA Thor with CUDA capability sm\_110 is not compatible · Issue #26791 · vllm-project/vllm · GitHub](https://github.com/vllm-project/vllm/issues/26791)

---

<div class="post-metadata">

**Author:** ![Jon](https://avatars.discourse-cdn.com/v4/letter/j/3bc359/32.png) [@Jon](https://discuss.vllm.ai/u/Jon)\
**Post date:** [November 14, 2025, 8:10am UTC](https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910/3 "2025-11-14T08:10:20Z")

</div>

我现在也在参考了，还不知道行不行，谢谢大佬

---

<div class="post-metadata">

**Author:** ![Jon](https://avatars.discourse-cdn.com/v4/letter/j/3bc359/32.png) [@Jon](https://discuss.vllm.ai/u/Jon)\
**Post date:** [November 14, 2025, 8:58am UTC](https://discuss.vllm.ai/t/cuda12-9-pytorch2-8-vllm0-11-0/1910/4 "2025-11-14T08:58:59Z")

</div>

大佬，可以帮忙看一下怎么解决吗  
(llm312) nvidia@localhost:~$ python -m vllm.entrypoints.openai.api\_server \

> ```auto
> --model /data/lqy/qwen/Qwen3-VL-8B-Instruct \
> --served-model-name Qwen3-VL-8B-Instruct \
> --host 0.0.0.0 \
> --port 9000 \
> --gpu-memory-utilization 0.7 \
> --skip-mm-profiling
> 
> ```

INFO 11-14 16:54:20 [**init**.py:216] Automatically detected platform cuda.  
Traceback (most recent call last):  
File “”, line 198, in \_run\_module\_as\_main  
File “”, line 88, in \_run\_code  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/entrypoints/openai/api\_server.py”, line 42, in  
from vllm.config import VllmConfig  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/config/ **init**.py”, line 34, in  
from vllm.config.lora import LoRAConfig  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/config/lora.py”, line 14, in  
from vllm.platforms import current\_platform  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/platforms/ **init**.py”, line 248, in **getattr**  
\_current\_platform = resolve\_obj\_by\_qualname(  
^^^^^^^^^^^^^^^^^^^^^^^^  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/utils/ **init**.py”, line 2680, in resolve\_obj\_by\_qualname  
module = importlib.import\_module(module\_name)  
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
File “/data/conda/envs/llm312/lib/python3.12/importlib/ **init**.py”, line 90, in import\_module  
return \_bootstrap.\_gcd\_import(name[level:], package, level)  
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^  
File “/data/conda/envs/llm312/lib/python3.12/site-packages/vllm/platforms/cuda.py”, line 18, in  
import vllm.\_C # noqa  
^^^^^^^^^^^^^^  
ImportError: libcudart.so.12: cannot open shared object file: No such file or directory  
(llm312) nvidia@localhost:~$
