# Vllm serve拉起服务，到最后报错了

**URL:** <https://discuss.vllm.ai/t/vllm-serve/2736>\
**Category:** General\
**Created:** [June 27, 2026, 9:36am UTC](https://discuss.vllm.ai/t/vllm-serve/2736 "2026-06-27T09:36:31Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![zZzbinn](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/zzzbinn/32/1504_2.png) [@zZzbinn](https://discuss.vllm.ai/u/zZzbinn)\
**Post date:** [June 27, 2026, 9:36am UTC](https://discuss.vllm.ai/t/vllm-serve/2736/1 "2026-06-27T09:36:31Z")

</div>

vllm serve /autodl-fs/data/weight/DeepSeek-V3.1/ --trust-remote-code --max-num-seqs 16 --max-model-len 8192 --max-num-batched-tokens 8192 --gpu-memory-utilization 0.95 --tensor-parallel-size 8 --distributed\_executor\_backend “mp” --no-enable-prefix-caching --enable-expert-parallel --no-async-scheduling --default-chat-template-kwargs ‘{“enable\_thinking”: false}’ --compilation\_config ‘{“cudagraph\_mode”: “FULL\_DECODE\_ONLY”}’

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [June 27, 2026, 9:36am UTC](https://discuss.vllm.ai/t/vllm-serve/2736/2 "2026-06-27T09:36:46Z")

</div>

Your command mostly uses valid vLLM serve arguments, but there are some issues:

- The correct flag is `--distributed-executor-backend`, not `--distributed_executor_backend`.
- The correct flag is `--compilation-config`, not `--compilation_config`.
- Use standard double quotes `"` for JSON and arguments, not curly quotes or single quotes.
- All other flags are supported and documented for distributed, expert parallel, and compilation config usage in vLLM serve. See [engine args](https://docs.vllm.ai/en/stable/serving/engine_args.html#searchbox) and [serve args](https://docs.vllm.ai/en/latest/serving/serve_args.html#cli-arguments).

Would you like a corrected command or more detail on any specific flag?

Sources:

- [engine args](https://docs.vllm.ai/en/stable/serving/engine_args.html#searchbox)
- [serve args](https://docs.vllm.ai/en/latest/serving/serve_args.html#cli-arguments)

Hint: Mention @Herald in the post for followups.
