# Command lines to run pp=2 on two nodes in vllm 0.10.x

**URL:** <https://discuss.vllm.ai/t/command-lines-to-run-pp-2-on-two-nodes-in-vllm-0-10-x/1986>\
**Category:** General\
**Created:** [November 20, 2025, 4:33am UTC](https://discuss.vllm.ai/t/command-lines-to-run-pp-2-on-two-nodes-in-vllm-0-10-x/1986 "2025-11-20T04:33:37Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![weiyshay](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/weiyshay/32/701_2.png) [@weiyshay](https://discuss.vllm.ai/u/weiyshay)\
**Post date:** [November 20, 2025, 4:33am UTC](https://discuss.vllm.ai/t/command-lines-to-run-pp-2-on-two-nodes-in-vllm-0-10-x/1986/1 "2025-11-20T04:33:37Z")

</div>

Hi,  
I am trying to run PP=2, tp=2 on two nodes using vllm 0.10.2, I read some articles that I can use ray to manage the cluster, however, it is also possible to run this using native vllm without ray.

I tried a few command lines, like ‘master-addr’, ‘master-port’ or ‘entrypoints.controller’ etc, but it tells that the parameters are not recognized.

I have two questions:  
#1: is it supported to run PP=2 without using ray on vllm 0.10.x?

#2: what are the correct parameters to run if yes? any document to the usage/command line?

Appreciate any comments.

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [November 20, 2025, 4:33am UTC](https://discuss.vllm.ai/t/command-lines-to-run-pp-2-on-two-nodes-in-vllm-0-10-x/1986/2 "2025-11-20T04:33:57Z")

</div>

1. Running pipeline parallelism (PP \> 1) across multiple nodes without Ray is **not supported** in vLLM 0.10.x. Multi-node distributed inference currently requires Ray as the runtime engine; native multiprocessing is only supported for single-node deployments. Parameters like `master-addr`, `master-port`, or `entrypoints.controller` are not recognized by vLLM and are not documented in the official sources for multi-node use without Ray. See [distributed serving docs](https://docs.vllm.ai/en/latest/serving/distributed_serving.html#running-vllm-on-multiple-nodes) and [parallelism scaling](https://docs.vllm.ai/en/latest/serving/parallelism_scaling/#__span-2-2).

2. For multi-node PP/TP, you must use Ray. The correct approach is to set up a Ray cluster (using the provided `run_cluster.sh` or KubeRay), then launch vLLM with `--tensor-parallel-size` and `--pipeline-parallel-size` as needed. There is no support for launching multi-node PP/TP with `torchrun` or similar native PyTorch methods in vLLM 0.10.x. See [parallelism scaling documentation](https://docs.vllm.ai/en/latest/serving/parallelism_scaling/#__span-2-2) and [github discussion](https://github.com/vllm-project/vllm/issues/3902).

Would you like more detail on how to set up Ray for multi-node vLLM?

Sources:

- [distributed serving docs](https://docs.vllm.ai/en/latest/serving/distributed_serving.html#running-vllm-on-multiple-nodes)
- [parallelism scaling](https://docs.vllm.ai/en/latest/serving/parallelism_scaling/#__span-2-2)
- [github discussion](https://github.com/vllm-project/vllm/issues/3902)

Hint: Mention @RunLLM in the post for followups.

---

<div class="post-metadata">

**Author:** ![weiyshay](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/weiyshay/32/701_2.png) [@weiyshay](https://discuss.vllm.ai/u/weiyshay)\
**Post date:** [November 20, 2025, 4:38am UTC](https://discuss.vllm.ai/t/command-lines-to-run-pp-2-on-two-nodes-in-vllm-0-10-x/1986/3 "2025-11-20T04:38:28Z")

</div>

yes, please detail on how to set up Ray for multi-node vLLM
