# Qwen 2.5 VL for videos

**URL:** <https://discuss.vllm.ai/t/qwen-2-5-vl-for-videos/1460>\
**Category:** General\
**Created:** [August 26, 2025, 9:35am UTC](https://discuss.vllm.ai/t/qwen-2-5-vl-for-videos/1460 "2025-08-26T09:35:56Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![KerolosAtef](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/kerolosatef/32/646_2.png) [@KerolosAtef](https://discuss.vllm.ai/u/KerolosAtef)\
**Post date:** [August 26, 2025, 9:35am UTC](https://discuss.vllm.ai/t/qwen-2-5-vl-for-videos/1460/1 "2025-08-26T09:35:56Z")

</div>

Could you share a simple video example of using Qwen 2.5 VL with vllm? What is the best configuration that can be used for videos to speed up the inference and optimize the memory consumption?

I tried this example which needed to be optimized  
```python `  
_from_ transformers _import_ AutoProcessor

_from_ vllm _import_ LLM, SamplingParams

_from_ qwen\_vl\_utils _import_ process\_vision\_info

MODEL\_PATH = “Qwen2.5-VL-7B-Instruct”

llm = LLM(

```
*model*=MODEL_PATH,

```

)

sampling\_params = SamplingParams(

```
*temperature*=0.1,

*top_p*=0.001,

*repetition_penalty*=1.05,

*max_tokens*=4096,

*stop_token_ids*=\[\],

```

)

video\_messages = [

```
{"role": "system", "content": "You are a helpful assistant."},

{"role": "user", "content": \[

        {"type": "text", "text": "Describe the video content in detail."},

        {

            "type": "video", 

            "video": "/home/jupyter/APIs/Main_endpoints/local_videos/episode_1.mp4",

            "total_pixels": 20480 \* 28 \* 28, "min_pixels": 16 \* 28 \* 28,

            *# "fps":0.1,*

            "nframes":32,

        }

    \]

},

```

]

_# Here we use video messages as a demonstration_

messages = video\_messages

processor = AutoProcessor.from\_pretrained(MODEL\_PATH)

prompt = processor.apply\_chat\_template(

```
messages,

*tokenize*=False,

*add_generation_prompt*=True,

```

)

image\_inputs, video\_inputs, video\_kwargs = process\_vision\_info(messages, _return\_video\_kwargs_=True)

print(“Video FPS:”, video\_kwargs.get(“fps”, “Not provided”))

mm\_data = {}

_if_ image\_inputs is not None:

```
mm_data\["image"\] = image_inputs

```

_if_ video\_inputs is not None:

```
mm_data\["video"\] = video_inputs

```

llm\_inputs = {

```
"prompt": prompt,

"multi_modal_data": mm_data,

*# FPS will be returned in video_kwargs*

"mm_processor_kwargs": video_kwargs,

```

}

outputs = llm.generate([llm\_inputs], _sampling\_params_=sampling\_params)

_for_ output _in_ outputs:

```
generated_text = output.outputs\[0\].text

print(generated_text)

print('--------------------------')

```

```

---

<div class="post-metadata">

**Author:** ![RunLLM](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/runllm/32/233_2.png) [@RunLLM](https://discuss.vllm.ai/u/RunLLM)\
**Post date:** [August 26, 2025, 9:36am UTC](https://discuss.vllm.ai/t/qwen-2-5-vl-for-videos/1460/2 "2025-08-26T09:36:38Z")

</div>

A simple, optimized video example for Qwen2.5-VL with vLLM is shown below. To speed up inference and reduce memory, set `limit_mm_per_prompt={"video": 1}` to limit video items per prompt, and use `mm_processor_kwargs` to restrict video resolution and frame count (e.g., `"max_pixels": 768*768`, `"nframes": 8`). This reduces the number of tokens and memory used. Also, set `gpu_memory_utilization` to a value like 0.8 and `enforce_eager=True` if you encounter CUDA graph memory issues. Example:

```python
from transformers import AutoProcessor
from vllm import LLM, SamplingParams
from qwen_vl_utils import process_vision_info

MODEL_PATH = "Qwen/Qwen2.5-VL-7B-Instruct"
VIDEO_PATH = "/home/jupyter/APIs/Main_endpoints/local_videos/episode_1.mp4"

llm = LLM(
    model=MODEL_PATH,
    gpu_memory_utilization=0.8,
    enforce_eager=True,
    limit_mm_per_prompt={"video": 1},
    mm_processor_kwargs={"max_pixels": 768*768, "nframes": 8, "fps": 1},
)

sampling_params = SamplingParams(
    temperature=0.1,
    top_p=0.001,
    repetition_penalty=1.05,
    max_tokens=1024,
    stop_token_ids=[],
)

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": [
        {"type": "text", "text": "Describe the video content in detail."},
        {"type": "video", "video": VIDEO_PATH, "total_pixels": 768*768, "nframes": 8, "fps": 1},
    ]},
]

processor = AutoProcessor.from_pretrained(MODEL_PATH)
prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs, video_kwargs = process_vision_info(messages, return_video_kwargs=True)
mm_data = {}
if video_inputs is not None:
    mm_data["video"] = video_inputs

llm_inputs = {
    "prompt": prompt,
    "multi_modal_data": mm_data,
    "mm_processor_kwargs": video_kwargs,
}

outputs = llm.generate([llm_inputs], sampling_params=sampling_params)
for output in outputs:
    print(output.outputs[0].text)

```

For more details and further memory optimization tips, see [vLLM multimodal docs](https://docs.vllm.ai/en/latest/features/multimodal_inputs.html#__span-2-38) and [conserving memory](https://docs.vllm.ai/en/latest/configuration/conserving_memory.html).

Would you like a deeper explanation of each configuration or more advanced optimization strategies?

Sources:

- [vLLM multimodal inputs documentation](https://docs.vllm.ai/en/latest/features/multimodal_inputs.html#__span-2-38)
- [vLLM conserving memory documentation](https://docs.vllm.ai/en/latest/configuration/conserving_memory.html)
- [Qwen2.5-VL video example PR](https://github.com/vllm-project/vllm/pull/21530)
