# Qwen3.5-27B-FP8 Speculative Decoding

**URL:** <https://discuss.vllm.ai/t/qwen3-5-27b-fp8-speculative-decoding/2447>\
**Category:** Speculative Decoding\
**Created:** [March 12, 2026, 1:39pm UTC](https://discuss.vllm.ai/t/qwen3-5-27b-fp8-speculative-decoding/2447 "2026-03-12T13:39:19Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![WittyExplorer](https://avatars.discourse-cdn.com/v4/letter/w/b5e925/32.png) [@WittyExplorer](https://discuss.vllm.ai/u/WittyExplorer)\
**Post date:** [March 12, 2026, 1:39pm UTC](https://discuss.vllm.ai/t/qwen3-5-27b-fp8-speculative-decoding/2447/1 "2026-03-12T13:39:19Z")

</div>

I want to test speculative decoding with Qwen3.5-27B-FB8. The only way that works for me is `{“method”: “mtp”, “num_speculative_tokens”:1}` . Increasing `num_speculative_tokens` to `2` results in error. Is that expected?

Looking at `eagle` it appears that one needs special eagle models for it to work. The links from documentation pointing to redhat and some other collection don’t have anything for Qwen3.5. The only model I could find is for Qwen-9B. Does it mean that currently it’s not possible to use `eagle`for my model?

Then I tried to use draft model. I tried `Qwen-9B` and `Qwen-4B`, however in both cases I got `RuntimeError: The size of tensor a (5120) must match the size of tensor b (4096) at non-singleton dimension 1` or similar error. And I also wasn’t able to find anything resembling a draft model on HF. Does it mean that using `draft model` is also not possible at the moment? Or is there some other model that could work as draft model?

Does it make sense to test any other methods?

With `MTP`my throughput benchmark result went from `120` requests per minute to `140`. Which is about 16%. Not bad for adding one config parameter but also not as good as 1.5x to 2x performance gains I saw online. Is that considered to be a good improvement or is there potential to get more?

The command I use to run `vllm` in a docker container:

```yaml
Qwen/Qwen3.5-27B-FP8
      --port 8000
      --enable-auto-tool-choice
      --tool-call-parser qwen3_coder
      --reasoning-parser qwen3
      --enable-prefix-caching
      --speculative-config '{"method": "draft_model", "num_speculative_tokens": 5, "model": "Qwen/Qwen3.5-4B"}'

```

---

<div class="post-metadata">

**Author:** ![benchislett](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/benchislett/32/380_2.png) [@benchislett](https://discuss.vllm.ai/u/benchislett)\
**Post date:** [March 12, 2026, 5:36pm UTC](https://discuss.vllm.ai/t/qwen3-5-27b-fp8-speculative-decoding/2447/3 "2026-03-12T17:36:36Z")

</div>

RunLLM is wrong here. You should be able to use it with MTP for multiple tokens, and that will be the most performant way to do it.

Please create a github issue and include the error log so we can triage

---

<div class="post-metadata">

**Author:** ![Jianhua-Cui](https://yyz2.discourse-cdn.com/flex036/user_avatar/discuss.vllm.ai/jianhua-cui/32/1233_2.png) [@Jianhua-Cui](https://discuss.vllm.ai/u/Jianhua-Cui)\
**Post date:** [April 11, 2026, 2:02am UTC](https://discuss.vllm.ai/t/qwen3-5-27b-fp8-speculative-decoding/2447/4 "2026-04-11T02:02:01Z")

</div>

Qwen3.5 uses hybrid linear attention throughout. Its `conv_states` and `recurrent_states` do not have a `sequence_length` dimension, so they cannot be selectively accepted the way a traditional KV cache can. In other words, if your draft contains 4 tokens but the target LLM only accepts the first 2 during parallel verification, we have no way to restore the corresponding `conv_states` and `recurrent_states`.

I think this may be the reason why only MTP-1 is supported.
