No multimodal on vLLM? Model Qwen/Qwen3.8-27B-FP8 is treated as multimodal but has no registered multimodal processor; running in text-only mode.

#146
by pabbbb - opened

Model treated as text only?

Qwen/Qwen3.8-27B-FP8 is treated as multimodal but has no registered multimodal processor; running in text-only mode.

using vLLM:

exec vllm serve Qwen/Qwen3.8-27B-FP8
--kv-cache-dtype fp8
--enable-auto-tool-choice
--tool-call-parser qwen3_coder
--reasoning-parser qwen3
--max-model-len 262144
--max-num-batched-tokens 8192
--gpu-memory-utilization 0.8
--mm-encoder-tp-mode data
--trust-remote-code
--speculative-config '{"method":"mtp","num_speculative_tokens":"3"}'

You need to use the nightly release, v0.27.1 does not have qwen3.8 support. See https://recipes.vllm.ai/Qwen/Qwen3.8-27B

Sign up or log in to comment