Instructions to use Qwen/Qwen3.5-0.8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen3.5-0.8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Qwen/Qwen3.5-0.8B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Qwen/Qwen3.5-0.8B") model = AutoModelForMultimodalLM.from_pretrained("Qwen/Qwen3.5-0.8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- AMD Developer Cloud
- Local Apps Settings
- vLLM
How to use Qwen/Qwen3.5-0.8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen3.5-0.8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.5-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Qwen/Qwen3.5-0.8B
- SGLang
How to use Qwen/Qwen3.5-0.8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.5-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.5-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen3.5-0.8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen3.5-0.8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Qwen/Qwen3.5-0.8B with Docker Model Runner:
docker model run hf.co/Qwen/Qwen3.5-0.8B
When vllm deployment is started using mtp config, an error will be reported. Please help!
Using vllm==0.18.1 and transformers==5.5.0, the vllm server of mtp still cannot be started. Is there any recommended environment configuration for successful startup?
However, configuring enforce_eager enables normal inference.
(EngineCore_DP0 pid=3924877) INFO 04-07 14:29:00 [kv_cache_utils.py:1319] Maximum concurrency for 8,196 tokens per request: 339.23x
Capturing CUDA graphs (mixed prefill-decode, PIECEWISE): 100%|█████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████████| 50/50 [00:01<00:00, 32.15it/s]
Capturing CUDA graphs (decode, FULL): 0%| | 0/50 [01:10<?, ?it/s]
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] EngineCore failed to start.
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] Traceback (most recent call last):
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/compilation/cuda_graph.py", line 281, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] output = self.runnable(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._call_impl(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return forward_call(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/model_executor/models/qwen3_5.py", line 738, in forward
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] hidden_states = self.language_model.model(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/compilation/decorators.py", line 402, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self.aot_compiled_fn(self, *args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/_dynamo/aot_compile.py", line 124, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self.fn(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/model_executor/models/qwen3_next.py", line 1132, in forward
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] def forward(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/compilation/caching.py", line 198, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self.optimized_call(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 936, in call_wrapped
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._wrapped_call(self, *args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 455, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] raise e
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 442, in __call__
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return super(self.cls, obj).__call__(*args, **kwargs) # type: ignore[misc]
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._call_impl(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return forward_call(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File ".50", line 208, in forward
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] submod_1 = self.submod_1(getitem, s59, getitem_1, getitem_2, getitem_3); getitem = getitem_1 = getitem_2 = submod_1 = None
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 936, in call_wrapped
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._wrapped_call(self, *args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 455, in call
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] raise e
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/fx/graph_module.py", line 442, in call
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return super(self.cls, obj).call(*args, **kwargs) # type: ignore[misc]
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1776, in _wrapped_call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._call_impl(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1787, in _call_impl
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return forward_call(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File ".52", line 5, in forward
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] gdn_attention_core = torch.ops.vllm.gdn_attention_core(mixed_qkv, b_1, a_1, core_attn_out, 'language_model.model.layers.0.linear_attn'); mixed_qkv = b_1 = a_1 = core_attn_out = gdn_attention_core = None
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/_ops.py", line 1209, in call
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return self._op(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/model_executor/models/qwen3_next.py", line 1451, in gdn_attention_core
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self._forward_core(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/model_executor/models/qwen3_next.py", line 683, in _forward_core
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] mixed_qkv_spec = causal_conv1d_update(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/model_executor/layers/mamba/ops/causal_conv1d.py", line 1196, in causal_conv1d_update
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] _causal_conv1d_update_kernel[grid](
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/triton/runtime/jit.py", line 370, in
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/triton/runtime/jit.py", line 743, in run
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] launch_metadata = kernel.launch_metadata(grid, stream, *bound_args.values())
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/triton/compiler/compiler.py", line 482, in launch_metadata
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self._init_handles()
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/triton/compiler/compiler.py", line 465, in _init_handles
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self.module, self.function, self.n_regs, self.n_spills, self.n_max_threads = driver.active.utils.load_binary(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] RuntimeError: Triton Error [CUDA]: operation not permitted when stream is capturing
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100]
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] During handling of the above exception, another exception occurred:
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100]
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] Traceback (most recent call last):
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/engine/core.py", line 1090, in run_engine_core
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] engine_core = EngineCoreProc(*args, engine_index=dp_rank, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/engine/core.py", line 834, in init
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] super().init(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/engine/core.py", line 120, in init
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] num_gpu_blocks, num_cpu_blocks, kv_cache_config = self._initialize_kv_caches(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/engine/core.py", line 279, in _initialize_kv_caches
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self.model_executor.initialize_from_config(kv_cache_configs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/executor/abstract.py", line 118, in initialize_from_config
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] compilation_times: list[float] = self.collective_rpc("compile_or_warm_up_model")
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/executor/uniproc_executor.py", line 76, in collective_rpc
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] result = run_method(self.driver_worker, method, args, kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/serial_utils.py", line 459, in run_method
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/worker/gpu_worker.py", line 530, in compile_or_warm_up_model
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] cuda_graph_memory_bytes = self.model_runner.capture_model()
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/tracing/otel.py", line 178, in sync_wrapper
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/worker/gpu_model_runner.py", line 5363, in capture_model
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self._capture_cudagraphs(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/worker/gpu_model_runner.py", line 5464, in _capture_cudagraphs
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] dummy_run(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/utils/_contextlib.py", line 124, in decorate_context
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] return func(*args, **kwargs)
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/v1/worker/gpu_model_runner.py", line 5002, in _dummy_run
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] outputs = self.model(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/vllm/compilation/cuda_graph.py", line 275, in call
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] with torch.cuda.graph(
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/cuda/graphs.py", line 268, in exit
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] self.cuda_graph.capture_end()
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] File "/root/paddlejob/workspace/env_run/ms-swift-main/hunyuanocr/lib/python3.10/site-packages/torch/cuda/graphs.py", line 130, in capture_end
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] super().capture_end()
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] torch.AcceleratorError: CUDA error: operation failed due to a previous error during capture
(EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] Search for cudaErrorStreamCaptureInvalidated' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information. (EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect. (EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] For debugging consider passing CUDA_LAUNCH_BLOCKING=1 (EngineCore_DP0 pid=3924877) ERROR 04-07 14:30:12 [core.py:1100] Compile with TORCH_USE_CUDA_DSA` to enable device-side assertions.