SmolVLA - LIBERO (GGUF for vla.cpp)

GGUF conversion of HuggingFaceVLA/smolvla_libero for inference with vla.cpp, a lightweight C++ inference engine for Vision-Language-Action models built on top of llama.cpp.

SmolVLA-450M is the golden reference architecture in vla.cpp - a SigLIP vision tower + SmolLM2 language backbone + flow-matching action expert. Every other arch in the engine is validated against the SmolVLA recipe.

Usage

Build vla-server from the vla.cpp repo, then:

# Terminal 1 - serve (use the CUDA build for inference)
./build-cuda/vla-server --bind tcp://*:5566 \
    smolvla-libero.gguf

# Terminal 2 - drive a LIBERO episode (inside the LIBERO uv venv)
python eval/client/run_sim_client_direct.py \
    --arch smolvla \
    --task libero_object --task-id 0 --n-episodes 10 \
    --vla-addr tcp://localhost:5566

smolvla uses --n-action-steps 1 (the client default).

Benchmark

Full libero_object sweep (10 tasks × 20 episodes = 200 episodes), vla-server + run_sim_client_direct.py:

Hardware n_act Success rate client/step client/call Peak mem
RTX 3060 (sm_86) 4 90.5% 28.16 ms 113 ms 1410 MiB VRAM
Jetson AGX Orin (sm_87) 4 90.5% 65.41 ms 262 ms 689 MiB RAM
Jetson Orin Nano 8 GB (sm_87) 4 88.0% 141.81 ms 567 ms 2031 MiB RAM

License

Weights follow the upstream license of HuggingFaceVLA/smolvla_libero (Apache-2.0). The vla.cpp conversion tooling and inference engine are MIT-licensed.

Downloads last month
171
GGUF
Model size
0.6B params
Architecture
smolvla
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Video Preview
loading

Model tree for vrfai/smolvla-libero-gguf

Collection including vrfai/smolvla-libero-gguf