SmolVLA - LIBERO (GGUF for vla.cpp)
GGUF conversion of HuggingFaceVLA/smolvla_libero
for inference with vla.cpp, a lightweight
C++ inference engine for Vision-Language-Action models built on top of
llama.cpp.
SmolVLA-450M is the golden reference architecture in vla.cpp - a SigLIP vision tower + SmolLM2 language backbone + flow-matching action expert. Every other arch in the engine is validated against the SmolVLA recipe.
Usage
Build vla-server from the vla.cpp repo,
then:
# Terminal 1 - serve (use the CUDA build for inference)
./build-cuda/vla-server --bind tcp://*:5566 \
smolvla-libero.gguf
# Terminal 2 - drive a LIBERO episode (inside the LIBERO uv venv)
python eval/client/run_sim_client_direct.py \
--arch smolvla \
--task libero_object --task-id 0 --n-episodes 10 \
--vla-addr tcp://localhost:5566
smolvla uses --n-action-steps 1 (the client default).
Benchmark
Full libero_object sweep (10 tasks × 20 episodes = 200 episodes), vla-server +
run_sim_client_direct.py:
| Hardware | n_act | Success rate | client/step | client/call | Peak mem |
|---|---|---|---|---|---|
| RTX 3060 (sm_86) | 4 | 90.5% | 28.16 ms | 113 ms | 1410 MiB VRAM |
| Jetson AGX Orin (sm_87) | 4 | 90.5% | 65.41 ms | 262 ms | 689 MiB RAM |
| Jetson Orin Nano 8 GB (sm_87) | 4 | 88.0% | 141.81 ms | 567 ms | 2031 MiB RAM |
License
Weights follow the upstream license of
HuggingFaceVLA/smolvla_libero
(Apache-2.0). The vla.cpp conversion tooling and inference engine are MIT-licensed.
- Downloads last month
- 171
We're not able to determine the quantization variants.
Model tree for vrfai/smolvla-libero-gguf
Base model
HuggingFaceTB/SmolLM2-360M