Instructions to use Lalakai/Hunter-1-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Lalakai/Hunter-1-4B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lalakai/Hunter-1-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Lalakai/Hunter-1-4B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Lalakai/Hunter-1-4B:Q4_K_M # Run inference directly in the terminal: llama cli -hf Lalakai/Hunter-1-4B:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Lalakai/Hunter-1-4B:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Lalakai/Hunter-1-4B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Lalakai/Hunter-1-4B:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Lalakai/Hunter-1-4B:Q4_K_M
Use Docker
docker model run hf.co/Lalakai/Hunter-1-4B:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Lalakai/Hunter-1-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Lalakai/Hunter-1-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Lalakai/Hunter-1-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Lalakai/Hunter-1-4B:Q4_K_M
- Ollama
How to use Lalakai/Hunter-1-4B with Ollama:
ollama run hf.co/Lalakai/Hunter-1-4B:Q4_K_M
- Unsloth Desktop
- Pi
How to use Lalakai/Hunter-1-4B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lalakai/Hunter-1-4B:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Lalakai/Hunter-1-4B:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Lalakai/Hunter-1-4B with Docker Model Runner:
docker model run hf.co/Lalakai/Hunter-1-4B:Q4_K_M
- Lemonade
How to use Lalakai/Hunter-1-4B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Lalakai/Hunter-1-4B:Q4_K_M
Run and chat with the model
lemonade run user.Hunter-1-4B-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Lalakai/Hunter-1-4B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lalakai/Hunter-1-4B:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Lalakai/Hunter-1-4B:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Lalakai/Hunter-1-4B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Lalakai/Hunter-1-4B:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Lalakai/Hunter-1-4B:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hunter 4B
Hunter 4B is a 4-billion-parameter conversational reasoning model from Lalakai Labs, built for local inference on consumer hardware. It works through problems step by step inside <thought>...</thought> traces before giving its final answer, and is tuned for reasoning, agentic coding, tool use, and cybersecurity analysis.
Weights are distributed in GGUF format for llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes. GGUF stores tensors only and does not use pickled code.
Highlights
- 4B parameters, optimized for reasoning, coding, and security tasks
- Visible reasoning in
<thought>traces for math, code, and multi-step problems - Q4_K_M quantization runs in roughly 3.5 GB of working memory at a 4K context
- Apache-2.0 license
Model details
| Developer | Lalakai Labs |
| Model type | Causal language model (conversational), GGUF |
| Parameters | 4B |
| Language | English |
| Version | v1.0 |
| License | Apache-2.0 |
Training data
Hunter 4B was trained on approximately 180,000 curated examples. Raw web scrapes were not used. Every example was curated, deduplicated (MinHash LSH), and formatted into a ChatML envelope with explicit reasoning traces.
| Category | Share | Focus |
|---|---|---|
| Complex reasoning | 40% | Step-by-step problem solving, math, distilled reasoning traces |
| Agentic coding | 25% | Code synthesis, debugging, syntax discipline |
| Multi-step tool use | 20% | Structured function calls, JSON schema adherence, API loops |
| World knowledge | 15% | General facts and language grounding |
A core focus of the mix is security-oriented reasoning, including vulnerability analysis, secure-code review, and adversarial thinking, alongside everyday coding and math.
Training method
Parameter-efficient QLoRA fine-tuning (rank 64), merged back to FP16, then quantized to Q4_K_M using an importance matrix (imatrix) so that the weights most important for reasoning are preserved under compression.
Files
| File | Size | Description |
|---|---|---|
Hunter-4B-v1.0-FP16.gguf |
~8.05 GB | Full-precision weights. Highest quality; needs ~10 GB of RAM or VRAM. |
Hunter-4B-v1.0-Q4_K_M.gguf |
~2.50 GB | 4-bit quantized weights. Recommended for most users. |
Hunter-4B-v1.0.imatrix.gguf |
~3.9 MB | Importance matrix used for quantization. Only needed to re-quantize. |
SHA256SUMS |
Checksums for verifying downloads. |
Verify your download before running:
sha256sum -c SHA256SUMS
Quick start
llama.cpp
llama-cli -m Hunter-4B-v1.0-Q4_K_M.gguf -p "Hello! Who are you?"
To serve over HTTP:
llama-server -m Hunter-4B-v1.0-Q4_K_M.gguf --port 8080
Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama(
model_path="Hunter-4B-v1.0-Q4_K_M.gguf",
n_ctx=4096,
verbose=False,
)
print(llm("Hello! Who are you?", max_tokens=256)["choices"][0]["text"])
Ollama
Create a Modelfile with the ChatML template and stop tokens:
FROM ./Hunter-4B-v1.0-Q4_K_M.gguf
TEMPLATE """<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
{{ .Response }}<|im_end|>"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.6
PARAMETER top_p 0.90
PARAMETER top_k 40
PARAMETER repeat_penalty 1.1
SYSTEM """You are Hunter-4B, a reasoning model developed by Lalakai Labs. Write out your reasoning inside <thought> ... </thought> tags before giving your final answer."""
ollama create hunter-4b -f Modelfile
ollama run hunter-4b
Hugging Face CLI
hf download Lalakai/Hunter-1-4B Hunter-4B-v1.0-Q4_K_M.gguf
Choosing a file
- Local chat on a typical laptop: Q4_K_M
- Highest output quality, if you have the memory: FP16
- Creating your own quantizations: FP16 plus the .imatrix file
Intended use
- Step-by-step explanations, math, and coding help
- Agentic coding and tool-use workflows on local hardware
- Summarization and structured multi-step analysis
- Security-oriented reasoning, such as vulnerability analysis and secure-code review, with human oversight
Evaluation
No formal benchmark results have been published for this model yet. They will be added to this card when available.
Limitations
- A 4B model will not match frontier-scale models on the hardest reasoning tasks or on niche knowledge. Use the FP16 weights and a larger context window (-c) when quality matters most.
- Like all language models, Hunter 4B can produce hallucinations, reflect biases in its training data, and generate unsafe content when prompted adversarially. Review outputs for your use case and apply your own safety filtering in production.
Responsible use
Use this model under human supervision and within applicable law and your organization's policies. Do not use it to target systems you are not authorized to test.
License
Released under the Apache License, Version 2.0 (https://www.apache.org/licenses/LICENSE-2.0).
Contact
Built by Lalakai Labs (https://x.com/LalakaiAI).
Acknowledgements
Hunter 4B is a fine-tune of Qwen 3.5 4B (https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0). We thank the Qwen team for the base model.
- Downloads last month
- 252
4-bit