Hunter 4B banner

Hunter 4B

Hunter 4B is a 4-billion-parameter conversational reasoning model from Lalakai Labs, built for local inference on consumer hardware. It works through problems step by step inside <thought>...</thought> traces before giving its final answer, and is tuned for reasoning, agentic coding, tool use, and cybersecurity analysis.

Weights are distributed in GGUF format for llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes. GGUF stores tensors only and does not use pickled code.

Highlights

  • 4B parameters, optimized for reasoning, coding, and security tasks
  • Visible reasoning in <thought> traces for math, code, and multi-step problems
  • Q4_K_M quantization runs in roughly 3.5 GB of working memory at a 4K context
  • Apache-2.0 license

Model details

Developer Lalakai Labs
Model type Causal language model (conversational), GGUF
Parameters 4B
Language English
Version v1.0
License Apache-2.0

Training data

Hunter 4B was trained on approximately 180,000 curated examples. Raw web scrapes were not used. Every example was curated, deduplicated (MinHash LSH), and formatted into a ChatML envelope with explicit reasoning traces.

Category Share Focus
Complex reasoning 40% Step-by-step problem solving, math, distilled reasoning traces
Agentic coding 25% Code synthesis, debugging, syntax discipline
Multi-step tool use 20% Structured function calls, JSON schema adherence, API loops
World knowledge 15% General facts and language grounding

A core focus of the mix is security-oriented reasoning, including vulnerability analysis, secure-code review, and adversarial thinking, alongside everyday coding and math.

Training method

Parameter-efficient QLoRA fine-tuning (rank 64), merged back to FP16, then quantized to Q4_K_M using an importance matrix (imatrix) so that the weights most important for reasoning are preserved under compression.

Files

File Size Description
Hunter-4B-v1.0-FP16.gguf ~8.05 GB Full-precision weights. Highest quality; needs ~10 GB of RAM or VRAM.
Hunter-4B-v1.0-Q4_K_M.gguf ~2.50 GB 4-bit quantized weights. Recommended for most users.
Hunter-4B-v1.0.imatrix.gguf ~3.9 MB Importance matrix used for quantization. Only needed to re-quantize.
SHA256SUMS Checksums for verifying downloads.

Verify your download before running:

sha256sum -c SHA256SUMS
Quick start
llama.cpp
llama-cli -m Hunter-4B-v1.0-Q4_K_M.gguf -p "Hello! Who are you?"
To serve over HTTP:
llama-server -m Hunter-4B-v1.0-Q4_K_M.gguf --port 8080
Python (llama-cpp-python)
from llama_cpp import Llama

llm = Llama(
    model_path="Hunter-4B-v1.0-Q4_K_M.gguf",
    n_ctx=4096,
    verbose=False,
)
print(llm("Hello! Who are you?", max_tokens=256)["choices"][0]["text"])
Ollama
Create a Modelfile with the ChatML template and stop tokens:
FROM ./Hunter-4B-v1.0-Q4_K_M.gguf

TEMPLATE """<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
{{ .Response }}<|im_end|>"""

PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.6
PARAMETER top_p 0.90
PARAMETER top_k 40
PARAMETER repeat_penalty 1.1

SYSTEM """You are Hunter-4B, a reasoning model developed by Lalakai Labs. Write out your reasoning inside <thought> ... </thought> tags before giving your final answer."""
ollama create hunter-4b -f Modelfile
ollama run hunter-4b
Hugging Face CLI
hf download Lalakai/Hunter-1-4B Hunter-4B-v1.0-Q4_K_M.gguf
Choosing a file
- Local chat on a typical laptop: Q4_K_M
- Highest output quality, if you have the memory: FP16
- Creating your own quantizations: FP16 plus the .imatrix file
Intended use
- Step-by-step explanations, math, and coding help
- Agentic coding and tool-use workflows on local hardware
- Summarization and structured multi-step analysis
- Security-oriented reasoning, such as vulnerability analysis and secure-code review, with human oversight
Evaluation
No formal benchmark results have been published for this model yet. They will be added to this card when available.
Limitations
- A 4B model will not match frontier-scale models on the hardest reasoning tasks or on niche knowledge. Use the FP16 weights and a larger context window (-c) when quality matters most.
- Like all language models, Hunter 4B can produce hallucinations, reflect biases in its training data, and generate unsafe content when prompted adversarially. Review outputs for your use case and apply your own safety filtering in production.
Responsible use
Use this model under human supervision and within applicable law and your organization's policies. Do not use it to target systems you are not authorized to test.
License
Released under the Apache License, Version 2.0 (https://www.apache.org/licenses/LICENSE-2.0).
Contact
Built by Lalakai Labs (https://x.com/LalakaiAI).
Acknowledgements
Hunter 4B is a fine-tune of Qwen 3.5 4B (https://huggingface.co/Qwen/Qwen3.5-4B) (Apache-2.0). We thank the Qwen team for the base model.
Downloads last month
252
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lalakai/Hunter-1-4B

Finetuned
Qwen/Qwen3.5-4B
Quantized
(513)
this model

Collection including Lalakai/Hunter-1-4B