Instructions to use AxiomicLabs/GPT-X3-150M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AxiomicLabs/GPT-X3-150M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AxiomicLabs/GPT-X3-150M", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AxiomicLabs/GPT-X3-150M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AxiomicLabs/GPT-X3-150M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AxiomicLabs/GPT-X3-150M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxiomicLabs/GPT-X3-150M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/AxiomicLabs/GPT-X3-150M
- SGLang
How to use AxiomicLabs/GPT-X3-150M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AxiomicLabs/GPT-X3-150M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxiomicLabs/GPT-X3-150M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AxiomicLabs/GPT-X3-150M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxiomicLabs/GPT-X3-150M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use AxiomicLabs/GPT-X3-150M with Docker Model Runner:
docker model run hf.co/AxiomicLabs/GPT-X3-150M
GPT-X3-150M is Axiomic Labs’ third-generation base model, built on the TX5 hybrid GDN2/gated GQA architecture and trained from scratch with TrainWork. It ranks #1 on the Open SLM Leaderboard and is the first model under 150M parameters to exceed an Intelligence Index of 30.
Performance
Data Efficiency
GPT-X3-150M outperforms Hugging Face's SmolLM2-135M on the Intelligence Index while using about 27× fewer training tokens (75B vs 2T).
Benchmark results and Intelligence Index methodology follow the Open SLM Leaderboard. Training compute is estimated as 6 × parameters × training tokens.
Relative Leaderboard Performance
| Organization | Model | Parameters | Int Index | HellaSwag | ARC-Easy | ARC-Challenge | PIQA | ArithMark-3 |
|---|---|---|---|---|---|---|---|---|
| Axiomic Labs | GPT-X3-150M | 149M | 33.05 | 46.16% | 59.64% | 31.31% | 71.33% | 50.90% |
| LiquidAI | LFM2.5-350M-Base | 354.5M | 30.50 | 44.98% | 68.86% | 40.96% | 67.68% | 35.90% |
| OpenCerebral | Boris-2-Preview-1 | 144M | 27.91 | 45.23% | 58.38% | 29.95% | 69.21% | 37.60% |
| Bench Labs | cagliostro-v3.5 | 146.35M | 27.49 | 43.41% | 53.75% | 29.35% | 68.06% | 45.30% |
| HuggingFace | SmolLM2-135M | 135M | 27.13 | 43.22% | 58.63% | 29.69% | 68.44% | 39.20% |
| Bench Labs | cagliostro-v3 | 146.35M | 26.55 | 42.51% | 54.88% | 28.75% | 67.46% | 43.70% |
| Axiomic Labs | GPT-X2.5-135M | 135M | 25.17 | 40.57% | 51.81% | 29.18% | 69.42% | 38.40% |
| Gemma-3-270M | 268.1M | 25.16 | 41.33% | 57.11% | 28.24% | 68.66% | 35.60% | |
| BananaMind | BananaMind-2-Pro | 139M | 24.96 | 42.78% | 53.58% | 27.82% | 67.52% | 38.20% |
| MobileLLM-R1-140M-base | 140M | 24.64 | 33.84% | 49.92% | 24.74% | 63.22% | 65.70% | |
| LiquidAI | LFM2.5-230M-Base | 229.7M | 23.96 | 39.39% | 59.55% | 35.41% | 63.60% | 37.80% |
| Axiomic Labs | GPT-X2-125M | 125M | 23.36 | 40.41% | 51.47% | 27.82% | 67.30% | 37.20% |
| Axiomic Labs | GPT-X-125M | 125M | 19.94 | 36.57% | 50.76% | 26.62% | 64.96% | 35.60% |
| OpenAI | GPT-2 Medium | 354.8M | 19.68 | 39.33% | 43.35% | 24.91% | 66.10% | 34.60% |
| SupraLabs | Supra2-100M-Base | 100.7M | 19.41 | 35.98% | 47.81% | 24.83% | 65.40% | 36.90% |
| OpenAI | GPT-2 (124M) | 124M | 13.58 | 31.26% | 39.35% | 22.35% | 62.08% | 35.70% |
The Intelligence Index chance-normalizes HellaSwag, combined ARC, PIQA, and ArithMark-3, then applies weights of 1.00, 1.00, 1.00, and 0.65 respectively.
What's New in X3 vs GPT-X2.5?
Model size and architecture
- 30 XGQA layers -> 24 hybrid layers (18 GDN2 + 6 gated GQA).
- Three GDN2 layers between each full-attention layer make the stack a 3:1 hybrid.
- XSA -> normalized, gated GQA with RoPE on 32/64 head dimensions.
- Adds query/key normalization and an output gate to X3's full-attention layers. Ablations showed XSA redundant with attention gating.
- 32,770 -> 65,536 vocabulary tokens; uses the LFM2.5 tokenizer.
- Larger tokenizer yields across the board improvements including speed, training efficiency, long context performance, etc.
Training and optimization
- 11-source progressive mix -> nine-source fixed mix.
- X3 keeps the source proportions constant throughout training as small models struggle to adapt to data changes.
- Next-token training -> six-token TST for the first 30%, then next-token training; 75B compute tokens and ~187.5B raw tokens total.
- TST predicts six target tokens per compute position during its phase.
- AdamW -> hybrid Muon + AdamW.
- Muon handles matrix weights; AdamW handles embeddings, norms, biases, and other non-matrix parameters.
Architecture
| Component | Details |
|---|---|
| Architecture | TX5 |
| Tokenizer | LFM2.5 65k (with BOS removed) |
| Layers | 24 |
| Training Tokens | 75B |
| Attention | GDN2:GQA |
| Attention Layout | 3:1 |
| QK Normalization | OffsetRMSNorm on queries and keys in GQA layers |
| Embeddings | Tied |
| Activation | SiLU |
| Normalization | RMSNorm |
| Dropout | 0 |
Config
vocab_size = 65,536 (LFM2.5 tokenizer no BOS)
n_layer = 24
n_embd = 576
block_size = 2048
gdn_head_dim = 64 (9 GDN2 heads)
gqa_head_dim = 64 (9 query heads)
gqa_n_kv_heads = 3 (3:1 grouped-query attention)
gqa_rotary_dim = 32
gdn_layers_per_gqa = 3 (GDN2:GQA = 3:1)
swiglu_hidden_ratio = 2.75 (1,584 units)
rope_theta = 100,000
Parameter Breakdown
| Component | Params |
|---|---|
| Token embeddings (65,536 x 576; LM head tied) | 37,748,736 |
| GDN2 attention blocks (18) | 38,632,482 |
| Gated GQA blocks (6) | 7,299,840 |
| SwiGLU MLPs (24) | 65,691,648 |
| Block RMSNorms | 27,648 |
| Final RMSNorm | 576 |
| Total | 149,400,930 |
Training
GPT-X3 is trained for 143,052 steps across two GPUs, with a global batch of 524,288 compute tokens per step. The run uses six-token TST for the first 30% of training, then switches to standard next-token prediction.
Data
The configured source folders are mixed from step 0 with one fixed set of weights; there are no later curriculum stages. The run uses 143,052 steps at a 524,288-token global batch (about 75.0004B compute tokens). Token-Superposition Training (TST) is enabled for the first 30% of the run with six-token bags, consuming about 135B raw tokens in that phase before standard autoregressive training resumes.
| Organization | Source | Weight |
|---|---|---|
| Axiomic Labs | SFTset-SLM | 1% |
| IFM | TXT360-v2 | 40% |
| OpenBMB | UltraMath | 10% |
| OpenBMB | UltraCode | 2% |
| Axiomic Labs | FactBase-DCLM | 1% |
| NVIDIA | ClimbMix | 43% |
| NVIDIA | OpenMathInstruct | 2% |
| Various | Misc. | 0.5% |
| Microsoft | OrcaMath | 0.5% |
Token Superposition Training (TST)
During TST, embeddings for each group of six input tokens are averaged into one model position, and the model predicts the six tokens in the next group.
| Training phase | Share | Objective | Compute tokens | Raw tokens |
|---|---|---|---|---|
| TST (~42,916 steps) | 30% | Predict the next group of six tokens | ~22.5B | ~135B |
| Autoregressive (~100,136 steps) | 70% | Predict the next token | ~52.5B | ~52.5B |
Optimization
- Optimizer: Hybrid Muon and AdamW. Muon handles 2D weight matrices (max LR 0.02, momentum 0.95, Nesterov, five Newton-Schulz steps); AdamW handles embeddings, normalization parameters, biases, and other non-matrix parameters (max LR 1.5e-3, betas 0.9/0.95).
- Weight decay: 0.01 on Muon parameters; AdamW's parameter group uses 0 weight decay.
- Learning-rate schedule: 2,000-step linear warmup, stable until 82.5% of training, then a
1-sqrtcooldown to zero. - Batch size: 524,288 compute tokens per optimizer step, distributed across two GPUs with gradient accumulation.
- Sequence length: 2,048 tokens.
- Precision and stability: bfloat16 mixed precision, global gradient-norm clipping at 1.0.
Hardware
- 2x RTX 3080 Ti
- Training time: ~330 hours
Usage
GPT-X3-150M is a base model for text completion. Give it a passage to continue, such as the beginning of a paragraph.
The following Transformers example is a draft for the released checkpoint; it requires the weights, tokenizer, and Transformers model code to be published first.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AxiomicLabs/GPT-X3-150M"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.float32,
).to("cuda").eval()
prompt = "Artificial intelligence is"
inputs = tokenizer(
prompt,
return_tensors="pt",
add_special_tokens=False,
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=120,
do_sample=True,
temperature=0.8,
top_p=0.95,
)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Citation
@misc{gptx3_2026,
title={GPT-X3-150M},
author={Axiomic Labs},
year={2026},
howpublished={\url{https://huggingface.co/AxiomicLabs/GPT-X3-150M}},
}
- Downloads last month
- 338


