X3 Banner

GPT-X3-150M is Axiomic Labs’ third-generation base model, built on the TX5 hybrid GDN2/gated GQA architecture and trained from scratch with TrainWork. It ranks #1 on the Open SLM Leaderboard and is the first model under 150M parameters to exceed an Intelligence Index of 30.

Performance

Data Efficiency

Intelligence Index vs Training Compute

GPT-X3-150M outperforms Hugging Face's SmolLM2-135M on the Intelligence Index while using about 27× fewer training tokens (75B vs 2T). Benchmark results and Intelligence Index methodology follow the Open SLM Leaderboard. Training compute is estimated as 6 × parameters × training tokens.

Relative Leaderboard Performance

Benchmark Performance vs Competitors

Organization Model Parameters Int Index HellaSwag ARC-Easy ARC-Challenge PIQA ArithMark-3
Axiomic Labs GPT-X3-150M 149M 33.05 46.16% 59.64% 31.31% 71.33% 50.90%
LiquidAI LFM2.5-350M-Base 354.5M 30.50 44.98% 68.86% 40.96% 67.68% 35.90%
OpenCerebral Boris-2-Preview-1 144M 27.91 45.23% 58.38% 29.95% 69.21% 37.60%
Bench Labs cagliostro-v3.5 146.35M 27.49 43.41% 53.75% 29.35% 68.06% 45.30%
HuggingFace SmolLM2-135M 135M 27.13 43.22% 58.63% 29.69% 68.44% 39.20%
Bench Labs cagliostro-v3 146.35M 26.55 42.51% 54.88% 28.75% 67.46% 43.70%
Axiomic Labs GPT-X2.5-135M 135M 25.17 40.57% 51.81% 29.18% 69.42% 38.40%
Google Gemma-3-270M 268.1M 25.16 41.33% 57.11% 28.24% 68.66% 35.60%
BananaMind BananaMind-2-Pro 139M 24.96 42.78% 53.58% 27.82% 67.52% 38.20%
Facebook MobileLLM-R1-140M-base 140M 24.64 33.84% 49.92% 24.74% 63.22% 65.70%
LiquidAI LFM2.5-230M-Base 229.7M 23.96 39.39% 59.55% 35.41% 63.60% 37.80%
Axiomic Labs GPT-X2-125M 125M 23.36 40.41% 51.47% 27.82% 67.30% 37.20%
Axiomic Labs GPT-X-125M 125M 19.94 36.57% 50.76% 26.62% 64.96% 35.60%
OpenAI GPT-2 Medium 354.8M 19.68 39.33% 43.35% 24.91% 66.10% 34.60%
SupraLabs Supra2-100M-Base 100.7M 19.41 35.98% 47.81% 24.83% 65.40% 36.90%
OpenAI GPT-2 (124M) 124M 13.58 31.26% 39.35% 22.35% 62.08% 35.70%

The Intelligence Index chance-normalizes HellaSwag, combined ARC, PIQA, and ArithMark-3, then applies weights of 1.00, 1.00, 1.00, and 0.65 respectively.


What's New in X3 vs GPT-X2.5?

Model size and architecture

  • 30 XGQA layers -> 24 hybrid layers (18 GDN2 + 6 gated GQA).
    • Three GDN2 layers between each full-attention layer make the stack a 3:1 hybrid.
  • XSA -> normalized, gated GQA with RoPE on 32/64 head dimensions.
    • Adds query/key normalization and an output gate to X3's full-attention layers. Ablations showed XSA redundant with attention gating.
  • 32,770 -> 65,536 vocabulary tokens; uses the LFM2.5 tokenizer.
    • Larger tokenizer yields across the board improvements including speed, training efficiency, long context performance, etc.

Training and optimization

  • 11-source progressive mix -> nine-source fixed mix.
    • X3 keeps the source proportions constant throughout training as small models struggle to adapt to data changes.
  • Next-token training -> six-token TST for the first 30%, then next-token training; 75B compute tokens and ~187.5B raw tokens total.
    • TST predicts six target tokens per compute position during its phase.
  • AdamW -> hybrid Muon + AdamW.
    • Muon handles matrix weights; AdamW handles embeddings, norms, biases, and other non-matrix parameters.

Architecture

Component Details
Architecture TX5
Tokenizer LFM2.5 65k (with BOS removed)
Layers 24
Training Tokens 75B
Attention GDN2:GQA
Attention Layout 3:1
QK Normalization OffsetRMSNorm on queries and keys in GQA layers
Embeddings Tied
Activation SiLU
Normalization RMSNorm
Dropout 0

Config

vocab_size             = 65,536    (LFM2.5 tokenizer no BOS)
n_layer                = 24
n_embd                 = 576
block_size             = 2048
gdn_head_dim           = 64        (9 GDN2 heads)
gqa_head_dim           = 64        (9 query heads)
gqa_n_kv_heads         = 3         (3:1 grouped-query attention)
gqa_rotary_dim         = 32
gdn_layers_per_gqa     = 3         (GDN2:GQA = 3:1)
swiglu_hidden_ratio    = 2.75      (1,584 units)
rope_theta             = 100,000

Parameter Breakdown

Component Params
Token embeddings (65,536 x 576; LM head tied) 37,748,736
GDN2 attention blocks (18) 38,632,482
Gated GQA blocks (6) 7,299,840
SwiGLU MLPs (24) 65,691,648
Block RMSNorms 27,648
Final RMSNorm 576
Total 149,400,930

Training

GPT-X3 is trained for 143,052 steps across two GPUs, with a global batch of 524,288 compute tokens per step. The run uses six-token TST for the first 30% of training, then switches to standard next-token prediction.

Data

The configured source folders are mixed from step 0 with one fixed set of weights; there are no later curriculum stages. The run uses 143,052 steps at a 524,288-token global batch (about 75.0004B compute tokens). Token-Superposition Training (TST) is enabled for the first 30% of the run with six-token bags, consuming about 135B raw tokens in that phase before standard autoregressive training resumes.

Token Superposition Training (TST)

During TST, embeddings for each group of six input tokens are averaged into one model position, and the model predicts the six tokens in the next group.

Training phase Share Objective Compute tokens Raw tokens
TST (~42,916 steps) 30% Predict the next group of six tokens ~22.5B ~135B
Autoregressive (~100,136 steps) 70% Predict the next token ~52.5B ~52.5B

Optimization

  • Optimizer: Hybrid Muon and AdamW. Muon handles 2D weight matrices (max LR 0.02, momentum 0.95, Nesterov, five Newton-Schulz steps); AdamW handles embeddings, normalization parameters, biases, and other non-matrix parameters (max LR 1.5e-3, betas 0.9/0.95).
  • Weight decay: 0.01 on Muon parameters; AdamW's parameter group uses 0 weight decay.
  • Learning-rate schedule: 2,000-step linear warmup, stable until 82.5% of training, then a 1-sqrt cooldown to zero.
  • Batch size: 524,288 compute tokens per optimizer step, distributed across two GPUs with gradient accumulation.
  • Sequence length: 2,048 tokens.
  • Precision and stability: bfloat16 mixed precision, global gradient-norm clipping at 1.0.

Hardware

  • 2x RTX 3080 Ti
  • Training time: ~330 hours

Usage

GPT-X3-150M is a base model for text completion. Give it a passage to continue, such as the beginning of a paragraph.

The following Transformers example is a draft for the released checkpoint; it requires the weights, tokenizer, and Transformers model code to be published first.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AxiomicLabs/GPT-X3-150M"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=torch.float32,
).to("cuda").eval()

prompt = "Artificial intelligence is"
inputs = tokenizer(
    prompt,
    return_tensors="pt",
    add_special_tokens=False,
).to(model.device)

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=120,
        do_sample=True,
        temperature=0.8,
        top_p=0.95,
    )

print(tokenizer.decode(output[0], skip_special_tokens=True))

Citation

@misc{gptx3_2026,
  title={GPT-X3-150M},
  author={Axiomic Labs},
  year={2026},
  howpublished={\url{https://huggingface.co/AxiomicLabs/GPT-X3-150M}},
}
Downloads last month
338
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train AxiomicLabs/GPT-X3-150M

Space using AxiomicLabs/GPT-X3-150M 1