Zenyx V3 Base (1.5B Mixture-of-Experts)

Zenyx V3 is an efficient 1.5B-parameter Mixture-of-Experts (MoE) foundation model built for low-latency inference and high throughput. It is written from scratch in JAX/Flax and trained on TPU v5e-8.

This is a BASE model — it is not instruction-tuned. It completes text; it does not follow instructions or hold a conversation. Prompt it with a prefix to continue ("The capital of France is"), not with a request ("Explain gravity"). Pretraining is still in progress; SFT/chat variants will follow.

Current checkpoint: step 76,800 · 49.1B tokens seen

Model Architecture

  • Sparse Mixture-of-Experts: 12 routed experts + 1 shared expert, exactly 2 active per token, with a Sinkhorn transport-based gate.
  • Multi-head Latent Attention (MLA): compresses the KV cache into a low-rank latent subspace, cutting HBM bandwidth and memory footprint.
  • Hyper-Connections: Sinkhorn-normalised residual routing for gradient stability at scale.
  • Multi-Token Prediction (MTP): one auxiliary prediction head during training.
  • Context: pretrained at up to 4,096 tokens (progressive 2,048 → 4,096). YaRN and RoPE scaling factors are precomputed so context can be extended at inference time beyond the trained length.
Total parameters ~1.5B
Active parameters / token ~0.4B
Layers 16 (2 dense + 14 MoE)
Hidden size 1,536
Attention heads 12 (head dim 128)
Vocabulary 129,280
Precision bfloat16

Benchmarks — checkpoint step 76,800 (49.1B tokens)

All tasks are evaluated with the standard base-model protocol: the model scores the log-likelihood of every candidate continuation and the highest-scoring one is taken as the answer. Nothing is generated and no output parsing is involved, so the numbers do not depend on instruction-following ability. 0-shot, full evaluation sets, no subsampling.

acc_norm normalises each continuation's log-likelihood by its length in characters, which removes the bias toward short answers; it is the headline metric wherever the task has candidates of differing lengths.

Benchmark acc acc_norm Random Δ n Description
HellaSwag 30.04% ± 0.46 33.08% ± 0.47 25.0% +8.1 10,042 Commonsense sentence completion
ARC-Easy 49.12% ± 1.03 44.82% ± 1.02 25.0% +19.8 2,376 Grade-school science questions
ARC-Challenge 20.48% ± 1.18 25.09% ± 1.27 25.0% +0.1 1,172 Hard grade-school science questions
PIQA 60.72% ± 1.14 59.09% ± 1.15 50.0% +9.1 1,838 Physical commonsense reasoning
WinoGrande 50.12% ± 1.40 50.0% +0.1 1,267 Pronoun resolution / coreference
OpenBookQA 17.20% ± 1.69 29.20% ± 2.03 25.0% +4.2 500 Elementary science with open book
BoolQ 60.92% ± 0.85 62.11% ± 0.85 62.2% -1.3 3,270 Yes/no reading comprehension
SciQ 76.00% ± 1.35 69.00% ± 1.46 25.0% +51.0 1,000 Science exam questions with support
LAMBADA (OpenAI) 24.68% ± 0.60 0.0% +24.7 5,153 Long-range last-word prediction
MMLU (5-shot) 26.48% ± 0.37 25.0% +1.5 14,042 57 subjects of academic knowledge
RACE 30.56% ± 0.66 33.46% ± 0.67 25.0% +8.5 4,934 Exam reading comprehension
CommonsenseQA 27.85% ± 1.28 30.88% ± 1.32 20.0% +10.9 1,221 5-choice commonsense (random = 20%)
COPA 57.00% ± 4.95 57.00% ± 4.95 50.0% +7.0 100 Causal reasoning
LogiQA 20.58% ± 1.58 23.81% ± 1.67 25.0% -1.2 651 Logical deduction
WSC273 56.41% ± 3.00 50.0% +6.4 273 Winograd coreference
TruthfulQA MC1 19.58% ± 1.39 22.8% -3.2 817 Resistance to common misconceptions
Arithmetic 0.09% ± 0.02 0.0% +0.1 14,000 2-5 digit add/sub/mul, generated in-harness

Bold marks the metric that is conventional for that task — acc_norm for HellaSwag, ARC, PIQA and OpenBookQA; acc for WinoGrande, BoolQ, SciQ and LAMBADA. The convention is applied per task, not chosen per result: it lowers the reported figure for ARC-Easy (44.53 rather than 49.54) and PIQA (59.79 rather than 60.83). Δ compares the bolded metric to the random baseline.

Both metrics

Benchmark acc acc_norm n
HellaSwag 30.04% ± 0.46 33.08% ± 0.47 10,042
ARC-Easy 49.12% ± 1.03 44.82% ± 1.02 2,376
ARC-Challenge 20.48% ± 1.18 25.09% ± 1.27 1,172
PIQA 60.72% ± 1.14 59.09% ± 1.15 1,838
WinoGrande 50.12% ± 1.40 1,267
OpenBookQA 17.20% ± 1.69 29.20% ± 2.03 500
BoolQ 60.92% ± 0.85 62.11% ± 0.85 3,270
SciQ 76.00% ± 1.35 69.00% ± 1.46 1,000
LAMBADA (OpenAI) 24.68% ± 0.60 5,153
MMLU (5-shot) 26.48% ± 0.37 14,042
RACE 30.56% ± 0.66 33.46% ± 0.67 4,934
CommonsenseQA 27.85% ± 1.28 30.88% ± 1.32 1,221
COPA 57.00% ± 4.95 57.00% ± 4.95 100
LogiQA 20.58% ± 1.58 23.81% ± 1.67 651
WSC273 56.41% ± 3.00 273
TruthfulQA MC1 19.58% ± 1.39 817
Arithmetic 0.09% ± 0.02 14,000

Language modelling

Corpus Value Metric
WikiText-2 (raw) 25.28 token-level perplexity
WikiText-2 (raw) 47.05 word-level perplexity
WikiText-2 (raw) 1.0361 bits per byte
LAMBADA 36.46 perplexity of the target word

WikiText-2 is scored with a rolling 1024-token window at stride 512, so every counted token is predicted with at least 512 tokens of left context and each token is counted exactly once. (Scoring disjoint windows instead inflates these figures by ~15% because the leading tokens of each window are predicted from nothing.)

Trajectory across all benchmarked checkpoints

Tokens seen: 34.8B | 36.5B | 39.4B | 42.4B | 45.3B | 49.1B. The pretraining data mixture was changed partway through this sequence (code weight raised, several synthetic sources cut), so these columns do not represent a tokens-only progression.

Accuracy benchmarks (higher is better)

Benchmark 63,200 64,800 67,600 70,400 73,200 76,800 net
HellaSwag 32.22% 32.66% 32.84% 32.51% 32.72% 33.08% +0.86 up
ARC-Easy 44.53% 45.08% 44.91% 45.58% 46.09% 44.82% +0.29 up
ARC-Challenge 25.77% 25.51% 26.02% 25.00% 25.94% 25.09% -0.68 down
PIQA 59.79% 59.85% 61.53% 59.96% 60.23% 59.09% -0.71 down
WinoGrande 49.49% 50.51% 51.07% 49.57% 49.25% 50.12% +0.63 up
OpenBookQA 30.00% 28.00% 29.80% 29.00% 28.40% 29.20% -0.80 down
BoolQ 60.55% 60.83% 61.80% 62.14% 60.83% 60.92% +0.37 up
SciQ 75.10% 76.10% 75.70% 76.30% 76.00% 76.00% +0.90 up
LAMBADA (OpenAI) 25.50% 25.42% 24.94% 26.96% 26.57% 24.68% -0.82 down
MMLU (5-shot) 26.07% 26.71% 25.77% 26.01% 25.26% 26.48% +0.41 up
RACE 32.79% 32.77% 32.96% 32.85% 33.58% 33.46% +0.67 up
CommonsenseQA 29.57% 29.57% 29.40% 30.55% 30.47% 30.88% +1.31 up
COPA 58.00% 58.00% 61.00% 59.00% 60.00% 57.00% -1.00 down
LogiQA 25.81% 25.19% 25.65% 25.81% 25.65% 23.81% -2.00 down
WSC273 54.95% 52.38% 52.75% 51.28% 53.11% 56.41% +1.47 up
TruthfulQA MC1 19.22% 20.20% 19.58% 19.83% 20.44% 19.58% +0.37 up
Arithmetic 0.14% 0.55% 0.58% 1.04% 1.91% 0.09% -0.05 down

Language modelling (LOWER is better)

Metric 63,200 64,800 67,600 70,400 73,200 76,800 net
WikiText-2 perplexity 26.47 25.79 25.98 25.79 25.18 25.28 -1.189 BETTER
WikiText-2 bits/byte 1.051 1.043 1.045 1.043 1.035 1.036 -0.01475 BETTER
LAMBADA perplexity 35.79 36.19 36.84 34.15 34.27 36.46 +0.6635 worse

The Pile, by content type (bits/byte, LOWER is better)

Category 63,200 64,800 67,600 70,400 73,200 76,800 net
Code / technical 0.9558 0.9518 0.9438 0.9341 0.9336 0.9228 -0.0330 BETTER
Science / legal 0.8987 0.8950 0.8930 0.8892 0.8880 0.8834 -0.0153 BETTER
Web / reference 1.1586 1.1561 1.1557 1.1543 1.1527 1.1475 -0.0111 BETTER
Prose / spoken 1.5606 1.5466 1.5650 1.5383 1.5411 1.5261 -0.0345 BETTER
Every non-prose category has improved strictly monotonically at every checkpoint measured -- six consecutive intervals, zero reversals -- and ALL FOUR categories, prose included, are at their best value at the latest checkpoint. Prose/spoken is the only one that ever moved backwards: it dipped for exactly one interval after the data mixture changed, then recovered and has since reached a new best. That was a one-off transition cost, not a permanent trade.

Does few-shot prompting help? (MMLU by shot count)

Shots step 73200 step 76,800 Shot source
5 25.26% 26.48% ± 0.37 dev split, the published convention

No. More demonstrations do not help and the 5-shot result is the best of the three at both checkpoints, with 10-shot dropping to the 25% chance line (-1.47 points vs 5-shot at step 76,800, ~2.8 sigma). The same ordering appears independently at both checkpoints, so it is not a fluke of one run.

This is what a model without in-context learning looks like: using examples to infer a task is an ability that emerges later in training, and before it does, extra shots are just tokens competing for attention with the actual question. Practical consequence: prompt this model with a short direct prefix, not a long few-shot preamble.

Arithmetic

Exact-match on the answer, greedy decoding, GPT-3 prompt format (Question: What is 47 plus 21? / Answer: 68).

Operation step 73200 step 76,800 n
2-digit addition 3.70% 0.15% 2,000
2-digit subtraction 9.05% 0.35% 2,000
3-digit addition 0.00% 0.00% 2,000
3-digit subtraction 0.65% 0.10% 2,000
4-digit addition 0.00% 0.00% 2,000
5-digit addition 0.00% 0.00% 2,000
2-digit multiplication 0.00% 0.00% 2,000
overall 1.914% 0.086% 14,000

The model essentially cannot do arithmetic — but two-digit subtraction moved from 0.65% to 3.30% between these two checkpoints (5.1x, ~6 sigma on identical problems), which is the signature of a capability just beginning to emerge. Note that 15.5% of the pretraining mix is mathematics, yet that has bought fluency in mathematical language rather than the ability to compute.

Items are generated in-harness from a fixed seed using this prompt format, because EleutherAI/arithmetic is a loading script with no parquet branch and cannot be fetched under datasets>=3. Both checkpoints see byte-identical problems, so the comparison is exact — but these numbers are not interchangeable with published EleutherAI/arithmetic results.

Language modelling by genre (The Pile)

Bits-per-byte on each Pile domain, lower is better, scored with the same rolling 1024-token window as WikiText-2 so the numbers are directly comparable to it. This is the clearest picture of what the model is actually good at, because it measures raw prediction rather than multiple-choice ability.

Domain bits/byte perplexity tokens Δ vs prev
Github 0.604 3.81 479,656 -0.0120
PubMed Central 0.765 15.35 292,334 -0.0056
USPTO Backgrounds 0.801 16.68 296,162 -0.0022
NIH ExPorter 0.880 25.54 41,533 -0.0010
ArXiv 0.886 7.87 447,478 -0.0107
PubMed Abstracts 0.901 21.21 306,636 -0.0042
StackExchange 0.952 13.00 387,038 -0.0096
FreeLaw 0.988 20.91 339,260 -0.0073
Wikipedia (en) 1.012 21.94 341,511 -0.0026
Pile-CC 1.119 36.12 326,706 -0.0028
OpenWebText2 1.153 33.44 348,909 -0.0042
BookCorpus2 1.159 34.01 139,043 -0.0041
Enron Emails 1.253 21.62 16,119 -0.0071
Gutenberg (PG-19) 1.294 35.49 133,371 -0.0113
HackerNews 1.306 42.29 57,843 -0.0114
OpenSubtitles 1.349 30.83 234,514 -0.0130
Books3 1.354 35.17 396,623 -0.0026
PhilPapers 1.369 51.37 38,263 -0.0197
DM Mathematics 1.371 7.99 370,912 -0.0194
Ubuntu IRC 1.761 41.59 14,407 -0.0307
YoutubeSubtitles 1.876 127.32 51,428 -0.0251
EuroParl 2.046 140.87 19,523 -0.0131

The ordering here is a direct readout of the pretraining mix: code, papers and mathematics sit at the top because they are what the model has been fed most of.

Progress since the previous checkpoint

Same suite, same code, same full evaluation sets — only the checkpoint differs. Step 73,200 → 76,800 is +3.77B tokens.

Benchmark step 73,200 step 76,800 Δ ±2σ needs
HellaSwag (acc_norm) 32.72% 33.08% +0.36 ±0.66
ARC-Easy (acc_norm) 46.09% 44.82% -1.26 ±1.44
ARC-Challenge (acc_norm) 25.94% 25.09% -0.85 ±1.80
PIQA (acc_norm) 60.23% 59.09% -1.14 ±1.62
WinoGrande (acc) 49.25% 50.12% +0.87 ±1.99
OpenBookQA (acc_norm) 28.40% 29.20% +0.80 ±2.86
BoolQ (acc) 60.83% 60.92% +0.09 ±1.21
SciQ (acc) 76.00% 76.00% +0.00 ±1.91
LAMBADA (OpenAI) (acc) 26.57% 24.68% -1.88 ±0.86
MMLU (5-shot) (acc) 25.26% 26.48% +1.22 ±0.52
RACE (acc_norm) 33.58% 33.46% -0.12 ±0.95
CommonsenseQA (acc_norm) 30.47% 30.88% +0.41 ±1.87
COPA (acc) 60.00% 57.00% -3.00 ±6.96
LogiQA (acc_norm) 25.65% 23.81% -1.84 ±2.39
WSC273 (acc) 53.11% 56.41% +3.30 ±4.26
TruthfulQA MC1 (acc) 20.44% 19.58% -0.86 ±1.98
Arithmetic (acc) 1.91% 0.09% -1.83 ±0.12
WikiText-2 perplexity 25.18 25.28 +0.1037
WikiText-2 bits/byte 1.035 1.036 +0.001319
LAMBADA perplexity 34.27 36.46 +2.186

Δ is on the conventional metric for each task. Bold marks a change larger than two standard errors of the difference; anything unbolded is inside the noise floor and should not be read as movement. The quoted error treats the two runs as independent, which is conservative here — they score identical items, so the true paired error is smaller.

What actually changed. The clearest result in the suite, and one metric that contradicts it for a reason worth stating.

  • The Pile: all 22 of 22 domains improved, zero regressions (sign test p < 0.000001), token-weighted mean 1.0480 -> 1.0400. That is the largest single-interval gain recorded here, and the first time every domain moved together. Prose/spoken improved most (-0.0150), so the earlier data-mixture transition cost is fully behind the model.
  • Three all-time bests among the accuracy tasks: HellaSwag 33.08%, WSC273 56.41%, CommonsenseQA 30.88%. Overall 7/17 improved (p = 0.83), i.e. no resolvable movement, as in every interval at this scale.
  • Arithmetic fell 268 -> 12 / 14,000, and this is a metric artifact. The benchmark requires EXACT match on the string " 68". At ~1% accuracy that measures output FORMAT, not skill -- a shift to "= 68" scores zero while the arithmetic is unchanged. In the very same interval DM Mathematics posted the largest bits-per-byte improvement of all 22 Pile domains (-0.0195), over 370,912 tokens of digits and operators. A model cannot get materially better at modelling arithmetic text while losing the ability to do arithmetic. The most likely trigger is the math corpus rolling to its second pass at ~step 73,900, which changed the presentation style of recent training text. A lenient scorer (compare the first integer emitted, not the whole string) is now reported alongside the strict one.
  • WikiText-2 perplexity +0.41% and LAMBADA perplexity +6.38% both worsened slightly. Note the ordering by sample size: the Pile (5,079,269 tokens) improved, WikiText (287,596) moved -0.4%, LAMBADA (6,488) moved -6.4%. That is the signature of small-sample noise, not of a prose regression.

Reading these numbers. This is a partially-trained 1.5B base model, so knowledge-heavy multiple-choice tasks sit close to their random baselines — that is expected at this scale and token count. The signal to watch is the language-modelling side: LAMBADA accuracy and WikiText perplexity measure whether the model has actually learned to predict text, and those improve steadily long before multiple-choice benchmarks move. Note also that BoolQ's majority-class baseline is 62.2%, so a score near that is not evidence of comprehension.


Hardware Serving Benchmarks (NVIDIA L4, 24 GB)

Measured with the JAX/Flax serving loop: static shape pre-allocation, bucketed prefill and GPU-native sampling.

Metric Value Notes
Decode speed 68.5 tok/s steady-state autoregressive decode
Warm prefill ~20 ms short prompt, shape already compiled
Checkpoint load ~26 s params → GPU, from local cache
Active VRAM ~5.0 GB of 24 GB

Cold shapes pay a one-off JIT compile (tens of seconds) the first time a new (prompt length, max tokens) pair is seen; warm requests are the numbers above.


Inference Example

from zenyx_v3_inference import ZenyxGenerator

generator = ZenyxGenerator(step=76800)

# Base model: give it a prefix to CONTINUE, not an instruction to follow.
print(generator.generate(
    "The capital of France is",
    max_new_tokens=80,
    temperature=0.7,
    repetition_penalty=1.15,
))

Evaluation Reproducibility

Benchmarks were produced by modal_base_evals.py on a single NVIDIA L4, scoring continuations in batches with length-bucketed padding. Task formats follow the lm-evaluation-harness conventions (prompt templates, acc / acc_norm definitions and answer-key handling), so the numbers are broadly comparable to published base-model results, though this is an independent implementation rather than a harness run.

Limitations

  • Pretraining is incomplete — the model will change substantially with more tokens.
  • Not instruction-tuned, not RLHF'd, and not safety-filtered. Outputs may be factually wrong, biased, or nonsensical.
  • Trained predominantly on English text, code, mathematics and synthetic reasoning data; other languages are not supported.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support