qwen3-0.6b-math-opd-frombase

Qwen3-0.6B-base post-trained by on-policy distillation (OPD) toward a frozen Qwen3-4B teacher on MATH, from qwen3-nano-math-reasoner.

Result β€” MATH-500 (full, greedy, HF runtime, max_new_tokens=2048)

Model Accuracy
Qwen3-0.6B base ~16%
this checkpoint (OPD-from-base, step 80) 28.8% (144/500)
SFT seed (Qwen3-235B offline distill) 37.8%

A ~1.8x rise over base. The student samples its own rollouts; the frozen Qwen3-4B teacher (shared tokenizer) scores every token; the student minimizes a per-token JSD to the teacher with a k3 trust-region anchor toward the base init.

Training

  • Init: Qwen3-0.6B base; Teacher: Qwen/Qwen3-4B (frozen)
  • Loss: per-token JSD(teacher||student) + k3 anchor + entropy floor
  • lr 2e-6, kl_coeff 0.05, num_rollouts 4, max_new_tokens 768
  • Early-stopped at step 80 β€” the accuracy curve is non-monotonic (14% -> 26% -> 30% peak -> 20%); past the peak the fragile base over-trains and degrades.
  • Trainer: src/qwen_opd.py (--student_checkpoint base).

Honest caveat

This mostly elicits and formats the latent math Qwen3's pretraining already deposited (reasoning stage + Qwen2.5-Math synthetic data), rather than teaching new capability β€” which is why it lands below the offline-SFT seed. Format is imperfect (rambles to the token cap).

Scratch-format .pth; load via the repo's load_hf_qwen_model.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for jeetganatra/qwen3-0.6b-math-opd-frombase

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1345)
this model