qwen3-0.6b-math-opd-frombase
Qwen3-0.6B-base post-trained by on-policy distillation (OPD) toward a frozen Qwen3-4B teacher on MATH, from qwen3-nano-math-reasoner.
Result β MATH-500 (full, greedy, HF runtime, max_new_tokens=2048)
| Model | Accuracy |
|---|---|
| Qwen3-0.6B base | ~16% |
| this checkpoint (OPD-from-base, step 80) | 28.8% (144/500) |
| SFT seed (Qwen3-235B offline distill) | 37.8% |
A ~1.8x rise over base. The student samples its own rollouts; the frozen Qwen3-4B teacher (shared tokenizer) scores every token; the student minimizes a per-token JSD to the teacher with a k3 trust-region anchor toward the base init.
Training
- Init: Qwen3-0.6B base; Teacher:
Qwen/Qwen3-4B(frozen) - Loss: per-token JSD(teacher||student) + k3 anchor + entropy floor
- lr 2e-6, kl_coeff 0.05, num_rollouts 4, max_new_tokens 768
- Early-stopped at step 80 β the accuracy curve is non-monotonic (14% -> 26% -> 30% peak -> 20%); past the peak the fragile base over-trains and degrades.
- Trainer:
src/qwen_opd.py(--student_checkpoint base).
Honest caveat
This mostly elicits and formats the latent math Qwen3's pretraining already deposited (reasoning stage + Qwen2.5-Math synthetic data), rather than teaching new capability β which is why it lands below the offline-SFT seed. Format is imperfect (rambles to the token cap).
Scratch-format .pth; load via the repo's load_hf_qwen_model.