icedduck/lab1-exRL_bin

coding stage ②-B: SFT + GRPO, binary reward, strong observation; pass@1 .446 with tests, .295 without, re-measured .459 / .314 (report §3.14, §3.15).

Part of LLM from scratch, to the aha moment, to a coding agent: fifteen experiments on four RTX 4090s. Code, report, the evaluation protocol and the scripts that produced this checkpoint: https://github.com/bethehand/lab1-llm-scratch-to-aha-moment-to-coding-agent

Evaluate it yourself from the repository root, for example python3 bench_observation_price.py --hf-dir icedduck/lab1-exRL_bin for the coding-line checkpoints. Weights are derived from Qwen/Qwen2.5-Coder-1.5B and remain under the Qwen license.

Downloads last month
6
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for icedduck/lab1-exRL_bin

Finetuned
(65)
this model