Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing Paper • 2607.07953 • Published 12 days ago • 14
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models Paper • 2607.05803 • Published 13 days ago • 10
LLM-as-a-Verifier: A General-Purpose Verification Framework Paper • 2607.05391 • Published 14 days ago • 15
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients Paper • 2606.18216 • Published Jun 16 • 63
Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning Paper • 2602.21103 • Published Jun 2 • 9
CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning Paper • 2605.28742 • Published May 27 • 4
Reinforcement Learning from Rich Feedback with Distributional DAgger Paper • 2606.05152 • Published Jun 3 • 3
On the Scaling of PEFT: Towards Million Personal Models of Trillion Parameters Paper • 2606.02437 • Published Jun 1 • 239
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Paper • 2605.29548 • Published May 28 • 12
SkillOpt: Executive Strategy for Self-Evolving Agent Skills Paper • 2605.23904 • Published May 22 • 257
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time Paper • 2604.11626 • Published Apr 13 • 103
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass Paper • 2604.10966 • Published Apr 13 • 12
Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation Paper • 2604.13010 • Published Apr 14 • 20
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 114