Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning Paper • 2609.03430 • Published 2 days ago • 125
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning Paper • 2609.00638 • Published 4 days ago • 70
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation Paper • 2608.24138 • Published 11 days ago • 13
Where to Look Matters: On-Policy Self-Distillation for Long-Video Understanding Paper • 2608.25356 • Published 10 days ago • 20
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 16 days ago • 274
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies Paper • 2607.03723 • Published Jul 4 • 5
Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling Paper • 2606.03102 • Published Jun 2 • 15
Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scaling Paper • 2605.27030 • Published May 26 • 33
Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Use Paper • 2605.14038 • Published May 13 • 15
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Paper • 2605.08083 • Published May 8 • 71
EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents Paper • 2605.13941 • Published May 13 • 25
G-Zero: Self-Play for Open-Ended Generation from Zero Data Paper • 2605.09959 • Published May 11 • 18
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Paper • 2605.09269 • Published May 10 • 6
Reinforcing Multimodal Reasoning Against Visual Degradation Paper • 2605.09262 • Published May 10 • 7
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling Paper • 2605.08083 • Published May 8 • 71
On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows Paper • 2605.06110 • Published May 7 • 17