DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation Paper • 2503.01622 • Published Mar 3, 2025
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy Paper • 2507.18392 • Published Jul 24, 2025 • 20
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration Paper • 2604.12843 • Published Apr 15 • 1
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks Paper • 2605.28556 • Published May 27 • 75
Bamba Collection Collection of Bamba - hybrid Mamba2 model architecture based models trained on open data • 7 items • Updated Mar 2 • 24