IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
Abstract
The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models.
A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.
Community
Excited to share IdeaAMBIG!
We ask a simple question: when an AI system is given a research idea, can it tell whether the idea is actually specified well enough to implement?
We introduce a benchmark of 660 evidence-grounded specification gaps and evaluate 13 LLMs on detecting, localizing, and resolving implementation-critical defects.
Our most striking result: LLMs are much better at fixing a gap than finding it. On real-world cases, the best model achieves only 9.6% defect recovery, but 80.6% clarification success once the defect is given. Providing the gold resolution raises downstream codification readiness from 14% to 98%.
This points to an important bottleneck for research agents: before asking models to implement an idea, we may first need them to reliably recognize what is missing, ambiguous, or inconsistent.
Would love to hear the community’s thoughts on this!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents (2026)
- ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents (2026)
- RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists (2026)
- From Discussion to Execution: Replicating Buggy and Correct Data Science Code (2026)
- SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents (2026)
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring (2026)
- SCIRIGOR:Evaluating Open-Ended Scientific Analysis Beyond Final Scores (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.10539 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper