Human Feedback Types
provisional (inferred)None explicit
No explicit feedback protocol extracted.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
HFEPX · Eval paper review
Ben Rank, Hardik Bhatnagar, Ameya Prabhu, Shira Eisenberg +3 more
Published
Mar 9, 2026
Citations
0
Trust level
Provisional
Usefulness score
Unavailable
Extraction confidence
0% (Provisional)
Derived from abstract and metadata only.
Signals refreshed
Mar 10, 2026
Signal extraction is still processing. This page currently shows metadata-first guidance until structured protocol fields are ready.
This page is a lightweight research summary built from the abstract and metadata while deeper extraction catches up.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A provisional background reference while structured extraction finishes.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
This page is still relying on abstract and metadata signals, not a fuller protocol read.
Eval-fit score is unavailable until extraction completes.
If you are doing eval pipeline work, start here
AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities. This raises a deeper question: can these systems extend their capabilities to automate AI research itself? In this paper, we explore post-training, the critical phase that turns base LLMs into useful assistants. We introduce PostTrainBench to benchmark how well LLM agents can perform post-training autonomously under bounded compute constraints (10 hours on one H100 GPU). We ask frontier agents (e.g., Claude Code with Opus 4.6) to optimize the performance of a base LLM on a particular benchmark (e.g., Qwen3-4B on AIME). Importantly, we do not provide any predefined strategies to the agents and instead give them full autonomy to find necessary information on the web, run experiments, and curate data. We find that frontier agents make substantial progress but generally lag behind instruction-tuned LLMs from leading providers: 23.2% for the best agent vs. 51.1% for official instruction-tuned models. However, agents can exceed instruction-tuned models in targeted scenarios: GPT-5.1 Codex Max achieves 89% on BFCL with Gemma-3-4B vs. 67% for the official model. We also observe several failure modes worth flagging. Agents sometimes engage in reward hacking: training on the test set, downloading existing instruction-tuned checkpoints instead of training their own, and using API keys they find to generate synthetic data without authorization. These behaviors are concerning and highlight the importance of careful sandboxing as these systems become more capable. Overall, we hope PostTrainBench will be useful for tracking progress in AI R&D automation and for studying the risks that come with it. Website and code are available at https://posttrainbench.com/.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
Tool Use evaluation
Includes extracted eval setup.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
Not reported
No explicit QC controls found.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
Not extracted
No benchmark anchors detected.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
Not extracted
No metric anchors detected.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
Unknown
Rater source not explicitly reported.
"AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities."
This page is using abstract-level cues only right now. Treat the signals below as provisional.
Evaluation fields are inferred from the abstract only.
AI agents have become surprisingly proficient at software engineering over the past year, largely due to improvements in reasoning capabilities.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.