Human Feedback Types
strongExpert Verification
Directly usable for protocol triage.
"Efficient reproduction of research papers is pivotal to accelerating scientific progress."
HFEPX · Eval paper review
Xuanle Zhao, Zilin Sang, Yuxuan Li, Qi Shi +6 more
Published
May 27, 2025
Citations
0
Trust level
Moderate
Usefulness score
50/100 (Medium)
Extraction confidence
60% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Apr 24, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Efficient reproduction of research papers is pivotal to accelerating scientific progress. However, the increasing complexity of proposed methods often renders reproduction a labor-intensive endeavor, necessitating profound domain expertise. To address this, we introduce the paper lineage, which systematically mines implicit knowledge from the cited literature. This algorithm serves as the backbone of our proposed \ours, a multi-agent framework designed to autonomously reproduce experimental code in a complete, end-to-end manner. To ensure code executability, \ours incorporates a sampling-based unit testing strategy for rapid validation. To assess reproduction capabilities, we introduce \ourbench, a benchmark featuring verified implementations, alongside comprehensive metrics for evaluating both reproduction and execution fidelity. Extensive evaluations on PaperBench and \ourbench demonstrate that \ours consistently surpasses existing baselines across all metrics. Notably, it yields substantial improvements in reproduction fidelity and final execution performance. The code is available at https://github.com/AI9Stars/AutoReproduce.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Expert Verification
Directly usable for protocol triage.
"Efficient reproduction of research papers is pivotal to accelerating scientific progress."
None explicit
Validate eval design from full paper text.
"Efficient reproduction of research papers is pivotal to accelerating scientific progress."
Not reported
No explicit QC controls found.
"Efficient reproduction of research papers is pivotal to accelerating scientific progress."
Ourbench, Paperbench
Useful for quick benchmark comparison.
"To assess reproduction capabilities, we introduce \ourbench, a benchmark featuring verified implementations, alongside comprehensive metrics for evaluating both reproduction and execution fidelity."
Not extracted
No metric anchors detected.
"Efficient reproduction of research papers is pivotal to accelerating scientific progress."
Domain Experts
Helpful for staffing comparability.
"However, the increasing complexity of proposed methods often renders reproduction a labor-intensive endeavor, necessitating profound domain expertise."
No metric terms were extracted from the available abstract.
Efficient reproduction of research papers is pivotal to accelerating scientific progress.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Expert Verification
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Ourbench, Paperbench
Metric reporting is present
No metric terms extracted.