Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing."
HFEPX · Eval paper review
Hovhannes Tamoyan, Sean Narenthiran, Erik Arakelyan, Mira Mezini +1 more
Published
Jun 23, 2026
Citations
0
Trust level
Moderate
Usefulness score
25/100 (Low)
Extraction confidence
55% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Jun 23, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A benchmark-and-metrics comparison anchor.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have emerged, yet are still evaluated as file retrieval rather than actionable diagnosis, producing locations without the diagnostic context a repair agent needs. We introduce SHERLOC (Structured Hypothesis-driven Exploration and Reasoning for Localization), a training-free framework pairing a reasoning LLM with compact repository tools and self-recovery, without fine-tuning or multi-agent orchestration. SHERLOC reaches state-of-the-art localization across model scales: 84.33% accuracy@1 on SWE-Bench Lite and 81.27% recall@1 on SWE-Bench Verified; at ~30B parameters, it matches or outperforms other agentic methods. Injecting our locations and diagnostic findings into repair agents yields, on average, +5.95 pp resolve rate on SWE-Bench Verified while cutting localization and total tokens by 36.7% and 23.1%.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing."
Automatic Metrics
Includes extracted eval setup.
"LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing."
Not reported
No explicit QC controls found.
"LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing."
SWE Bench, SWE Bench Lite, SWE Bench Verified
Useful for quick benchmark comparison.
"SHERLOC reaches state-of-the-art localization across model scales: 84.33% accuracy@1 on SWE-Bench Lite and 81.27% recall@1 on SWE-Bench Verified; at ~30B parameters, it matches or outperforms other agentic methods."
Accuracy, Recall, Recall@1
Useful for evaluation criteria comparison.
"SHERLOC reaches state-of-the-art localization across model scales: 84.33% accuracy@1 on SWE-Bench Lite and 81.27% recall@1 on SWE-Bench Verified; at ~30B parameters, it matches or outperforms other agentic methods."
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: SWE-bench, SWE-bench Lite, SWE-bench Verified
Metric reporting is present
Detected: accuracy, recall, recall@1