Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear."
HFEPX · Eval paper review
Aofan Liu, Jingxiang Meng
Published
Apr 24, 2026
Citations
0
Trust level
Low
Usefulness score
0/100 (Low)
Extraction confidence
30% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Apr 24, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
Background context only.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
This paper looks adjacent to evaluation work, but not like a strong protocol reference.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear. We frame self-correction as a cybernetic feedback loop in which the same language model serves as both controller and plant, and use a two-state Markov model over {Correct, Incorrect} to operationalize a simple deployment diagnostic: iterate only when ECR/EIR > Acc/(1 - Acc). In this view, EIR functions as a stability margin and prompting functions as lightweight controller design. Across 7 models and 3 datasets (GSM8K, MATH, StrategyQA), we find a sharp near-zero EIR threshold (<= 0.5%) separating beneficial from harmful self-correction. Only o3-mini (+3.4 pp, EIR = 0%), Claude Opus 4.6 (+0.6 pp, EIR ~ 0.2%), and o4-mini (+/-0 pp) remain non-degrading; GPT-5 degrades by -1.8 pp. A verify-first prompt ablation provides causal evidence that this threshold is actionable through prompting alone: on GPT-4o-mini it reduces EIR from 2% to 0% and turns -6.2 pp degradation into +0.2 pp (paired McNemar p < 10^-4), while producing little change on already-sub-threshold models. ASC further illustrates the stopping trade-off: it halts harmful refinement but incurs a 3.8 pp confidence-elicitation cost. Overall, the paper argues that self-correction should be treated not as a default behavior, but as a control decision governed by measurable error dynamics.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear."
None explicit
Validate eval design from full paper text.
"Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear."
Not reported
No explicit QC controls found.
"Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear."
GSM8K
Useful for quick benchmark comparison.
"Across 7 models and 3 datasets (GSM8K, MATH, StrategyQA), we find a sharp near-zero EIR threshold (<= 0.5%) separating beneficial from harmful self-correction."
Not extracted
No metric anchors detected.
"Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear."
No metric terms were extracted from the available abstract.
Iterative self-correction is widely used in agentic LLM systems, but when repeated refinement helps versus hurts remains unclear.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: GSM8K
Metric reporting is present
No metric terms extracted.