Human Feedback Types
strongCritique Edit
Directly usable for protocol triage.
"Language-model agents can improve after failure or carry text across episodes without revising what counts as success."
HFEPX · Eval paper review
Guodong Xu
Published
Aug 21, 2026
Citations
0
Trust level
Moderate
Usefulness score
50/100 (Medium)
Extraction confidence
55% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Aug 21, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Secondary protocol comparison source
Use if you need
A concrete protocol example with enough signal to inform rater workflow design.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The abstract does not clearly describe the evaluation setup.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrower attribution problem of criterion revision: when criterion K0 accepts an outcome violating a broader commitment B, what observations justify saying that the system formed and persistently used K1? We require five non-compensatory conditions: criterion-failure detection, a model-emitted proposal, new-episode transfer, intervention sensitivity on the claimed carrier, and preservation. We evaluate CMB-0.1 on twelve cross-domain cases and four arms: stateless inference, append-only history, model-generated but harness-committed state, and evaluator-written oracle state. Seven mechanism fixtures yield 84 deterministic scorer trials; four local quantized artifacts yield 96 calls and 192 model-case-arm trials. No model trial satisfies all five conditions, but this zero does not establish general capability absence. Eleven calls remain invalid after one retry; several commitments disclose the target distinction; the harness performs commits; deletion reuses a stateless call; and conflict changes multiple factors. Qwen2.5-7B answers every transfer and preservation item without revision state, exposing zero-state reconstruction. These failures make CMB-0.1 an instrument-calibration result rather than a model ranking. We derive a prospective, trace-anchored CMB-0.4 protocol requiring concealed transfer, explicit WRITE/NO-WRITE/ESCALATE actions, a separately logged policy-selected commit, matched interventions, repeated hidden items, and a frozen executable oracle. It is a successor design, not a completed confirmatory result. The paper contributes a measurement chain, an empirical diagnosis of its first implementation, and a more discriminating protocol for future tests of criterion revision.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Critique Edit
Directly usable for protocol triage.
"Language-model agents can improve after failure or carry text across episodes without revising what counts as success."
None explicit
Validate eval design from full paper text.
"Language-model agents can improve after failure or carry text across episodes without revising what counts as success."
Calibration
Calibration/adjudication style controls detected.
"These failures make CMB-0.1 an instrument-calibration result rather than a model ranking."
Not extracted
No benchmark anchors detected.
"Language-model agents can improve after failure or carry text across episodes without revising what counts as success."
Not extracted
No metric anchors detected.
"Language-model agents can improve after failure or carry text across episodes without revising what counts as success."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Language-model agents can improve after failure or carry text across episodes without revising what counts as success.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Critique Edit
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
Detected: Calibration
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.