Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Before letting an agent operate over real context, can you prove it used the right evidence?"
HFEPX · Eval paper review
Jeffrey Flynt
Published
Jun 22, 2026
Citations
0
Trust level
Moderate
Usefulness score
27/100 (Low)
Extraction confidence
50% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Jul 2, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access. In one case study, two frontier LLM judges scored a plausible agent response 0.85 and higher. But the trace told a different story: the agent had never retrieved the artifact its answer depended on, yielding a GroundEval score of 0.000. We introduce GroundEval, a judge-free framework for evaluating agents against grounded, time-bounded, and access-controlled evidence. GroundEval uses a domain configuration to generate questions, lets the agent choose how to answer, and then scores both the final answer and the recorded trajectory that produced it. The benchmark targets three failures that LLM-as-judge evaluation struggles to detect: whether an agent checked before claiming absence, reasoned only from evidence available to the actor at the relevant time, and used the correct causal mechanism rather than a plausible one. These correspond to three tracks: Silence, Perspective, and Counterfactual. GroundEval exposes when plausible answers rest on invalid evidence paths, and produces structured per-question diagnostics that pair tool activity with the agent's turn-level narration, making each score inspectable rather than merely reported. Our case studies suggest this failure mode is common rather than exceptional, one that final-answer and judge-based evaluation cannot detect by construction.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Before letting an agent operate over real context, can you prove it used the right evidence?"
Llm As Judge
Includes extracted eval setup.
"Before letting an agent operate over real context, can you prove it used the right evidence?"
Not reported
No explicit QC controls found.
"Before letting an agent operate over real context, can you prove it used the right evidence?"
Groundeval
Useful for quick benchmark comparison.
"GroundEval turns that question into a deterministic test of what the agent searched, fetched, cited, and was permitted to access."
Not extracted
No metric anchors detected.
"Before letting an agent operate over real context, can you prove it used the right evidence?"
No metric terms were extracted from the available abstract.
Before letting an agent operate over real context, can you prove it used the right evidence?
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Llm As Judge
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Groundeval
Metric reporting is present
No metric terms extracted.