Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly."
HFEPX · Eval paper review
Yubo Li, Yidi Miao, Ramayya Krishnan, Rema Padman
Published
Sep 22, 2026
Citations
0
Trust level
Moderate
Usefulness score
47/100 (Medium)
Extraction confidence
55% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Sep 29, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the exact study setup in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly."
Llm As Judge, Automatic Metrics
Includes extracted eval setup.
"LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly."
Adjudication
Calibration/adjudication style controls detected.
"Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic."
Not extracted
No benchmark anchors detected.
"LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly."
Accuracy
Useful for evaluation criteria comparison.
"With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly."
No benchmark or dataset names were extracted from the available abstract.
LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Llm As Judge, Automatic Metrics
Quality control reporting appears
Detected: Adjudication
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy