Human Feedback Types
strongRubric Rating
Directly usable for protocol triage.
"Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation."
HFEPX · Eval paper review
Jessica Huynh, Alfredo Gomez, Athiya Deviyani, Renee Shelby +2 more
Published
May 7, 2026
Citations
0
Trust level
Moderate
Usefulness score
77/100 (High)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
May 7, 2026
This paper has strong direct human-feedback and evaluation protocol signal and is suitable as a primary eval pipeline reference.
Use this for comparison and orientation, not as your only source.
Best use
Primary benchmark and eval reference
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Use this as a primary source when designing or comparing eval protocols.
If you are doing eval pipeline work, start here
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation. However, there is limited statistical analysis of how modifications in a rubric presented to both humans and autoraters affect their score agreement. Rubrics that ask for an overall or \emph{holistic} judgment - for example, rating the ``quality'' of an essay - may be inconsistently interpreted due to the complexity or subjectivity of the criteria. Conversely, rubrics can ask for \emph{analytic} judgments, which decompose assessment criteria - for example, ``quality'' into ``fluency'' and ``organization''. While these rubrics can be edited to improve the individual accuracy of both human and automated scoring, this approach may result in disagreement between the two scores, or with the associated holistic judgment. Designing and deploying reliable autoraters requires understanding not just the relationship between human and autorater annotations but how that relationship changes as holistic or analytic judgments are elicited. The results indicate that rubric edits providing representative examples and additional context, and reducing positional bias in the rubric increased human-autorater agreement, while higher rubric complexity and conservative aggregation methods tended to decrease it. The findings from the automatic essay scoring and instruction-following evaluation domains suggest that practitioners should carefully analyze domain- and rubric-specific performance to move towards higher human-autorater agreement.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Rubric Rating
Directly usable for protocol triage.
"Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation."
Llm As Judge, Automatic Metrics
Includes extracted eval setup.
"Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation."
Not reported
No explicit QC controls found.
"Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation."
Not extracted
No benchmark anchors detected.
"Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation."
Accuracy, Agreement
Useful for evaluation criteria comparison.
"However, there is limited statistical analysis of how modifications in a rubric presented to both humans and autoraters affect their score agreement."
No benchmark or dataset names were extracted from the available abstract.
Autoraters, also referred to as LLM-as-judges, are increasingly used for evaluation and automated content moderation.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Rubric Rating
Evaluation mode is explicit
Detected: Llm As Judge, Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy, agreement