Human Feedback Types
strongPairwise Preference, Critique Edit
Directly usable for protocol triage.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
HFEPX · Eval paper review
Wei Xia, Jin Wu, Haoran Shi, Xiangyu Wang +1 more
Published
Jun 18, 2026
Citations
0
Trust level
Moderate
Usefulness score
57/100 (Medium)
Extraction confidence
65% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Jun 18, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The abstract does not clearly name benchmarks or metrics.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback. However, existing approaches often treat scoring and feedback as separate components: neural scoring models provide limited interpretability, while Large Language Model (LLM)-based feedback is typically insensitive to learners proficiency levels. To address this fragmentation, this work proposes PsyScore, a psychometrically-aware framework that integrates diagnostic assessment with instructional scaffolding through a shared latent ability representation. PsyScore comprises three key modules: a Trait-Adaptive Neural IRT Scorer that incorporates the Graded Partial Credit Model (GPCM) into a neural architecture, enabling the precise estimation of student ability while maintaining psychometric interpretability, a ZPD-Scaffolded Feedback Generator, which conditions multi-agent feedback strategies on the diagnosed ability parameter to adapt instructional focus across different proficiency levels, and a Multi-Perspective Feedback Evaluation Strategy that assesses feedback quality via pairwise preference judgements and student revision simulations. Experiments on the ASAP++ dataset demonstrate that PsyScore achieves competitive scoring performance while providing more pedagogically aligned feedback.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference, Critique Edit
Directly usable for protocol triage.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
Simulation Env
Includes extracted eval setup.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
Not reported
No explicit QC controls found.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
Not extracted
No benchmark anchors detected.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
Not extracted
No metric anchors detected.
"Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Effective Automated Essay Scoring (AES) are expected to support both reliable assessment and actionable instructional feedback.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference, Critique Edit
Evaluation mode is explicit
Detected: Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.