Human Feedback Types
strongRubric Rating
Directly usable for protocol triage.
"Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety."
HFEPX · Eval paper review
Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
Published
Jan 21, 2026
Citations
0
Trust level
High
Usefulness score
65/100 (Medium)
Extraction confidence
80% (High)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Aug 26, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this as a practical starting point for protocol research, then validate against the original paper.
Best use
Secondary protocol comparison source
Use if you need
A benchmark-and-metrics comparison anchor.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety. These risks are hard to assess: subtle clinical errors are often missed by generic metrics and LLM judges using general criteria, while expert-authored fine-grained rubrics are expensive and difficult to scale. In this paper, we propose a retrieval-augmented multi-agent framework for automatically generating instance-specific evaluation rubrics. Our approach grounds evaluation in authoritative medical evidence by decomposing retrieved content into atomic facts and synthesizing them with user interaction constraints to form fine-grained evaluation criteria. Evaluated on HealthBench and LLMEval-Med, our framework achieves Clinical Intent Alignment (CIA) scores of 50.20% and 31.90%, significantly outperforming the GPT-4o baseline and showing consistent improvements across English and Chinese medical benchmarks. In discriminative tests on HealthBench, our rubrics achieve a 7.8% point higher win rate than GPT-4o and increase the mean score difference from 4.972 to 8.658. Ablation studies further show that individual components contribute differently across datasets, with interaction-intent modeling providing the most consistent contribution to clinical-criterion coverage. Beyond evaluation, our rubrics guide response refinement, improving response quality by 9.2%. These results suggest that automated, knowledge-grounded rubric generation provides a scalable foundation for evaluating and improving medical LLMs. The code is available at https://anonymous.4open.science/r/Automated-Rubric-Generation-E716/.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Rubric Rating
Directly usable for protocol triage.
"Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety."
Automatic Metrics
Includes extracted eval setup.
"Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety."
Not reported
No explicit QC controls found.
"Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety."
Healthbench, Llmeval
Useful for quick benchmark comparison.
"Evaluated on HealthBench and LLMEval-Med, our framework achieves Clinical Intent Alignment (CIA) scores of 50.20% and 31.90%, significantly outperforming the GPT-4o baseline and showing consistent improvements across English and Chinese medical benchmarks."
Win rate
Useful for evaluation criteria comparison.
"In discriminative tests on HealthBench, our rubrics achieve a 7.8% point higher win rate than GPT-4o and increase the mean score difference from 4.972 to 8.658."
Domain Experts
Helpful for staffing comparability.
"These risks are hard to assess: subtle clinical errors are often missed by generic metrics and LLM judges using general criteria, while expert-authored fine-grained rubrics are expensive and difficult to scale."
Large Language Models (LLMs) are increasingly used for clinical decision support, where hallucinations and unsafe suggestions may pose direct risks to patient safety.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Rubric Rating
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Healthbench, Llmeval
Metric reporting is present
Detected: win rate