Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
HFEPX · Eval paper review
Bingquan Zhang, Xiaoxiao Liu, Yuchi Wang, Lei Zhou +2 more
Published
Nov 12, 2025
Citations
0
Trust level
Low
Usefulness score
15/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Mar 23, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale. Although large language model (LLM)-based virtual standardized patients (VSPs) have been proposed as an alternative, their behavior remains unstable and lacks rigorous comparison with human standardized patients. We propose EasyMED, a multi-agent VSP framework that separates case-grounded information disclosure from response generation to support stable, inquiry-conditioned patient behavior. We also introduce SPBench, a human-grounded benchmark with eight expert-defined criteria for interaction-level evaluation. Experiments show that EasyMED more closely matches human SP behavior than existing VSPs, particularly in case consistency and controlled disclosure. A four-week controlled study further demonstrates learning outcomes comparable to human SP training, with stronger early gains for novice learners and improved flexibility, psychological safety, and cost efficiency.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
Automatic Metrics
Includes extracted eval setup.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
Not reported
No explicit QC controls found.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
Not extracted
No benchmark anchors detected.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
Not extracted
No metric anchors detected.
"Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale."
Domain Experts
Helpful for staffing comparability.
"We also introduce SPBench, a human-grounded benchmark with eight expert-defined criteria for interaction-level evaluation."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Standardized patients (SPs) are indispensable for clinical skills training but remain expensive and difficult to scale.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.