Human Feedback Types
partialRubric Rating
Directly usable for protocol triage.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
HFEPX · Eval paper review
Adrian Arnaiz-Rodriguez, Miguel Baidal, Erik Derner, Jenn Layton Annable +4 more
Published
Sep 29, 2025
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Apr 8, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health. Despite their support capabilities, safe detection and response to crises such as suicidal ideation and self-harm are still unclear, hindered by the lack of unified crisis taxonomies and clinical evaluation standards. We address this by creating: (1) a taxonomy of six crisis categories; (2) a dataset of over 2,000 inputs from 12 mental health datasets, classified into these categories; and (3) a clinical response assessment protocol. We also use LLMs to identify crisis inputs and audit five models for response safety and appropriateness. First, we built a clinical-informed crisis taxonomy and evaluation protocol. Next, we curated 2,252 relevant examples from over 239,000 user inputs, then tested three LLMs for automatic classification. In addition, we evaluated five models for the appropriateness of their responses to a user's crisis, graded on a 5-point Likert scale from harmful (1) to appropriate (5). While some models respond reliably to explicit crises, risks still exist. Many outputs, especially in self-harm and suicidal categories, are inappropriate or unsafe. Different models perform variably; some, like gpt-5-nano and deepseek-v3.2-exp, have low harm rates, but others, such as gpt-4o-mini and grok-4-fast, generate more unsafe responses. All models struggle with indirect signals, default replies, and context misalignment. These results highlight the urgent need for better safeguards, crisis detection, and context-aware responses in LLMs. They also show that alignment and safety practices, beyond scale, are crucial for reliable crisis support. Our taxonomy, datasets, and evaluation methods support ongoing AI mental health research, aiming to reduce harm and protect vulnerable users.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Rubric Rating
Directly usable for protocol triage.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
None explicit
Validate eval design from full paper text.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
Not reported
No explicit QC controls found.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
Not extracted
No benchmark anchors detected.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
Not extracted
No metric anchors detected.
"Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Large language model-powered chatbots have transformed how people seek information, especially in high-stakes contexts like mental health.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Rubric Rating
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.