Human Feedback Types
strongRed Team
Directly usable for protocol triage.
"The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions."
HFEPX · Eval paper review
Qingyang Xu, Yaling Shen, Stephanie Fong, Zimu Wang +6 more
Published
Apr 6, 2026
Citations
0
Trust level
Moderate
Usefulness score
67/100 (Medium)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Apr 6, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions. A key challenge is distinguishing therapeutic empathy from maladaptive validation, where supportive responses may inadvertently reinforce harmful beliefs or behaviors in multi-turn conversations. This risk is largely overlooked by existing red-teaming frameworks, which focus mainly on generic harms or optimization-based attacks. To address this gap, we introduce Personality-based Client Simulation Attack (PCSA), the first red-teaming framework that simulates clients in psychological counseling through coherent, persona-driven client dialogues to expose vulnerabilities in psychological safety alignment. Experiments on seven general and mental health-specialized LLMs show that PCSA substantially outperforms four competitive baselines. Perplexity analysis and human inspection further indicate that PCSA generates more natural and realistic dialogues. Our results reveal that current LLMs remain vulnerable to domain-specific adversarial tactics, providing unauthorized medical advice, reinforcing delusions, and implicitly encouraging risky actions.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Red Team
Directly usable for protocol triage.
"The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions."
Simulation Env
Includes extracted eval setup.
"The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions."
Not reported
No explicit QC controls found.
"The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions."
Not extracted
No benchmark anchors detected.
"The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions."
Perplexity
Useful for evaluation criteria comparison.
"Perplexity analysis and human inspection further indicate that PCSA generates more natural and realistic dialogues."
No benchmark or dataset names were extracted from the available abstract.
The increasing use of large language models (LLMs) in mental healthcare raises safety concerns in high-stakes therapeutic interactions.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Red Team
Evaluation mode is explicit
Detected: Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: perplexity