Human Feedback Types
partialRed Team
Directly usable for protocol triage.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
HFEPX · Eval paper review
Xinpeng Wang, Mingyang Wang, Yihong Liu, Hinrich Schütze +1 more
Published
May 22, 2025
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Feb 25, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety. Recent research has revealed that refusal behavior can be mediated by a single direction in activation space, enabling targeted interventions to bypass refusals. While this is primarily demonstrated in an English-centric context, appropriate refusal behavior is important for any language, but poorly understood. In this paper, we investigate the refusal behavior in LLMs across 14 languages using PolyRefuse, a multilingual safety dataset created by translating malicious and benign English prompts into these languages. We uncover the surprising cross-lingual universality of the refusal direction: a vector extracted from English can bypass refusals in other languages with near-perfect effectiveness, without any additional fine-tuning. Even more remarkably, refusal directions derived from any safety-aligned language transfer seamlessly to others. We attribute this transferability to the parallelism of refusal vectors across languages in the embedding space and identify the underlying mechanism behind cross-lingual jailbreaks. These findings provide actionable insights for building more robust multilingual safety defenses and pave the way for a deeper mechanistic understanding of cross-lingual vulnerabilities in LLMs.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Red Team
Directly usable for protocol triage.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
None explicit
Validate eval design from full paper text.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
Not reported
No explicit QC controls found.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
Not extracted
No benchmark anchors detected.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
Not extracted
No metric anchors detected.
"Refusal mechanisms in large language models (LLMs) are essential for ensuring safety."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Refusal mechanisms in large language models (LLMs) are essential for ensuring safety.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Red Team
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.