Human Feedback Types
provisional (inferred)None explicit
No explicit feedback protocol extracted.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
HFEPX · Eval paper review
Xinhe Wang, Katia Sycara, Yaqi Xie
Published
Apr 27, 2026
Citations
0
Trust level
Provisional
Usefulness score
Unavailable
Extraction confidence
0% (Provisional)
Derived from abstract and metadata only.
Signals refreshed
Apr 27, 2026
Signal extraction is still processing. This page currently shows metadata-first guidance until structured protocol fields are ready.
This page is a lightweight research summary built from the abstract and metadata while deeper extraction catches up.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A provisional background reference while structured extraction finishes.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
This page is still relying on abstract and metadata signals, not a fuller protocol read.
Eval-fit score is unavailable until extraction completes.
If you are doing eval pipeline work, start here
Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's intent. It has been found that this binary training regime often leads to brittleness, since the user intent cannot reliably be evaluated, especially if the attacker obfuscates their intent, and also makes the system seem unhelpful. In response, frontier models, such as GPT-5, have shifted from refusal-based safeguards to safe completion, that aims to maximize helpfulness while obeying safety constraints. However, safe completion could be exploited when a user pretends their intention is benign. Specifically, this intent inversion would be effective in multi-turn conversation, where the attacker has multiple opportunities to reinforce their deceptively benign intent. In this work, we introduce a novel multi-turn jailbreaking method that exploits this vulnerability. Our approach gradually builds conversational trust by simulating benign-seeming intentions and by exploiting the consistency property of the model, ultimately guiding the target model toward harmful, detailed outputs. Most crucially, our approach also uncovered an additional class of model vulnerability that we call para-jailbreaking that has been unnoticed up to now. Para-jailbreaking describes the situation where the model may not reveal harmful direct reply to the attack query, however the information that it reveals is nevertheless harmful. Our contributions are threefold. First, it achieves high success rates against frontier models including GPT-5-thinking and Claude-Sonnet-4.5. Second, our approach revealed and addressed para-jailbreaking harmful output. Third, experiments on multimodal VLM models showed that our approach outperformed state-of-the-art models.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
None explicit
Validate eval design from full paper text.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
Not reported
No explicit QC controls found.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
Not extracted
No benchmark anchors detected.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
Not extracted
No metric anchors detected.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
Unknown
Rater source not explicitly reported.
"Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking."
This page is using abstract-level cues only right now. Treat the signals below as provisional.
Evaluation fields are inferred from the abstract only.
Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.