Human Feedback Types
provisional (inferred)None explicit
No explicit feedback protocol extracted.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
HFEPX · Eval paper review
Vamshi Krishna Bonagiri, Ponnurangam Kumaragurum, Khanh Nguyen, Benjamin Plaut
Published
Oct 18, 2025
Citations
0
Trust level
Provisional
Usefulness score
Unavailable
Extraction confidence
0% (Provisional)
Derived from abstract and metadata only.
Signals refreshed
Jun 26, 2026
Signal extraction is still processing. This page currently shows metadata-first guidance until structured protocol fields are ready.
This page is a lightweight research summary built from the abstract and metadata while deeper extraction catches up.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A provisional background reference while structured extraction finishes.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
This page is still relying on abstract and metadata signals, not a fuller protocol read.
Eval-fit score is unavailable until extraction completes.
If you are doing eval pipeline work, start here
As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical. While uncertainty quantification is well-studied for single-turn tasks, multi-turn agentic scenarios with real-world tool access present unique challenges where uncertainties and ambiguities compound, leading to severe or catastrophic risks beyond traditional text generation failures. We propose using "quitting" as a simple yet effective behavioral mechanism for LLM agents to recognize and withdraw from situations where they lack confidence. Leveraging the ToolEmu framework, we conduct a systematic evaluation of quitting behavior across 12 state-of-the-art LLMs. Our results demonstrate a highly favorable safety-helpfulness trade-off: agents prompted to quit with explicit instructions improve safety by an average of +0.39 on a 0-3 scale across all models (+0.64 for proprietary models), while maintaining a negligible average decrease of -0.03 in helpfulness. Our analysis demonstrates that simply adding explicit quit instructions proves to be a highly effective safety mechanism that can immediately be deployed in existing agent systems, and establishes quitting as an effective first-line defense mechanism for autonomous agents in high-stakes applications.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
None explicit
Validate eval design from full paper text.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
Not reported
No explicit QC controls found.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
Not extracted
No benchmark anchors detected.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
Not extracted
No metric anchors detected.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
Unknown
Rater source not explicitly reported.
"As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical."
This page is using abstract-level cues only right now. Treat the signals below as provisional.
Evaluation fields are inferred from the abstract only.
As Large Language Model (LLM) agents increasingly operate in complex environments with real-world consequences, their safety becomes critical.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.