Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states."
HFEPX · Eval paper review
Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue +5 more
Published
Aug 12, 2026
Citations
0
Trust level
Moderate
Usefulness score
27/100 (Low)
Extraction confidence
50% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Aug 12, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states."
Simulation Env
Includes extracted eval setup.
"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states."
Not reported
No explicit QC controls found.
"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states."
Toolhazard Bench
Useful for quick benchmark comparison.
"Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks."
Not extracted
No metric anchors detected.
"Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states."
No metric terms were extracted from the available abstract.
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Toolhazard-Bench
Metric reporting is present
No metric terms extracted.