Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior."
HFEPX · Eval paper review
Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu +6 more
Published
Aug 12, 2026
Citations
0
Trust level
Moderate
Usefulness score
27/100 (Low)
Extraction confidence
50% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Aug 12, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior."
Simulation Env
Includes extracted eval setup.
"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior."
Not reported
No explicit QC controls found.
"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior."
Cooperbench
Useful for quick benchmark comparison.
"We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $τ^2$-bench, and CooperBench."
Not extracted
No metric anchors detected.
"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior."
No metric terms were extracted from the available abstract.
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Cooperbench
Metric reporting is present
No metric terms extracted.