Human Feedback Types
strongDemonstrations
Directly usable for protocol triage.
"These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations."
HFEPX · Eval paper review
Zelai Xu, Ruize Zhang, Chao Yu, Huining Yuan +8 more
Published
Feb 4, 2025
Citations
0
Trust level
Moderate
Usefulness score
77/100 (High)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Feb 26, 2026
This paper has strong direct human-feedback and evaluation protocol signal and is suitable as a primary eval pipeline reference.
Use this for comparison and orientation, not as your only source.
Best use
Primary protocol reference for eval design
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Use this as a primary source when designing or comparing eval protocols.
If you are doing eval pipeline work, start here
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence. In this paper, we present VolleyBots, a novel robot sports testbed where multiple drones cooperate and compete in the sport of volleyball under physical dynamics. VolleyBots integrates three features within a unified platform: competitive and cooperative gameplay, turn-based interaction structure, and agile 3D maneuvering. These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations. We provide a comprehensive suite of tasks ranging from single-drone drills to multi-drone cooperative and competitive tasks, accompanied by baseline evaluations of representative reinforcement learning (RL), multi-agent reinforcement learning (MARL) and game-theoretic algorithms. Simulation results show that on-policy RL methods outperform off-policy methods in single-agent tasks, but both approaches struggle in complex tasks that combine motion control and strategic play. We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy. To highlight VolleyBots' sim-to-real potential, we further demonstrate the zero-shot deployment of a policy trained entirely in simulation on real-world drones.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Demonstrations
Directly usable for protocol triage.
"These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations."
Automatic Metrics, Simulation Env
Includes extracted eval setup.
"Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence."
Not reported
No explicit QC controls found.
"Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence."
Not extracted
No benchmark anchors detected.
"Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence."
Win rate
Useful for evaluation criteria comparison.
"We additionally design a hierarchical policy which achieves 69.5% win rate against the strongest baseline in the 3 vs 3 task, demonstrating its potential for tackling the complex interplay between low-level control and high-level strategy."
Domain Experts
Helpful for staffing comparability.
"These intertwined features yield a complex problem combining motion control and strategic play, with no available expert demonstrations."
No benchmark or dataset names were extracted from the available abstract.
Robot sports, characterized by well-defined objectives, explicit rules, and dynamic interactions, present ideal scenarios for demonstrating embodied intelligence.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Demonstrations
Evaluation mode is explicit
Detected: Automatic Metrics, Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: win rate