Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment."
HFEPX · Eval paper review
Subba Reddy Oota, Francisco Herrera, Jordi Cabot Sagrera, Marcos López de Prado +1 more
Published
Sep 29, 2026
Citations
0
Trust level
Moderate
Usefulness score
27/100 (Low)
Extraction confidence
55% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Sep 29, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A benchmark-and-metrics comparison anchor.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three consistent patterns. First, generic Plan+ReAct often fails to preserve declared planning structure, especially for longer plans: across three benchmarks, only (22)--(45%) of trajectories preserve it, whereas pattern-specific executors enforce the intended structure. Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark. Third, the largest gains come from execution: pattern-specific executors improve task success from (0.48) to (0.92) on ALFWorld and from (0.36) to (0.44) on SWE-bench Verified over Plan+ReAct. Current LLMs, however, do not reliably select the strongest mode for each task, although few-shot examples improve selection in some benchmark--model combinations. Overall, reliable agent planning requires both effective mode selection and faithful execution: routing substantially closes the execution gap, while task-specific mode selection remains open.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment."
Simulation Env
Includes extracted eval setup.
"Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment."
Not reported
No explicit QC controls found.
"Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment."
ALFWorld, SWE Bench, SWE Bench Verified
Useful for quick benchmark comparison.
"Second, planning-mode effectiveness varies across environments and models: Search performs best on ALFWorld, Hierarchical on SWE-bench, and the strongest pattern can vary across models within the same benchmark."
Task success
Useful for evaluation criteria comparison.
"Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures."
Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Simulation Env
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: ALFWorld, SWE-bench, SWE-bench Verified
Metric reporting is present
Detected: task success