Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs)."
HFEPX · Eval paper review
Ziyan Wang, Zheng Wang, Xingwei Qu, Qi Cheng +4 more
Published
Oct 5, 2025
Citations
0
Trust level
Low
Usefulness score
25/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Mar 17, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs). Yet on-policy algorithms such as Group Relative Policy Optimization (GRPO) often suffer in early training: noisy gradients from low-quality rollouts lead to unstable updates and inefficient exploration. We introduce Slow-Fast Policy Optimization (SFPO), a simple yet efficient framework to address the above limitations via decomposing each step into three stages: a short fast trajectory of inner steps on the same batch, a reposition mechanism to control off-policy drift, and a final slow correction. This reposition-before-update design preserves the objective and rollout process unchanged, making SFPO plug-compatible with existing policy-gradient pipelines. Extensive experiments demonstrate that SFPO consistently improves stability, reduces number of rollouts, and accelerates convergence of reasoning RL training. Specifically, it outperforms GRPO by up to 2.80 points in average on math reasoning benchmarks. It also achieves up to 4.93\texttimes{} fewer rollouts and an up to 4.19\texttimes{} reduction in wall-clock time to match GRPO's best accuracy. Project website is available at https://slow-fast-po.github.io/.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs)."
Automatic Metrics
Includes extracted eval setup.
"Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs)."
Not reported
No explicit QC controls found.
"Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs)."
Not extracted
No benchmark anchors detected.
"Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs)."
Accuracy
Useful for evaluation criteria comparison.
"It also achieves up to 4.93\texttimes{} fewer rollouts and an up to 4.19\texttimes{} reduction in wall-clock time to match GRPO's best accuracy."
No benchmark or dataset names were extracted from the available abstract.
Reinforcement learning (RL) has become central to enhancing reasoning in large language models (LLMs).
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy