Skip to content
OpenTrain AIFor AI Companies

Offline Reinforcement Learning for LLM Multi-Step Reasoning

Published Dec 1, 2024
arXiv PDF
Researcher verdict
Starting point
Use as implementation starting point
Benchmark evidence
Missing
Not verified yet
Time to first repro
A few hours
Fast first run
Risk flags
0
None detected

Results and benchmarks

Freshness tier: hot
Offline Reinforcement Learning for LLM Multi-Step Reasoning presents a reinforcement learning method.

Implementation

Best maintained implementation now

Recommended
Confidence: High
Reproducibility: Strong

jwhj/OREO

116 stars · 5 forks · Last push Jan 21, 2025 · Apache-2.0 license

  • License
  • CI
  • Dependencies
  • Docker

Official implementation from Papers with Code · Repository link is mentioned in the paper metadata · Community adoption signal (116 stars)

Why this implementation
Confidence: high

jwhj/oreo is the strongest maintained implementation based on ranking signals. CI workflows are present. License is declared (Apache-2.0).

Open jwhj/oreo
Reproduction risks
  • No repository-level red flags were detected, but paper-specific preprocessing and hyperparameter details may still be under-specified.
  • Selected jwhj/oreo as the strongest maintained implementation for new work.
  • Includes CI workflow signals.
  • Includes dependency/environment manifest signals.
  • Repository activity is within the last 24 months.

Compare implementation paths

Compare maintenance quality, reproducibility coverage, and evidence confidence before choosing a reproduction baseline.

jwhj/oreo
best maintained
Maintenance
Stale
Confidence
High
Reproducibility
Strong
Stars
116
Last push
Jan 21, 2025 (582d)

Official implementation from Papers with Code · Repository link is mentioned in the paper metadata

  • No push in 12+ months
  • No tagged releases
zhaoolee/garss
alternative
Maintenance
Active
Confidence
Low
Reproducibility
Limited
Stars
1,426
Last push
Aug 23, 2026 (2d)

Community adoption signal (1426 stars)

  • No Docker setup
  • Dependency manifest missing
  • Low confidence match
jwhj/OREO
alternative
Maintenance
Stale
Confidence
Low
Reproducibility
Strong
Stars
116
Last push
Jan 21, 2025 (582d)

Matched via arXiv identifier search · Community adoption signal (116 stars)

  • No push in 12+ months
  • No tagged releases
  • Low confidence match

Reproduction readiness

Time to first repro: hours
Last checked: Aug 24, 2026

Setup required

Dependencies pinned, manual setup needed

  • jwhj/oreo has pyproject.toml but requires manual environment setup.
  • Last push was 582 days ago, so expect possible dependency version conflicts.
Open jwhj/oreo

Quick start

git clone https://github.com/jwhj/oreo.git
pip install -e .

Repositories and ecosystem

No additional verified repositories beyond the primary recommendation.

Hugging Face artifacts

No direct paper-linked artifacts were found. Showing strongest curated related artifacts for faster exploration.

Research context

Tasks

None detected

Methods

Reinforcement learning

Domains

Large Language Models

Evaluation and human feedback data

Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.

Open in HFEPX
Explore similar papers

Jump to Paper2Code search queries derived from this paper's research context.

Data includes links from Papers with Code ( CC-BY-SA-4.0 ).