Skip to content
OpenTrain AIFor AI Companies

Thinking Before Thinking: Scaling Agentic Inference Through Meta-Reasoning

Paras Dahal, Anton Bakhtin, Taco Cohen, Zhengxing Chen, Carole-Jean Wu +7 morePublished Sep 29, 2026
arXiv PDF
Researcher verdict
Context only
Use as context only
Benchmark evidence
Missing
Not verified yet
Time to first repro
A few days
Plan setup time
Risk flags
2
Review before use

Abstract

Domain fit: AI-adjacent · Paper appears method- or tooling-adjacent to AI workflows with partial ecosystem coverage.

As agents take on longer and more complex problems, controlling the execution becomes a task in its own right. Each step in the run brings new control choices, like which partial work to build on, whether to start fresh, or when to stop. We introduce agentic meta-reasoning, an inference-time harness that makes these choices an explicit and structured reasoning process. Workers carry out the task-level computation, while a controller consolidates what the run has established, explores next options, assesses what each option is worth under the remaining budget, and dispatches the chosen work with context drawn from persistent memory. Between decisions the controller carries only a compact account of the run rather than replaying its full history. Our baselines span production coding agents and research harnesses, together with a Direct Control Agent using the same workers and compute budget allowance. On ProgramBench, which tests long-horizon agentic capability through program reconstruction, meta-reasoning achieves 71.5% with GPT-5.5 against 58.0% for Codex; with Opus 4.8 it achieves 67.2% against 65.5% for Claude Code. On the other benchmarks, spanning abstract reasoning, multi-domain long-horizon reasoning, and proof generation, it gains between 3.6 and 4.2 points over direct control, averaged across three frontier models. It keeps improving over the tested budget ranges where direct control plateaus, though its overhead can hurt at small budgets. Artifact-graph analysis reveals more reuse of earlier work, higher coverage of correct solutions in most settings, and nonuniform gains in final selection. These results indicate that spending computation on structured control becomes more important as agents scale to longer runs.

Results and benchmarks

Freshness tier: warm
As agents take on longer and more complex problems, controlling the execution becomes a task in its own right.

Implementation

No direct implementation yet

Maintained implementation evidence is not confirmed for this paper yet.

Use the implementation status and reproduction sections for the current action plan.

Implementation evidence summary
Confidence: low

Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.

Reproduction risks
  • Estimate is based on paper-only reproduction flow

Reproduction readiness

Time to first repro: days
Last checked: Sep 30, 2026

No repo

No verified implementation available

  • No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.

Hardware requirements

  • Expect multi-day setup/compute for meaningful reproduction based on current guidance.

Hugging Face artifacts

No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.

Tip: start with models, then check datasets and spaces if you need evaluation data or demos.

Research context

Tasks

Agentic tool use

Methods

Agentic systems

Domains

AI Agents

Evaluation and human feedback data

Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.

Open in HFEPX
Explore similar papers

Jump to Paper2Code search queries derived from this paper's research context.