Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Math

Math evaluation papers that call for domain expertise or specialist review (123 papers).

Papers in tag: 123

Need Math evaluators for your project?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (12)
  • Llm As Judge (1)

Human Feedback Types

  • Pairwise Preference (6)
  • Red Team (3)
  • Critique Edit (2)

Required Expertise

  • Math (20)
  • Coding (6)
  • Law (3)
Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao · Aug 20, 2026 · Citations: 0

Pairwise Preference Llm As Judge MathCoding
  • LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
  • The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop.
Index SLM Technical Report

Tianjiao Li, Lusheng Zhang, Shien He, Xiaojing Liu, Tianxing Yan, Mengran Yu · Jul 10, 2026 · Citations: 0

Pairwise Preference MathCoding
  • The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with an…
  • On a suite of standard benchmarks covering examination, reasoning, mathematics, and code, Index-1.9B-Base attains an average score of 64.92, competitive with or exceeding open models of several times its size.
Online Safety Monitoring for LLMs

Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick · Jul 2, 2026 · Citations: 0

Red Team Math
  • Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical.
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026 · Citations: 0

Demonstrations Automatic Metrics MathCoding
  • We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
  • In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning

Sangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung · Jun 29, 2026 · Citations: 0

Pairwise Preference Math
  • Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while…
  • However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem.
LatentRevise: Learning from Zero-Hit Reasoning

Yiqiu Guo, Xueting Han, Qi Jia, Guangtao Zhai, Jing Bai · Jun 29, 2026 · Citations: 0

Critique Edit Math
  • Used as training data, these trajectories improve SFT and RLVR on math benchmarks over standard baselines.
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov, Martin Vechev · Jun 29, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Math
  • As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
  • Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning

Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Math
  • Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
  • Trained on GSM8K, Cliff-DPO improves accuracy across benchmarks by up to +6.6.
Blockwise Policy-Drift Gating for On-Policy Distillation

Liwen Zheng, Haiyun Jiang · Jun 23, 2026 · Citations: 0

Automatic Metrics Math
  • In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric.
  • On Teacher-TopK/LSM, Block64 gives the best four-benchmark mean pass@8 among trained students.
SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?

Jiamin Chen, Yidi Wu, Qiexiang Wang, Qianben Chen, Yuchen Li, Yansen Zhang · May 28, 2026 · Citations: 0

Pairwise Preference Automatic Metrics MathCoding
  • Therefore, we present Seeded Elimination with Adaptive LLM-as-a-Meta-Judge, a self-improving evaluation protocol for extracting latent ranking signal from saturated benchmarks.
  • We evaluate SEAL on multiple saturated benchmarks covering code generation, mathematical reasoning, knowledge-intensive question answering, and tool-use agent task completion.
Scaling Laws for Agent Harnesses via Effective Feedback Compute

Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che · May 28, 2026 · Citations: 0

Automatic Metrics MathLaw
  • Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair.
  • Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS.
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu, Fan Zhang · May 21, 2026 · Citations: 0

Expert Verification Automatic Metrics MathCoding
  • In this paper, we present Maestro (Multimodal Agent for Expert-Skill Targeted Reinforced Orchestration), a Reinforcement Learning (RL)-driven orchestration framework that reframes heterogeneous multimodal tasks as a sequential…
  • We evaluate Maestro across ten representative multimodal benchmarks spanning mathematical reasoning, chart understanding, high-resolution perception, and domain-specific analysis.
Efficient Agentic Reasoning Through Self-Regulated Simulative Planning

Mingkai Deng, Jinyu Hou, Lara Sá Neves, Varad Pimpalkhute, Taylor W. Killian, Zhengzhong Liu · May 21, 2026 · Citations: 0

Automatic Metrics Math
  • To test this, we develop SR^2AM (Self-Regulated Simulative Reasoning Agentic LLM), realizing both as distinct stages within an LLM's chain-of-thought, with the LLM as world model.
  • Across math, science, tabular analysis, and web information seeking, v0.1-8B and v1.0-30B achieve Pass@1 competitive with 120-355B and 685B-1T parameter systems respectively, while v1.0-30B uses 25.8-95.3% fewer reasoning tokens than…