Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang · Aug 21, 2026 · Citations: 0
Tag: Math
Math evaluation papers that call for domain expertise or specialist review (123 papers).
Papers in tag: 123
Need Math evaluators for your project?
Post a Job →Research Utility Snapshot
Evaluation Modes
- Automatic Metrics (12)
- Llm As Judge (1)
Human Feedback Types
- Pairwise Preference (6)
- Red Team (3)
- Critique Edit (2)
Required Expertise
- Math (20)
- Coding (6)
- Law (3)
Seongjae Kang, Taehyung Yu, Sung Ju Hwang · Aug 20, 2026 · Citations: 0
- Customer-service LLM agents must follow organizational policy when acting on a user's behalf.
- Runtime safeguards can intervene on risky actions, but action-local checks do not guide an agent through a multi-step procedure.
Bin Zhu, Yi Xie, Yanghui Rao · Aug 20, 2026 · Citations: 0
- LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
- The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop.
Haonan He, Xinyue Fan · Aug 20, 2026 · Citations: 0
- For instance, LoRA-GA^2 surpasses the leading baseline by an average of 0.66 points on the GLUE benchmark, and outperforms the strongest baseline by 1.03 points on GSM8K and 0.87 points on HumanEval, respectively.
Tianjiao Li, Lusheng Zhang, Shien He, Xiaojing Liu, Tianxing Yan, Mengran Yu · Jul 10, 2026 · Citations: 0
- The series comprises four models: Index-1.9B-Base, a foundation model with 1.9 billion non-embedding parameters pre-trained on 2.8 trillion predominantly Chinese and English tokens; Index-1.9B-Pure, a control variant trained with an…
- On a suite of standard benchmarks covering examination, reasoning, mathematics, and code, Index-1.9B-Base attains an average score of 64.92, competitive with or exceeding open models of several times its size.
Mona Schirmer, Metod Jazbec, Alexander Timans, Christian Naesseth, Maja Waldron, Eric Nalisnick · Jul 2, 2026 · Citations: 0
- Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical.
Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026 · Citations: 0
- We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
- In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
Sangmook Lee, Minbeom Kim, Jeonghye Kim, Dohyung Kim, Sojeong Rhee, Kyomin Jung · Jun 29, 2026 · Citations: 0
- Using a human-calibrated LLM judge framework, we show that prior diversity measures are unreliable proxies for approach-level diversity, and this mismatch carries over to diversity-aware RLVR, where target metrics are preserved while…
- However, optimizing an LLM judge diversity reward during training causes the policy to exploit judge-specific preferences rather than broaden its approaches, leaving direct optimization of approach-level diversity as an open problem.
Yiqiu Guo, Xueting Han, Qi Jia, Guangtao Zhai, Jing Bai · Jun 29, 2026 · Citations: 0
- Used as training data, these trajectories improve SFT and RLVR on math benchmarks over standard baselines.
Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov, Martin Vechev · Jun 29, 2026 · Citations: 0
- As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
- Importantly, we show that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrieval benchmarks.
Jaeyong Ko, Pilsung Kang, Yukyung Lee · Jun 24, 2026 · Citations: 0
- Across seven models and three mathematical reasoning benchmarks (GSM1K, MATH500, AIME 2025), cliff tokens act as failure triggers; deleting the first cliff token and resampling recovers pass@64 to 1.0, while keeping it limits recovery to…
- Trained on GSM8K, Cliff-DPO improves accuracy across benchmarks by up to +6.6.
Yutong Yin, Mingyu Jin, Jin Pan, Changyi Yang, Zijie Xia, Dhruv Pai · Jun 24, 2026 · Citations: 0
- On mathematical reasoning benchmarks, LBR improves both Pass@1 and Pass@32 over discrete chain-of-thought, vanilla discrete-token RLVR, and RL-compatible soft-token branching baselines.
Hyunji Nam, Keertana Chidambaram, Dorottya Demszky, Natasha Jaques · Jun 23, 2026 · Citations: 0
Liwen Zheng, Haiyun Jiang · Jun 23, 2026 · Citations: 0
- In a six-variant Qwen3 math reasoning benchmark with a uniform 200-step training budget for all trained variants, we use pass@8 as the primary problem-level solve-rate metric.
- On Teacher-TopK/LSM, Block64 gives the best four-benchmark mean pass@8 among trained students.
Fiona Y. Wang, Markus J. Buehler · May 31, 2026 · Citations: 0
- We develop a category-theoretic account of agentic discovery for materials science.
Jiamin Chen, Yidi Wu, Qiexiang Wang, Qianben Chen, Yuchen Li, Yansen Zhang · May 28, 2026 · Citations: 0
- Therefore, we present Seeded Elimination with Adaptive LLM-as-a-Meta-Judge, a self-improving evaluation protocol for extracting latent ranking signal from saturated benchmarks.
- We evaluate SEAL on multiple saturated benchmarks covering code generation, mathematical reasoning, knowledge-intensive question answering, and tool-use agent task completion.
Shahinul Hoque, Jinghuai Zhang, Jinyuan Sun, Fnu Suya · May 28, 2026 · Citations: 0
Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, Wanxiang Che · May 28, 2026 · Citations: 0
- Agent harnesses shape language-model performance by controlling tool use, feedback, verification, memory, and repair.
- Across synthetic, real, held-out, and prospective evaluations, EFC-based coordinates outperform raw-compute baselines and SAS.
Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu, Fan Zhang · May 21, 2026 · Citations: 0
- In this paper, we present Maestro (Multimodal Agent for Expert-Skill Targeted Reinforced Orchestration), a Reinforcement Learning (RL)-driven orchestration framework that reframes heterogeneous multimodal tasks as a sequential…
- We evaluate Maestro across ten representative multimodal benchmarks spanning mathematical reasoning, chart understanding, high-resolution perception, and domain-specific analysis.
Mingkai Deng, Jinyu Hou, Lara Sá Neves, Varad Pimpalkhute, Taylor W. Killian, Zhengzhong Liu · May 21, 2026 · Citations: 0
- To test this, we develop SR^2AM (Self-Regulated Simulative Reasoning Agentic LLM), realizing both as distinct stages within an LLM's chain-of-thought, with the LLM as world model.
- Across math, science, tabular analysis, and web information seeking, v0.1-8B and v1.0-30B achieve Pass@1 competitive with 120-355B and 685B-1T parameter systems respectively, while v1.0-30B uses 25.8-95.3% fewer reasoning tokens than…