- Aligning Multimodal Sequential Recommendations via Robust Direct Preference Optimization with Sparse MoE
Hejin Huang, Jusheng Zhang, Kaitong Cai, Jian Wang, Rong Pan · Mar 31, 2026 · Citations: 0
Automatic Metrics General
Preference-based alignment objectives have been widely adopted, from RLHF-style pairwise learning in large language models to emerging applications in recommender systems.
- QED-Nano: Teaching a Tiny Model to Prove Hard Theorems
LM-Provers, Yuxiao Qu, Amrith Setlur, Jasper Dekoninck, Edward Beeching · Apr 6, 2026 · Citations: 0
Automatic Metrics MathCoding
To support further research on open mathematical reasoning, we release the full QED-Nano pipeline, including the QED-Nano and QED-Nano-SFT models, the FineProofs-SFT and FineProofs-RL datasets, and the training and evaluation code.
- CAMEL: Confidence-Gated Reflection for Reward Modeling
Zirui Zhu, Hailun Xu, Yang Luo, Yong Liu, Kanchan Sarkar · Feb 24, 2026 · Citations: 0
Automatic Metrics General
Building on this insight, we propose CAMEL, a confidence-gated reflection framework that performs a lightweight single-token preference decision first and selectively invokes reflection only for low-confidence instances.
- Distilling Feedback into Memory-as-a-Tool
Víctor Gallego · Jan 9, 2026 · Citations: 0
Automatic Metrics General
We propose a framework that amortizes the cost of inference-time reasoning by converting transient critiques into retrievable guidelines, through a file-based memory system and agent-controlled tool calls.
- S0 Tuning: Zero-Overhead Adaptation of Hybrid Recurrent-Attention Models
Jack Young · Apr 1, 2026 · Citations: 0
Automatic Metrics MathCoding
Using roughly 48 execution-verified HumanEval training solutions, tuning a single initial state matrix per recurrent layer, with zero inference overhead, outperforms LoRA by +10.8 pp (p < 0.001) on HumanEval.
- Learning When to Sample: Confidence-Aware Self-Consistency for Efficient LLM Chain-of-Thought Reasoning
Juming Xiong, Kevin Guo, Congning Ni, Chao Yan, Katherine Brown · Mar 9, 2026 · Citations: 0
Automatic Metrics Math
Recent self-consistency-based approaches further improve accuracy but require sampling and aggregating multiple reasoning trajectories, leading to substantial additional computational overhead.
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Moments
Yuquan Wang, Mi Zhang, Yining Wang, Geng Hong, Mi Wen · Aug 6, 2025 · Citations: 0
General
It injects timely safety aha moments during the reasoning process to guide the model towards harmless yet helpful reasoning.
- GLiGuard: Schema-Conditioned Classification for LLM Safeguard
Urchade Zaratiana, Mary Newhauser, George Hurn-Maloney, Ash Lewis · May 8, 2026 · Citations: 0
Automatic Metrics Coding
Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimensions.
- $\texttt{YC-Bench}$: Benchmarking AI Agents for Long-Term Planning and Consistent Execution
Muyu He, Adit Jain, Anand Kumar, Vincent Tu, Soumyadeep Bakshi · Apr 1, 2026 · Citations: 0
Automatic Metrics General
As LLM agents tackle increasingly complex tasks, a critical question is whether they can maintain strategic coherence over long horizons: planning under uncertainty, learning from delayed feedback, and adapting when early mistakes compound.
- Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization
Qianben Chen, Tianrui Qin, King Zhu, Qiexiang Wang, Chengjun Yu · Feb 26, 2026 · Citations: 0
Automatic Metrics General
Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios.
- Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
Amin Karimi Monsefi, Dominic Culver, Nikhil Bhendawade, Manuel R. Ciosici, Yizhe Zhang · May 8, 2026 · Citations: 0
Automatic Metrics General
Each training trajectory is built through a chain of blind stochastic jumps with no evaluation of sequence quality; a single bad decision at an early midpoint propagates through subsequent steps, yet the student must imitate the result.
- Luna-2: Scalable Single-Token Evaluation with Small Language Models
Vatsal Goel, Rishon Dsouza, Nikhil Ega, Amey Ramesh Rambatla, Rob Friel · Feb 20, 2026 · Citations: 0
Llm As JudgeAutomatic Metrics General
We present Luna-2, a novel architecture that leverages decoder-only small language models (SLMs) into a deterministic evaluation model to reliably compute complex task-specific LLMAJ metrics (e.g.
- TabAgent: A Framework for Replacing Agentic Generative Components with Tabular-Textual Classifiers
Ido Levy, Eilam Shapira, Yinon Goldshtein, Avi Yaeli, Nir Mashkif · Feb 18, 2026 · Citations: 0
Automatic Metrics General
We propose TabAgent, a framework for replacing generative decision components in closed-set selection tasks with a compact textual-tabular classifier trained on execution traces.
- AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering
Yuxin Wang, Jiahao Lu, Qifeng Wu, Shicheng Fang, Chuanyuan Tan · May 29, 2026 · Citations: 0
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
Yuxuan Ye, Raul Santos-Rodriguez, Edwin Simpson · May 28, 2026 · Citations: 0
- Training Deliberative Monitors for Black-Box Scheming Detection
Aditya Sinha, Akshat Naik, Victor Gillioz, Simon Storf, Kilian Merkelbach · May 28, 2026 · Citations: 0
- Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspective
Shenghao Ye, Yuxiang Wang, Yu Guo, Dong Jin, Shuangwu Chen · May 28, 2026 · Citations: 0
- Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
Sterling Huang, Abigayle Brown, Jiyoo Noh, Jiakang Xu, Wantong Huo · May 18, 2026 · Citations: 0
- FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing
Zihan Tang, Leqi Shen, Hui Chen, Ao Wang, Ben Wan · May 17, 2026 · Citations: 0
- When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviews
Sandeep Kumar, Yash Kamdar, Abid Hossain, Bharti Kumari, Tanik Saikh · May 11, 2026 · Citations: 0
- PaT: Planning-after-Trial for Efficient Test-Time Code Generation
Youngsik Yoon, Sungjae Lee, Seockbean Song, Siwei Wang, Wei Chen · May 8, 2026 · Citations: 0
- A Few Good Clauses: Comparing LLMs vs Domain-Trained Small Language Models on Structured Contract Extraction
Nicole Lincoln, Nick Whitehouse, Jaron Mar, Rivindu Perera · May 7, 2026 · Citations: 0
- ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesis
Atharva Naik, Yash Mathur, Prakam, Carolyn Rose, David Mortensen · May 6, 2026 · Citations: 0
- RAG over Thinking Traces Can Improve Reasoning Tasks
Negar Arabzadeh, Wenjie Ma, Sewon Min, Matei Zaharia · May 5, 2026 · Citations: 0
- Length Value Model: Scalable Value Pretraining for Token-Level Length Modeling
Zhen Zhang, Changyi Yang, Zijie Xia, Zhen Yang, Chengzhi Liu · Apr 29, 2026 · Citations: 0
- Anchor-and-Resume Concession Under Dynamic Pricing for LLM-Augmented Freight Negotiation
Hoang Nguyen, Lu Wang, Marta Gaia Bras · Apr 22, 2026 · Citations: 0
- Decoding Text Spans for Efficient and Accurate Named-Entity Recognition
Andrea Maracani, Savas Ozkan, Junyi Zhu, Sinan Mutlu, Mete Ozay · Apr 22, 2026 · Citations: 0
- Stochasticity in Tokenisation Improves Robustness
Sophie Steger, Rui Li, Sofiane Ennadir, Anya Sims, Arno Solin · Apr 17, 2026 · Citations: 0
- Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization
Haochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu, Enyan Dai · Apr 16, 2026 · Citations: 0
- DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines
Gabriel Pimenta de Freitas Cardoso, Caio Lucas da Silva Chacon, Jonas Felipe da Fonseca Oliveira, Paulo Henrique de Medeiros Araujo · Apr 15, 2026 · Citations: 0
- Robust Explanations for User Trust in Enterprise NLP Systems
Guilin Zhang, Kai Zhao, Jeffrey Friedman, Xu Chu, Amine Anoun · Apr 13, 2026 · Citations: 0
- Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang · Apr 3, 2026 · Citations: 0
- Ensemble Self-Training for Unsupervised Machine Translation
Ido Aharon, Jonathan Shaki, Sarit Kraus · Mar 17, 2026 · Citations: 0
- Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
Disha Sheshanarayana, Rajat Subhra Pal, Manjira Sinha, Tirthankar Dasgupta · Mar 16, 2026 · Citations: 0
- Adaptive Vision-Language Model Routing for Computer Use Agents
Xunzhuo Liu, Bowei He, Xue Liu, Andy Luo, Haichen Zhang · Mar 13, 2026 · Citations: 0
- Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning
Junseok Kim, Nakyeong Yang, Kyungmin Min, Kyomin Jung · Jan 6, 2026 · Citations: 0
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards
Zhilin Wang, Jiaqi Zeng, Olivier Delalleau, Ellie Evans, Daniel Egert · Sep 25, 2025 · Citations: 0