- Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization
Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou · Apr 8, 2026 · Citations: 0
Human EvalAutomatic Metrics General
Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
- SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov · Jun 29, 2026 · Citations: 0
Automatic Metrics Math
As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
- MemRerank: Preference Memory for Personalized Product Reranking
Zhiyuan Peng, Xuyang Wu, Huaixiao Tou, Yi Fang, Yu Gong · Mar 31, 2026 · Citations: 0
Automatic Metrics General
LLM-based shopping agents increasingly rely on long purchase histories and multi-turn interactions for personalization, yet naively appending raw history to prompts is often ineffective due to noise, length, and relevance mismatch.
- Multi-Agent Environments for Vehicle Routing Problems
Ricardo Gama, Ricardo Cunha, Daniel Fuertes, Carlos R. del-Blanco, Hugo L. Fernandes · Nov 21, 2024 · Citations: 0
Simulation Env Coding
Here, we propose MAEnvs4VRP library, a unified framework for multi-agent vehicle routing environments that supports classical, dynamic, stochastic, and multi-task problem variants within a single modular design.
- CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
Ekaterina Trofimova, Emil Sataev, Abhijit Singh Jowhari · Aug 23, 2024 · Citations: 0
Automatic Metrics Coding
Evaluations on diverse scientific papers demonstrate CodeRefine's ability to improve code implementation from the paper, potentially accelerating the adoption of cutting-edge algorithms in real-world applications.
- DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
Dana Rezazadegan, Atie Kia, Phongpadid Nandavong, Dominique Carlon, Jeremy Nguyen · Jun 26, 2026 · Citations: 0
- Compact Geometric Representations of Hierarchies
Prashant Gokhale, Piotr Indyk, Yuhao Liu, Sandeep Silwal, Tony Chang Wang · Jun 16, 2026 · Citations: 0
- Efficient Benchmarking Is Just Feature Selection and Multiple Regression
Sam Bowyer, Acyr Locatelli, Kris Cao · May 25, 2026 · Citations: 0
- Qwen Goes Brrr: Off-the-Shelf RAG for Ukrainian Multi-Domain Document Understanding
Anton Bazdyrev, Ivan Bashtovyi, Ivan Havlytskyi, Oleksandr Kharytonov, Artur Khodakovskyi · May 11, 2026 · Citations: 0
- The Cost of Context: Mitigating Textual Bias in Multimodal Retrieval-Augmented Generation
Hoin Jung, Xiaoqian Wang · May 7, 2026 · Citations: 0
- Learning When to Remember: Risk-Sensitive Contextual Bandits for Abstention-Aware Memory Retrieval in LLM-Based Coding Agents
Mehmet Iscan · Apr 30, 2026 · Citations: 0
- Efficient Rationale-based Retrieval: On-policy Distillation from Generative Rerankers based on JEPA
Teng Chen, Sheng Xu, Feixiang Guo, Xiaoyu Wang, Qingqing Gu · Apr 25, 2026 · Citations: 0
- Self-Aware Vector Embeddings for Retrieval-Augmented Generation: A Neuroscience-Inspired Framework for Temporal, Confidence-Weighted, and Relational Knowledge
Naizhong Xu · Apr 22, 2026 · Citations: 0
- RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration
Fabian Ridder, Laurin Lessel, Malte Schilling · Apr 17, 2026 · Citations: 0
- Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
Manan Gupta, Dhruv Kumar · Apr 16, 2026 · Citations: 0
- Better and Worse with Scale: How Contextual Entrainment Diverges with Model Size
Dikshant Kukreja, Kshitij Sah, Gautam Gupta, Avinash Anand, Rajiv Ratn Shah · Apr 14, 2026 · Citations: 0
- Forgetting to Witness: Efficient Federated Unlearning and Its Visible Evaluation
Houzhe Wang, Xiaojie Zhu, Chi Chen · Apr 6, 2026 · Citations: 0
- Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional non-robustness and how to fix them
Ole Delzer, Sidney Bender · Apr 6, 2026 · Citations: 0
- Abnormal Head Movements in Neurological Conditions: A Knowledge-Based Dataset with Application to Cervical Dystonia
Saja Al-Dabet, Sherzod Turaev, Nazar Zaki · Apr 2, 2026 · Citations: 0
- Decidable By Construction: Design-Time Verification for Trustworthy AI
Houston Haynes · Mar 26, 2026 · Citations: 0
- Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA
Saahil Mathur, Ryan David Rittner, Vedant Ajit Thakur, Daniel Stuart Schiff, Tunazzina Islam · Mar 25, 2026 · Citations: 0
- L2GTX: From Local to Global Time Series Explanations
Ephrem Tibebe Mekonnen, Luca Longo, Lucas Rizzo, Pierpaolo Dondio · Mar 13, 2026 · Citations: 0
- CIRCUS: Circuit Consensus under Uncertainty via Stability Ensembles
Swapnil Parekh · Feb 28, 2026 · Citations: 0
- Scaling Search Relevance: Augmenting App Store Ranking with LLM-Generated Judgments
Evangelia Christakopoulou, Vivekkumar Patel, Hemanth Velaga, Sandip Gaikwad · Feb 26, 2026 · Citations: 0
- InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs
Lv Tang, Tianyi Zheng, Bo Li, Xingyu Li · Feb 2, 2026 · Citations: 0
- From Domains to Instances: Dual-Granularity Data Synthesis for LLM Unlearning
Xiaoyu Xu, Minxin Du, Zitong Li, Zi Liang, Zhibiao Guo · Jan 7, 2026 · Citations: 0
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment
Priyanka Mudgal · Nov 6, 2025 · Citations: 0
- Patent Representation Learning via Self-supervision
You Zuo, Kim Gerdes, Eric Villemonte de La Clergerie, Benoît Sagot · Nov 3, 2025 · Citations: 0
- Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe Adaptation
Tianyi Hu, Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich · Jul 29, 2025 · Citations: 0
- From Time Series Analysis to Question Answering: A Survey in the LLM Era
Wei Li, Zhe Xie, Yuxuan Liang, Xinli Hao, Yunyao Cheng · Jun 13, 2025 · Citations: 0