- Preference Shapes Relevance: Cross-component Hierarchical Semantic Alignment for Personalized Generative Retrieval
Gaoming Zhang, Angqing Jiang, Jianchun Song, Kena Qi, Dayao Chen · Aug 31, 2026 · Citations: 0
Coding
To address these challenges, we propose Cross-component Hierarchical semantic Alignment for Personalized generative retrieval (CHAP), a novel personalized GR framework from a hierarchical perspective.
- Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking
Xiaoyang Chen, Jie Liu, Haijin Liang, Haibo Shi, Jin Ma · Aug 31, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- ShardMemo: Scope-Before-Routing for Agentic Memory Retrieval
Yang Zhao, Chengxiao Dai, Mengying Kou, Yue Xiu, Dusit Niyato · Jan 29, 2026 · Citations: 0
Automatic Metrics Law
We present SHARD-MEMO, an agentic memory system built on scope-before-routing: metadata predicates first identify the admissible shards, and a learned router then selects a small number of them for shard-local approximate nearest neighbor…
- When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
Kazuki Nakayashiki · Aug 26, 2026 · Citations: 0
General
An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer authoritative record.
- CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval
Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu · Aug 26, 2026 · Citations: 0
Simulation Env Coding
Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem.
- ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains
Jinpu Jiang, Xuan Wu, Wenhao Song, Bo Yang, You Zhou · Aug 26, 2026 · Citations: 0
Automatic Metrics General
To address this limitation, we propose ReliableRAG, which, to the best of our knowledge, is the first reliability-driven framework that mitigates deceptive misinformation in multi-hop QA through fine-grained evaluation of individual…
- Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson · Aug 21, 2026 · Citations: 0
Automatic Metrics Coding
We propose an Evaluation Agent, middleware that combines Natural Language Inference (NLI) factual verification, a five-signal poison detector with relevance-weighted aggregation, and a Trust Index T = 0.4 F + 0.35 C + 0.25 (1 - P ) with a…
- Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance
Sergiy Lunyakin · Aug 21, 2026 · Citations: 0
Automatic Metrics General
An evaluation on FinanceBench (145 questions) compares KDAF against zero-context inference, BM25, concept-weighted lexical retrieval, and ungrounded graph traversal.
- Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents
Yiyang Feng, Biddut Sarker Bijoy, Niranjan Balasubramanian, Jiawei Zhou · Aug 20, 2026 · Citations: 0
Automatic Metrics Coding
Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience.
- GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval
Ernest Beta, Odysseas S. Chlapanis, Dimitrios Galanis, Ion Androutsopoulos · Aug 19, 2026 · Citations: 0
Automatic Metrics LawMultilingual
We introduce GreekBarRetrieval, a public retrieval benchmark derived from, and complementing GreekBarBench, which did not include retrieval.
- EvoSelect: Data-Efficient LLM Evolution for Targeted Task Adaptation
Ting-Wei Li, Sirui Chen, Jiaru Zou, Yingbing Huang, Tianxin Wei · Apr 28, 2026 · Citations: 0
Automatic Metrics General
Such adaptation often requires iteratively improving the model toward a targeted task, yet collecting high-quality human-labeled data to support this process is costly and difficult to scale.
- Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization
Aarik Gulaya · May 27, 2026 · Citations: 0
Automatic Metrics General
We evaluate the Specification on a prototype benchmark of held-out behavioral predictions scored by a calibrated 5-judge LLM panel.
- CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
Enhan Li, Junhao He, Hongyang Du · Aug 13, 2026 · Citations: 0
General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- GEM: A Generative Embedding Model Bridging Reasoning and Retrieval
Zhili Shen, Craig Macdonald · Aug 13, 2026 · Citations: 0
Automatic Metrics Coding
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Le Zhang, Ke Sun · Aug 12, 2026 · Citations: 0
Automatic Metrics General
We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent.
- CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation
Zhipeng Song, Yizhi Zhou, Xiangyu Kong, Jiulong Jiao, Xuezhou Ye · May 6, 2026 · Citations: 0
Automatic Metrics General
CAR converts these confidence changes into coarse precedence constraints and returns the feasible ranking with minimum Kendall distance from the baseline, preserving existing pairwise preferences unless generator-side evidence supports…
- QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
Yilin Liu, Rui Meng, Wangze Ni, Jianxin Yan, Heng Cao · Aug 12, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning
Yuhang Wu, Xiangqing Shen, Fanfan Wang, Cangqi Zhou, Zhen Wu · Apr 2, 2026 · Citations: 0
Automatic Metrics General
However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process.
- mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health
Yi Ren · Jun 28, 2026 · Citations: 0
Automatic Metrics Medicine
Medical question-answering benchmarks rarely cover the maternal, neonatal, child, and reproductive-health questions a nurse-midwife asks, and, to our knowledge, no public chunk-level relevance benchmark exists for maternal-health guideline…
- PACE: A Proxy for Agentic Capability Evaluation
Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja, Jiayi Geng · Jul 2, 2026 · Citations: 0
Automatic Metrics Coding
We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict model performances on agentic benchmarks.
- AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations
Javier Irigoyen, Roberto Daza, Francisco Jurado, Julian Fierrez, Ruben Tolosana · Jul 2, 2026 · Citations: 0
Automatic Metrics General
For each question, the dataset includes an explanation written by a human teacher alongside 11 explanations generated by LLM-simulated teacher profiles associated with distinct pedagogical risks.
- Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory
ZhiShu Jiang, Haibo Liu, Xin Shen, Guanqiang QI, Chenxi Miao · May 28, 2026 · Citations: 0
General
Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled for the right user.
- Evergreen: Efficient Claim Verification for Semantic Aggregates
Alexander W. Lee, Benjamin Han, Shayak Sen, Sam Yeom, Ugur Cetintemel · Apr 28, 2026 · Citations: 0
Automatic Metrics General
On a benchmark of production-inspired workloads over restaurant review and customer support datasets, Evergreen's optimized configurations occupy the entire cost-quality Pareto frontier.
- What Survives Into Context: A Diagnostic for Budget-Constrained Multi-Hop RAG and When Submodular Evidence Packing Improves It
Ananto Nayan Bala · Jul 1, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolio
Anzhe Xie, Weihang Su, Yujia Zhou, Yiqun Liu, Qingyao Ai · Jun 15, 2026 · Citations: 0
Automatic Metrics General
Its structured, verifiable workflow makes it an ideal substrate for evaluating systematic scientific reasoning, yet existing benchmarks lack ground truth across the full retrieval-screening-synthesis pipeline.
- DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasks
Romain Karpinsky, Julien Mozziconacci, Mickaël Delcey · Jun 29, 2026 · Citations: 0
Automatic Metrics General
However, systematic benchmark comparisons across these methods remain scarce.
- Information Dynamics of Language Communication
Leonardo S. Goodall, Andrea I. Luppi, Pedro A. M. Mediano · Jun 29, 2026 · Citations: 0
Medicine
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
Nikolay Georgiev, Maria Drencheva, Kseniia Ibragimova, Ivo Petrov, Dimitar I. Dimitrov · Jun 29, 2026 · Citations: 0
Automatic Metrics Math
As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databases, theorem libraries, and educational resources.
- Improving Answer Extraction in Context-based Question Answering Systems Using LLMs
Hafez Abdelghaffar, Ahmed Alansary, Ali Hamdi · Jun 4, 2026 · Citations: 0
Automatic Metrics General
Our methodology involves fine-tuning a pre-trained LLM on a benchmark QA dataset to improve its contextual comprehension and answer extraction capabilities.
- Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation
Ahmed Alansary, Molham Mohamed, Ali Hamdi · Jun 3, 2026 · Citations: 0
Automatic Metrics Medicine
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization
Ahmed Alansary, Ali Hamdi · Jun 3, 2026 · Citations: 0
Automatic Metrics General
The generated summaries are then evaluated using automatic evaluation metrics that capture both lexical similarity and semantic relevance.
- Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents
Yuxin Wang, Paul Thomas, Zhiwei Yu, Yuan Gao, Saeed Hassanpour · Jun 24, 2026 · Citations: 0
Automatic Metrics General
Specifically, how they shape an agent's responses under varying conversational contexts and whether they lead to substantively different response behaviors.
- PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptation
Kirill Dubovikov, Omar El Mansouri, Hachem Madmoun, Yanda Li, Sandeep Kumar · Jun 23, 2026 · Citations: 0
Automatic Metrics General
Reranker adaptation improves the public Earth Science benchmark by 44% relative and a six-task reasoning-intensive panel by 23%.
- AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression
Yijing Chen, Wenhui Tan, Xiaoyi Yu, Yuyue Wang, Xin Cheng · Jun 23, 2026 · Citations: 0
Automatic Metrics General
Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively.
- FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection
Shaswata Mitra, Subash Neupane, Martin Duclos, Sudip Mittal, Aritran Piplai · Aug 26, 2025 · Citations: 0
Automatic Metrics General
To address these challenges, we introduce FALCON, an agentic framework for CTI-grounded rule retrieval, generation, and validation.
- Telenor Nordics Customer Service self-help corpus
Mike Riess · May 26, 2026 · Citations: 0
Automatic Metrics Multilingual
The documents have been sourced from the public self-help pages of four Nordic telecommunications operators and subsequently filtered for person-identifiable information and relevance through a combined LLM and human annotation pipeline.
- Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models
Darrien McKenzie, Nicklas Hansen, Xiaolong Wang · Jun 18, 2026 · Citations: 0
Automatic Metrics General
Empirically, we find that different sampling strategies induce non-trivial tradeoffs between productivity (learning signal), diversity (coverage of the task manifold), and utility (evaluation relevance).
- MemRerank: Preference Memory for Personalized Product Reranking
Zhiyuan Peng, Xuyang Wu, Huaixiao Tou, Yi Fang, Yu Gong · Mar 31, 2026 · Citations: 0
Automatic Metrics General
LLM-based shopping agents increasingly rely on long purchase histories and multi-turn interactions for personalization, yet naively appending raw history to prompts is often ineffective due to noise, length, and relevance mismatch.
- Evaluation of Automatic Speech Recognition Using Generative Large Language Models
Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek · Apr 23, 2026 · Citations: 0
Automatic Metrics General
Embedding-based semantic metrics are better correlated with human perception, but decoder-based Large Language Models (LLMs) remain underexplored for this task.
- Beyond Transcripts: A Renewed Perspective on Audio Chaptering
Fabian Retkowski, Maike Züfle, Thai Binh Nguyen, Jan Niehues, Alexander Waibel · Feb 9, 2026 · Citations: 0
Automatic Metrics General
Despite its relevance, research remains limited and text-based, leaving key questions unresolved about leveraging audio information, handling ASR errors, and transcript-free evaluation.
- SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extraction
Jiawei He, Mengyu Shi, Jiawei Liu, Dong Sun, Chunrong Fang · May 22, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Evaluation of Chunking Strategies for Effective Text Embedding in Low-Resource Language on Agricultural Documents
Sovandara Chhoun, Pichdara Po, Sereiwathna Ros, Wan-Sup Cho, Saksonita Khoeurn · May 21, 2026 · Citations: 0
Automatic Metrics Multilingual
For evaluation, we perform 5-fold cross-validation over 18 question-answer pairs.
- A Comparative Study of Language Models for Khmer Retrieval-Augmented Question Answering
Sereiwathna Ros, Phannet Pov, Ratanaktepi Chhor, Kimleang Ly, Wan-Sup Cho · May 21, 2026 · Citations: 0
Automatic Metrics General
We conduct a two-phase comparative evaluation.
- Causal Intervention-Based Memory Selection for Long-Horizon LLM Agents
Saksham Sahai Srivastava · May 17, 2026 · Citations: 0
Automatic Metrics Coding
Long-horizon LLM agents rely on persistent memory to support interactions across sessions, yet existing memory systems often retrieve context using semantic similarity or broad history inclusion, treating retrieved memories as uniformly…
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuning
Manish Nagaraj, Sakshi Choudhary, Utkarsh Saxena, Deepak Ravikumar, Kaushik Roy · Oct 8, 2025 · Citations: 0
General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution
Qinhao Chen, Linyang He, Nima Mesgarani · Apr 18, 2026 · Citations: 0
Automatic Metrics General
PIE connects Pruning, automatic Interpretation, and interpretation Evaluation, establishing a comprehensive benchmarking environment to systematically measure behavioral fidelity and downstream interpretability under pruning.
- Attribution-Guided Pruning for Insight and Control: Circuit Discovery and Targeted Correction in Small-scale LLMs
Sayed Mohammad Vakilzadeh Hatefi, Maximilian Dreyer, Reduan Achtibat, Patrick Kahardipraja, Thomas Wiegand · Jun 16, 2025 · Citations: 0
Coding
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Screening Is Enough
Ken M. Nakanishi · Apr 1, 2026 · Citations: 0
General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval
Chunyu Li, Jingyi Kang, Ding Chen, Mengyuan Zhang, Jiajun Shen · May 7, 2026 · Citations: 0
Automatic Metrics General
In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory.
- Correct Is Not Enough: Training Reasoning Planners with Executor-Grounded Rewards
Tianyang Han, Hengyu Shi, Junjie Hu, Xu Yang, Zhiling Wang · May 5, 2026 · Citations: 0
Automatic Metrics MathLaw
Extensive experiments on code and math benchmarks show that this executor-grounded reasoning reward improves the two-stage planner-executor system over execution-only training, suggesting that reasoning supervision should evaluate not only…
- Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics
Dominik Dahlem, Diego Maniloff, Mac Misiura · May 6, 2026 · Citations: 0
Automatic Metrics General
The resulting two-axis diagnostic (φ for capacity, G for direction) yields a falsifiable polarity prediction: bottleneck- and diffuse-dominated benchmarks should exhibit opposite polarity.
- Sub-Token Routing in LoRA for Adaptation and Query-Aware KV Compression
Wei Jiang, Wei Wang · Apr 23, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
Yilun Zhao, Jinbiao Wei, Tingyu Song, Siyue Zhang, Chen Zhao · May 5, 2026 · Citations: 0
Automatic Metrics General
This capability is increasingly important for agentic search systems, where retrievers must provide complementary evidence across iterative search and synthesis.
- Reproducing Complex Set-Compositional Information Retrieval
Vincent Degenhart, Dewi Timman, Arjen P. de Vries, Faegheh Hasibi, Mohanna Hoveyda · May 5, 2026 · Citations: 0
Automatic Metrics Coding
We conduct a reproducibility study to benchmark major retrieval families and reasoning-targeted methods on QUEST and QUEST+Variants, and introduce LIMIT+, a controlled benchmark where relevance depends on arbitrary attribute predicates and…
- Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
Negar Arabzadeh, Andrew Drozdov, Michael Bendersky, Matei Zaharia · Apr 24, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- LLMs as Assessors: Right for the Right Reason?
Sourav Saha, Mandar Mitra, Aditya Dutta · Jan 13, 2026 · Citations: 0
Automatic Metrics General
A good deal of recent research has focused on how Large Language Models (LLMs) may be used as judges in place of humans to evaluate the quality of the output produced by various text / image processing systems.
- AgentSearchBench: A Benchmark for AI Agent Search in the Wild
Bin Wu, Arastun Mammadli, Xiaoyu Zhang, Emine Yilmaz · Apr 24, 2026 · Citations: 0
Automatic Metrics Coding
The rapid growth of AI agent ecosystems is transforming how complex tasks are delegated and executed, creating a new challenge of identifying suitable agents for a given task.
- How Hard is it to Decide if a Fact is Relevant to a Query?
Meghyn Bienvenu, Diego Figueira, Pierre Lafourcade · Apr 24, 2026 · Citations: 0
Automatic Metrics General
Relevance has already been shown to be harder than query evaluation: namely, Σ^p_2-complete for CQs, even over a binary signature.
- Context-Fidelity Boosting: Enhancing Faithful Generation through Watermark-Inspired Decoding
Weixu Zhang, Fanghua Ye, Qiang Gao, Jian Li, Haolun Wu · Apr 24, 2026 · Citations: 0
Automatic Metrics General
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
- Handling Missing Modalities in Multimodal Survival Prediction for Non-Small Cell Lung Cancer
Filippo Ruffini, Camillo Maria Caruso, Claudia Tacconi, Lorenzo Nibid, Francesca Miccolis · Jan 15, 2026 · Citations: 0
MedicineMultilingual
Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.