Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Pairwise Preference

Pairwise Preference papers with explicit human-feedback protocol signal (436 papers).

Papers in tag: 436

Running a Pairwise Preference study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (10)
  • Llm As Judge (3)
  • Human Eval (1)

Human Feedback Types

  • Pairwise Preference (20)
  • Expert Verification (1)
  • Rubric Rating (1)

Required Expertise

  • General (13)
  • Coding (4)
  • Medicine (2)
Localize-Then-Decide Guarantees for LLM Judgments

Xinyu Li, Yi Zhou, Guanqun Cao, Zeyu Fu, Tianjin Huang, Gaojie Jin · Aug 26, 2026 · Citations: 0

Pairwise Preference General
  • Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging.
  • Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans.
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu · Aug 26, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored.
  • This paper presents a comparative study of both families for ASR evaluation.
Enhancing LLMs in Predictive Political QA with Semi-Structured Data

Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang · Aug 21, 2026 · Citations: 0

Pairwise Preference Simulation Env General
  • We identify two complementary signals for predictive political QA: actor stances that capture issue-specific preferences, and high-order structure signals that capture indirect dependencies among political actors.
Personalized Privacy Control in LLMs via Attention Head Intervention

Junseok Kim, Nakyeong Yang, Kyomin Jung · Aug 21, 2026 · Citations: 0

Pairwise Preference General
  • The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns.
  • To address this limitation, we introduce personalized privacy, which incorporates user-specific disclosure preferences into privacy control.
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu · Aug 21, 2026 · Citations: 0

Pairwise PreferenceExpert Verification Llm As JudgeAutomatic Metrics Medicine
  • While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format.
  • To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter…
Source-Free MT Evaluation Is Not MT Evaluation

Baban Gain, Ramakrishna Appicharla, Asif Ekbal · Aug 21, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Multilingual
  • Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often correlate less well with human judgments.
  • As a result, source-free, reference-based evaluation has become the practical norm, even though it is unfaithful to the definition of translation adequacy and unfair to systems whose outputs preserve the source meaning while differing from…
LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

Matan Rusanovsky, Yoav Miron, Roy Uziel, Omer Belhasin, Ran Zilberstein, Maor Ashkenazi · Aug 20, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • Over the vanilla DFlash drafter, LiLiCorr raises acceptance length on every benchmark by 9 to 19%, while its scoring head accounts for about 2.8% of the per-block latency.
  • Against DFlash and two concurrent methods that also restore coherence at draft time, LiLiCorr delivers the highest throughput in 70 of 72 settings: nine benchmarks at two target sizes under greedy and temperature-one decoding, and a…
When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave · Aug 20, 2026 · Citations: 0

Pairwise Preference General
  • To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned…
  • Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability…
OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models

Nikita Khudov · Aug 20, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Coding
  • We introduce OenoBench, a wine-domain knowledge benchmark of 3,266 multiple-choice questions across six pillars (regions, grape varieties, viticulture, winemaking, producers, business) and four difficulty tiers.
  • Evaluating sixteen frontier configurations, we find: (i) overall accuracy spans 53%-84%, led by o3 at 83.6%; (ii) reasoning-mode lift concentrates in DeepSeek R1 (+6.8pp) and is absent in Claude Opus and Gemini Pro; (iii) Anthropic shows…
Stopping and Routing LLM Judge Panels

Bin Zhu, Yi Xie, Yanghui Rao · Aug 20, 2026 · Citations: 0

Pairwise Preference Llm As Judge MathCoding
  • LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
  • The deployment question is not only which judge is best, but which judges should be called, on which examples, and when panel construction should stop.
PersonalBench: Measuring the Authorship Gap in LLM Personalization

Yash Ganpat Sawant · Aug 20, 2026 · Citations: 0

Pairwise Preference Llm As JudgeAutomatic Metrics General
  • Personalized text generation aims to make LLMs write in a specific individual's style, yet existing benchmarks measure task accuracy or preference alignment rather than whether the model's output actually resembles the target author's…
  • We introduce PersonalBench, a benchmark that evaluates inference-time personalization methods through three independent lenses: LUAR (a trained authorship verification model), an LLM-as-judge, and automated stylometrics.
The Asymmetric Harms of LLM Compression

Yuan Wu, Mairui Li, Lesia Semenova, Chudi Zhong · Aug 20, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • Finally, we demonstrate that stable aggregate bias scores can conceal substantial, opposing shifts in stereotypical preferences across demographic subgroups.
  • Together, these findings reveal asymmetric behavioral changes that aggregate performance measures fail to capture, highlighting the need for granular evaluation of compressed models before deployment.
PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment

Jiawei Feng, Jiancan Wu, Xingyu Zhu, Junkang Wu, Xiang Wang, Xiangnan He · Aug 20, 2026 · Citations: 0

Pairwise Preference General
  • Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences.
  • To address these challenges, we propose Perception-Enhanced Alignment DPO (PEA-DPO), a framework for multimodal LLMs alignment, which explicitly leverages visual preference signals to overcome visual insensitivity.
HARP: Hierarchical Adaptive Ranking with Preference-Adaptive Fusion for Query-Based CVE Prioritization

Haochen Liu, Zhengzhang Chen, Haoyu Wang, Yanchi Liu, Jundong Li, Haifeng Chen · Aug 19, 2026 · Citations: 0

Pairwise Preference General
  • Vulnerability prioritization is inherently preference dependent, since the same CVE can receive different remediation priority under different operational preference scenarios.
  • In practice, organizations already operate under a preference scenario, but this preference is often implicit and difficult to express as a written prompt instruction, while triage queries usually do not encode it.
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation

Cedar Site Bai, Zhenyu Liao, Duanshun Li, Sheikh Sarwar, Huiyuan Chen, Yuan Chen · Aug 16, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • However, guiding multi-turn interactions to elicit user preferences effectively remains challenging.
  • Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li · Aug 13, 2026 · Citations: 0

Pairwise Preference Human Eval Coding
  • In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on rollout feedback.
  • Across seven controlled code-agent-model configurations, integrating the learned DesignHarness consistently improves performance, increasing the average PosterBench Score from 54.99 to 67.39 (+12.4%).
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

Chenrun Wang, Mingxuan Zhu, Tiancheng Huang, Wenjie Li, Yujie Zhang, Zichen Zhu · Aug 13, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions.
  • In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation.
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang · Aug 13, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Medicine
  • Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment.
  • We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports.