Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Automatic Metrics

Automatic Metrics evaluation setups appearing in the current HFEPX corpus (2567 papers).

Papers in tag: 2,567

Running a Automatic Metrics study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (20)
  • Llm As Judge (1)

Human Feedback Types

  • Critique Edit (2)
  • Pairwise Preference (1)

Required Expertise

  • General (16)
  • Coding (3)
  • Math (2)
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic · Aug 26, 2026 · Citations: 0

Critique Edit Llm As JudgeAutomatic Metrics General
  • Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts.
  • On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to…
Learning New Facts with QLoRA: An Acquisition-Retention Frontier

Estelle Zheng, Sébastien Warichet, Emmanuel Helbert, Christophe Cerisara · Aug 26, 2026 · Citations: 0

Automatic Metrics MathCoding
  • We study factual acquisition in a controlled OpenStreetMap-derived benchmark where Qwen3-4B must acquire anonymized geographic associations while retaining unrelated capabilities.
  • Low-rank QLoRA preserves out-of-domain (OOD) performance but acquires fewer facts, whereas higher ranks improve same-fact paraphrase generalization at an increasing cost in performance on unrelated benchmarks.
Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

Raúl Vázquez, Aman Sinha, Chuyuan Li, Claudio Savelli, Eduardo Calò, Emilio Raimond · Aug 26, 2026 · Citations: 0

Automatic Metrics General
  • Building on the recently introduced SHEEP dataset, designed for long-term evaluation across model generations, the task invites participants to detect and classify fine-grained hallucination spans in image-conditioned text generation (VQA,…
  • The evaluation uses a five-class taxonomy of hallucinations spanning four languages: Chinese, English, French, and Italian.
Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking

Zhexi Feng, Wuxi Chen, Bingrui Zhang · Aug 26, 2026 · Citations: 0

Automatic Metrics General
  • An ICE-specific reversal appears in a released 301-question NQ-open DPR-BERT pipeline: its average-confidence baseline improves instance-level calibration error by 0.045 under exact match but worsens it by 0.074 under human correctness,…
  • TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, maintains at least nominal coverage, narrows intervals, and repairs confidence subject to a base-rate deployment…
AWM: Answerable Working Memory for Long-Document VQA Agents

Dongzhuoran Zhou, Yuqicheng Zhu, Yule Liu, Zhen Yang, Rui Lu, Yuxiao Dong · Aug 26, 2026 · Citations: 0

Automatic Metrics General
  • Long-document visual question answering increasingly relies on VLM agents that retrieve candidate pages, inspect page images, write findings to working memory, and synthesize answers.
  • Working memory should carry answer-supporting evidence across page inspections for later grounded answering, yet existing evaluation mainly checks final-answer correctness and evidence-page access.
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek, Shiran Liu · Aug 26, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored.
  • This paper presents a comparative study of both families for ASR evaluation.
MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

Jiaxin Yuan, Connor Martinez Lockhart, Xiaoyu Liu, Jiaqi Wang, Chenghao Deng, Xiayimei Han · Aug 26, 2026 · Citations: 0

Automatic Metrics Math
  • Formal theorem proving enables machine-verifiable evaluation of mathematical reasoning, yet existing benchmarks often emphasize aggregate proof accuracy, concentrate on a narrow range of mathematics, and provide limited evidence of…
  • We introduce MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics.