CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era

Zhengqing Yuan, Kaiwen Shi, Zheyuan Zhang, Lichao Sun, Nitesh V. Chawla, Yanfang Ye · Feb 26, 2026 · Citations: 0

Abstract

Scientific research relies on accurate citation for attribution and integrity, yet large language models (LLMs) introduce a new risk: fabricated references that appear plausible but correspond to no real publications. Such hallucinated citations have already been observed in submissions and accepted papers at major machine learning venues, exposing vulnerabilities in peer review. Meanwhile, rapidly growing reference lists make manual verification impractical, and existing automated tools remain fragile to noisy and heterogeneous citation formats and lack standardized evaluation. We present the first comprehensive benchmark and detection framework for hallucinated citations in scientific writing. Our multi-agent verification pipeline decomposes citation checking into claim extraction, evidence retrieval, passage matching, reasoning, and calibrated judgment to assess whether a cited source truly supports its claim. We construct a large-scale human-validated dataset across domains and define unified metrics for citation faithfulness and evidence alignment. Experiments with state-of-the-art LLMs reveal substantial citation errors and show that our framework significantly outperforms prior methods in both accuracy and interpretability. This work provides the first scalable infrastructure for auditing citations in the LLM era and practical tools to improve the trustworthiness of scientific references.

HFEPX Relevance Assessment

This paper has direct human-feedback and/or evaluation protocol signal and is likely useful for eval pipeline design.

Eval-Fit Score

25/100 • Low

Treat as adjacent context, not a core eval-method reference.

Human Feedback Signal

Not explicit in abstract metadata

Evaluation Signal

Detected

HFEPX Fit

High-confidence candidate

If you are doing eval pipeline work, start here:

Human Eval Hub LLM-as-Judge Hub Pairwise Preference Hub Tool-Use Eval Hub

Human Data Lens

Uses human feedback: No
Feedback types: None
Rater population: Unknown
Unit of annotation: Unknown
Expertise required: General
Extraction source: Runtime deterministic fallback

Evaluation Lens

Evaluation modes: Automatic Metrics
Agentic eval: Multi Agent
Quality controls: Not reported
Confidence: 0.45
Flags: ambiguous, runtime_fallback_extraction

Protocol And Measurement Signals

Benchmarks / Datasets

No benchmark or dataset names were extracted from the available abstract.

Reported Metrics

accuracyfaithfulness

Research Brief

Deterministic synthesis

Meanwhile, rapidly growing reference lists make manual verification impractical, and existing automated tools remain fragile to noisy and heterogeneous citation formats and lack standardized evaluation. HFEPX signals include Automatic Metrics, Multi Agent with confidence 0.45. Updated from current HFEPX corpus.

Generated Mar 3, 2026, 8:34 PM · Grounded in abstract + metadata only

Key Takeaways

Meanwhile, rapidly growing reference lists make manual verification impractical, and existing automated tools remain fragile to noisy and heterogeneous citation formats and lack…
We present the first comprehensive benchmark and detection framework for hallucinated citations in scientific writing.

Researcher Actions

Treat this as method context, then pivot to protocol-specific HFEPX hubs.
Identify benchmark choices from full text before operationalizing conclusions.
Validate metric comparability (accuracy, faithfulness).

Caveats

Generated from title, abstract, and extracted metadata only; full-paper implementation details are not parsed.
Extraction confidence is probabilistic and should be validated for critical decisions.

Recommended Queries

human-eval protocol design agent eval benchmark comparison inter-rater agreement adjudication

Research Summary

Contribution Summary

Meanwhile, rapidly growing reference lists make manual verification impractical, and existing automated tools remain fragile to noisy and heterogeneous citation formats and lack standardized evaluation.
We present the first comprehensive benchmark and detection framework for hallucinated citations in scientific writing.
Our multi-agent verification pipeline decomposes citation checking into claim extraction, evidence retrieval, passage matching, reasoning, and calibrated judgment to assess whether a cited source truly supports its claim.

Why It Matters For Eval

Meanwhile, rapidly growing reference lists make manual verification impractical, and existing automated tools remain fragile to noisy and heterogeneous citation formats and lack standardized evaluation.
We present the first comprehensive benchmark and detection framework for hallucinated citations in scientific writing.

Researcher Checklist

Gap: Human feedback protocol is explicit

No explicit human feedback protocol detected.
Pass: Evaluation mode is explicit

Detected: Automatic Metrics
Gap: Quality control reporting appears

No calibration/adjudication/IAA control explicitly detected.
Gap: Benchmark or dataset anchors are present

No benchmark/dataset anchor extracted from abstract.
Pass: Metric reporting is present

Detected: accuracy, faithfulness

Related Papers

Papers are ranked by protocol overlap, extraction signal alignment, and semantic proximity.

A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruning Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systems Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
Hierarchical LLM-Based Multi-Agent Framework with Prompt Optimization for Multi-Robot Task Planning Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
Reshaping MOFs text mining with a dynamic multi-agents framework of large language model Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
SAMAS: A Spectrum-Guided Multi-Agent System for Achieving Style Fidelity in Literary Translation Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
The Emergence of Lab-Driven Alignment Signatures: A Psychometric Framework for Auditing Latent Bias and Compounding Risk in Generative AI Protocol Overlap

Citations: 0 Relevance: 4.60 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
- Shared metric mentions
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning Protocol Overlap

Citations: 0 Relevance: 3.70 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archives Protocol Overlap

Citations: 0 Relevance: 3.70 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup
An Agentic System for Rare Disease Diagnosis with Traceable Reasoning Protocol Overlap

Citations: 0 Relevance: 3.70 Shared tag: Multi Agent
- Shared HFEPX protocol tags
- Aligned agent-evaluation setup

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote