Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

HFEPX Metric Hub

Relevance Metric Papers

Updated from current HFEPX corpus (2026-09-02). This page tracks 60 papers for Relevance.

Read Full Context

Updated from current HFEPX corpus (2026-09-02). This page tracks 60 papers for Relevance. Use it to compare how relevance is measured across human feedback and evaluation studies.

Papers: 60 Last published: Aug 31, 2026 Global RSS

When This Metric Page Is Useful

Useful for shortlisting papers and comparing how this metric is used across different eval setups. Quality band: High .

Metric Coverage

100.0%

60 sampled papers include metric names.

Benchmark Anchoring

28.3%

Papers with explicit dataset/benchmark anchors for fair comparison.

Quality Controls

10.0%

6 papers report calibration/adjudication/IAA controls.

  • 60 papers are not low-signal flagged in this sample.
  • Use the protocol matrix below to avoid comparing metrics across incompatible eval setups.

Recommended next step: Use the top metric-reliable papers first, then compare benchmark context in the matrix before drawing conclusions.

Main limitation: Benchmark coverage is good enough for side-by-side comparison, but you should still verify definitions in the source papers.

What This Metric Page Tells You

What This Metric Page Tells You

  • Use this page to compare how relevance is operationalized across benchmarks and rater setups.
Metric Notes (Expanded)

Metric-Driven Protocol Takeaways

  • Relevance is often paired with automatic_metrics, simulation_env.

Metric Interpretation

  • relevance: 60 papers
  • accuracy: 14 papers
  • recall: 9 papers
  • cost: 8 papers

Benchmark Context

  • HotpotQA: 3 papers
  • FEVER: 2 papers
  • Agentsearchbench: 1 papers

Start Here (Metric-Reliable First 6)

Ranked for metric reporting completeness and comparability.

Metric Protocol Matrix (Top 10)

Compare metric, benchmark, and evaluation context side by side.

Paper Metrics Benchmarks Eval Modes Quality Controls
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Aug 26, 2026

Recall, Cost ALFWorld Simulation Env Calibration
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

Aug 21, 2026

Accuracy, F1 TruthfulQA, FEVER Automatic Metrics Calibration
mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

Jun 28, 2026

Relevance Mamabench, Mamaretrieval Automatic Metrics Not reported
PACE: A Proxy for Agentic Capability Evaluation

Jul 2, 2026

Accuracy, Spearman GAIA, SWE Bench Automatic Metrics Not reported
AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

Jul 2, 2026

Precision, Relevance ScienceQA, Airiskeval Automatic Metrics Not reported
Beyond Polarization: The Generative Constraint of Chain-of-Thought in Pointwise Reranking

Aug 31, 2026

Accuracy, Relevance Not reported Automatic Metrics Calibration
Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance

Aug 21, 2026

Accuracy, F1 Financebench Automatic Metrics Not reported
GreekBarRetrieval: A Benchmark for Greek Statutory Retrieval

Aug 19, 2026

Recall, Ndcg Greekbarretrieval, Greekbarbench Automatic Metrics Not reported
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory

Aug 12, 2026

Accuracy, Cost Egor1 Bench Automatic Metrics Not reported
CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation

May 6, 2026

F1, Ndcg NQ, HotpotQA Automatic Metrics Not reported
How To Use This Page

Checklist

  • Gap: Human feedback

    Human feedback is present in 12 of 60 papers.

  • Gap: Quality controls

    Quality controls is present in 6 of 60 papers.

  • Moderate: Benchmarks

    Benchmarks is present in 17 of 60 papers.

  • Strong: Metrics

    Metrics is present in 60 of 60 papers.

  • Gap: Known rater population

    Known rater population is present in 5 of 60 papers.

  • Moderate: Known annotation unit

    Known annotation unit is present in 19 of 60 papers.

Strengths

  • Metrics is present in 60 of 60 papers.

Known Gaps

  • Human feedback is present in 12 of 60 papers.
  • Quality controls is present in 6 of 60 papers.
  • Known rater population is present in 5 of 60 papers.

Suggested Next Analyses

  • Review the most recent relevance papers first, then compare benchmark context before reusing the metric.

Recommended Queries

Known Limitations
  • This synthetic persisted page is generated from extraction data because the cached metric payload was missing for relevance.
Coverage Snapshot

Top Metrics

  • Relevance (60)
  • Accuracy (14)
  • Recall (9)
  • Cost (8)

Evaluation Modes

  • Automatic Metrics (50)
  • Simulation Env (1)

Top Benchmarks

  • HotpotQA (3)
  • FEVER (2)
  • Agentsearchbench (1)
  • Airiskeval (1)

Agentic Mix

  • None (55)
  • Long Horizon (5)

Top Papers Reporting This Metric

Related Metrics And Hubs