Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

HFEPX Archive Slice

HFEPX Daily Papers for 2026-06-18

Daily archive slice for 2026-06-18 from the HFEPX corpus. Updated from current HFEPX corpus (2026-09-18); covers 60 papers from 2026-06-18.

Papers: 60 Last published: Jun 18, 2026 Global RSS

Researcher Quick Triage

Use this archive page for time-slice monitoring (what changed in evaluation methods, metrics, and protocol quality this period). Quality band: High .

High-Signal Coverage

100.0%

60 / 60 papers are not low-signal flagged.

Benchmark Anchors

20.0%

Papers with benchmark/dataset mentions in extraction output.

Metric Anchors

46.7%

Papers with reported metric mentions in extraction output.

  • 4 papers report explicit quality controls for this archive period.
  • Prioritize papers with both benchmark and metric anchors for reliable longitudinal comparisons.

Primary action: Use this slice for trend comparison: review top papers first, then validate shifts in the protocol matrix.

Get this digest every Friday →

Subscribe

Why This Time Slice Matters

  • Use this archive slice to monitor protocol drift and shifts in evaluation methods over 2026-06-18.

Protocol Takeaways For This Period

  • Evaluation modes for this slice cluster around automatic_metrics, simulation_env.

Start Here (Highest-Signal Papers In This Slice)

Ranked by protocol completeness and evidence density for faster period-over-period review.

Protocol Matrix (Top 10)

Quickly compare method ingredients across this archive slice.

Paper Eval Modes Benchmarks Metrics Quality Controls
Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systems

Jun 18, 2026

Automatic Metrics Herabench Cost, Token cost Not reported
Source-Grounded Data Generation for Text-to-JSON Learning

Jun 18, 2026

Automatic Metrics Stage Eval Accuracy, Exact match Not reported
GEMS: Geometric Constraints Enable Multi-Semantic Superposition in LLMs

Jun 18, 2026

Automatic Metrics GSM8K Accuracy, Perplexity Not reported
CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis

Jun 18, 2026

Automatic Metrics Wikisplitbench, Claimdecompbench Accuracy, F1 Not reported
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning

Jun 18, 2026

Automatic Metrics CommonsenseQA Accuracy, Cost Not reported
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA

Jun 18, 2026

Automatic Metrics ChartQA Accuracy Not reported
NEST: Narrative Event Structures in Time for Long Video Understanding

Jun 18, 2026

Automatic Metrics Needle In A Haystack F1 Not reported
Toward Calibrated Mixture-of-Experts Under Distribution Shift

Jun 18, 2026

Automatic Metrics Not reported Accuracy Calibration
Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Users

Jun 18, 2026

Automatic Metrics Not reported Accuracy Not reported
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse

Jun 18, 2026

Automatic Metrics Not reported Accuracy Calibration
Researcher Workflow (Detailed)

Checklist

  • Gap: Human feedback

    Human feedback is present in 11 of 60 papers.

  • Gap: Quality controls

    Quality controls is present in 4 of 60 papers.

  • Gap: Benchmarks

    Benchmarks is present in 12 of 60 papers.

  • Moderate: Metrics

    Metrics is present in 28 of 60 papers.

  • Gap: Known rater population

    Known rater population is present in 3 of 60 papers.

  • Moderate: Known annotation unit

    Known annotation unit is present in 13 of 60 papers.

Known Gaps

  • Human feedback is present in 11 of 60 papers.
  • Quality controls is present in 4 of 60 papers.
  • Benchmarks is present in 12 of 60 papers.

Suggested Next Analyses

  • Compare 2026-06-18 against neighboring archive slices to flag protocol drift.

Recommended Queries

Known Limitations
  • This synthetic archive page is generated on-demand from extraction data because no cached payload was available for 2026-06-18.
Research Utility Snapshot (Detailed)

Evaluation Modes

  • Automatic Metrics (23)
  • Simulation Env (3)
  • Llm As Judge (1)

Top Metrics

  • Accuracy (12)
  • Cost (7)
  • F1 (6)
  • Recall (3)

Top Benchmarks

  • ALFWorld (1)
  • ChartQA (1)
  • Claimdecompbench (1)
  • Combeval (1)

Quality Controls

  • Calibration (4)

Papers In This Archive Slice

Recent Archive Slices