Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Expert Verification

Expert Verification papers with explicit human-feedback protocol signal (130 papers).

Papers in tag: 130

Running a Expert Verification study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (14)
  • Llm As Judge (1)

Human Feedback Types

  • Expert Verification (20)
  • Critique Edit (1)
  • Pairwise Preference (1)

Required Expertise

  • Medicine (13)
  • Coding (5)
  • Law (3)
Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy

Afonso Baldo, Hugo Pitorro, Areti Vassilopoulos, Anabela C. Areias, Maya D'Eon, Fabíola Costa · Aug 21, 2026 · Citations: 0

Expert Verification Automatic Metrics General
  • We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that…
  • Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models.
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu · Aug 21, 2026 · Citations: 0

Pairwise PreferenceExpert Verification Llm As JudgeAutomatic Metrics Medicine
  • While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format.
  • To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter…
ContractScrub: A benchmark for final review of legal contracts

Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke, Nabeel Seedat, Andrew M. Bean · Aug 20, 2026 · Citations: 0

Expert Verification Automatic Metrics Law
  • Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted.
  • We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect…
Mint-Agent: Introducing Finance-Native Agentic Foundation Models

Mint-Agent Team, Kun Wang, Gavin Zhang, Yaze Geng, Lei Tang, Yaoyang Yi · Aug 17, 2026 · Citations: 0

Expert Verification Automatic Metrics General
  • We present Mint-Agent, a family of finance-native agentic models designed around these two scales of financial intelligence.
  • Across professional financial benchmarks, our models demonstrate two defining strengths: (1) Reliability: Mint-Ag achieves 98.33% on RFC-Bench, surpassing GPT-5.6-Sol and Claude-Opus-4.8 by 3.66 and 3.00 points; and (2) Executability:…
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem · Aug 13, 2026 · Citations: 0

Expert Verification MedicineCoding
  • We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning.
  • MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution.
Principal Trait Analysis: Towards Deriving "Skills" in Human-AI Collaboration

Hunter McNichols, Kai Du, Andrew Lan · Aug 11, 2026 · Citations: 0

Expert Verification Automatic Metrics Coding
  • Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration.
  • We evaluate PTA on two human-AI collaborative coding datasets, an educational setting (students working with an AI tutor) and a professional setting (developers working with an AI coding agent).
ADVENT: LLM-Driven Automatic Predicate Invention for ILP

Tingting Yu, Pei-Cing Huang, Chan Hsu, Chan-Tung Ku, Yihuang Kang · Jul 2, 2026 · Citations: 0

Expert Verification Automatic Metrics Coding
  • Experiments on nine poker-hand concepts across seven LLMs show that LLM-driven PI achieves 58% success rate where ILP alone fails entirely, formal verification raises this to 80%, and the knowledge pool yields gains up to +31 percentage…
FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

Zhiyun Zhang, Liwen Sun, Xiang Qian, Chenyan Xiong · Jul 1, 2026 · Citations: 0

Rubric RatingExpert Verification Automatic Metrics MedicineCoding
  • Across seven medical benchmarks, FaithMed improves over agentic-search baselines (+9% on average) and outcome-only RL (+5.8%), while raising average evidence-based medicine rubric scores over agentic-search Qwen3 baselines (+15.5%).
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

William Philipp, Finn Fassbender, Thorsten Langer, Martje Pauly, Rebecca Herzog, Alexander Baumann · Jul 1, 2026 · Citations: 0

Expert Verification Automatic Metrics Medicine
  • Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches.
  • We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large…
How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification

Ali H. Lazem, William J. Teahan · Jun 28, 2026 · Citations: 0

Expert Verification Automatic Metrics Medicine
  • Analysing the complete output of a multi-agent clinical extraction pipeline applied to 167,034 patient narratives, 2.51 billion generated tokens across the ten text-bearing channels of an eleven-channel pipeline, we introduce…
  • In a controlled downstream test, de-duplicating the corpus before adaptation improved a clinical encoder on external disease-recognition benchmarks at equal token budget, robustly across adaptation depths and replicated on a second…
MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization

Aueaphum Aueawatthanaphisut · Jun 18, 2026 · Citations: 0

Expert Verification Medicine
  • The framework coordinates specialized agents for clinical text, longitudinal EHR, medical imaging, physiological sensor signals, guideline retrieval, uncertainty auditing, and referral planning.
  • We also outline a real-data evaluation design using public and credentialed clinical datasets spanning EHR, radiology, ECG, ICU time series, and referral-proxy outcomes.
Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text

Priyanshi Garg, Ishita Rao, Jieqiong Ding, Amandalynne Paullada · Jun 17, 2026 · Citations: 0

Expert Verification Medicine
  • We show how governance constraints, ICD-based cohort selection, single-annotator labeling, and hospital-stay-level aggregation produce labels that reflect clinician-documented judgments, treat suicidality as a bounded episode, and assume…
eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini · Jun 10, 2026 · Citations: 0

Expert VerificationCritique Edit Medicine
  • Finally, we propose CRF-filling as a novel structured information extraction benchmark, and provide zero-shot baseline resulting from Gemma-27B and MedGemma-27B.
MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

Valentina Bui Muti, Eugénie Dulout, Ziquan Fu · May 28, 2026 · Citations: 0

Expert Verification Automatic Metrics Medicine
  • We introduce a reusable pipeline for generating terminology-grounded HL7 FHIR R4 bundles from unstructured text, enabling controllable evaluation of clinical decision support systems over structured inputs.
  • Evaluation on MedCase-Structured reveals consistently lower diagnostic accuracy for LLMs on structured FHIR inputs than with plain text, highlighting the importance of deployment-aligned benchmarking.
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

Jinyang Wu, Guocheng Zhai, Ruihan Jin, Yuhao Shen, Zhengxi Lu, Fan Zhang · May 21, 2026 · Citations: 0

Expert Verification Automatic Metrics MathCoding
  • In this paper, we present Maestro (Multimodal Agent for Expert-Skill Targeted Reinforced Orchestration), a Reinforcement Learning (RL)-driven orchestration framework that reframes heterogeneous multimodal tasks as a sequential…
  • We evaluate Maestro across ten representative multimodal benchmarks spanning mathematical reasoning, chart understanding, high-resolution perception, and domain-specific analysis.
Safety and accuracy follow different scaling laws in clinical large language models

Sebastian Wind, Tri-Thien Nguyen, Jeta Sopa, Mahshad Lotfinia, Sebastian Bickelhaup, Michael Uder · May 5, 2026 · Citations: 0

Expert Verification Automatic Metrics LawMedicine
  • We introduce SaFE-Scale, a framework for measuring how clinical LLM safety changes across model scale, evidence quality, retrieval strategy, context exposure, and inference-time compute.
  • To instantiate this framework, we introduce RadSaFE-200, a Radiology Safety-Focused Evaluation benchmark of 200 multiple-choice questions with clinician-defined clean evidence, conflict evidence, and option-level labels for high-risk error,…