Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Medicine

Medicine evaluation papers that call for domain expertise or specialist review (167 papers).

Papers in tag: 167

Need Medicine evaluators for your project?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (13)
  • Llm As Judge (4)
  • Simulation Env (2)

Human Feedback Types

  • Rubric Rating (6)
  • Expert Verification (5)
  • Pairwise Preference (4)

Required Expertise

  • Medicine (20)
  • Coding (4)
  • Multilingual (1)
Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Rishiraj Sengupta, Sotiris Chatzimiltis, Mohammad Shojafar, Xiatian Zhu · Aug 21, 2026 · Citations: 0

Pairwise PreferenceExpert Verification Llm As JudgeAutomatic Metrics Medicine
  • While existing benchmarks rely on restrictive MCQs with fixed answer keys, this paper evaluates 5G domain understanding and fault analysis in a free-text generation format.
  • To address this we evaluate three lightweight LLMs, Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, on free-text 5G domain knowledge and fault-analysis tasks across three benchmarks, TeleQNA ORAN FT, 5G-Faults FT, and TeleInter…
HealMed: Multilingual Evaluation of Large Language Models in Medicine

Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin · Aug 20, 2026 · Citations: 0

Critique Edit MedicineMultilingual
  • We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine.
  • The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions.
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination

Saisha Shetty, Satvik Tripathi, Austin Lin, Colin Zhao, Theodore Kim, Don Enwerem · Aug 13, 2026 · Citations: 0

Expert Verification MedicineCoding
  • We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministic multi-agent orchestration for clinical reasoning.
  • MARC coordinates role-specialized agents for extraction, reasoning, answer generation, and evaluation, with explicit context passing and traceable intermediate outputs, enabling stage-wise failure attribution.
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives

Chengyang He, Tahreem Arif, Marko Zivkovic, Lijing Wang, Yue Ning, Ping Wang · Aug 13, 2026 · Citations: 0

Pairwise Preference Automatic Metrics Medicine
  • Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and causality assessment.
  • We conduct evaluation on MedTempo, a new benchmark of 5,347 vaccine adverse-event narratives spanning three COVID-19 vaccine types, with expert-validated temporal stage annotations for 3,166 reports.
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora · Aug 12, 2026 · Citations: 0

Rubric Rating Llm As JudgeAutomatic Metrics Medicine
  • On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude…
  • VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower.
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

Lior Baruch, Moshe Butman, Kfir Bar, Doron Friedman · Aug 12, 2026 · Citations: 0

Pairwise Preference Simulation Env Medicine
  • This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead.
  • Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets.
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou · Aug 12, 2026 · Citations: 0

Rubric Rating Medicine
  • Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.
  • Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures…
Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology

Del Coburn, Scott Sanner, Dan Silver · Aug 11, 2026 · Citations: 0

Automatic Metrics Medicine
  • We introduce Social Chain of Thought (SCoT),a multi-round pipeline for medical differential diagnosis that structures multi-agent interaction as a deliberative framework for collabora.
  • Evaluating SCoT against single-agent baselines, one-agent pipeline ablations, and best-of-n scaling, we show that its recall advantage is not reproduced by monolithic inference alone.
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui · Aug 3, 2026 · Citations: 0

Llm As JudgeAutomatic Metrics Medicine
  • However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question.
  • To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies…
ZenGen: Social Mind for LLMs

ZenGen Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin · Jul 26, 2026 · Citations: 0

Rubric Rating Automatic Metrics Medicine
  • For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms.
  • Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band.
Token Reduction Is Not Cost Reduction

Sarel Weinberger, Amir Hozez · Jul 13, 2026 · Citations: 0

Automatic Metrics MedicineCoding
  • Token-reduction tools for coding agents are often evaluated by the number of tokens they remove, but token count alone does not determine end-to-end inference cost.
  • We evaluate three token-reduction approaches against an unmodified Claude Code baseline across controlled coding tasks, measuring provider-billed cost, task success, cache traffic, and agent behavior.
FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

Zhiyun Zhang, Liwen Sun, Xiang Qian, Chenyan Xiong · Jul 1, 2026 · Citations: 0

Rubric RatingExpert Verification Automatic Metrics MedicineCoding
  • Across seven medical benchmarks, FaithMed improves over agentic-search baselines (+9% on average) and outcome-only RL (+5.8%), while raising average evidence-based medicine rubric scores over agentic-search Qwen3 baselines (+15.5%).
Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

William Philipp, Finn Fassbender, Thorsten Langer, Martje Pauly, Rebecca Herzog, Alexander Baumann · Jul 1, 2026 · Citations: 0

Expert Verification Automatic Metrics Medicine
  • Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches.
  • We introduce MedQADE, the first standardised open-response clinical benchmark for German, a major clinical language lacking native evaluation infrastructure, comprising 3,800 items annotated by ten practising physicians and nine Large…
How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification

Ali H. Lazem, William J. Teahan · Jun 28, 2026 · Citations: 0

Expert Verification Automatic Metrics Medicine
  • Analysing the complete output of a multi-agent clinical extraction pipeline applied to 167,034 patient narratives, 2.51 billion generated tokens across the ten text-bearing channels of an eleven-channel pipeline, we introduce…
  • In a controlled downstream test, de-duplicating the corpus before adaptation improved a clinical encoder on external disease-recognition benchmarks at equal token budget, robustly across adaptation depths and replicated on a second…
mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

Yi Ren · Jun 28, 2026 · Citations: 0

Rubric Rating Automatic Metrics Medicine
  • Medical question-answering benchmarks rarely cover the maternal, neonatal, child, and reproductive-health questions a nurse-midwife asks, and, to our knowledge, no public chunk-level relevance benchmark exists for maternal-health guideline…
  • We release two benchmarks that fill these gaps.
A French OSCE Dialogue Dataset and Controllable Virtual Patient System for Clinical Training

Doria Bonzi, Tom Bourgeade, Fabrice Lefèvre, Irina Illina · Jun 26, 2026 · Citations: 0

Llm As JudgeSimulation Env Medicine
  • However, training is often limited by the low availability of human standardized patients, motivating the development of realistic virtual patients (VPs).
  • Additionally, we propose a multi-level evaluation framework assessing patient simulation quality, student performance, and linguistic quality, using an LLM-as-a-Judge approach.