Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Rubric Rating

Rubric Rating papers with explicit human-feedback protocol signal (126 papers).

Papers in tag: 126

Running a Rubric Rating study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (11)
  • Llm As Judge (4)
  • Simulation Env (1)

Human Feedback Types

  • Rubric Rating (20)
  • Pairwise Preference (2)
  • Critique Edit (1)

Required Expertise

  • General (9)
  • Coding (6)
  • Medicine (6)
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench

Praveen Reddy, Charuta Mandke, Suvrankar Datta, Sarah Khan, Siddharth Reddy Anthireddy, Shitij Arora · Aug 12, 2026 · Citations: 0

Rubric Rating Llm As JudgeAutomatic Metrics Medicine
  • On 4,023 English-language HealthBench questions (80.5% of the benchmark), scored with a GPT-4.1 judge, VITA ranked first with 51.9% of possible rubric points, ahead of GPT-5.4 (46.1%), o4-mini (44.3%), Gemini 3.1 Pro (42.6%), and Claude…
  • VITA's advantages in accuracy and completeness persisted under the neutral judge; its communication scores were lower.
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam, Shravan Mohan, Yakov Gazman, Oded Vainas · Aug 12, 2026 · Citations: 0

Rubric Rating Llm As Judge General
  • Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention.
  • Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average…
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust · Aug 12, 2026 · Citations: 0

Rubric Rating Llm As Judge Coding
  • We introduce FrontierFinance, a fully open benchmark of 220 expert-crafted queries and 11,543 source-attributed rubrics spanning six crucial use cases across the full investor workflow.
  • Evaluating frontier models and agent systems under a common harness restricted to publicly available data, we find that the tool harness, not the model alone, strongly shapes quality and efficiency; that Samaya's in-house system leads at…
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL

Minglai Yang, Xinyu Guo, Utkarsh Tyagi, Mian Zhang, Razvan Dumitru, Sunjie Hou · Aug 12, 2026 · Citations: 0

Rubric Rating Medicine
  • Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models on tasks with no deterministic answer.
  • Comparing no dropout against dropout at 30% and 50% on both benchmark pairs, dropout raises the OOD gold score at every matched checkpoint (+1 to +2 points on HealthBench-Hard, +6 to +7 points on ResearchQA), lowers the two hacking measures…
ZenGen: Social Mind for LLMs

ZenGen Team, Ao Xiang, Bi Jingping, Chen Jiahui, Chen Lehan, Chen Yilin · Jul 26, 2026 · Citations: 0

Rubric Rating Automatic Metrics Medicine
  • For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms.
  • Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band.
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira · Jul 2, 2026 · Citations: 0

Rubric Rating Automatic Metrics General
  • Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014).
  • These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted…
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He, Sulong Xu, Simiu Gu · Jul 2, 2026 · Citations: 0

Rubric Rating Automatic Metrics General
  • We introduce SkillCoach, a self-evolving rubric framework for evaluating and enhancing agentic skill-use.
  • Experiments show that evolved rubrics substantially improve evaluation quality, expose failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering for enhancing agentic skill-use.
FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoning

Zhiyun Zhang, Liwen Sun, Xiang Qian, Chenyan Xiong · Jul 1, 2026 · Citations: 0

Rubric RatingExpert Verification Automatic Metrics MedicineCoding
  • Across seven medical benchmarks, FaithMed improves over agentic-search baselines (+9% on average) and outcome-only RL (+5.8%), while raising average evidence-based medicine rubric scores over agentic-search Qwen3 baselines (+15.5%).
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives

Aayush Aluru, Chloe Ho, Muhammad Hammouri, Kerry Luo, Myra Malik, Ryan Lagasse · Jul 1, 2026 · Citations: 0

Pairwise PreferenceRubric Rating General
  • MAGNET, a multi-agent goal-driven narrative engine for storytelling, generates stories with persona-grounded character agents that propose actions based on a shared world state and evolving story goals, while ATLAS is a graph-based pipeline…
  • At 100 pages, MAGNET reduced annotations and hallucinations by 41 and 50%, respectively, compared to the single model baseline and by 34 and 45%, respectively, compared to IBSEN, with pairwise rubric evaluation showing similar results.
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho · Jun 30, 2026 · Citations: 0

Rubric RatingCritique Edit Automatic Metrics General
  • UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield…
Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

Yangda Peng, Yunjia Qi, Hao Peng, Haotian Xia, Guanzhong He, Xintong Shi · Jun 29, 2026 · Citations: 0

Rubric Rating Llm As JudgeAutomatic Metrics Coding
  • Rubric-based scoring has become a widely used paradigm in model evaluation, typically with LLM-as-a-Judge (LaaJ) for rubric scoring.
  • We introduce RuVerBench, the first benchmark for assessing LaaJ reliability in rubric verification for agentic scenarios.
mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

Yi Ren · Jun 28, 2026 · Citations: 0

Rubric Rating Automatic Metrics Medicine
  • Medical question-answering benchmarks rarely cover the maternal, neonatal, child, and reproductive-health questions a nurse-midwife asks, and, to our knowledge, no public chunk-level relevance benchmark exists for maternal-health guideline…
  • We release two benchmarks that fill these gaps.
Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B

Zhaoyang Li, Ruijie Zhang, Jiaqi Liu, Zhaoji Sun · Jun 27, 2026 · Citations: 0

Rubric Rating General
  • Agricultural applications, however, are domain-specific, region-dependent, time-sensitive, and safety-critical.
  • Without data governance, expert evaluation, and evidence constraints, an agricultural assistant mayproduce unreliable advice on crop diseases, pesticide use, fertilization, or policy interpretation.To avoid presenting unverified simulated…
The Verification Horizon: No Silver Bullet for Coding Agent Rewards

Binghai Wang, Chenlong Zhang, Dayiheng Liu, Jiajun Zhang, Jiawei Chen, Mingze Li · Jun 24, 2026 · Citations: 0

Rubric Rating Automatic Metrics Coding
  • For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabilities and engineering harnesses grow more sophisticated, generating complex candidate solutions is no longer difficult --…
  • Every verifier we can build is only a proxy for human intent, never the intent itself.
Qwen-AgentWorld: Language World Models for General Agents

Yuxin Zuo, Zikai Xiao, Li Sheng, Fei Huang, Jianhong Tu, Yuxuan Liu · Jun 23, 2026 · Citations: 0

Rubric Rating Simulation Env Coding
  • We introduce Qwen-AgentWorld-35B-A3B and Qwen-AgentWorld-397B-A17B, the first language world models capable of simulating agentic environments covering 7 domains via long chain-of-thought reasoning.
  • Leveraging more than 10M environment interaction trajectories of 7 domains in real-world environments, we develop Qwen-AgentWorld through a three-stage training pipeline: CPT injects general-purpose world modeling capabilities from the…