Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Critique Edit

Critique Edit papers with explicit human-feedback protocol signal (93 papers).

Papers in tag: 93

Running a Critique Edit study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (6)
  • Llm As Judge (1)
  • Simulation Env (1)

Human Feedback Types

  • Critique Edit (20)
  • Pairwise Preference (3)
  • Expert Verification (1)

Required Expertise

  • General (11)
  • Medicine (4)
  • Coding (2)
Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic · Aug 26, 2026 · Citations: 0

Critique Edit Llm As JudgeAutomatic Metrics General
  • Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts.
  • On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to…
Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

Sahil Mishra, Niranjan Rajeev, Tanmoy Chakraborty · Aug 21, 2026 · Citations: 0

Critique Edit Automatic Metrics General
  • Here, we introduce Tree-of-Concerns, a multi-agent framework that deploys specialized skeptic personas, each operating through a category-specific analytical lens, as parallel debate trees to extract unstated limitations from scientific…
  • Through experiments on ToC-Bench, our benchmark of 414 research papers with 1,905 unstated limitations, sourced from reviewer-reported weaknesses and follow-up citation critiques, we demonstrate that ToC improves precision by 79% and…
HealMed: Multilingual Evaluation of Large Language Models in Medicine

Yingjian Chen, Fan Gao, Sherry T. Tong, Haoyu Zhang, Aosong Feng, Kevin W. Jin · Aug 20, 2026 · Citations: 0

Critique Edit MedicineMultilingual
  • We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language models in medicine.
  • The benchmark was developed over two years by 23 physicians and medical experts based across nine countries and regions.
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Zhuoyang Qian, Biao Wu, Yiran Wang, Chris D Yan, Desan Dai, Liangwei Zheng · Aug 12, 2026 · Citations: 0

Critique Edit Coding
  • We present Spark-to-Paper, an end-to-end research paper generation system implemented as thirteen composable skills inside an existing coding assistant, without requiring a separate agent platform or orchestration service.
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang · Aug 12, 2026 · Citations: 0

Pairwise PreferenceCritique Edit Automatic Metrics General
  • Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences.
  • On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%.
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

Alicia Guerra, Yibo Hu · Aug 5, 2026 · Citations: 0

Critique Edit General
  • Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly.
  • We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled…
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

Shayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho · Jun 30, 2026 · Citations: 0

Rubric RatingCritique Edit Automatic Metrics General
  • UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield…
LatentRevise: Learning from Zero-Hit Reasoning

Yiqiu Guo, Xueting Han, Qi Jia, Guangtao Zhai, Jing Bai · Jun 29, 2026 · Citations: 0

Critique Edit Math
  • Used as training data, these trajectories improve SFT and RLVR on math benchmarks over standard baselines.
Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Models

Sajjad Ghiasvand, Maryam Amirizaniani, Haniyeh Ehsani Oskouie, Mahnoosh Alizadeh, Ramtin Pedarsani · Jun 29, 2026 · Citations: 0

Critique Edit General
  • Open-ended aesthetic critique is a challenge for multimodal large language models (MLLMs): unlike multiple-choice aesthetic benchmarks, it has no single correct answer, and most aesthetic evaluation has measured models against numeric…
  • We evaluate MLLM critiques against ranked human references and ask whether they are close to human ones.
LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

Thi Huyen Nguyen, Zahra Ahmadi · Jun 23, 2026 · Citations: 0

Critique Edit General
  • The rapid growth of scientific submissions has pushed traditional peer review toward its scalability limits, motivating the exploration of large language models (LLMs) as intelligent automated evaluation assistants.
  • We present a structured taxonomy of modeling approaches (including prompt-based, supervised, retrieval-augmented, and alignment-optimized approaches), and synthesize empirical findings across existing benchmarks.
Do Thinking Tokens Help with Safety?

Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora · Jun 23, 2026 · Citations: 0

Critique Edit Automatic Metrics General
  • Today's reasoning models use thinking tokens to attain stronger performance on benchmarks than their instruction-tuned counterparts.
  • It is also generally believed that this more "deliberative" mode should improve alignment and safety, by providing the model a safe space to consider whether its planned answer to a request violates its safety principles.
Posterior Refinement: Fast Language Generation via Any-Order Flow Maps

Manan Agarwal, Sheel Shah, Chanhyuk Lee, Jaehoon Yoo, Jerry Huang, Seunghoon Hong · Jun 23, 2026 · Citations: 0

Critique Edit General
  • Across diverse benchmarks, we demonstrate that FMLM+ with Posterior Refinement improves the speed--quality tradeoff over both MDM and FMLM families, providing a scalable foundation for high-fidelity language modeling.
Self-Preference Is Weak or Absent in Verifiable Instruction-Following Revision: A Four-Model Test Under Genuine Authorship

William Guey, Pierrick Bougault · Jun 18, 2026 · Citations: 0

Pairwise PreferenceCritique Edit Law
  • Across four mid-tier model families and 85 author-versus-fresh comparisons, we find no detectable self-preference: authors reject verified-good fixes to their own drafts at essentially the same rate as fresh models judging the same drafts…
  • The one robust observation is qualitative: when authors do reject a verified-good fix, 97% of their stated reasons are flaw-catching rather than preference, that is, about the character of rejections, not an elevated rate.
eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini · Jun 10, 2026 · Citations: 0

Expert VerificationCritique Edit Medicine
  • Finally, we propose CRF-filling as a novel structured information extraction benchmark, and provide zero-shot baseline resulting from Gemma-27B and MedGemma-27B.
CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generation

Wenhan Xiao, Ziwei Zhang, Chuanyue Yu, Xingcheng Fu, Qingyun Sun, Runhua Xu · May 28, 2026 · Citations: 0

Critique Edit MedicineCoding
  • To learn these capabilities, we design two reward functions: Conservative Judgement Alignment (CJA) first encourages calibrated high-level judgements while mitigating the over-aggressive phenomenon, whereas Diagnostic Quality Alignment…
  • Experiments across five QA benchmarks show that CRITIC-R1 consistently improves answer quality over strong RAG baselines.