Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: General

General papers in the current HFEPX explorer (790 papers).

Papers in tag: 790

Running a General study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (11)
  • Llm As Judge (1)
  • Simulation Env (1)

Human Feedback Types

  • Pairwise Preference (8)
  • Rubric Rating (4)
  • Critique Edit (2)

Required Expertise

  • General (20)
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation

Ofir Ben Shoham, Shrutendra Harsola, Vignesh Subrahmaniam, Shravan Mohan, Yakov Gazman, Oded Vainas · Aug 12, 2026 · Citations: 0

Rubric Rating Llm As Judge General
  • Our reward is an LLM-as-a-judge rubric that scores each recommendation across multiple binary dimensions of advice quality, augmented with a safety gate for harm prevention.
  • Since LLM-based evaluation alone cannot confirm whether improvements reflect genuine business value rather than adaptation to the judge, we complement it with a judge-independent audit based on a standard doubly-robust Conditional Average…
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques

Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang · Aug 12, 2026 · Citations: 0

Pairwise PreferenceCritique Edit Automatic Metrics General
  • Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences.
  • On the out-of-domain Music Arena benchmark with 733 preference pairs, it achieves the highest accuracy of 71.35%.
Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study

Simone Mungari · Aug 12, 2026 · Citations: 0

Pairwise Preference General
  • As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest.
  • We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Nimet Beyza Bozdag, Emre Can Acikgoz, Gokhan Tur, Dilek Hakkani-Tür · Aug 12, 2026 · Citations: 0

Automatic Metrics General
  • As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.
  • We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction.
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs

Vu Duc Anh, Nhat M. Hoang, Do Xuan Long, Cong-Duy Nguyen, Ponhvoan Srey, Luu Anh Tuan · Aug 12, 2026 · Citations: 0

Pairwise Preference General
  • The first stage strengthens step-level reasoning via step-level preference optimization, while the second stage explicitly trains models to self-verify and self-correct.
  • Comprehensive in-domain and out-of-domain evaluations across multiple LLMs demonstrate that SFS-DPO and SFS-DPO-R consistently outperform prior step-level training baselines.
Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment

Haokai Zhao, Yunze Xiao, Weihao Xuan, Flora Salim, Benjamin Tag, Aditya Joshi · Aug 12, 2026 · Citations: 0

Pairwise Preference General
  • Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences.
  • However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour.
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models

Si'an Xie, Jiaxun Liu, Biao Yang, Wei Yuan, Fan Yang, Tingting Gao · Aug 11, 2026 · Citations: 0

Automatic Metrics General
  • We introduce MPAR-Bench, a bilingual English-Chinese benchmark that isolates reasoning breadth through multi-point associative reasoning.
  • We construct 1,000 items using a multi-agent clue-generation pipeline, embedding-based diversity filtering, and human verification.
The Evaluator Is Part of the Experiment: Measuring Open-Ended LLM Conformity

Alicia Guerra, Yibo Hu · Aug 5, 2026 · Citations: 0

Critique Edit General
  • Open-ended revisions require a different measurement strategy because answer quality is graded, latent, and judged imperfectly.
  • We introduce an experimental protocol implemented across a pooled main peer-condition corpus and separately constructed decomposition corpora, allowing us to separate ordinary re-answering, candidate-content exposure, a bundled…
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li · Jul 2, 2026 · Citations: 0

Simulation Env General
  • Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress.
  • We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget.
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach

Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira · Jul 2, 2026 · Citations: 0

Rubric Rating Automatic Metrics General
  • Gemini~3.0 Pro with rubric-guided prompting achieved the highest human-AI agreement (ICC(3,1) = 0.888, MAE = 0.10, Bland-Altman bias = -0.014).
  • These results show that question complexity is a reliable predictor of the difficulty LLMs face in grading accurately, and they establish a principled, taxonomy-based framework for determining which questions are suitable for AI-assisted…
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao · Jul 2, 2026 · Citations: 0

Automatic Metrics General
  • Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
  • A public online benchmark of frontier LLMs on the same game reports zero wins at the lowest difficulty across five configurations, and the developer-reported human win rate at the same difficulty is 16%; the task is hard but not saturated.
Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism

Minjong Cheon · Jul 2, 2026 · Citations: 0

Pairwise Preference General
  • Crucially, this robustness does not transfer: it attenuates across domains and, in the safety-critical vaccine domain, can reverse, with myth-rebuttal weakening under skeptical pressure.
  • We synthesize these into a four-way taxonomy separating active from accidental robustness, and argue that behavioral evaluation alone cannot distinguish a model that resists skepticism because it understands the signal from one that only…
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He, Sulong Xu, Simiu Gu · Jul 2, 2026 · Citations: 0

Rubric Rating Automatic Metrics General
  • We introduce SkillCoach, a self-evolving rubric framework for evaluating and enhancing agentic skill-use.
  • Experiments show that evolved rubrics substantially improve evaluation quality, expose failures hidden by final accuracy, and provide stronger supervision signals than outcome-only filtering for enhancing agentic skill-use.
On the Limits of Steering Vectors for Preference-Aligned Generation

Melanie Subbiah, Zara Hall, Kathleen McKeown · Jul 2, 2026 · Citations: 0

Pairwise Preference Automatic Metrics General
  • Using the PLUME writing personalization benchmark, we extract steering vectors for a range of preferences and evaluate them on summarization and email-writing tasks across two open-source models (Qwen2.5-7B-Instruct and…
  • Taken together, our results suggest that steering vectors face meaningful limits as a general-purpose tool for preference alignment.