Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubric

Ruipeng Jia, Yunyi Yang, Yuxin Wu, Yongbo Gai, Siyuan Tao, Mengyu Zhou, Jianhe Lin, Xiaoxi Jiang, Guanjun Jiang · Feb 15, 2026 · Citations: 0

General Llm As Judge Pairwise Preference Rubric Rating

Open arXiv RSS feed

Abstract

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking in open-ended alignment. We argue that robust alignment for non-verifiable tasks is fundamentally a principle generalization problem: reward should not be a learned function internalized into a judge, but an explicit reasoning process executed under inspectable principles. To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR) and lightweight Pointwise Verifiable Rubrics (PVRs), which provide both hard-constraint guardrails and verifiable reward components when ground-truth or programmatic checks are available. OpenRS uses an explicit meta-rubric -- a constitution-like specification that governs how rubrics are instantiated, weighted, and enforced -- and instantiates adaptive rubrics on the fly by conditioning on the semantic differences between two candidate responses. It then performs criterion-wise pairwise comparisons and aggregates criterion-level preferences externally, avoiding pointwise weighted scalarization while improving discriminability in open-ended settings. To keep principles consistent yet editable across various domains, we introduce a two-level meta-rubric refinement pipeline (automated evolutionary refinement for general principles and a reproducible human-in-the-loop procedure for domain principles), complemented with pointwise verifiable rubrics that act as both guardrails against degenerate behaviors and a source of verifiable reward for objective sub-tasks. Finally, we instantiate OpenRS as reward supervision in pairwise RL training.

HFEPX Relevance Assessment

This paper has direct human-feedback and/or evaluation protocol signal and is likely useful for eval pipeline design.

Eval-Fit Score

57/100 • Medium

Useful as a secondary reference; validate protocol details against neighboring papers.

Human Feedback Signal

Detected

Evaluation Signal

Detected

HFEPX Fit

High-confidence candidate

If you are doing eval pipeline work, start here:

Human Eval Hub LLM-as-Judge Hub Pairwise Preference Hub Tool-Use Eval Hub

Human Data Lens

Uses human feedback: Yes
Feedback types: Pairwise Preference, Rubric Rating
Rater population: Unknown
Unit of annotation: Pairwise
Expertise required: General
Extraction source: Persisted extraction

Evaluation Lens

Evaluation modes: Llm As Judge
Agentic eval: None
Quality controls: Not reported
Confidence: 0.65
Flags: None

Protocol And Measurement Signals

Benchmarks / Datasets

No benchmark or dataset names were extracted from the available abstract.

Reported Metrics

No metric terms were extracted from the available abstract.

Research Brief

Deterministic synthesis

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking in open-ended alignment. HFEPX signals include Pairwise Preference, Rubric Rating, Llm As Judge with confidence 0.65. Updated from current HFEPX corpus.

Generated Mar 2, 2026, 9:56 PM · Grounded in abstract + metadata only

Key Takeaways

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking…
To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR)…

Researcher Actions

Compare its human-feedback setup against pairwise and rubric hubs.
Identify benchmark choices from full text before operationalizing conclusions.
Verify metric definitions before comparing against your eval pipeline.

Caveats

Generated from title, abstract, and extracted metadata only; full-paper implementation details are not parsed.
Extraction confidence is probabilistic and should be validated for critical decisions.

Recommended Queries

llm-as-judge calibration pairwise preference data quality inter-rater agreement adjudication

Research Summary

Contribution Summary

Scalar reward models compress multi-dimensional human preferences into a single opaque score, creating an information bottleneck that often leads to brittleness and reward hacking in open-ended alignment.
To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR) and lightweight Pointwise Verifiable Rubrics (PVRs), which…
To keep principles consistent yet editable across various domains, we introduce a two-level meta-rubric refinement pipeline (automated evolutionary refinement for general principles and a reproducible human-in-the-loop procedure for domain…

Why It Matters For Eval

To operationalize this view, we present the Open Rubric System (OpenRS), a plug-and-play, rubrics-based LLM-as-a-Judge framework built around Pairwise Adaptive Meta-Rubrics (PAMR) and lightweight Pointwise Verifiable Rubrics (PVRs), which…
To keep principles consistent yet editable across various domains, we introduce a two-level meta-rubric refinement pipeline (automated evolutionary refinement for general principles and a reproducible human-in-the-loop procedure for domain…

Researcher Checklist

Pass: Human feedback protocol is explicit

Detected: Pairwise Preference, Rubric Rating
Pass: Evaluation mode is explicit

Detected: Llm As Judge
Gap: Quality control reporting appears

No calibration/adjudication/IAA control explicitly detected.
Gap: Benchmark or dataset anchors are present

No benchmark/dataset anchor extracted from abstract.
Gap: Metric reporting is present

No metric terms extracted.

Related Papers

Papers are ranked by protocol overlap, extraction signal alignment, and semantic proximity.

HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialogue Protocol Overlap

Citations: 0 Relevance: 12.00 Shared tag: Pairwise PreferenceShared tag: Rubric RatingShared tag: Llm As Judge
- Shared 3 HFEPX protocol tags
- Aligned human feedback protocol
- Aligned evaluation mode
LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering Protocol Overlap

Citations: 0 Relevance: 8.20 Shared tag: Pairwise PreferenceShared tag: Rubric Rating
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Protocol Overlap

Citations: 0 Relevance: 8.20 Shared tag: Pairwise PreferenceShared tag: Rubric Rating
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languages Protocol Overlap

Citations: 0 Relevance: 8.20 Shared tag: Pairwise PreferenceShared tag: Rubric Rating
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
Multi-Agent Comedy Club: Investigating Community Discussion Effects on LLM Humor Generation Protocol Overlap

Citations: 0 Relevance: 8.20 Shared tag: Pairwise PreferenceShared tag: Rubric Rating
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
The Subjectivity of Respect in Police Traffic Stops: Modeling Community Perspectives in Body-Worn Camera Footage Protocol Overlap

Citations: 0 Relevance: 8.20 Shared tag: Pairwise PreferenceShared tag: Rubric Rating
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions Protocol Overlap

Citations: 0 Relevance: 7.90 Shared tag: Rubric RatingShared tag: Llm As Judge
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
- Aligned evaluation mode
Small Reward Models via Backward Inference Protocol Overlap

Citations: 0 Relevance: 7.90 Shared tag: Rubric RatingShared tag: Llm As Judge
- Shared 2 HFEPX protocol tags
- Aligned human feedback protocol
- Aligned evaluation mode
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment Protocol Overlap

Citations: 0 Relevance: 4.10 Shared tag: Pairwise Preference
- Shared HFEPX protocol tags
- Aligned human feedback protocol
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment Protocol Overlap

Citations: 0 Relevance: 4.10 Shared tag: Pairwise Preference
- Shared HFEPX protocol tags
- Aligned human feedback protocol
Balancing Multiple Objectives in Urban Traffic Control with Reinforcement Learning from AI Feedback Protocol Overlap

Citations: 0 Relevance: 4.10 Shared tag: Pairwise Preference
- Shared HFEPX protocol tags
- Aligned human feedback protocol
Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languages Protocol Overlap

Citations: 0 Relevance: 4.10 Shared tag: Pairwise Preference
- Shared HFEPX protocol tags
- Aligned human feedback protocol

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote