Human Feedback Types
strongExpert Verification
Directly usable for protocol triage.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
HFEPX · Eval paper review
Yejin Bang, Kirsty Fielding, Brandan Oliver, Brian Birke +2 more
Published
Aug 20, 2026
Citations
0
Trust level
Moderate
Usefulness score
65/100 (Medium)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Aug 20, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs. Contract ``scrubbing,'' the final review of transactional agreements for errors and inconsistencies, is a particularly suitable task for automation, because it is routine, painstaking work requiring detailed attention to long documents. Scrubbing also seems to align naturally with the general capabilities expected of frontier LLMs around long-context reasoning, consistency checking, and named entity recognition (NER). Despite the economic value and potential for automation, no formal evaluations of LLMs performing contract scrubbing have been conducted. We introduce ContractScrub, the first benchmark designed to evaluate contract scrubbing capabilities, comprising contracts hand-crafted by experienced lawyers over diverse error categories such as misuse of defined terms, incorrect references, and inconsistent language. Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Expert Verification
Directly usable for protocol triage.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
Automatic Metrics
Includes extracted eval setup.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
Not reported
No explicit QC controls found.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
Not extracted
No benchmark anchors detected.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
Recall
Useful for evaluation criteria comparison.
"Frontier models perform surprisingly poorly with only one model reaching 0.75 macro average recall despite strong performance on seemingly related general benchmarks, demonstrating the practical limits of current models and the importance of narrowly targeted, domain-specific benchmarks for measuring real-world impact."
Domain Experts
Helpful for staffing comparability.
"Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs."
No benchmark or dataset names were extracted from the available abstract.
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Expert Verification
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: recall