Human Feedback Types
partialPairwise Preference
Directly usable for protocol triage.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
HFEPX · Eval paper review
Oscar Sainz, Naiara Perez, Julen Etxaniz, Joseba Fernandez de Landa +8 more
Published
Jun 9, 2025
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Mar 13, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages. In this paper, we explore alternatives to conventional instruction adaptation pipelines in low-resource scenarios. We assume a realistic scenario for low-resource languages, where only the following are available: corpora in the target language, existing open-weight multilingual base and instructed backbone LLMs, and synthetically generated instructions sampled from the instructed backbone. We present a comprehensive set of experiments for Basque that systematically study different combinations of these components evaluated on benchmarks and human preferences from 1,680 participants. Our conclusions show that target language corpora are essential, with synthetic instructions yielding robust models, and, most importantly, that using as backbone an instruction-tuned model outperforms using a base non-instructed model. Scaling up to Llama 3.1 Instruct 70B as backbone, our model comes near frontier models of much larger sizes for Basque, without using any Basque instructions. We release code, models, instruction datasets, and human preferences to support full reproducibility in future research on low-resource language adaptation. https://github.com/hitz-zentroa/latxa-instruct
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference
Directly usable for protocol triage.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
None explicit
Validate eval design from full paper text.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
Not reported
No explicit QC controls found.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
Not extracted
No benchmark anchors detected.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
Not extracted
No metric anchors detected.
"Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Instructing language models with user intent requires large instruction datasets, which are only available for a limited set of languages.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.