Human Feedback Types
strongExpert Verification
Directly usable for protocol triage.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
HFEPX · Eval paper review
Ikram Belmadani, Oumaima El Khettari, Pacôme Constant dit Beaufils, Benoit Favre +1 more
Published
Mar 6, 2026
Citations
0
Trust level
High
Usefulness score
67/100 (Medium)
Extraction confidence
75% (High)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Mar 6, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this as a practical starting point for protocol research, then validate against the original paper.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Secondary protocol comparison source
Use if you need
A concrete protocol example with enough signal to inform rater workflow design.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The abstract does not clearly name benchmarks or metrics.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts. Yet, in specialized fields such as medicine, the scarcity of high-quality French instruction data limits effective supervision. To address this gap, we introduce MedInjection-FR, a large-scale French biomedical instruction dataset comprising 571K instruction-response pairs drawn from three complementary sources: native, synthetic, and translated data. We design a controlled experimental framework to systematically assess how data provenance affects instruction tuning, using Qwen-4B-Instruct fine-tuned across seven configurations combining these sources. Results show that native data yield the strongest performance, while mixed setups, particularly native and translated, provide complementary benefits. Synthetic data alone remains less effective but contributes positively when balanced with native supervision. Evaluation on open-ended QA combines automatic metrics, LLM-as-a-judge assessment, and human expert review; although LLM-based judgments correlate best with human ratings, they show sensitivity to verbosity. These findings highlight that data authenticity and diversity jointly shape downstream adaptation and that heterogeneous supervision can mitigate the scarcity of native French medical instructions.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Expert Verification
Directly usable for protocol triage.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
Llm As Judge
Includes extracted eval setup.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
Adjudication
Calibration/adjudication style controls detected.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
Not extracted
No benchmark anchors detected.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
Not extracted
No metric anchors detected.
"Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts."
Domain Experts
Helpful for staffing comparability.
"Evaluation on open-ended QA combines automatic metrics, LLM-as-a-judge assessment, and human expert review; although LLM-based judgments correlate best with human ratings, they show sensitivity to verbosity."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Instruction tuning has become essential for adapting large language models (LLMs) to follow domain-specific prompts.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Expert Verification
Evaluation mode is explicit
Detected: Llm As Judge
Quality control reporting appears
Detected: Adjudication
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.