Human Feedback Types
partialPairwise Preference
Directly usable for protocol triage.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
HFEPX · Eval paper review
Dilara Torunoğlu-Selamet, Dogukan Arslan, Rodrigo Wilkens, Wei He +74 more
Published
Jan 13, 2026
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Feb 24, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community. As such, they constitute an interesting challenge for assessing the linguistic (and to some extent cultural) capabilities of NLP systems. In this paper, we present XMPIE, a parallel multilingual and multimodal dataset of potentially idiomatic expressions. The dataset, containing 34 languages and over ten thousand items, allows comparative analyses of idiomatic patterns among language-specific realisations and preferences in order to gather insights about shared cultural aspects. This parallel dataset allows to evaluate model performance for a given PIE in different languages and whether idiomatic understanding in one language can be transferred to another. Moreover, the dataset supports the study of PIEs across textual and visual modalities, to measure to what extent PIE understanding in one modality transfers or implies in understanding in another modality (text vs. image). The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors. The result is a high-quality benchmark for evaluating multilingual and multimodal idiomatic language understanding.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference
Directly usable for protocol triage.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
None explicit
Validate eval design from full paper text.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
Not reported
No explicit QC controls found.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
Not extracted
No benchmark anchors detected.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
Not extracted
No metric anchors detected.
"Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community."
Domain Experts
Helpful for staffing comparability.
"The data was created by language experts, with both textual and visual components crafted under multilingual guidelines, and each PIE is accompanied by five images representing a spectrum from idiomatic to literal meanings, including semantically related and random distractors."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Potentially idiomatic expressions (PIEs) construe meanings inherently tied to the everyday experience of a given language community.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.