Human Feedback Types
partialCritique Edit
Directly usable for protocol triage.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
HFEPX · Eval paper review
Leonor Veloso, Lea Hirlimann, Lucija Mihić Zidar, Philipp Wicke +2 more
Published
Sep 22, 2025
Citations
0
Trust level
Low
Usefulness score
40/100 (Low)
Extraction confidence
45% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Aug 13, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Background context only
Use if you need
Background context only.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The available metadata is too thin to trust this as a primary source.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process queer language incorrectly, e.g., they misclassify it as hate speech or generate negative responses. To address this problem, we propose Slaying, the first real-world dataset of English queer slang. Slaying is community-validated, and includes over 500 queer slang terms that pertain to more than 20 queer subcommunities. We argue that queer language data resources have great potential in NLP -- e.g., as components of large pretraining corpora and as the basis for benchmarks -- and can improve queer users' experience of NLP systems. We leverage Slaying for two novel findings in support of this argument: (i) For a number of language models, we show that they are unbiased towards the queer community, but at the same time unable to process its language, i.e., absence of representation bias does not entail the absence of linguistic bias. (ii) Model performance on queer slang varies across queer subcommunities; it is generally worse for slang pertaining to African-American and Latine communities. These findings are relevant for both the queer NLP and the broader ML communities. Slaying is available to the public, and open to future revisions and extensions. Warning: This paper contains profane and potentially offensive language.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Critique Edit
Directly usable for protocol triage.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
None explicit
Validate eval design from full paper text.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
Not reported
No explicit QC controls found.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
Not extracted
No benchmark anchors detected.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
Not extracted
No metric anchors detected.
"Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Critique Edit
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.