Human Feedback Types
provisional (inferred)None explicit
No explicit feedback protocol extracted.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
HFEPX · Eval paper review
Dongxin Guo, Jikun Wu, Siu Ming Yiu
Published
May 10, 2026
Citations
0
Trust level
Provisional
Usefulness score
Unavailable
Extraction confidence
0% (Provisional)
Derived from abstract and metadata only.
Signals refreshed
May 10, 2026
Signal extraction is still processing. This page currently shows metadata-first guidance until structured protocol fields are ready.
This page is a lightweight research summary built from the abstract and metadata while deeper extraction catches up.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A provisional background reference while structured extraction finishes.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
This page is still relying on abstract and metadata signals, not a fuller protocol read.
Eval-fit score is unavailable until extraction completes.
If you are doing eval pipeline work, start here
Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces. We present QD-LLM, a framework for parameter-efficient neuroevolution that evolves prompt embeddings, compact neural interfaces (~32K parameters) that steer generation in frozen LLMs (70B+ parameters), within a Quality-Diversity (QD) optimization framework. Our contributions: (1) evolved prompt embeddings via gradient-free optimization enabling behavioral steering without model fine-tuning; (2) hybrid behavior characterization combining semantic and explicit features with formal coverage bounds (Theorem 1) under validated near-independence (NMI $= 0.08 \pm 0.02$); (3) co-evolutionary variation operators including targeted behavioral mutation via finite-difference gradient estimation. On HumanEval (164 problems), MBPP, and creative writing benchmarks, QD-LLM achieves 46.4% higher coverage and 41.4% higher QD-Score than QDAIF ($p<0.001$, 30 runs, Vargha-Delaney $A=0.94$). We demonstrate downstream utility: diverse archives improve test generation (34% more edge cases) and fine-tuning data quality (8.3% accuracy gain). We validate across open-source LLMs (Llama-3-70B, Mistral-Large) with full embedding access, establishing prompt embedding evolution as an effective paradigm bridging neuroevolution and modern LLMs.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
Automatic metrics
Includes extracted eval setup.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
Not reported
No explicit QC controls found.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
Not extracted
No benchmark anchors detected.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
Accuracy
Useful for evaluation criteria comparison.
"We demonstrate downstream utility: diverse archives improve test generation (34% more edge cases) and fine-tuning data quality (8.3% accuracy gain)."
Unknown
Rater source not explicitly reported.
"Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces."
This page is using abstract-level cues only right now. Treat the signals below as provisional.
Evaluation fields are inferred from the abstract only.
Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.