Human Feedback Types
provisional (inferred)None explicit
No explicit feedback protocol extracted.
"Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan."
HFEPX · Eval paper review
Geonhui Jang, Dongyoon Han, YoungJoon Yoo
Published
Apr 16, 2026
Citations
0
Trust level
Provisional
Usefulness score
Unavailable
Extraction confidence
0% (Provisional)
Derived from abstract and metadata only.
Signals refreshed
Apr 16, 2026
Signal extraction is still processing. This page currently shows metadata-first guidance until structured protocol fields are ready.
This page is a lightweight research summary built from the abstract and metadata while deeper extraction catches up.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A provisional background reference while structured extraction finishes.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
This page is still relying on abstract and metadata signals, not a fuller protocol read.
Eval-fit score is unavailable until extraction completes.
If you are doing eval pipeline work, start here
Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan. Existing approaches augment reasoning steps or inject specific structure into how models think, but leave scattered problem conditions unchanged. Inspired by the way humans organize fragmented information into coherent explanations, we propose StoryCoder, a narrative reformulation framework that transforms code generation questions into coherent natural language narratives, providing richer contextual structure than simple rephrasings. Each narrative consists of three components: a task overview, constraints, and example test cases, guided by the selected algorithm and genre. Experiments across 11 models on HumanEval, LiveCodeBench, and CodeForces demonstrate consistent improvements, with an average gain of 18.7% in zero-shot pass@10. Beyond accuracy, our analyses reveal that narrative reformulation guides models toward correct algorithmic strategies, reduces implementation errors, and induces a more modular code structure. The analyses further show that these benefits depend on narrative coherence and genre alignment, suggesting that structured problem representation is important for code generation regardless of model scale or architecture. Our code is available at https://github.com/gu-ni/StoryCoder.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan."
Automatic metrics
Includes extracted eval setup.
"Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan."
Not reported
No explicit QC controls found.
"Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan."
LiveCodeBench
Useful for quick benchmark comparison.
"Experiments across 11 models on HumanEval, LiveCodeBench, and CodeForces demonstrate consistent improvements, with an average gain of 18.7% in zero-shot pass@10."
Accuracy
Useful for evaluation criteria comparison.
"Beyond accuracy, our analyses reveal that narrative reformulation guides models toward correct algorithmic strategies, reduces implementation errors, and induces a more modular code structure."
Unknown
Rater source not explicitly reported.
"Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan."
This page is using abstract-level cues only right now. Treat the signals below as provisional.
Evaluation fields are inferred from the abstract only.
Effective code generation requires both model capability and a problem representation that carefully structures how models reason and plan.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.