Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks."
HFEPX · Eval paper review
Yutao Hou, Zeguan Xiao, Fei Yu, Yihan Jiang +5 more
Published
Jun 5, 2025
Citations
0
Trust level
Low
Usefulness score
0/100 (Low)
Extraction confidence
25% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Apr 24, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
Background context only.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
This paper looks adjacent to evaluation work, but not like a strong protocol reference.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks. However, these models exhibit unexpected brittleness, often failing on simple variations of the same underlying task. Existing robustness evaluations predominantly rely on hand-crafted templates or a limited set of perturbation rules. Consequently, such approaches lack the adaptability to probe latent vulnerabilities unique to specific models and remain susceptible to data contamination. To address this, we propose the Math Stress Tester (MaSTer), an automated framework inspired by software stress testing. MaSTer generates adversarial variants via a multi-round rewrite-verify loop, ensuring semantic consistency while successfully inducing model failure. Our framework generates benchmark variants dynamically for each LLM, thus minimizing the risk of data contamination. Experiments on GSM8K and MATH-500 demonstrate the effectiveness of MaSTer on mathematical tasks. Additionally, we validate the framework's extensibility to non-mathematical tasks, highlighting its broad applicability. Furthermore, we demonstrate that the synthesized variants generated by MaSTer can be utilized as a fine-tuning dataset to significantly enhance the model's robustness.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks."
None explicit
Validate eval design from full paper text.
"Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks."
Not reported
No explicit QC controls found.
"Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks."
MATH 500, GSM8K
Useful for quick benchmark comparison.
"Experiments on GSM8K and MATH-500 demonstrate the effectiveness of MaSTer on mathematical tasks."
Not extracted
No metric anchors detected.
"Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks."
No metric terms were extracted from the available abstract.
Large Language Models (LLMs) have demonstrated remarkable capabilities in various reasoning-intensive tasks.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
No clear evaluation mode extracted.
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: MATH-500, GSM8K
Metric reporting is present
No metric terms extracted.