Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
HFEPX · Eval paper review
Hongbo Ma, Sansheng Cao, Jiajun Fan, Bangji Yang +1 more
Published
Sep 29, 2026
Citations
0
Trust level
Low
Usefulness score
0/100 (Low)
Extraction confidence
35% (Low)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Sep 29, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this as background context only. Do not make protocol decisions from this page alone.
All signals on this page are inferred from the abstract only and may be inaccurate. Do not use this page as a primary protocol reference.
Best use
Background context only
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
This paper looks adjacent to evaluation work, but not like a strong protocol reference.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost. We find that the core of reasoning capacity lies in the Thinking model's weight component within the null space of a projection defined by the corresponding Non-thinking model's dominant singular directions, and removing the subspace component can largely improve reasoning efficiency without hurting the accuracy gained during thinking-mode post-training. Unlike existing efforts that mostly operate within the dominant subspace, we are the first to unveil the critical role of the null space and harness it for model optimization. Motivated by this finding, we propose Spectral Null-Space Swap ($S^3$), a training-free composition of paired Non-thinking and Thinking checkpoints. Our method keeps the Non-thinking model inside its own dominant subspace and takes the Thinking checkpoint outside it, improving reasoning efficiency while maintaining accuracy. We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains. $S^3$ establishes new empirical Pareto Frontiers among training-free model composition strategies: across all settings, it reduces inference token overhead by an average of 27.4% compared to full Thinking models while simultaneously improving overall task accuracy by 1.0 percentage point (e.g., yielding +8.3% accuracy on HMMT25 alongside a 33.0% token speedup). We further use attention entropy for explanation and find that the retained component produces more concentrated attention, and we use a simplified analytical model about optimization to demonstrate why null-space can effectively reduce attention entropy, thereby improving the efficiency of reasoning.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
Automatic Metrics
Includes extracted eval setup.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
Not reported
No explicit QC controls found.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
Not extracted
No benchmark anchors detected.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
Accuracy, Token cost
Useful for evaluation criteria comparison.
"LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost."
Domain Experts
Helpful for staffing comparability.
"We extensively evaluate $S^3$ on 2B-30B dense and mixture-of-experts (MoE) architectures spanning 28 evaluation environments across mathematical, multimodal, and audio reasoning domains."
No benchmark or dataset names were extracted from the available abstract.
LLMs trained with Chain-of-thought excel in reasoning capability, but often come with excessive token cost.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy, token cost