Human Feedback Types
strongPairwise Preference
Directly usable for protocol triage.
"Large language model (LLM) agents are moving beyond prompting alone."
HFEPX · Eval paper review
Pengcheng Jiang, Jiacheng Lin, Zhiyi Shi, Zifeng Wang +30 more
Published
Dec 18, 2025
Citations
0
Trust level
Moderate
Usefulness score
55/100 (Medium)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Mar 9, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Use this page for context, then validate protocol choices against stronger HFEPX references before implementation decisions.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Read the full paper before copying any benchmark, metric, or protocol choices.
Main weakness
The abstract does not clearly name benchmarks or metrics.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Large language model (LLM) agents are moving beyond prompting alone. ChatGPT marked the rise of general-purpose LLM assistants, DeepSeek showed that on-policy reinforcement learning with verifiable rewards can improve reasoning and tool use, and OpenClaw highlights a newer direction in which agents accumulate persistent memory and reusable skills. Yet the research landscape remains fragmented across post-training, retrieval, memory, and skill systems. This survey studies these developments under a single notion of \emph{adaptation}: improving an agent, its tools, or their interaction after pretraining. We organize the field with a four-paradigm framework spanning agent adaptation and tool adaptation. On the agent side, A1 (tool-execution-signaled) and A2 (agent-output-signaled) improve the agent itself through supervised fine-tuning, preference optimization, and reinforcement learning with verifiable rewards. On the tool side, T1 (agent-agnostic) provides reusable pre-trained modules any agent can call, while T2 (agent-supervised) uses the agent's outputs to train memory systems, skill libraries, or lightweight subagents. Using this framework, we review post-training methods, adaptive memory architectures, and agent skills; compare their trade-offs in cost, flexibility, and generalization; and summarize evaluation practices across deep research, software development, computer use, and drug discovery. We conclude by outlining open problems in agent-tool co-adaptation, continual learning, safety, and efficient deployment.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Pairwise Preference
Directly usable for protocol triage.
"Large language model (LLM) agents are moving beyond prompting alone."
Automatic Metrics
Includes extracted eval setup.
"Large language model (LLM) agents are moving beyond prompting alone."
Not reported
No explicit QC controls found.
"Large language model (LLM) agents are moving beyond prompting alone."
Not extracted
No benchmark anchors detected.
"Large language model (LLM) agents are moving beyond prompting alone."
Not extracted
No metric anchors detected.
"Large language model (LLM) agents are moving beyond prompting alone."
No benchmark or dataset names were extracted from the available abstract.
No metric terms were extracted from the available abstract.
Large language model (LLM) agents are moving beyond prompting alone.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Pairwise Preference
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
No metric terms extracted.