Human Feedback Types
strongDemonstrations
Directly usable for protocol triage.
"The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task."
HFEPX · Eval paper review
Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado +5 more
Published
Oct 4, 2026
Citations
0
Trust level
High
Usefulness score
65/100 (Medium)
Extraction confidence
80% (High)
Derived from extracted protocol signals and abstract evidence.
Rater population
Not reported
Signals refreshed
Oct 6, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this as a practical starting point for protocol research, then validate against the original paper.
Best use
Secondary protocol comparison source
Use if you need
A benchmark-and-metrics comparison anchor.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) Non-Replayability, since past activity cannot be replayed to evaluate skill updates; and (3) Interleaved Trajectories, since logs may mix several tasks without marking their boundaries. To address these, we introduce TeleTune, a framework for learning a textual skill library from offline logs without recorded goals, cannot be replayed during optimization, and may interleave tasks. TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress. The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task. At test time, the agent is provided with the learned library and the workflow-based retrieved demonstrations. Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination. We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%. Under the heaviest perturbation of the WorkArena training data,TeleTune keeps the highest average success rate at 68.5%, 6.3% above the strongest baseline. Our analyses show (1) skill optimization and workflow-based retrieval are complementary, (2) optimizing on fixed logs costs 5 to 75 times fewer tokens than validating the same edits with live episodes, (3) skill-guided progress tracks the live success rate.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Demonstrations
Directly usable for protocol triage.
"The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task."
Automatic Metrics
Includes extracted eval setup.
"Computer-use agents need to capture procedural knowledge of how people use software."
Not reported
No explicit QC controls found.
"Computer-use agents need to capture procedural knowledge of how people use software."
Mind2Web, WorkArena
Useful for quick benchmark comparison.
"Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination."
Accuracy, Success rate
Useful for evaluation criteria comparison.
"TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress."
Computer-use agents need to capture procedural knowledge of how people use software.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Demonstrations
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Mind2Web, WorkArena
Metric reporting is present
Detected: accuracy, success rate