Ziyu Chen, Yilun Zhao, Jiashuo Sun, Yiling Ma, Manasi Patwardhan, Arman Cohan · Oct 6, 2026 · Citations: 0
Tag: Demonstrations
Demonstrations papers in the current HFEPX explorer (108 papers).
Papers in tag: 108
Running a Demonstrations study?
Post a Job →Research Utility Snapshot
Evaluation Modes
- Automatic Metrics (8)
- Simulation Env (1)
Human Feedback Types
- Demonstrations (20)
- Pairwise Preference (1)
Required Expertise
- General (12)
- Math (5)
- Coding (3)
Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang · Oct 6, 2026 · Citations: 0
- To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by…
- Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses.
Xi Ding, Naichen Shi, Jiawei Zhang · Oct 6, 2026 · Citations: 0
Venkata M Sangaraju, Sudhir Vissa · Oct 5, 2026 · Citations: 0
- Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance…
- Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column.
Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang · Oct 5, 2026 · Citations: 0
- LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done.
- We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives.
Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal · Oct 4, 2026 · Citations: 0
- Computer-use agents need to capture procedural knowledge of how people use software.
- We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%.
Songtao Li, Yijia Zhang, Shidi Zhang, Jianyuan Yuan, Fengyu Zhang, Hongfei Lin · Oct 2, 2026 · Citations: 0
- To address these limitations, we propose GAMA, a guideline-augmented multi-agent framework for schema-as-code BioNER.
Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao · Sep 29, 2026 · Citations: 0
- For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction.
- Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks.
Huan Rong, Chao Yin, Anouar Imel, Yijie Xia, Tinghuai Ma · Sep 29, 2026 · Citations: 0
- In this way, the safety issues arising in AD can be mitigated through constrained actions.
Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio · Sep 29, 2026 · Citations: 0
- Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training.
- Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics.
Jacob Epifano · Sep 29, 2026 · Citations: 0
Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge · Sep 29, 2026 · Citations: 0
- The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark.
- These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
Botong Zhao, Fang Yu, Tim, Senhua Zhu, Xinyuan Chen, Yue Lu · Aug 31, 2026 · Citations: 0
- Across the three simulation benchmarks, \method achieves the strongest overall performance while preserving the direct actor's online execution path.
Simeng Zhang, Yilong Chen, Wenyuan Zhang, Zhenyu Zhang, Yao Chen, Junyuan Shang · Aug 21, 2026 · Citations: 0
Utkarsh Bahuguna · Aug 11, 2026 · Citations: 0
- On the full GPQA Diamond benchmark (198 graduate-level science questions), majority voting reduces per-problem accuracy on a majority of problems for two instruction-tuned models from different families: 56.6% of problems for Qwen2.5-7B and…
Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma, Yi Feng · Aug 9, 2026 · Citations: 0
- Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding.
- Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains.
Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026 · Citations: 0
- We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
- In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
Tao Feng, Xinke Jiang, Chao Wu · Jun 29, 2026 · Citations: 0
- Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memory, when…
- Experiments on multiple benchmarks show that KbSD consistently improves both task accuracy and hallucination mitigation over strong baselines, with the largest gains appearing in the challenging quadrants where sparse rewards are least…
Xiao You, Tianwei Yan, Shan Zhao · Jun 28, 2026 · Citations: 0
Alex Kwon · Jun 28, 2026 · Citations: 0
- LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps trust.
- We show this rewriting manufactures confidence: across our constructed agent settings, a casual, hedged remark becomes a confident, dated assertion the agent then obeys like a verified fact, granting every above-clearance request it faces.