Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Demonstrations

Demonstrations papers in the current HFEPX explorer (108 papers).

Papers in tag: 108

Running a Demonstrations study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Automatic Metrics (8)
  • Simulation Env (1)

Human Feedback Types

  • Demonstrations (20)
  • Pairwise Preference (1)

Required Expertise

  • General (12)
  • Math (5)
  • Coding (3)
Sherpa: Teaching LLMs to Teach Adaptively

Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang · Oct 6, 2026 · Citations: 0

Pairwise PreferenceDemonstrations Math
  • To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by…
  • Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses.
Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents

Venkata M Sangaraju, Sudhir Vissa · Oct 5, 2026 · Citations: 0

Demonstrations Math
  • Enterprise AI agents that share a memory store face two unaddressed risks: sensitive data can leak through legitimately computed results the requester could not derive, and departments can silently compute a same-named key performance…
  • Existing agent-memory systems (e.g., MemGPT, Zep, A-MEM) gate retrieval by content, ownership, and role, not derivation, missing a cached insight that embeds a forbidden column.
Recursive Video In-Context Learning for Agentic Robot

Wenrui Bao, Xinxin Liu, Bingxin Xu, Yuzhang Shang · Oct 5, 2026 · Citations: 0

Demonstrations General
  • LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done.
  • We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives.
TeleTune: Evolving Agent Skills From Offline Telemetry

Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal · Oct 4, 2026 · Citations: 0

Demonstrations Automatic Metrics General
  • Computer-use agents need to capture procedural knowledge of how people use software.
  • We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%.
Skill-Space Shooting for Autonomous Robot Policy Improvement

Zihang Rui, Renhao Wang, Haoxu Huang, Yang Gao · Sep 29, 2026 · Citations: 0

Demonstrations General
  • For this improvement to scale across tasks, it must make effective use of experience without requiring human demonstration of each correction.
  • Recent agentic systems offer a way to reduce this reliance on human effort by using foundation models to autonomously compose learned behaviors to complete tasks.
Pixels to Keys: Exploring Spatial and Motion Cues in Gameplay Inverse Dynamics

Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio · Sep 29, 2026 · Citations: 0

Demonstrations Automatic Metrics General
  • Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrations, but they rarely include player inputs for training.
  • Our experiments on Trackmania highlight the importance of factors like the model architecture and motion flow extraction in preprocessing, while also showing the limits of evaluation through unbalanced metrics.
Can Language Models Learn to Forecast Stock Prices

Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge · Sep 29, 2026 · Citations: 0

Demonstrations Automatic Metrics MathCoding
  • The resulting AURA-4B more than doubles the starting direction--magnitude score, from 20.94 to 43.31, and is comparable to frontier language models on this benchmark.
  • These results show that post-training can substantially improve financial forecasting performance, together with changes in how the model investigates the market, on this outcome-selected benchmark.
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Changhao Xiang, Shilin Zhang, Zheng Ma, Kanzhi Cheng, Ruize Ma, Yi Feng · Aug 9, 2026 · Citations: 0

Demonstrations Law
  • Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding.
  • Using this framework, we construct OpenVisTool-42K, a dataset spanning five visual reasoning domains, together with OpenVisTool-Bench, a benchmark covering the same domains.
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations

Mehul Damani, Isha Puri, Idan Shenfeld, Jacob Andreas · Jul 1, 2026 · Citations: 0

Demonstrations Automatic Metrics MathCoding
  • We propose an adversarial generator-discriminator framework that augments verifiable rewards with a learned signal from human demonstrations.
  • In story generation, our method significantly improves win rate while producing stories that are diverse and more human-like.
KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Search

Tao Feng, Xinke Jiang, Chao Wu · Jun 29, 2026 · Citations: 0

Demonstrations Automatic Metrics General
  • Agentic search equips large language models with dynamic retrieval abilities, but existing reinforcement learning methods remain limited by reward sparsity in knowledge boundary calibration -- deciding when to trust parametric memory, when…
  • Experiments on multiple benchmarks show that KbSD consistently improves both task accuracy and hallucination mitigation over strong baselines, with the largest gains appearing in the challenging quadrants where sparse rewards are least…
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts

Alex Kwon · Jun 28, 2026 · Citations: 0

Demonstrations General
  • LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation into stored "facts" that later steps trust.
  • We show this rewriting manufactures confidence: across our constructed agent settings, a casual, hedged remark becomes a confident, dated assertion the agent then obeys like a verified fact, granting every above-clearance request it faces.