Skip to content
OpenTrain AIFor AI Companies
← Back to explorer

Tag: Simulation Env

Simulation Env evaluation setups appearing in the current HFEPX corpus (279 papers).

Papers in tag: 279

Running a Simulation Env study?

Post a Job →

Research Utility Snapshot

Evaluation Modes

  • Simulation Env (20)
  • Automatic Metrics (3)

Human Feedback Types

  • Pairwise Preference (2)

Required Expertise

  • General (14)
  • Coding (4)
  • Medicine (3)
CaSKG: Counterfactual-Causal Skill Graphs for Scalable Agent Skill Retrieval

Zhiyuan Li, Linyuan Gao, Xuechun Ding, Hongwei Chen, Yuan Wu, Yi Chang · Aug 26, 2026 · Citations: 0

Simulation Env Coding
  • Reusable skill libraries allow large language model (LLM) agents to reuse procedural knowledge across tasks, but they also turn memory access into a challenging retrieval problem.
  • CaSKG first builds a high-recall directed candidate graph from semantic, lexical, input/output, and structural evidence, with repair evidence and an optional LLM judge further refining candidate scores.
Enhancing LLMs in Predictive Political QA with Semi-Structured Data

Yinan Liu, Zihan Zhou, Zichun Jin, Xinyu Wang, Bin Wang, Xiaochun Yang · Aug 21, 2026 · Citations: 0

Pairwise Preference Simulation Env General
  • We identify two complementary signals for predictive political QA: actor stances that capture issue-specific preferences, and high-order structure signals that capture indirect dependencies among political actors.
Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay

Haiyue Zhang · Aug 20, 2026 · Citations: 0

Simulation Env General
  • Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals used to train LLM agents -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's…
  • A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory).
Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories

Hexi Wang, Yujia Zhou, Bangde Du, Weihang Su, Xinyuan Cao, Qingyi Pan · Aug 20, 2026 · Citations: 0

Simulation Env General
  • Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed.
  • Inspired by complementary memory systems, we propose LifeMem, a longitudinal memory framework that combines structured life-event retrieval with agent-specific parametric memory for experience integration.
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories

Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu · Aug 13, 2026 · Citations: 0

Simulation Env General
  • Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed.
  • We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memory and formulate an evaluation framework that holds candidate retrieval, target state, model, decoding, and tool budget fixed while varying…
PatientAct: Theory-Grounded Mental Health Client Simulation

Sahand Sabour, TszYam NG, Yaqian Chen, Guanqun Bi, Jialu Zhao, Minlie Huang · Aug 13, 2026 · Citations: 0

Simulation Env MedicineCoding
  • Abstract shows limited direct human-feedback or evaluation-protocol detail; use as adjacent methodological context.
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL

Simon Yu, Nicholas Tomlin, Marwa Abdulhai, Ximing Lu, Derek Chong, Abe Hou · Aug 12, 2026 · Citations: 0

Simulation Env General
  • Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior.
  • Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users.
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations

Lior Baruch, Moshe Butman, Kfir Bar, Doron Friedman · Aug 12, 2026 · Citations: 0

Pairwise Preference Simulation Env Medicine
  • This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead.
  • Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets.
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Yutao Mou, Pengfei Yang, Zhe Yin, Zhangchi Xue, Xiaotian Luan, Dingyao Yu · Aug 12, 2026 · Citations: 0

Simulation Env General
  • Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states.
  • To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute.
Diagnosis Before Recovery: Turning Agent Failures into Selective Self-Correction

Pan Wang, Yihao Hu, Hang Wang, Zirui Lv, Xin Zhang, Jianshe Li · Aug 12, 2026 · Citations: 0

Simulation Env MedicineCoding
  • Coding agents benefit from this property because compilers, tests, and execution traces turn many failures into typed recovery signals, but broad language-agent tasks often expose only a coarse task failure.
  • This creates a tension for generic recovery playbooks: they broaden the agent's context precisely when the system needs a narrower repair interface, mixing incompatible signals for invalid actions, missing procedures, and strict-format…
Will Scaling Improve Social Simulation with LLMs?

Caleb Ziems, William Held, Su Doga Karaca, David Grusky, Tatsunori Hashimoto, Diyi Yang · Jul 2, 2026 · Citations: 0

Automatic MetricsSimulation Env Law
  • We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal…
  • Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU.
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du, Jiacheng Chen, Tianle Li · Jul 2, 2026 · Citations: 0

Simulation Env General
  • Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended software-engineering progress.
  • We introduce Autonomous Policy Evolution, a controlled evaluation setting in which a harness-model agent repeatedly edits an executable policy system under a fixed interaction budget.
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation

Peng Yun, Shouwang Huang, Hao Li, Jinxi Li, Jianan Wang, Bo Yang · Jul 2, 2026 · Citations: 0

Simulation Env General
  • We introduce PhysMani-Bench, a dynamic manipulation benchmark with 16 tasks, and demonstrate a superior success rate over strong baselines in both simulation and real-world robot experiments.
PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation

Junhao Chen, Xiang Li, Mingjin Chen, Boran Zhang, Henghaofan Zhang, Yibin Xu · Jul 2, 2026 · Citations: 0

Simulation Env Coding
  • We present PairCoder, which grounds review in the toolchain and realizes it as two agent pair programming: a Driver agent writes the program, a Navigator agent reviews it against verification evidence (diagnostics, execution results, and…
  • Across 17 public benchmarks and seven models from three vendors, PairCoder improves essentially every benchmark whose artifact is verifiable, on full official metric suites rather than execution alone (for example, Blender scene…
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

Muhammad Usman Safder, Ayesha Gull, Rania Elbadry, Fan Zhang, Yankai Chen, Xueqing Peng · Jun 30, 2026 · Citations: 0

Simulation Env General
  • Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout…
  • To measure MSD objectively, we introduce FinPersona-Bench, a simulation benchmark in which a synthetic market decouples observable price from hidden fundamental value, enabling falsifiable evaluation across three failure modes: trading…
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation

Songjun Tu, Chengdong Xu, Qichao Zhang, Yiwen Ma, Yaocheng Zhang, Linjing Li · Jun 28, 2026 · Citations: 0

Simulation Env General
  • Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not oracular: they may help in one state while misleading the same policy in another.
  • We introduce UCOB, a framework for learning to utilize and evolve agentic skills via credit-aware on-policy bidirectional self-distillation.