300 canonical paper links on this archive page.
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessionsarxiv-2606.23654 Sparse Blocked context onlyJun 22, 2026
- GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agentsarxiv-2606.24551 Sparse Blocked context onlyJun 22, 2026
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Viewsarxiv-2606.23557 Sparse Blocked context onlyJun 22, 2026
- VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instructarxiv-2606.23543 Sparse Blocked context onlyJun 22, 2026
- Self-Compacting Language Model Agentsarxiv-2606.23525 Sparse Blocked context onlyJun 22, 2026
- UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observationarxiv-2606.23503 Curated Related Blocked context onlyJun 22, 2026
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interactionarxiv-2606.23449 Sparse Blocked context onlyJun 22, 2026
- ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Modelsarxiv-2606.23404 Sparse Blocked context onlyJun 22, 2026
- Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?arxiv-2606.23189 Sparse Blocked context onlyJun 22, 2026
- Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluationarxiv-2606.23127 Sparse Blocked context onlyJun 22, 2026
- Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latentsarxiv-2606.23085 Sparse Blocked context onlyJun 22, 2026
- When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agentsarxiv-2606.22936 Sparse Blocked context onlyJun 22, 2026
- Libretto: Giving LLM Agents a Sense of Musical Structurearxiv-2606.22708 Sparse Blocked context onlyJun 21, 2026
- PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Modelsarxiv-2606.22540 Sparse Blocked context onlyJun 21, 2026
- OpenBioRQ: Unsolved Biomedical Research Questions for Agentsarxiv-2606.21959 Sparse Blocked context onlyJun 20, 2026
- A Verifiable Search Is Not a Learnable Chain-of-Thoughtarxiv-2606.21884 Sparse Blocked context onlyJun 20, 2026
- PrivacyAlign: Contextual Privacy Alignment for LLM Agentsarxiv-2606.21710 Sparse Blocked context onlyJun 19, 2026
- Improving Text-to-Music Generation with Human Preference Rewardsarxiv-2606.21670 Sparse Blocked context onlyJun 19, 2026
- UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gatingarxiv-2606.21661 Sparse Blocked context onlyJun 19, 2026
- Counsel: A Meta-Evaluation Dataset for Agentic Tasksarxiv-2606.21627 Sparse Blocked context onlyJun 19, 2026
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streamsarxiv-2606.21337 Sparse Blocked context onlyJun 19, 2026
- Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approacharxiv-2606.21215 Sparse Blocked context onlyJun 19, 2026
- BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discoveryarxiv-2606.20997 Sparse Blocked context onlyJun 19, 2026
- Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York Cityarxiv-2606.20980 Direct Blocked context onlyJun 18, 2026
- CogniRoute: Learning to Route Social Evidence in Omni-Modal Modelsarxiv-2606.20970 Sparse Blocked context onlyJun 18, 2026
- Grouped Query Experts: Mixture-of-Experts on GQA Self-Attentionarxiv-2606.20945 Sparse Blocked context onlyJun 18, 2026
- Vesta: A Generalist Embodied Reasoning Modelarxiv-2606.20905 Sparse Blocked context onlyJun 18, 2026
- JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoisingarxiv-2606.20563 Sparse Blocked context onlyJun 18, 2026
- Current World Models Lack a Persistent State Corearxiv-2606.20545 Sparse Blocked context onlyJun 18, 2026
- Toward Calibrated Mixture-of-Experts Under Distribution Shiftarxiv-2606.20544 Sparse Blocked context onlyJun 18, 2026
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligencearxiv-2606.20515 Sparse Blocked context onlyJun 18, 2026
- World Action Models: A Surveyarxiv-2606.20781 Sparse Blocked context onlyJun 18, 2026
- Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipearxiv-2606.20381 Sparse Blocked context onlyJun 18, 2026
- Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptationarxiv-2606.20196 Sparse Blocked context onlyJun 18, 2026
- ReNikud: Audio-Supervised Hebrew Grapheme-to-Phoneme Conversionarxiv-2606.20179 Sparse Blocked context onlyJun 18, 2026
- EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policiesarxiv-2606.20092 Sparse Blocked context onlyJun 18, 2026
- When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agentsarxiv-2606.20023 Sparse Blocked context onlyJun 18, 2026
- ENPIRE: Agentic Robot Policy Self-Improvement in the Real Worldarxiv-2606.19980 Sparse Blocked context onlyJun 18, 2026
- MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimizationarxiv-2606.19930 Sparse Blocked context onlyJun 18, 2026
- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Managementarxiv-2606.19926 Curated Related Blocked context onlyJun 18, 2026
- Light-weight Pronunciation Assessment via Discrete Speech Token Surprisalarxiv-2606.19910 Sparse Blocked context onlyJun 18, 2026
- Leverage Is Not Reach: A Control-Window Law for Single-Neuron Steering in Language Modelsarxiv-2606.19831 Sparse Blocked context onlyJun 18, 2026
- JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Enginesarxiv-2606.19830 Sparse Blocked context onlyJun 18, 2026
- Clusters are All You Need: Pre-Training the Tsetlin Machine with Semantic Clusters from Language Models for Interpretabilityarxiv-2606.19815 Sparse Blocked context onlyJun 18, 2026
- Think Again or Think Longer? Selective Verification for Budget-Aware Reasoningarxiv-2606.19808 Sparse Blocked context onlyJun 18, 2026
- AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QAarxiv-2606.19782 Sparse Blocked context onlyJun 18, 2026
- Benchmarking Agentic Review Systemsarxiv-2606.19749 Sparse Blocked context onlyJun 18, 2026
- Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settingsarxiv-2606.19744 Sparse Blocked context onlyJun 18, 2026
- SAGE-OPD: Selective Agent-Guided Intervention for Multi-Turn On-Policy Distillationarxiv-2606.19659 Sparse Blocked context onlyJun 17, 2026
- Toten: A Knowledge-Based System For Structure-Preserving Representation Of Physical Quantities And Technical Notation In Brazilian Portuguesearxiv-2606.19626 Sparse Blocked context onlyJun 17, 2026
- Where Does Social Reasoning Come From? Capability Provenance in Language Modelsarxiv-2606.19625 Sparse Blocked context onlyJun 17, 2026
- FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelinesarxiv-2606.19605 Sparse Blocked context onlyJun 17, 2026
- Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Whyarxiv-2606.19602 Sparse Blocked context onlyJun 17, 2026
- LaViSA: A Language and Vision Structural Ambiguity Benchmarkarxiv-2606.19552 Sparse Blocked context onlyJun 17, 2026
- DeXposure-Claw: An Agentic System for DeFi Risk Supervisionarxiv-2606.19501 Sparse Blocked context onlyJun 17, 2026
- LooseControlVideo: Directorial Video Control using Spatial Blockingarxiv-2606.19495 Sparse Blocked context onlyJun 17, 2026
- Characterizing Narrative Content in Web-scale LLM Pretraining Dataarxiv-2606.19468 Sparse Blocked context onlyJun 17, 2026
- Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Modelsarxiv-2606.19297 Sparse Blocked context onlyJun 17, 2026
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performancearxiv-2606.19195 Curated Related Blocked context onlyJun 17, 2026
- OpenRath: Session-Centered Runtime State for Agent Systemsarxiv-2606.19409 Sparse Blocked context onlyJun 17, 2026
- The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RLarxiv-2606.19162 Sparse Blocked context onlyJun 17, 2026
- Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teamsarxiv-2606.19111 Sparse Blocked context onlyJun 17, 2026
- RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agentsarxiv-2606.19047 Sparse Blocked context onlyJun 17, 2026
- Enhancing Multilingual Reasoning via Steerable Model Mergingarxiv-2606.19002 Sparse Blocked context onlyJun 17, 2026
- Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancementarxiv-2606.18953 Sparse Blocked context onlyJun 17, 2026
- SenFlow: Inter-Sentence Flow Modeling for AI-Generated Text Detection in Hybrid Documentsarxiv-2606.18946 Sparse Blocked context onlyJun 17, 2026
- Physics-IQ Verifiedarxiv-2606.18943 Sparse Blocked context onlyJun 17, 2026
- ESBMC-GraphPLC: Formal Verification of Graphical PLCopen XML Ladder Diagram Programs Using SMT-Based Model Checkingarxiv-2606.18941 Sparse Blocked context onlyJun 17, 2026
- Improving Medical Communication using Rubric-Guided Counterfactual Recommendationsarxiv-2606.18889 Sparse Blocked context onlyJun 17, 2026
- Externalizing Research Synthesis and Validation in AI Scientists through a Research Harnessarxiv-2606.18874 Sparse Blocked context onlyJun 17, 2026
- Aligning Implied Statements for Implicit Hate Speech Generalizability with Context-Bounded Semi-hard Negative Miningarxiv-2606.18852 Sparse Blocked context onlyJun 17, 2026
- WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agentsarxiv-2606.18847 Sparse Blocked context onlyJun 17, 2026
- Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistancearxiv-2606.18613 Sparse Blocked context onlyJun 17, 2026
- CEO-Bench: Can Agents Play the Long Game?arxiv-2606.18543 Sparse Blocked context onlyJun 16, 2026
- MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrievalarxiv-2606.18508 Sparse Blocked context onlyJun 16, 2026
- EgoCS-400K: An Egocentric Gameplay Dataset for World Modelsarxiv-2606.18180 Sparse Blocked context onlyJun 16, 2026
- LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AIarxiv-2606.18021 Sparse Blocked context onlyJun 16, 2026
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?arxiv-2606.17861 Sparse Blocked context onlyJun 16, 2026
- MaineCoon: Pursuing A Real-Time Audio-Visual Social World Modelarxiv-2606.17800 Curated Related Blocked context onlyJun 16, 2026
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memoryarxiv-2606.17730 Sparse Blocked context onlyJun 16, 2026
- OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillationarxiv-2606.17628 Sparse Blocked context onlyJun 16, 2026
- ProCUA-SFT Technical Reportarxiv-2606.17321 Sparse Blocked context onlyJun 15, 2026
- MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revisionarxiv-2606.17162 Sparse Blocked context onlyJun 15, 2026
- Benchmarking LLM Agents on Meta-Analysis Articles from Nature Portfolioarxiv-2606.17041 Sparse Blocked context onlyJun 15, 2026
- TuneJury: An Open Metric for Improving Music Generation Preference Alignmentarxiv-2606.17006 Sparse Blocked context onlyJun 15, 2026
- OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendationarxiv-2606.16838 Sparse Blocked context onlyJun 15, 2026
- GD$^2$PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimizationarxiv-2606.16771 Sparse Blocked context onlyJun 15, 2026
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agentsarxiv-2606.16748 Sparse Blocked context onlyJun 15, 2026
- Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agentsarxiv-2606.16682 Sparse Blocked context onlyJun 15, 2026
- CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economiesarxiv-2606.16613 Sparse Blocked context onlyJun 15, 2026
- BadWorld: Adversarial Attacks on World Modelsarxiv-2606.16519 Sparse Blocked context onlyJun 15, 2026
- PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memoryarxiv-2606.16449 Sparse Blocked context onlyJun 15, 2026
- Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillationarxiv-2606.16429 Sparse Blocked context onlyJun 15, 2026
- A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimizationarxiv-2606.16154 Sparse Blocked context onlyJun 15, 2026
- Calibrated Triage, Not Autonomy: Confidence Estimation for Medical Vision-Language Modelsarxiv-2606.15910 Sparse Blocked context onlyJun 14, 2026
- Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Timearxiv-2606.15631 Sparse Blocked context onlyJun 14, 2026
- CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?arxiv-2606.15300 Sparse Blocked context onlyJun 13, 2026
- Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusionarxiv-2606.15236 Sparse Blocked context onlyJun 13, 2026
- MotionVLA: Vision-Language-Action Model for Humanoid Motionarxiv-2606.15142 Curated Related Blocked context onlyJun 13, 2026
- MVEB: Massive Video Embedding Benchmarkarxiv-2606.14958 Sparse Blocked context onlyJun 12, 2026
- Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansionarxiv-2606.14885 Sparse Blocked context onlyJun 12, 2026
- AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimizationarxiv-2606.14694 Sparse Blocked context onlyJun 12, 2026
- LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversationsarxiv-2606.14600 Sparse Blocked context onlyJun 12, 2026
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actionsarxiv-2606.14832 Sparse Blocked context onlyJun 12, 2026
- Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environmentsarxiv-2606.14397 Sparse Blocked context onlyJun 12, 2026
- Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networksarxiv-2606.14819 Sparse Blocked context onlyJun 12, 2026
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundryarxiv-2606.14249 Sparse Blocked context onlyJun 12, 2026
- ViT-Up: Faithful Feature Upsampling for Vision Transformersarxiv-2606.14024 Curated Related Blocked context onlyJun 12, 2026
- Self-Evolving Visual Questionerarxiv-2606.13929 Sparse Blocked context onlyJun 11, 2026
- The Price of Anarchy in Disaggregated Inferencearxiv-2606.17081 Sparse Blocked context onlyJun 11, 2026
- $μ_0$: A Scalable 3D Interaction-Trace World Modelarxiv-2606.13769 Sparse Blocked context onlyJun 11, 2026
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoningarxiv-2606.13673 Sparse Blocked context onlyJun 11, 2026
- $\texttt{WEAVER}$, Better, Faster, Longer: An Effective World Model for Robotic Manipulationarxiv-2606.13672 Sparse Blocked context onlyJun 11, 2026
- EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discoveryarxiv-2606.13662 Sparse Blocked context onlyJun 11, 2026
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillationarxiv-2606.13657 Sparse Blocked context onlyJun 11, 2026
- Surflo: Consistent 3D Surface Flow Model with Global Statearxiv-2606.13644 Sparse Blocked context onlyJun 11, 2026
- LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratoriesarxiv-2606.13578 Sparse Blocked context onlyJun 11, 2026
- MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffoldarxiv-2606.13376 Sparse Blocked context onlyJun 11, 2026
- VideoMDM: Towards 3D Human Motion Generation From 2D Supervisionarxiv-2606.13364 Sparse Blocked context onlyJun 11, 2026
- Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agentsarxiv-2606.13174 Sparse Blocked context onlyJun 11, 2026
- Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learningarxiv-2606.13106 Sparse Blocked context onlyJun 11, 2026
- Benchmarking AI Agents for Addressing Scientific Challenges Across Scalesarxiv-2606.12736 Sparse Blocked context onlyJun 10, 2026
- Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behaviorarxiv-2606.12730 Sparse Blocked context onlyJun 10, 2026
- From AGI to ASIarxiv-2606.12683 Sparse Blocked context onlyJun 10, 2026
- Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agentsarxiv-2606.12674 Sparse Blocked context onlyJun 10, 2026
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillationarxiv-2606.12575 Sparse Blocked context onlyJun 10, 2026
- Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMsarxiv-2606.12385 Sparse Blocked context onlyJun 10, 2026
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusionarxiv-2606.12303 Sparse Blocked context onlyJun 10, 2026
- FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agentsarxiv-2606.12087 Sparse Blocked context onlyJun 10, 2026
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinementarxiv-2606.11926 Sparse Blocked context onlyJun 10, 2026
- Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoningarxiv-2606.11683 Sparse Blocked context onlyJun 10, 2026
- TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Searcharxiv-2606.11662 Sparse Blocked context onlyJun 10, 2026
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligencearxiv-2606.14777 Sparse Blocked context onlyJun 10, 2026
- Forecasting Future Behavior as a Learning Taskarxiv-2606.11445 Sparse Blocked context onlyJun 9, 2026
- Data Journalist Agent: Transforming Data into Verifiable Multimodal Storiesarxiv-2606.11176 Sparse Blocked context onlyJun 9, 2026
- WorldOlympiad: Can Your World Model Survive a Triathlon?arxiv-2606.11129 Sparse Blocked context onlyJun 9, 2026
- IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoderarxiv-2606.11096 Sparse Blocked context onlyJun 9, 2026
- Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Modelsarxiv-2606.11025 Sparse Blocked context onlyJun 9, 2026
- RedAct: Redacting Agent Capability Traces for Procedural Skill Protectionarxiv-2606.10813 Sparse Blocked context onlyJun 9, 2026
- When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Modelsarxiv-2606.10740 Sparse Blocked context onlyJun 9, 2026
- DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratcharxiv-2606.10728 Sparse Blocked context onlyJun 9, 2026
- Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillationarxiv-2606.11270 Sparse Blocked context onlyJun 9, 2026
- $τ$-Rec: A Verifiable Benchmark for Agentic Recommender Systemsarxiv-2606.10156 Sparse Blocked context onlyJun 8, 2026
- iMaC: Translating Actions into Motion and Contact Images for Embodied World Modelsarxiv-2606.09813 Sparse Blocked context onlyJun 8, 2026
- Echo-Memory: A Controlled Study of Memory in Action World Modelsarxiv-2606.09803 Sparse Blocked context onlyJun 8, 2026
- End-to-End Context Compression at Scalearxiv-2606.09659 Sparse Blocked context onlyJun 8, 2026
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfacesarxiv-2606.09426 Sparse Blocked context onlyJun 8, 2026
- Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memoryarxiv-2606.09365 Sparse Blocked context onlyJun 8, 2026
- PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignmentarxiv-2606.09348 Sparse Blocked context onlyJun 8, 2026
- Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoningarxiv-2606.09290 Sparse Blocked context onlyJun 8, 2026
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attentionarxiv-2606.09079 Curated Related Blocked context onlyJun 8, 2026
- Beyond Scalar Rewards by Internalizing Reasoning into Score Distributionsarxiv-2606.09076 Sparse Blocked context onlyJun 8, 2026
- Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loopsarxiv-2606.08960 Sparse Blocked context onlyJun 8, 2026
- AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Modelsarxiv-2606.08952 Sparse Blocked context onlyJun 8, 2026
- PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleafarxiv-2606.08857 Sparse Blocked context onlyJun 7, 2026
- SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision Historyarxiv-2606.08671 Sparse Blocked context onlyJun 7, 2026
- PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systemsarxiv-2606.08481 Sparse Blocked context onlyJun 7, 2026
- Light-WAM: Efficient World Action Models with State-Fusion Action Decodingarxiv-2606.08242 Sparse Blocked context onlyJun 6, 2026
- MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learningarxiv-2606.08039 Sparse Blocked context onlyJun 6, 2026
- TBD-VLA: Temporal Block Diffusion Vision Language Action Modelarxiv-2606.07895 Sparse Blocked context onlyJun 5, 2026
- The Cold-Start Safety Gap in LLM Agentsarxiv-2606.07867 Sparse Blocked context onlyJun 5, 2026
- UniSHARP: Universal Sharp Monocular View Synthesisarxiv-2606.07514 Sparse Blocked context onlyJun 5, 2026
- Streaming Video Generation with Streaming Force Controlarxiv-2606.07508 Sparse Blocked context onlyJun 5, 2026
- Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoningarxiv-2606.07436 Sparse Blocked context onlyJun 5, 2026
- Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skillsarxiv-2606.07412 Sparse Blocked context onlyJun 5, 2026
- Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Testsarxiv-2606.07379 Sparse Blocked context onlyJun 5, 2026
- SWE-Explore: Benchmarking How Coding Agents Explore Repositoriesarxiv-2606.07297 Sparse Blocked context onlyJun 5, 2026
- dots.tts Technical Reportarxiv-2606.07080 Sparse Blocked context onlyJun 5, 2026
- SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gatingarxiv-2606.07074 Sparse Blocked context onlyJun 5, 2026
- Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Commentsarxiv-2606.06843 Sparse Blocked context onlyJun 5, 2026
- Regret Minimization with Adaptive Opponents in Repeated Gamesarxiv-2606.06486 Sparse Blocked context onlyJun 4, 2026
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulatorsarxiv-2606.06476 Sparse Blocked context onlyJun 4, 2026
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agentsarxiv-2606.06087 Sparse Blocked context onlyJun 4, 2026
- Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agentsarxiv-2606.06036 Sparse Blocked context onlyJun 4, 2026
- OPRD: On-Policy Representation Distillationarxiv-2606.06021 Sparse Blocked context onlyJun 4, 2026
- Robots Need More than VLA and World Modelsarxiv-2606.06556 Sparse Blocked context onlyJun 4, 2026
- Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systemsarxiv-2606.05985 Sparse Blocked context onlyJun 4, 2026
- When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agentsarxiv-2606.05806 Sparse Blocked context onlyJun 4, 2026
- DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Modelsarxiv-2606.05758 Sparse Blocked context onlyJun 4, 2026
- Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learningarxiv-2606.05645 Sparse Blocked context onlyJun 4, 2026
- AsyncWebRL: Efficient Multi-Step RL for Visual Web Agentsarxiv-2606.05597 Sparse Blocked context onlyJun 4, 2026
- SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variationsarxiv-2606.05563 Sparse Blocked context onlyJun 4, 2026
- Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Modelsarxiv-2606.05531 Sparse Blocked context onlyJun 4, 2026
- Personal AI Agent for Camera Roll VQAarxiv-2606.05275 Sparse Blocked context onlyJun 3, 2026
- Streaming Communication in Multi-Agent Reasoningarxiv-2606.05158 Sparse Blocked context onlyJun 3, 2026
- Reinforcement Learning from Rich Feedback with Distributional DAggerarxiv-2606.05152 Sparse Blocked context onlyJun 3, 2026
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?arxiv-2606.05080 Sparse Blocked context onlyJun 3, 2026
- DAR: Deontic Reasoning with Agentic Harnessesarxiv-2606.05009 Sparse Blocked context onlyJun 3, 2026
- M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasksarxiv-2606.05008 Sparse Blocked context onlyJun 3, 2026
- TIDE: Proactive Multi-Problem Discovery via Template-Guided Iterationarxiv-2606.04743 Sparse Blocked context onlyJun 3, 2026
- MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generationarxiv-2606.04688 Sparse Blocked context onlyJun 3, 2026
- Why Muon Outperforms Adam: A Curvature Perspectivearxiv-2606.04662 Sparse Blocked context onlyJun 3, 2026
- GENEB: Why Genomic Models Are Hard to Comparearxiv-2606.04525 Sparse Blocked context onlyJun 3, 2026
- The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?arxiv-2606.04455 Sparse Blocked context onlyJun 3, 2026
- A Cookbook of 3D Vision: Data, Learning Paradigms, and Applicationarxiv-2606.04291 Sparse Blocked context onlyJun 2, 2026
- MAOAM: Unified Object and Material Selection with Vision-Language Modelsarxiv-2606.04880 Sparse Blocked context onlyJun 2, 2026
- Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Modelsarxiv-2606.03988 Sparse Blocked context onlyJun 2, 2026
- AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generationarxiv-2606.03972 Sparse Blocked context onlyJun 2, 2026
- GridVQA-X: A Framework for Evaluating Multimodal Explainability Methodsarxiv-2606.14740 Sparse Blocked context onlyJun 2, 2026
- Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Studyarxiv-2606.04056 Sparse Blocked context onlyJun 2, 2026
- Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusionarxiv-2606.14732 Sparse Blocked context onlyJun 2, 2026
- SkillHarness: Harnessing Safe Skills for Computer-Use Agentsarxiv-2606.20636 Sparse Blocked context onlyJun 2, 2026
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulationarxiv-2606.03159 Sparse Blocked context onlyJun 2, 2026
- Neural Networks Provably Learn Spectral Representations for Group Compositionarxiv-2606.02993 Sparse Blocked context onlyJun 2, 2026
- Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactionsarxiv-2606.02859 Sparse Blocked context onlyJun 1, 2026
- $Ψ$-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialoguesarxiv-2606.02754 Sparse Blocked context onlyJun 1, 2026
- Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Modelsarxiv-2606.02580 Sparse Blocked context onlyJun 1, 2026
- SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Constructionarxiv-2606.02540 Sparse Blocked context onlyJun 1, 2026
- AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agentsarxiv-2606.02461 Sparse Blocked context onlyJun 1, 2026
- TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generationarxiv-2606.02320 Sparse Blocked context onlyJun 1, 2026
- Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectoriesarxiv-2606.02060 Sparse Blocked context onlyJun 1, 2026
- Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labelingarxiv-2606.02004 Sparse Blocked context onlyJun 1, 2026
- MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?arxiv-2606.01993 Sparse Blocked context onlyJun 1, 2026
- AutoMedBench: Towards Medical AutoResearch with Agentic AI Modelsarxiv-2606.01961 Sparse Blocked context onlyJun 1, 2026
- PlatonicNav: Unveiling Semantic Correspondence in Navigation with Platonic Topological Mapsarxiv-2606.01788 Sparse Blocked context onlyJun 1, 2026
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systemsarxiv-2606.01779 Sparse Blocked context onlyJun 1, 2026
- Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streamsarxiv-2606.01770 Sparse Blocked context onlyJun 1, 2026
- TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RLarxiv-2606.01599 Sparse Blocked context onlyJun 1, 2026
- Multi-Agent Computer Usearxiv-2606.01533 Sparse Blocked context onlyJun 1, 2026
- Joint Agent Memory and Exploration Learning via Novelty Signalsarxiv-2606.01528 Sparse Blocked context onlyJun 1, 2026
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verificationarxiv-2606.01476 Sparse Blocked context onlyMay 31, 2026
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Modelsarxiv-2606.01462 Sparse Blocked context onlyMay 31, 2026
- Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligencearxiv-2606.01444 Sparse Blocked context onlyMay 31, 2026
- ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formatsarxiv-2606.01348 Sparse Blocked context onlyMay 31, 2026
- CARVE: Certified Affordable Repair of Vetoed Maneuvers via Envelopes for Interactive Drivingarxiv-2606.02641 Sparse Blocked context onlyMay 31, 2026
- HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papersarxiv-2606.01132 Sparse Blocked context onlyMay 31, 2026
- 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Codearxiv-2606.01057 Sparse Blocked context onlyMay 31, 2026
- Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacherarxiv-2606.01000 Sparse Blocked context onlyMay 31, 2026
- MBench: A Comprehensive Benchmark on Memory Capability for Video World Modelsarxiv-2606.00793 Sparse Blocked context onlyMay 30, 2026
- Confidence-Adaptive SwiGLU for Mixture-of-Expertsarxiv-2606.00761 Sparse Blocked context onlyMay 30, 2026
- OCC-RAG: Optimal Cognitive Core for Faithful Question Answeringarxiv-2606.00683 Sparse Blocked context onlyMay 30, 2026
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgmentarxiv-2606.00644 Sparse Blocked context onlyMay 30, 2026
- Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMsarxiv-2606.00477 Sparse Blocked context onlyMay 30, 2026
- Masking Stale Observations Helps Search Agents -- Until It Doesn't: A Regime Map and Its Mechanismarxiv-2606.00408 Sparse Blocked context onlyMay 29, 2026
- AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agentsarxiv-2606.24893 Sparse Blocked context onlyMay 29, 2026
- Representation Forcing for Bottleneck-Free Unified Multimodal Modelsarxiv-2605.31604 Sparse Blocked context onlyMay 29, 2026
- SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Modelsarxiv-2605.31597 Sparse Blocked context onlyMay 29, 2026
- SVI-Bench: A Dynamic Microworld for Strategic Video Intelligencearxiv-2605.31529 Sparse Blocked context onlyMay 29, 2026
- How can embedding models bind concepts?arxiv-2605.31503 Sparse Blocked context onlyMay 29, 2026
- Scalable Inference-Time Annealing with Surrogate Likelihood Estimatorsarxiv-2605.31498 Sparse Blocked context onlyMay 29, 2026
- PaintBench: Deterministic Evaluation of Precise Visual Editingarxiv-2606.00188 Sparse Blocked context onlyMay 29, 2026
- DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memoryarxiv-2605.31336 Sparse Blocked context onlyMay 29, 2026
- Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasionarxiv-2605.31170 Sparse Blocked context onlyMay 29, 2026
- Trust-Region Behavior Blending for On-Policy Distillationarxiv-2605.31159 Sparse Blocked context onlyMay 29, 2026
- Light Interaction: Training-Free Inference Acceleration for Interactive Video World Modelsarxiv-2605.31158 Sparse Blocked context onlyMay 29, 2026
- SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenesarxiv-2605.31148 Sparse Blocked context onlyMay 29, 2026
- Task-Focused Memorization for Multimodal Agentsarxiv-2605.31075 Sparse Blocked context onlyMay 29, 2026
- GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restorationarxiv-2605.31039 Sparse Blocked context onlyMay 29, 2026
- PEEK: Picking Essential frames via Efficient Knowledge distillationarxiv-2605.31029 Sparse Blocked context onlyMay 29, 2026
- SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialoguearxiv-2605.30993 Sparse Blocked context onlyMay 29, 2026
- Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelismarxiv-2605.30852 Sparse Blocked context onlyMay 29, 2026
- LLM Anonymization Against Agentic Re-Identificationarxiv-2605.30848 Sparse Blocked context onlyMay 29, 2026
- PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Sayarxiv-2606.00152 Sparse Blocked context onlyMay 29, 2026
- Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defensearxiv-2605.30837 Sparse Blocked context onlyMay 29, 2026
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPOarxiv-2605.30789 Sparse Blocked context onlyMay 29, 2026
- Codifying the Judge: Scalable Evaluation via Program Distillationarxiv-2607.22561 Sparse Blocked context onlyMay 29, 2026
- Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputsarxiv-2605.30611 Sparse Blocked context onlyMay 28, 2026
- Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gesturesarxiv-2605.30608 Sparse Blocked context onlyMay 28, 2026
- Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?arxiv-2605.30557 Sparse Blocked context onlyMay 28, 2026
- MAAT: Multi-phase Adapter-Aware Targeted Unlearningarxiv-2605.30514 Sparse Blocked context onlyMay 28, 2026
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMsarxiv-2605.30501 Sparse Blocked context onlyMay 28, 2026
- LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysisarxiv-2605.30434 Sparse Blocked context onlyMay 28, 2026
- DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representationarxiv-2605.30350 Sparse Blocked context onlyMay 28, 2026
- NeuROK: Generative 4D Neural Object Kinematicsarxiv-2605.30347 Sparse Blocked context onlyMay 28, 2026
- YoCausal: How Far is Video Generation from World Model? A Causality Perspectivearxiv-2605.30346 Sparse Blocked context onlyMay 28, 2026
- Locally Coherent, Globally Incoherent: Bounding Compositional Incoherence in Multi-Component LLM Agentsarxiv-2605.30335 Sparse Blocked context onlyMay 28, 2026
- Colored Noise Diffusion Samplingarxiv-2605.30332 Sparse Blocked context onlyMay 28, 2026
- Resolution Diagnostics for Paired LLM Evaluationarxiv-2605.30315 Sparse Blocked context onlyMay 28, 2026
- Self-Trained Verification for Training- and Test-Time Self-Improvementarxiv-2605.30290 Sparse Blocked context onlyMay 28, 2026
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodimentsarxiv-2605.30280 Sparse Blocked context onlyMay 28, 2026
- LoMo: Local Modality Substitution for Deeper Vision-Language Fusionarxiv-2605.30265 Sparse Blocked context onlyMay 28, 2026
- VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agentsarxiv-2605.30256 Sparse Blocked context onlyMay 28, 2026
- CommunityFact: A Dynamic, Multilingual, Multi-domain Benchmark for Misinformation Detection in the Wildarxiv-2605.30241 Sparse Blocked context onlyMay 28, 2026
- ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Researcharxiv-2606.07591 Sparse Blocked context onlyMay 28, 2026
- Why Far Looks Up: Probing Spatial Representation in Vision-Language Modelsarxiv-2605.30161 Sparse Blocked context onlyMay 28, 2026
- Do Proactive Agents Really Need an LLM to Decide When to Wake and What to Anchor?arxiv-2605.30152 Sparse Blocked context onlyMay 28, 2026
- CCS: Clinical Consensus Selection for Radiology Report Generationarxiv-2605.30131 Sparse Blocked context onlyMay 28, 2026
- PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understandingarxiv-2605.30126 Sparse Blocked context onlyMay 28, 2026
- DirectorBench: Diagnosing Long-Form Video Generation with Personalized Multi-Agent Evaluationarxiv-2605.30090 Sparse Blocked context onlyMay 28, 2026
- Multimodal Music Recommendation System using LLMsarxiv-2606.00125 Sparse Blocked context onlyMay 28, 2026
- Conformal Certification of Reasoning Trace Prefixesarxiv-2605.30085 Sparse Blocked context onlyMay 28, 2026
- Native Audio-Visual Alignment for Generationarxiv-2605.30073 Sparse Blocked context onlyMay 28, 2026
- Audio Jailbreaks in Large Audio-Language Models: Taxonomy, Attack-Defense Analysis, and Cost-Aware Evaluationarxiv-2605.30031 Sparse Blocked context onlyMay 28, 2026
- Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmasarxiv-2605.30003 Sparse Blocked context onlyMay 28, 2026
- Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillationarxiv-2605.29992 Curated Related Blocked context onlyMay 28, 2026
- MIC: Maximizing Informational Capacity in Adaptive Representations via Isotropic Subspace Alignmentarxiv-2605.29987 Sparse Blocked context onlyMay 28, 2026
- MuPHI: Learning Implicit Multimodal Harm Reasoning via Semantically Grounded Reward Optimizationarxiv-2605.29951 Sparse Blocked context onlyMay 28, 2026
- Does The Way You Plan Matter? An Empirical Study of Planning Representations for LLM Web Agentsarxiv-2605.29927 Sparse Blocked context onlyMay 28, 2026
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Securityarxiv-2605.29801 Sparse Blocked context onlyMay 28, 2026
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panelsarxiv-2605.29800 Sparse Blocked context onlyMay 28, 2026
- Scaling Laws for Agent Harnesses via Effective Feedback Computearxiv-2605.29682 Sparse Blocked context onlyMay 28, 2026
- UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agentsarxiv-2605.29534 Sparse Blocked context onlyMay 28, 2026
- Xetrieval: Mechanistically Explaining Dense Retrievalarxiv-2605.29507 Sparse Blocked context onlyMay 28, 2026
- AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modelingarxiv-2605.29488 Sparse Blocked context onlyMay 28, 2026
- Honest Lying: Understanding Memory Confabulation in Reflexive Agentsarxiv-2605.29463 Sparse Blocked context onlyMay 28, 2026
- Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agentsarxiv-2605.29447 Sparse Blocked context onlyMay 28, 2026
- Learning Design Skills as Memory Policies for Agentic Photonic Inverse Designarxiv-2605.29421 Sparse Blocked context onlyMay 28, 2026
- Revisiting Observation Reduction for Web Agents: Comprehensive Evaluation with a Lightweight Frameworkarxiv-2605.29397 Sparse Blocked context onlyMay 28, 2026
- Offloading Score: Measuring AI Reliance Through Counterfactual Workflowsarxiv-2605.29392 Sparse Blocked context onlyMay 28, 2026
- Draft-OPD: On-Policy Distillation for Speculative Draft Modelsarxiv-2605.29343 Sparse Blocked context onlyMay 28, 2026