300 canonical paper links on this archive page.
- Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoningarxiv-2608.02831 Sparse Blocked context onlyAug 3, 2026
- Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Research Agentsarxiv-2608.02751 Sparse Blocked context onlyAug 3, 2026
- Quo Vadis, World Modeling?arxiv-2608.02713 Sparse Blocked context onlyAug 3, 2026
- CAPEval: A Decoupled Caption Evaluation across Understanding and Generationarxiv-2608.02589 Sparse Blocked context onlyAug 3, 2026
- UEmbed: Unified Sparse and Dense Multimodal Embeddingsarxiv-2608.02583 Curated Related Blocked context onlyAug 3, 2026
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Dataarxiv-2608.02580 Sparse Blocked context onlyAug 3, 2026
- SWE-Touch: Benchmarking Coding Agents When Users Touch the Codearxiv-2608.02499 Sparse Blocked context onlyAug 3, 2026
- InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesisarxiv-2608.02437 Sparse Blocked context onlyAug 3, 2026
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experiencearxiv-2608.02392 Sparse Blocked context onlyAug 3, 2026
- SKT: Skill-Use Training at Scale via Verified Synthetic Data Generationarxiv-2608.02287 Sparse Blocked context onlyAug 3, 2026
- PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputsarxiv-2608.02218 Sparse Blocked context onlyAug 3, 2026
- Douyin Multimodal Embedding Model Technical Reportarxiv-2608.02148 Sparse Blocked context onlyAug 3, 2026
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skillsarxiv-2608.01851 Sparse Blocked context onlyAug 3, 2026
- DAPD: Dual-Anchored Policy Distillationarxiv-2608.01735 Sparse Blocked context onlyAug 3, 2026
- Progressive Agent Skill Generation via Reinforcement Learningarxiv-2608.01678 Sparse Blocked context onlyAug 3, 2026
- Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AIarxiv-2608.01462 Sparse Blocked context onlyAug 2, 2026
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Spacearxiv-2608.01397 Sparse Blocked context onlyAug 2, 2026
- Prompt-Induced Waste in Coding Agents: Reasoning, Effort, Harness Design, and End-to-End Costarxiv-2608.01347 Sparse Blocked context onlyAug 2, 2026
- RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Evictionarxiv-2608.01247 Sparse Blocked context onlyAug 2, 2026
- Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systemsarxiv-2608.14667 Sparse Blocked context onlyAug 2, 2026
- HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inferencearxiv-2608.00577 Sparse Blocked context onlyAug 1, 2026
- Scaling Properties of Text Conditioning in Visual Generationarxiv-2607.29679 Sparse Blocked context onlyJul 31, 2026
- RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systemsarxiv-2607.29241 Sparse Blocked context onlyJul 31, 2026
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failuresarxiv-2607.28802 Sparse Blocked context onlyJul 30, 2026
- AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesisarxiv-2607.28618 Sparse Blocked context onlyJul 30, 2026
- QQWorld: Quantile-Quantile Matching for World Model Regularizationarxiv-2607.28415 Sparse Blocked context onlyJul 30, 2026
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledgerarxiv-2607.28374 Sparse Blocked context onlyJul 30, 2026
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadowarxiv-2607.28362 Sparse Blocked context onlyJul 30, 2026
- Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts Routingarxiv-2607.28308 Sparse Blocked context onlyJul 30, 2026
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agentsarxiv-2607.28227 Sparse Blocked context onlyJul 30, 2026
- Echoverse: Deep, Evolving Environments for Training Computer-Use Agents at Scalearxiv-2607.28074 Sparse Blocked context onlyJul 30, 2026
- Flux-OPD: On-Policy Distillation with Evolving Contextsarxiv-2607.28022 Sparse Blocked context onlyJul 30, 2026
- Complementary Matrix-Gated QKAN Fast-Weight Programmers for Quantum Dynamics Forecastingarxiv-2607.27945 Sparse Blocked context onlyJul 30, 2026
- FinanceHarness: Autonomous Financial Deep Research Frameworkarxiv-2607.27853 Sparse Blocked context onlyJul 30, 2026
- Articulated Object Reconstruction from Rest-State Observationarxiv-2607.27749 Sparse Blocked context onlyJul 30, 2026
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzlesarxiv-2607.27670 Sparse Blocked context onlyJul 30, 2026
- Harness-G: A Graph-Structured Harness for Search Agentsarxiv-2607.27652 Sparse Blocked context onlyJul 30, 2026
- Mental World Modelingarxiv-2607.27201 Sparse Blocked context onlyJul 29, 2026
- HumanCLAW: Can Vision-Language Models Act Through a Body?arxiv-2607.27180 Sparse Blocked context onlyJul 29, 2026
- LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generationarxiv-2608.00079 Sparse Blocked context onlyJul 29, 2026
- ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figuresarxiv-2607.26848 Sparse Blocked context onlyJul 29, 2026
- Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainabilityarxiv-2607.26637 Sparse Blocked context onlyJul 29, 2026
- Pass the Baton: Trajectory-Relayed On-Policy Distillationarxiv-2607.26057 Sparse Blocked context onlyJul 28, 2026
- Parallel Decoding Distillation for Fast Image and Video Generationarxiv-2607.26004 Curated Related Blocked context onlyJul 28, 2026
- GPT-Red: Automated Red Teaming via Self-Play at Scalearxiv-2607.26115 Sparse Blocked context onlyJul 28, 2026
- Shieldstralarxiv-2607.25857 Sparse Blocked context onlyJul 28, 2026
- CodeNib: A Multi-View Data System for Serving Repository Context to Coding Agentsarxiv-2607.25431 Sparse Blocked context onlyJul 28, 2026
- Temporal-Distance JEPA: Plan-Aware Representation Learning for Latent World Model Predictive Controlarxiv-2607.25337 Sparse Blocked context onlyJul 28, 2026
- CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisitionarxiv-2607.25294 Sparse Blocked context onlyJul 28, 2026
- AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognitionarxiv-2607.25289 Sparse Blocked context onlyJul 28, 2026
- OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysisarxiv-2607.25108 Sparse Blocked context onlyJul 27, 2026
- Data Pyramid for Embodied Manipulationarxiv-2607.24744 Sparse Blocked context onlyJul 27, 2026
- Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillationarxiv-2607.24731 Sparse Blocked context onlyJul 27, 2026
- DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipesarxiv-2607.24516 Sparse Blocked context onlyJul 27, 2026
- OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generationarxiv-2607.23855 Curated Related Blocked context onlyJul 26, 2026
- DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identificationarxiv-2607.23822 Sparse Blocked context onlyJul 26, 2026
- A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Foreverarxiv-2607.23806 Sparse Blocked context onlyJul 26, 2026
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KVarxiv-2607.23693 Sparse Blocked context onlyJul 26, 2026
- Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checkingarxiv-2607.23514 Sparse Blocked context onlyJul 26, 2026
- SceneActBench: Can Agents Act on the 3D Scenes They See?arxiv-2607.22393 Sparse Blocked context onlyJul 24, 2026
- Projection Pursuit CPCANet for Domain Generalizationarxiv-2607.22117 Sparse Blocked context onlyJul 24, 2026
- Spectral Prior for Reducing Exposure Bias in Diffusion Modelsarxiv-2607.22091 Sparse Blocked context onlyJul 24, 2026
- LAMAR: An Open Language-Aware Multilingual Alignment Rerankerarxiv-2607.22042 Curated Related Blocked context onlyJul 24, 2026
- AREX: Towards a Recursively Self-Improving Agent for Deep Researcharxiv-2607.21461 Sparse Blocked context onlyJul 23, 2026
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Constructionarxiv-2607.20911 Sparse Blocked context onlyJul 23, 2026
- Is Deep Research Reliable? Misleading Knowledge Induces False Conclusionsarxiv-2607.20891 Direct Blocked context onlyJul 23, 2026
- LLMs Get Lost in Evolving User Intentarxiv-2607.20734 Sparse Blocked context onlyJul 22, 2026
- Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanationsarxiv-2607.20379 Sparse Blocked context onlyJul 22, 2026
- SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodimentsarxiv-2607.20207 Sparse Blocked context onlyJul 22, 2026
- ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Modelsarxiv-2607.20092 Sparse Blocked context onlyJul 22, 2026
- ReferTrack: Referring Then Tracking for Embodied Visual Trackingarxiv-2607.20061 Sparse Blocked context onlyJul 22, 2026
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operationsarxiv-2607.19865 Sparse Blocked context onlyJul 22, 2026
- How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHFarxiv-2607.19712 Sparse Blocked context onlyJul 22, 2026
- Multimodal Speaker Verification as a Threat to Speaker Anonymizationarxiv-2607.19636 Sparse Blocked context onlyJul 22, 2026
- Masked Visual Actions for Unified World Modelingarxiv-2607.19343 Sparse Blocked context onlyJul 21, 2026
- Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completenessarxiv-2607.19322 Sparse Blocked context onlyJul 21, 2026
- FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documentsarxiv-2607.19238 Sparse Blocked context onlyJul 21, 2026
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPUarxiv-2607.19191 Sparse Blocked context onlyJul 21, 2026
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challengesarxiv-2607.19011 Sparse Blocked context onlyJul 21, 2026
- Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timingarxiv-2607.18934 Sparse Blocked context onlyJul 21, 2026
- HPD-Parsing: Hierarchical Parallel Document Parsingarxiv-2607.18839 Sparse Blocked context onlyJul 21, 2026
- Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learningarxiv-2607.18722 Sparse Blocked context onlyJul 21, 2026
- Generative World Renderer at the Speed of Playarxiv-2607.18703 Sparse Blocked context onlyJul 21, 2026
- AutoIndex: Learning Representation Programs for Retrievalarxiv-2607.18603 Sparse Blocked context onlyJul 21, 2026
- AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Researcharxiv-2608.11216 Sparse Blocked context onlyJul 20, 2026
- HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancementarxiv-2607.18217 Sparse Blocked context onlyJul 20, 2026
- FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applicationsarxiv-2607.18171 Sparse Blocked context onlyJul 20, 2026
- O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoningarxiv-2607.18142 Sparse Blocked context onlyJul 20, 2026
- Enhancing Rubric-based RL via Self-Distillationarxiv-2607.18082 Sparse Blocked context onlyJul 20, 2026
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Modelarxiv-2607.17977 Sparse Blocked context onlyJul 20, 2026
- ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Videoarxiv-2607.17790 Sparse Blocked context onlyJul 20, 2026
- Uncovering Latent Reasoning Strategies in Language Modelsarxiv-2607.17674 Sparse Blocked context onlyJul 20, 2026
- EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary Worldarxiv-2607.17250 Sparse Blocked context onlyJul 19, 2026
- Nonuniformity Principle in Human-AI Coworkingarxiv-2607.16530 Sparse Blocked context onlyJul 17, 2026
- Interactive Training 2: Auditable Control Plane for Live Model Trainingarxiv-2607.18314 Sparse Blocked context onlyJul 17, 2026
- Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligencearxiv-2607.16401 Sparse Blocked context onlyJul 17, 2026
- When Does Muon Help Agentic Reinforcement Learning?arxiv-2607.16169 Sparse Blocked context onlyJul 17, 2026
- JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Modelsarxiv-2607.16074 Sparse Blocked context onlyJul 17, 2026
- Loop the Loopies!arxiv-2607.16051 Sparse Blocked context onlyJul 17, 2026
- In the Driver's Seat: A Multi-Company Study on the Reality of Autonomous Driving System Testingarxiv-2607.15820 Sparse Blocked context onlyJul 17, 2026
- Recursive Harness Self-Improvementarxiv-2607.15524 Sparse Blocked context onlyJul 17, 2026
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalationarxiv-2607.15434 Sparse Blocked context onlyJul 16, 2026
- RoboTTT: Context Scaling for Robot Policiesarxiv-2607.15275 Sparse Blocked context onlyJul 16, 2026
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agentsarxiv-2607.15263 Sparse Blocked context onlyJul 16, 2026
- BadWAM: When World-Action Models Dream Right but Act Wrongarxiv-2607.15207 Sparse Blocked context onlyJul 16, 2026
- Video = World + Event Streamarxiv-2607.15038 Sparse Blocked context onlyJul 16, 2026
- LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budgetarxiv-2607.14952 Sparse Blocked context onlyJul 16, 2026
- Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheelarxiv-2607.14431 Sparse Blocked context onlyJul 15, 2026
- Cura 1T: Specialized Model for Agentic Healthcarearxiv-2607.15314 Sparse Blocked context onlyJul 15, 2026
- From Pixels to States: Rethinking Interactive World Models as Game Enginesarxiv-2607.14076 Sparse Blocked context onlyJul 15, 2026
- Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Codearxiv-2607.13921 Sparse Blocked context onlyJul 15, 2026
- Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learningarxiv-2607.14183 Sparse Blocked context onlyJul 15, 2026
- SPyCE: Skill-Policy Co-evolution for Multimodal Agentsarxiv-2607.13854 Sparse Blocked context onlyJul 15, 2026
- Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignmentarxiv-2607.13429 Sparse Blocked context onlyJul 15, 2026
- AffectFlow-DINO: Uncertainty-Aware Multi-Task Affect Estimation via Conditional Rectified Flowarxiv-2607.13250 Sparse Blocked context onlyJul 14, 2026
- ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillationarxiv-2607.13124 Sparse Blocked context onlyJul 14, 2026
- UniVR: Thinking in Visual Space for Unified Visual Reasoningarxiv-2607.12800 Sparse Blocked context onlyJul 14, 2026
- Tracing Agentic Failure from the Flow of Successarxiv-2607.12747 Sparse Blocked context onlyJul 14, 2026
- KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skillarxiv-2607.12625 Sparse Blocked context onlyJul 14, 2026
- Self-Improvements in Modern Agentic Systems: A Surveyarxiv-2607.13104 Sparse Blocked context onlyJul 14, 2026
- Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Modelsarxiv-2607.12463 Sparse Blocked context onlyJul 14, 2026
- Let RGB Be the Language of Visionarxiv-2607.12450 Sparse Blocked context onlyJul 14, 2026
- LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakesarxiv-2607.12310 Sparse Blocked context onlyJul 14, 2026
- Token Reduction Is Not Cost Reductionarxiv-2607.12161 Sparse Blocked context onlyJul 13, 2026
- Vinci2: Providing Proactive Assistance in Continuous Egocentric Videosarxiv-2607.11523 Sparse Blocked context onlyJul 13, 2026
- Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientistsarxiv-2607.11079 Sparse Blocked context onlyJul 13, 2026
- A Vocabulary for Multi-Agent Automated Research Systemsarxiv-2607.22682 Sparse Blocked context onlyJul 13, 2026
- SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learningarxiv-2607.10966 Sparse Blocked context onlyJul 13, 2026
- PanoWorld: Real-World Panoramic Generationarxiv-2607.09661 Sparse Blocked context onlyJul 10, 2026
- OpenLongTail: Generative Scaling of Long-Tail Driving Dataarxiv-2607.09655 Sparse Blocked context onlyJul 10, 2026
- A Sovereign, Open-Source Foundation Model for German and Englisharxiv-2607.09424 Sparse Blocked context onlyJul 10, 2026
- LongE2V: Long-Horizon Event-based Video Reconstruction, Prediction, and Frame Interpolation with Video Diffusion Modelsarxiv-2607.08770 Sparse Blocked context onlyJul 9, 2026
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generationarxiv-2607.08758 Sparse Blocked context onlyJul 9, 2026
- Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuningarxiv-2607.08393 Sparse Blocked context onlyJul 9, 2026
- Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Modelsarxiv-2607.08317 Sparse Blocked context onlyJul 9, 2026
- Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routingarxiv-2607.07953 Sparse Blocked context onlyJul 8, 2026
- DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environmentarxiv-2607.07820 Sparse Blocked context onlyJul 8, 2026
- Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoningarxiv-2607.07708 Sparse Blocked context onlyJul 8, 2026
- Sparse Delta Memory: Scaling the State of Linear RNNs through Sparsityarxiv-2607.07386 Sparse Blocked context onlyJul 8, 2026
- Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPEarxiv-2607.07740 Sparse Blocked context onlyJul 8, 2026
- Flow-ERD: Agent-type Aware Flow Matching with Entropy-Regularized Distillation for Diverse Traffic Simulationarxiv-2607.06957 Sparse Blocked context onlyJul 8, 2026
- WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligencearxiv-2607.06838 Sparse Blocked context onlyJul 7, 2026
- Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized Policiesarxiv-2607.06815 Sparse Blocked context onlyJul 7, 2026
- SPEAR: A Simulator for Photorealistic Embodied AI Researcharxiv-2607.06701 Sparse Blocked context onlyJul 7, 2026
- From Foundation to Application: Improving VLA Models in Practicearxiv-2607.06403 Sparse Blocked context onlyJul 7, 2026
- VaseMuseum: Digital Intelligent Museum for Ancient Greek Potteryarxiv-2607.06374 Sparse Blocked context onlyJul 7, 2026
- SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Reviewarxiv-2607.06065 Sparse Blocked context onlyJul 7, 2026
- RoboTALES: Learning Reasoning-Guided Robot Policies via Task-Aligned Simulated Futuresarxiv-2607.06018 Sparse Blocked context onlyJul 7, 2026
- Imagined Rollouts are Kinematic, Not Dynamic: A Diagnosis of Long-Horizon World-Model Failurearxiv-2607.05966 Sparse Blocked context onlyJul 7, 2026
- Where to cut, how deep: BPE and Unigram-LM on chemistry SMILESarxiv-2607.05691 Sparse Blocked context onlyJul 6, 2026
- Weak-to-Strong Generalization via Direct On-Policy Distillationarxiv-2607.05394 Sparse Blocked context onlyJul 6, 2026
- Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Modelsarxiv-2607.05390 Sparse Blocked context onlyJul 6, 2026
- Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generationarxiv-2607.05382 Sparse Blocked context onlyJul 6, 2026
- PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Spacearxiv-2607.05373 Sparse Blocked context onlyJul 6, 2026
- GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasksarxiv-2607.05369 Sparse Blocked context onlyJul 6, 2026
- Multiplayer Interactive World Models with Representation Autoencodersarxiv-2607.05352 Sparse Blocked context onlyJul 6, 2026
- TREK: Distill to Explore, Reinforce to Refinearxiv-2607.05339 Sparse Blocked context onlyJul 6, 2026
- EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environmentsarxiv-2607.05155 Sparse Blocked context onlyJul 6, 2026
- KVpop -- Key-Value Cache Compression with Predictive Online Pruningarxiv-2607.05061 Sparse Blocked context onlyJul 6, 2026
- Trust Region Policy Distillationarxiv-2607.04751 Sparse Blocked context onlyJul 6, 2026
- CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestrationarxiv-2607.05465 Sparse Blocked context onlyJul 6, 2026
- Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrievalarxiv-2607.04605 Sparse Blocked context onlyJul 6, 2026
- Wan-Streamer v0.2: Higher Resolution, Same Latencyarxiv-2607.04443 Sparse Blocked context onlyJul 5, 2026
- AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memesarxiv-2607.04410 Sparse Blocked context onlyJul 5, 2026
- Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMsarxiv-2607.03936 Sparse Blocked context onlyJul 4, 2026
- MentalThink: Shaping Thoughts in Mental SVG Worldarxiv-2607.03530 Sparse Blocked context onlyJul 3, 2026
- Perceptual Flow Matching for Few-Step Generative Modelingarxiv-2607.03524 Sparse Blocked context onlyJul 3, 2026
- Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignmentarxiv-2607.02920 Sparse Blocked context onlyJul 3, 2026
- Gemma 4 Technical Reportarxiv-2607.02770 Sparse Blocked context onlyJul 2, 2026
- Online Safety Monitoring for LLMsarxiv-2607.02510 Sparse Blocked context onlyJul 2, 2026
- What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debatesarxiv-2607.02507 Sparse Blocked context onlyJul 2, 2026
- Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robotsarxiv-2607.02501 Sparse Blocked context onlyJul 2, 2026
- Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignmentarxiv-2607.02471 Sparse Blocked context onlyJul 2, 2026
- TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolutionarxiv-2607.02469 Sparse Blocked context onlyJul 2, 2026
- Language Models as Measurement Apparatus for Culturearxiv-2607.02459 Sparse Blocked context onlyJul 2, 2026
- EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environmentsarxiv-2607.02440 Sparse Blocked context onlyJul 2, 2026
- ACID: Action Consistency via Inverse Dynamics for Planning with World Modelsarxiv-2607.02403 Sparse Blocked context onlyJul 2, 2026
- Optimizing Visual Generative Models via Distribution-wise Rewardsarxiv-2607.02291 Sparse Blocked context onlyJul 2, 2026
- AGVBench: A Reliability-Oriented Benchmark of Data Augmentation for Vein Recognitionarxiv-2607.02271 Sparse Blocked context onlyJul 2, 2026
- AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Modelsarxiv-2607.02269 Sparse Blocked context onlyJul 2, 2026
- AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agentsarxiv-2607.02255 Sparse Blocked context onlyJul 2, 2026
- PACE: A Proxy for Agentic Capability Evaluationarxiv-2607.02032 Sparse Blocked context onlyJul 2, 2026
- PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulationarxiv-2607.01938 Sparse Blocked context onlyJul 2, 2026
- TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27Barxiv-2607.01927 Sparse Blocked context onlyJul 2, 2026
- Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Draftersarxiv-2607.01893 Sparse Blocked context onlyJul 2, 2026
- SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Usearxiv-2607.01874 Sparse Blocked context onlyJul 2, 2026
- On the Limits of Steering Vectors for Preference-Aligned Generationarxiv-2607.01802 Sparse Blocked context onlyJul 2, 2026
- Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verificationarxiv-2607.01793 Sparse Blocked context onlyJul 2, 2026
- Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproductionarxiv-2607.01764 Sparse Blocked context onlyJul 2, 2026
- Multi-Head Recurrent Memory Agentsarxiv-2607.01523 Sparse Blocked context onlyJul 1, 2026
- RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluationarxiv-2607.01388 Sparse Blocked context onlyJul 1, 2026
- AutoMem: Automated Learning of Memory as a Cognitive Skillarxiv-2607.01224 Sparse Blocked context onlyJul 1, 2026
- Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?arxiv-2607.01211 Sparse Blocked context onlyJul 1, 2026
- Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrationsarxiv-2607.01181 Sparse Blocked context onlyJul 1, 2026
- AGC-Bench: Measuring Artificial General Creativityarxiv-2607.01152 Sparse Blocked context onlyJul 1, 2026
- ESC: Emotional Self-Correction for Reliable Vision-Language Modelsarxiv-2607.02089 Sparse Blocked context onlyJul 1, 2026
- Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documentsarxiv-2607.00895 Sparse Blocked context onlyJul 1, 2026
- Self-Evolving Agents with Anytime-Valid Certificatesarxiv-2607.00871 Sparse Blocked context onlyJul 1, 2026
- CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Modelsarxiv-2607.00862 Sparse Blocked context onlyJul 1, 2026
- MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmarkarxiv-2607.00724 Sparse Blocked context onlyJul 1, 2026
- Self-conditioned Flow Map Language Models via Fixed-point Flowsarxiv-2607.00714 Sparse Blocked context onlyJul 1, 2026
- Domain Arithmetic: One-Shot VLA Adaptation under Environmental Shiftsarxiv-2607.00666 Sparse Blocked context onlyJul 1, 2026
- Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explicationsarxiv-2607.00661 Sparse Blocked context onlyJul 1, 2026
- ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Servingarxiv-2607.00466 Sparse Blocked context onlyJul 1, 2026
- MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Moleculesarxiv-2607.00464 Sparse Blocked context onlyJul 1, 2026
- VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinementarxiv-2607.00446 Sparse Blocked context onlyJul 1, 2026
- When Classic Cache Policies Fail: Learning-Augmented Replacement for Semantic Retrieval Buffersarxiv-2607.00394 Sparse Blocked context onlyJul 1, 2026
- TRACE: State-Aware Query Processing over Temporal Evidence Graphs for Conversational Dataarxiv-2607.00339 Sparse Blocked context onlyJul 1, 2026
- Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditionsarxiv-2607.00304 Sparse Blocked context onlyJul 1, 2026
- Wake up for Touch! Mask-isolated Tactile Alignment Learning in MLLMsarxiv-2607.00302 Sparse Blocked context onlyJul 1, 2026
- EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systemsarxiv-2607.00297 Sparse Blocked context onlyJul 1, 2026
- ASPIRE: Agentic /Skills Discovery for Roboticsarxiv-2607.00272 Sparse Blocked context onlyJun 30, 2026
- PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seekingarxiv-2607.00115 Sparse Blocked context onlyJun 30, 2026
- QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agentsarxiv-2606.32034 Sparse Blocked context onlyJun 30, 2026
- TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learningarxiv-2606.32017 Sparse Blocked context onlyJun 30, 2026
- InstanceControl: Controllable Complex Image Generation without Instance Labelingarxiv-2606.31924 Sparse Blocked context onlyJun 30, 2026
- MemLearner: Learning to Query Context memory for Video World Modelsarxiv-2606.31734 Sparse Blocked context onlyJun 30, 2026
- Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialoguearxiv-2606.31719 Sparse Blocked context onlyJun 30, 2026
- AutoTrainess: Teaching Language Models to Improve Language Models Autonomouslyarxiv-2606.31551 Sparse Blocked context onlyJun 30, 2026
- Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answeringarxiv-2606.31432 Sparse Blocked context onlyJun 30, 2026
- 3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidancearxiv-2606.31329 Sparse Blocked context onlyJun 30, 2026
- BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decodingarxiv-2606.31315 Sparse Blocked context onlyJun 30, 2026
- AtomiMed: Hierarchical Atomic Fact-Checking for Universal Clinical-Aware Medical Report Evaluationarxiv-2606.31292 Sparse Blocked context onlyJun 30, 2026
- FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Modelarxiv-2606.31247 Sparse Blocked context onlyJun 30, 2026
- Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Carearxiv-2606.31036 Sparse Blocked context onlyJun 30, 2026
- AVTok: 1D Unified Tokenization for Holistic Audio-Video Generationarxiv-2606.30811 Sparse Blocked context onlyJun 29, 2026
- Morphing into Hybrid Attention Modelsarxiv-2606.30562 Sparse Blocked context onlyJun 29, 2026
- BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Languagearxiv-2606.30319 Sparse Blocked context onlyJun 29, 2026
- The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstructionarxiv-2606.30308 Sparse Blocked context onlyJun 29, 2026
- DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Modelarxiv-2606.30292 Sparse Blocked context onlyJun 29, 2026
- When Is a Draft Accepted? A Theory of Acceptance in Speculative Decodingarxiv-2606.30265 Sparse Blocked context onlyJun 29, 2026
- SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generationarxiv-2606.30201 Sparse Blocked context onlyJun 29, 2026
- DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoningarxiv-2606.30189 Sparse Blocked context onlyJun 29, 2026
- Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Mattersarxiv-2606.30128 Sparse Blocked context onlyJun 29, 2026
- Automating the Design of Embodied Agent Architecturesarxiv-2606.30111 Sparse Blocked context onlyJun 29, 2026
- Illuminating Unified Multimodal Model for Free-form Interleaved Text-Image Generationarxiv-2606.30054 Curated Related Blocked context onlyJun 29, 2026
- Monte Carlo Energy Aggregation for Mobile 3D Gaussian Splattingarxiv-2606.30017 Sparse Blocked context onlyJun 29, 2026
- SWE-Together: Evaluating Coding Agents in Interactive User Sessionsarxiv-2606.29957 Sparse Blocked context onlyJun 29, 2026
- Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?arxiv-2606.29920 Sparse Blocked context onlyJun 29, 2026
- MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluationarxiv-2606.29914 Sparse Blocked context onlyJun 29, 2026
- SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailingarxiv-2606.29887 Sparse Blocked context onlyJun 29, 2026
- A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agentsarxiv-2606.29719 Sparse Blocked context onlyJun 29, 2026
- Geometric Stability of Neural Population Codes: Regional Variation, Behavioral Relevance, and Circuit Dependencearxiv-2606.29655 Sparse Blocked context onlyJun 28, 2026
- Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languagesarxiv-2606.29649 Sparse Blocked context onlyJun 28, 2026
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resourcesarxiv-2606.29538 Sparse Blocked context onlyJun 28, 2026
- UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillationarxiv-2606.29502 Sparse Blocked context onlyJun 28, 2026
- Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillationarxiv-2606.29464 Sparse Blocked context onlyJun 28, 2026
- Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agentsarxiv-2606.29459 Sparse Blocked context onlyJun 28, 2026
- Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoningarxiv-2606.29425 Sparse Blocked context onlyJun 28, 2026
- The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scalingarxiv-2606.29278 Sparse Blocked context onlyJun 28, 2026
- PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agentsarxiv-2606.29225 Sparse Blocked context onlyJun 28, 2026
- Selective Memory Retention for Long-Horizon LLM Agentsarxiv-2606.29178 Sparse Blocked context onlyJun 28, 2026
- ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphsarxiv-2606.29067 Sparse Blocked context onlyJun 27, 2026
- Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification Systemarxiv-2606.28953 Sparse Blocked context onlyJun 27, 2026
- When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scalingarxiv-2606.28661 Sparse Blocked context onlyJun 27, 2026
- SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervisionarxiv-2606.28562 Sparse Blocked context onlyJun 26, 2026
- A Gravitational Interpretation of Fine-Tuning Reversionarxiv-2606.28525 Sparse Blocked context onlyJun 26, 2026
- NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learningarxiv-2606.27771 Sparse Blocked context onlyJun 26, 2026
- DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forumsarxiv-2606.27619 Sparse Blocked context onlyJun 26, 2026
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environmentsarxiv-2606.27537 Sparse Blocked context onlyJun 25, 2026
- DanceOPD: On-Policy Generative Field Distillationarxiv-2606.27377 Sparse Blocked context onlyJun 25, 2026
- Hallucination in World Models is Predictable and Preventablearxiv-2606.27326 Sparse Blocked context onlyJun 25, 2026
- Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)arxiv-2606.27163 Sparse Blocked context onlyJun 25, 2026
- How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoringarxiv-2606.26979 Sparse Blocked context onlyJun 25, 2026
- To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repairarxiv-2606.26978 Sparse Blocked context onlyJun 25, 2026
- Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoEarxiv-2606.26938 Sparse Blocked context onlyJun 25, 2026
- HyperDFlash: Hyper-Connection-Aligned Block Speculative Decoding with Gated Residual Reductionarxiv-2606.26744 Sparse Blocked context onlyJun 25, 2026
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editingarxiv-2606.26740 Sparse Blocked context onlyJun 25, 2026
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv-2606.26300 Sparse Blocked context onlyJun 24, 2026
- Fast LeWorldModelarxiv-2606.26217 Sparse Blocked context onlyJun 24, 2026
- Autodata: An agentic data scientist to create high quality synthetic dataarxiv-2606.25996 Sparse Blocked context onlyJun 24, 2026
- Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimizationarxiv-2606.25656 Sparse Blocked context onlyJun 24, 2026
- BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agentsarxiv-2606.25556 Sparse Blocked context onlyJun 24, 2026
- Fully Differentiable Neural Forced Alignment via Soft Dynamic Programmingarxiv-2606.25460 Sparse Blocked context onlyJun 24, 2026
- The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithmsarxiv-2606.25450 Sparse Blocked context onlyJun 24, 2026
- The Interplay of Harness Design and Post-Training in LLM Agentsarxiv-2606.25447 Sparse Blocked context onlyJun 24, 2026
- Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agentsarxiv-2606.25361 Sparse Blocked context onlyJun 24, 2026
- RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memoryarxiv-2606.25206 Sparse Blocked context onlyJun 23, 2026
- Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVRarxiv-2606.25178 Sparse Blocked context onlyJun 23, 2026
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Modelsarxiv-2606.25041 Sparse Blocked context onlyJun 23, 2026
- InSight: Self-Guided Skill Acquisition via Steerable VLAsarxiv-2606.24884 Sparse Blocked context onlyJun 23, 2026
- FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generationarxiv-2606.24876 Sparse Blocked context onlyJun 23, 2026
- World Value Models for Robotic Manipulationarxiv-2606.24742 Sparse Blocked context onlyJun 23, 2026
- MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recoveryarxiv-2606.24595 Sparse Blocked context onlyJun 23, 2026
- Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrievalarxiv-2606.24976 Sparse Blocked context onlyJun 23, 2026
- Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matchingarxiv-2606.24457 Sparse Blocked context onlyJun 23, 2026
- Trimming the Long-Tail of Visual World Modeling Evaluationarxiv-2606.24256 Sparse Blocked context onlyJun 23, 2026
- FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planningarxiv-2606.24231 Sparse Blocked context onlyJun 23, 2026
- MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrievalarxiv-2606.24200 Sparse Blocked context onlyJun 23, 2026
- BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasksarxiv-2606.24162 Sparse Blocked context onlyJun 23, 2026
- MedBench v5: A Dynamic, Process-Oriented, and Hallucination-Aware Benchmark for Clinical Multimodal Modelsarxiv-2606.24155 Sparse Blocked context onlyJun 23, 2026
- Metis: Bridging Text and Code Memory for Self-Evolving Agentsarxiv-2606.24151 Sparse Blocked context onlyJun 23, 2026
- ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detectionarxiv-2606.24112 Sparse Blocked context onlyJun 23, 2026
- Blockwise Policy-Drift Gating for On-Policy Distillationarxiv-2606.24084 Sparse Blocked context onlyJun 23, 2026
- RoPE-Aware Bit Allocation for KV-Cache Quantizationarxiv-2606.24033 Sparse Blocked context onlyJun 23, 2026
- TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMsarxiv-2607.22639 Sparse Blocked context onlyJun 22, 2026
- ChartWalker: Benchmarking the Cross-Chart RAG Taskarxiv-2606.23997 Sparse Blocked context onlyJun 22, 2026
- Learning to Trigger: Reinforcement Learning at the Large Hadron Colliderarxiv-2606.23993 Sparse Blocked context onlyJun 22, 2026
- ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generationarxiv-2606.23835 Sparse Blocked context onlyJun 22, 2026
- Semantic Browsing: Controllable Diversity for Image Generationarxiv-2606.23679 Sparse Blocked context onlyJun 22, 2026