300 canonical paper links on this archive page.
- Rethinking Stepwise Model Routing: A Cost-Efficient Table Reasoning Perspectivearxiv-2605.29319 Sparse Blocked context onlyMay 28, 2026
- CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrievalarxiv-2605.29271 Sparse Blocked context onlyMay 28, 2026
- Guidance Contrastive Token Credit Assignment for Discrete Policy Optimizationarxiv-2605.29198 Sparse Blocked context onlyMay 28, 2026
- The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressurearxiv-2605.29087 Sparse Blocked context onlyMay 27, 2026
- FRAPPE: Full Input, Residual Output Autoencoding with Projection Pursuit Encoderarxiv-2605.28992 Sparse Blocked context onlyMay 27, 2026
- Self-Improving Language Models with Bidirectional Evolutionary Searcharxiv-2605.28814 Sparse Blocked context onlyMay 27, 2026
- Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agentsarxiv-2605.28775 Sparse Blocked context onlyMay 27, 2026
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoningarxiv-2605.28774 Curated Related Blocked context onlyMay 27, 2026
- CubePart: An Open-Vocabulary Part-Controllable 3D Generatorarxiv-2605.28763 Sparse Blocked context onlyMay 27, 2026
- CORE: Contrastive Reflection Enables Rapid Improvements in Reasoningarxiv-2605.28742 Sparse Blocked context onlyMay 27, 2026
- LACUNA: Safe Agents as Recursive Program Holesarxiv-2605.28617 Sparse Blocked context onlyMay 27, 2026
- Compositional Text-to-Image Generation Via Region-aware Bimodal Direct Preference Optimizationarxiv-2605.28615 Sparse Blocked context onlyMay 27, 2026
- A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarksarxiv-2605.28556 Sparse Blocked context onlyMay 27, 2026
- GEM: Generative Supervision Helps Embodied Intelligencearxiv-2605.28548 Sparse Blocked context onlyMay 27, 2026
- On Compositional Learning Behaviours in Formal Mathematicsarxiv-2605.28512 Sparse Blocked context onlyMay 27, 2026
- Review Arcade: On the Human Alignment and Gameability of LLM Reviewsarxiv-2605.28897 Sparse Blocked context onlyMay 27, 2026
- Category-Level 3D Correspondence in Camera Space via Morphable Object Priorsarxiv-2605.28257 Sparse Blocked context onlyMay 27, 2026
- AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?arxiv-2605.28255 Sparse Blocked context onlyMay 27, 2026
- Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Modelsarxiv-2605.28132 Sparse Blocked context onlyMay 27, 2026
- Ask Now, Use Later: Benchmarking the Proactivity Gap in Long-Lived LLM Agentsarxiv-2605.28108 Sparse Blocked context onlyMay 27, 2026
- ResearchMath-14K: Scaling Research-Level Mathematics via Agentsarxiv-2605.28003 Sparse Blocked context onlyMay 27, 2026
- ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversationsarxiv-2605.27908 Sparse Blocked context onlyMay 27, 2026
- PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraftarxiv-2605.27762 Sparse Blocked context onlyMay 26, 2026
- SkillGrad: Optimizing Agent Skills Like Gradient Descentarxiv-2605.27760 Sparse Blocked context onlyMay 26, 2026
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player?arxiv-2605.27367 Sparse Blocked context onlyMay 26, 2026
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decodingarxiv-2605.27365 Curated Related Blocked context onlyMay 26, 2026
- Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Modelsarxiv-2605.27311 Sparse Blocked context onlyMay 26, 2026
- How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoningarxiv-2605.27310 Sparse Blocked context onlyMay 26, 2026
- Gemini Embedding 2: A Native Multimodal Embedding Model from Geminiarxiv-2605.27295 Sparse Blocked context onlyMay 26, 2026
- Less is More: Early Stopping Rollout for On-Policy Distillationarxiv-2605.27028 Sparse Blocked context onlyMay 26, 2026
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillationarxiv-2605.26844 Sparse Blocked context onlyMay 26, 2026
- Beyond Holistic Models: Systematic Component-level Benchmarking of Deep Multivariate Time-Series Forecastingarxiv-2605.26562 Sparse Blocked context onlyMay 26, 2026
- Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalizationarxiv-2605.26457 Sparse Blocked context onlyMay 26, 2026
- Advancing Creative Physical Intelligence in Large Multimodal Modelsarxiv-2605.26396 Sparse Blocked context onlyMay 25, 2026
- ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidencearxiv-2605.26340 Sparse Blocked context onlyMay 25, 2026
- Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systemsarxiv-2605.26302 Sparse Blocked context onlyMay 25, 2026
- LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AVarxiv-2605.26244 Sparse Blocked context onlyMay 25, 2026
- Helix4D: Complex 4D Mesh Generationarxiv-2605.26109 Sparse Blocked context onlyMay 25, 2026
- On-Policy Adversarial Flow Distillation for Autoregressive Video Generationarxiv-2605.26105 Sparse Blocked context onlyMay 25, 2026
- CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientistsarxiv-2605.26029 Sparse Blocked context onlyMay 25, 2026
- SemBridge: Language Transfer in Sparse Encoders via Multilingual Semantic Bridgesarxiv-2605.26002 Sparse Blocked context onlyMay 25, 2026
- RocketSmith: Agentic Additive Manufacturing of High-Powered Rocketsarxiv-2606.00097 Sparse Blocked context onlyMay 25, 2026
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluationarxiv-2605.25874 Curated Related Blocked context onlyMay 25, 2026
- Rethinking VLM Representation for VLA Initializationarxiv-2605.25802 Sparse Blocked context onlyMay 25, 2026
- StreamChar: Long-Horizon Streaming Character Audio-Video Generation with Decoupled Orchestrationarxiv-2605.25659 Sparse Blocked context onlyMay 25, 2026
- ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancementarxiv-2605.25569 Sparse Blocked context onlyMay 25, 2026
- Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoningarxiv-2605.25437 Sparse Blocked context onlyMay 25, 2026
- CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillationarxiv-2605.25378 Sparse Blocked context onlyMay 25, 2026
- Toward Native Multimodal Modeling: A Roadmaparxiv-2605.25343 Sparse Blocked context onlyMay 25, 2026
- Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generationarxiv-2605.25220 Sparse Blocked context onlyMay 24, 2026
- DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMsarxiv-2605.25188 Sparse Blocked context onlyMay 24, 2026
- WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Modelsarxiv-2605.25077 Sparse Blocked context onlyMay 24, 2026
- Your Embedding Model is SMARTer Than You Thinkarxiv-2605.24938 Sparse Blocked context onlyMay 24, 2026
- PANDO: Efficient Multimodal AI Agents via Online Skill Distillationarxiv-2605.24785 Sparse Blocked context onlyMay 24, 2026
- Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Frameworkarxiv-2605.24661 Sparse Blocked context onlyMay 23, 2026
- AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoningarxiv-2605.24486 Sparse Blocked context onlyMay 23, 2026
- Benchmarking Composed Image Retrieval for Applied Earth Observationarxiv-2605.24442 Sparse Blocked context onlyMay 23, 2026
- QUEST: Training Frontier Deep Research Agents with Fully Synthetic Tasksarxiv-2605.24218 Sparse Blocked context onlyMay 22, 2026
- SPACENUM: Revisiting Spatial Numerical Understanding in VLMsarxiv-2605.23898 Sparse Blocked context onlyMay 22, 2026
- From Activation to Causality: Discovery of Causal Visual Representations in the Human Brainarxiv-2605.23895 Sparse Blocked context onlyMay 22, 2026
- HorizonStream: Long-Horizon Attention for Streaming 3D Reconstructionarxiv-2605.23889 Sparse Blocked context onlyMay 22, 2026
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Modelsarxiv-2605.23699 Sparse Blocked context onlyMay 22, 2026
- One-Forcing: Towards Stable One-Step Autoregressive Video Generationarxiv-2605.23458 Sparse Blocked context onlyMay 22, 2026
- Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolutionarxiv-2605.23264 Sparse Blocked context onlyMay 22, 2026
- Foundation Protocol: A Coordination Layer for Agentic Societyarxiv-2605.23218 Sparse Blocked context onlyMay 22, 2026
- AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discoveryarxiv-2605.23204 Sparse Blocked context onlyMay 22, 2026
- EMMA: Extracting Multiple physical parameters from Multimodal Dataarxiv-2605.24047 Sparse Blocked context onlyMay 21, 2026
- MotiMotion: Motion-Controlled Video Generation with Visual Reasoningarxiv-2605.22818 Sparse Blocked context onlyMay 21, 2026
- Evaluating Commercial AI Chatbots as News Intermediariesarxiv-2605.22785 Sparse Blocked context onlyMay 21, 2026
- Diversed Model Discovery via Structured Table Discoveryarxiv-2605.22766 Sparse Blocked context onlyMay 21, 2026
- The Distillation Game: Adaptive Attacks & Efficient Defensesarxiv-2605.22737 Sparse Blocked context onlyMay 21, 2026
- ChronoMedKG: A Temporally-Grounded Biomedical Knowledge Graph and Benchmark for Clinical Reasoningarxiv-2605.22734 Sparse Blocked context onlyMay 21, 2026
- Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generatorsarxiv-2605.22717 Sparse Blocked context onlyMay 21, 2026
- Forecasting Scientific Progress with Artificial Intelligencearxiv-2605.22681 Sparse Blocked context onlyMay 21, 2026
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasksarxiv-2605.22535 Sparse Blocked context onlyMay 21, 2026
- Search-E1: Self-Distillation Drives Self-Evolution in Search-Augmented Reasoningarxiv-2605.22511 Sparse Blocked context onlyMay 21, 2026
- Reflecti-Mate: A Conversational Agent for Adaptive Decision-Making Support Through System 1 and System 2 Thinkingarxiv-2605.22509 Sparse Blocked context onlyMay 21, 2026
- Pattern-and-root inflectional morphology: the Arabic broken pluralarxiv-2605.22310 Sparse Blocked context onlyMay 21, 2026
- GHI: Graphormer over Conditioned Hypergraph Incidence for Aspect-Based Sentiment Analysisarxiv-2605.22228 Sparse Blocked context onlyMay 21, 2026
- Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?arxiv-2605.22170 Sparse Blocked context onlyMay 21, 2026
- One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systemsarxiv-2605.22144 Sparse Blocked context onlyMay 21, 2026
- Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Lawsarxiv-2605.21803 Sparse Blocked context onlyMay 20, 2026
- X-Token: Projection-Guided Cross-Tokenizer Knowledge Distillationarxiv-2605.21699 Sparse Blocked context onlyMay 20, 2026
- GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillationarxiv-2605.21605 Curated Related Blocked context onlyMay 20, 2026
- Lens: Rethinking Training Efficiency for Foundational Text-to-Image Modelsarxiv-2605.21573 Sparse Blocked context onlyMay 20, 2026
- Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoningarxiv-2605.21488 Sparse Blocked context onlyMay 20, 2026
- PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objectsarxiv-2605.21572 Sparse Blocked context onlyMay 20, 2026
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agentsarxiv-2605.21384 Sparse Blocked context onlyMay 20, 2026
- SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Researcharxiv-2605.22878 Sparse Blocked context onlyMay 20, 2026
- RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolutionarxiv-2605.21195 Sparse Blocked context onlyMay 20, 2026
- Decoupling Communication from Policy: Robust MARL under Bandwidth Constraintsarxiv-2605.21085 Sparse Blocked context onlyMay 20, 2026
- FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matchingarxiv-2605.20910 Sparse Blocked context onlyMay 20, 2026
- Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignmentarxiv-2605.20834 Sparse Blocked context onlyMay 20, 2026
- SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASRarxiv-2605.20712 Sparse Blocked context onlyMay 20, 2026
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientistsarxiv-2605.20668 Sparse Blocked context onlyMay 20, 2026
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learningarxiv-2605.20342 Sparse Blocked context onlyMay 19, 2026
- Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVRarxiv-2605.20164 Sparse Blocked context onlyMay 19, 2026
- Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decodingarxiv-2605.20104 Sparse Blocked context onlyMay 19, 2026
- AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaborationarxiv-2605.20025 Sparse Blocked context onlyMay 19, 2026
- OpenComputer: Verifiable Software Worlds for Computer-Use Agentsarxiv-2605.19769 Sparse Blocked context onlyMay 19, 2026
- SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objectsarxiv-2605.19587 Sparse Blocked context onlyMay 19, 2026
- GoLongRL: Capability-Oriented Long Context Reinforcement Learning with Multitask Alignmentarxiv-2605.19577 Curated Related Blocked context onlyMay 19, 2026
- CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editingarxiv-2605.19484 Sparse Blocked context onlyMay 19, 2026
- CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimizationarxiv-2605.19436 Sparse Blocked context onlyMay 19, 2026
- Generative Recursive Reasoningarxiv-2605.19376 Sparse Blocked context onlyMay 19, 2026
- Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstructionarxiv-2605.19354 Sparse Blocked context onlyMay 19, 2026
- WavFlow: Audio Generation in Waveform Spacearxiv-2605.18749 Sparse Blocked context onlyMay 18, 2026
- DexHoldem: Playing Texas Hold'em with Dexterous Embodied Systemarxiv-2605.18727 Sparse Blocked context onlyMay 18, 2026
- Semantic Generative Tuning for Unified Multimodal Modelsarxiv-2605.18714 Sparse Blocked context onlyMay 18, 2026
- AI for Auto-Research: Roadmap & User Guidearxiv-2605.18661 Sparse Blocked context onlyMay 18, 2026
- Forecasting Downstream Performance of LLMs With Proxy Metricsarxiv-2605.18607 Sparse Blocked context onlyMay 18, 2026
- MINTEval: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systemsarxiv-2605.18565 Sparse Blocked context onlyMay 18, 2026
- Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamicsarxiv-2605.18549 Sparse Blocked context onlyMay 18, 2026
- Code-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code Synthesisarxiv-2605.18451 Sparse Blocked context onlyMay 18, 2026
- NEWTON: Agentic Planning for Physically Grounded Video Generationarxiv-2605.18396 Sparse Blocked context onlyMay 18, 2026
- Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videosarxiv-2605.18233 Sparse Blocked context onlyMay 18, 2026
- Scalable Environments Drive Generalizable Agentsarxiv-2605.18181 Sparse Blocked context onlyMay 18, 2026
- Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routersarxiv-2605.18106 Sparse Blocked context onlyMay 18, 2026
- Semantic Reranking at Inference Time for Hard Examples in Rhetorical Role Labelingarxiv-2605.18007 Sparse Blocked context onlyMay 18, 2026
- PanoWorld: A Generative Spatial World Model for Consistent Whole-House Panorama Synthesisarxiv-2605.17916 Sparse Blocked context onlyMay 18, 2026
- Ethical Hyper-Velocity (EHV): A Provably Deterministic Governance-Aware JIT Compiler Architecture for Agentic Systemsarxiv-2605.17909 Sparse Blocked context onlyMay 18, 2026
- Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agentsarxiv-2605.17830 Sparse Blocked context onlyMay 18, 2026
- LatentUMM: Dual Latent Alignment for Unified Multimodal Modelsarxiv-2605.17766 Sparse Blocked context onlyMay 18, 2026
- Harnessing LLM Agents with Skill Programsarxiv-2605.17734 Sparse Blocked context onlyMay 18, 2026
- Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Modelsarxiv-2605.17672 Sparse Blocked context onlyMay 17, 2026
- Causal Intervention-Based Memory Selection for Long-Horizon LLM Agentsarxiv-2605.17641 Sparse Blocked context onlyMay 17, 2026
- SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screeningarxiv-2605.17610 Sparse Blocked context onlyMay 17, 2026
- Don't Guess, Just Ask: Resolving Ambiguity in Referring Segmentation via Multi-turn Clarificationarxiv-2605.17531 Sparse Blocked context onlyMay 17, 2026
- SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineeringarxiv-2605.17526 Sparse Blocked context onlyMay 17, 2026
- Trust No Tool: Evaluating and Defending LLM Agents under Untrusted Tool Feedbackarxiv-2605.17453 Sparse Blocked context onlyMay 17, 2026
- ContraFix: Skill-Enhanced Contrastive Runtime Analysis for Vulnerability Repairarxiv-2605.17450 Sparse Blocked context onlyMay 17, 2026
- Self-Improving CAD Generation Agents with Finite Element Analysis as Feedbackarxiv-2605.17448 Sparse Blocked context onlyMay 17, 2026
- Medical Context Distorts Decisions in Clinical Vision Language Modelsarxiv-2605.17436 Sparse Blocked context onlyMay 17, 2026
- Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaborationarxiv-2605.17423 Sparse Blocked context onlyMay 17, 2026
- NewsLens: A Multi-Agent Framework for Adversarial News Bias Navigationarxiv-2605.17364 Sparse Blocked context onlyMay 17, 2026
- HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Predictionarxiv-2605.17355 Sparse Blocked context onlyMay 17, 2026
- A2RBench: An Automatic Paradigm for Formally Verifiable Abstract Reasoning Benchmark Generationarxiv-2605.17278 Sparse Blocked context onlyMay 17, 2026
- Geometric Phase Transition Enables Extreme Hippocampal Memory Capacityarxiv-2605.17199 Sparse Blocked context onlyMay 16, 2026
- PluRule: A Benchmark for Moderating Pluralistic Communities on Social Mediaarxiv-2605.17187 Sparse Blocked context onlyMay 16, 2026
- Responsible Agentic AI Requires Explicit Provenancearxiv-2605.17169 Sparse Blocked context onlyMay 16, 2026
- Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languagesarxiv-2605.17152 Sparse Blocked context onlyMay 16, 2026
- UCSF-PDGM-VQA: Visual Question Answering dataset for brain tumor MRI interpretationarxiv-2605.17140 Sparse Blocked context onlyMay 16, 2026
- Scale Determines Whether Language Models Organize Representation Geometry for Predictionarxiv-2605.17084 Sparse Blocked context onlyMay 16, 2026
- TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Rewardarxiv-2607.21606 Sparse Blocked context onlyMay 16, 2026
- Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBencharxiv-2605.17079 Sparse Blocked context onlyMay 16, 2026
- S-Bus: Automatic Read-Set Reconstruction for Multi-Agent LLM State Coordinationarxiv-2605.17076 Sparse Blocked context onlyMay 16, 2026
- 1GC-7RC: One Graphic Card -- Seven Research Challenges! How Good Are AI Agents at Doing Your Job?arxiv-2605.17046 Sparse Blocked context onlyMay 16, 2026
- Algorithmic Cultivation: How Social Media Feeds Shape User Languagearxiv-2605.17010 Sparse Blocked context onlyMay 16, 2026
- TOBench: A Task-Oriented Omni-Modal Benchmark for Real-World Tool-Using Agentsarxiv-2605.16909 Sparse Blocked context onlyMay 16, 2026
- The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidencearxiv-2605.16895 Sparse Blocked context onlyMay 16, 2026
- NGM: A Plug-and-Play Training-Free Memory Module for LLMsarxiv-2605.16893 Sparse Blocked context onlyMay 16, 2026
- EmoMind: Decoding Affective Captions from Human Brain fMRIarxiv-2605.16739 Sparse Blocked context onlyMay 16, 2026
- DexJoCo: A Benchmark and Toolkit for Task-Oriented Dexterous Manipulation on MuJoCoarxiv-2605.16257 Sparse Blocked context onlyMay 15, 2026
- VideoSeeker: Incentivizing Instance-level Video Understanding via Native Agentic Tool Invocationarxiv-2605.16079 Sparse Blocked context onlyMay 15, 2026
- WorldAct: Activating Monolithic 3D Worlds into Interactive-Ready Object-Centric Scenesarxiv-2605.15843 Sparse Blocked context onlyMay 15, 2026
- FashionChameleon: Towards Real-Time and Interactive Human-Garment Video Customizationarxiv-2605.15824 Sparse Blocked context onlyMay 15, 2026
- ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Botharxiv-2605.15198 Sparse Blocked context onlyMay 14, 2026
- FutureSim: Replaying World Events to Evaluate Adaptive Agentsarxiv-2605.15188 Sparse Blocked context onlyMay 14, 2026
- Quantitative Video World Model Evaluation for Geometric-Consistencyarxiv-2605.15185 Sparse Blocked context onlyMay 14, 2026
- Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Videoarxiv-2605.15182 Sparse Blocked context onlyMay 14, 2026
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memoryarxiv-2605.15128 Sparse Blocked context onlyMay 14, 2026
- Learning from Language Feedback via Variational Policy Distillationarxiv-2605.15113 Sparse Blocked context onlyMay 14, 2026
- Beyond Individual Intelligence: Surveying Collaboration, Failure Attribution, and Self-Evolution in LLM-based Multi-Agent Systemsarxiv-2605.14892 Sparse Blocked context onlyMay 14, 2026
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Tracesarxiv-2605.14786 Sparse Blocked context onlyMay 14, 2026
- ViMU: Benchmarking Video Metaphorical Understandingarxiv-2605.14607 Sparse Blocked context onlyMay 14, 2026
- LiSA: Lifelong Safety Adaptation via Conservative Policy Inductionarxiv-2605.14454 Sparse Blocked context onlyMay 14, 2026
- FrontierSmith: Synthesizing Open-Ended Coding Problems at Scalearxiv-2605.14445 Sparse Blocked context onlyMay 14, 2026
- Nexus : An Agentic Framework for Time Series Forecastingarxiv-2605.14389 Sparse Blocked context onlyMay 14, 2026
- LLM-based Detection of Manipulative Political Narrativesarxiv-2605.14354 Sparse Blocked context onlyMay 14, 2026
- Minimal-Intervention KV Retention via Set-Conditioned Diversityarxiv-2605.14292 Sparse Blocked context onlyMay 14, 2026
- KVPO: ODE-Native GRPO for Autoregressive Video Alignment via KV Semantic Explorationarxiv-2605.14278 Sparse Blocked context onlyMay 14, 2026
- Auditing Agent Harness Safetyarxiv-2605.14271 Sparse Blocked context onlyMay 14, 2026
- BOOKMARKS: Efficient Active Storyline Memory for Role-playingarxiv-2605.14169 Sparse Blocked context onlyMay 13, 2026
- CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curvesarxiv-2605.14068 Sparse Blocked context onlyMay 13, 2026
- SPIN: Structural LLM Planning via Iterative Navigation for Industrial Tasksarxiv-2605.14051 Sparse Blocked context onlyMay 13, 2026
- HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Expertsarxiv-2605.13997 Sparse Blocked context onlyMay 13, 2026
- EVA-Bench: A New End-to-end Framework for Evaluating Voice Agentsarxiv-2605.13841 Sparse Blocked context onlyMay 13, 2026
- MinT: Managed Infrastructure for Training and Serving Millions of LLMsarxiv-2605.13779 Sparse Blocked context onlyMay 13, 2026
- RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Dataarxiv-2605.13775 Sparse Blocked context onlyMay 13, 2026
- FrameSkip: Learning from Fewer but More Informative Frames in VLA Trainingarxiv-2605.13757 Sparse Blocked context onlyMay 13, 2026
- Learning POMDP World Models from Observations with Language-Model Priorsarxiv-2605.13740 Sparse Blocked context onlyMay 13, 2026
- MMSkills: Towards Multimodal Skills for General Visual Agentsarxiv-2605.13527 Sparse Blocked context onlyMay 13, 2026
- Inducing Overthink: Hierarchical Genetic Algorithm-based DoS Attack on Black-Box Large Language Reasoning Modelsarxiv-2605.13338 Sparse Blocked context onlyMay 13, 2026
- Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scalingarxiv-2605.13301 Sparse Blocked context onlyMay 13, 2026
- Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generationarxiv-2605.12975 Sparse Blocked context onlyMay 13, 2026
- WriteSAE: Sparse Autoencoders for Recurrent Statearxiv-2605.12770 Sparse Blocked context onlyMay 12, 2026
- DocAtlas: Multilingual Document Understanding Across 80+ Languagesarxiv-2605.12623 Sparse Blocked context onlyMay 12, 2026
- LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleaguesarxiv-2605.12493 Sparse Blocked context onlyMay 12, 2026
- Beyond GRPO and On-Policy Distillation: An Empirical Sparse-to-Dense Reward Principle for Language-Model Post-Trainingarxiv-2605.12483 Sparse Blocked context onlyMay 12, 2026
- MEME: Multi-entity & Evolving Memory Evaluationarxiv-2605.12477 Sparse Blocked context onlyMay 12, 2026
- Reward Hacking in Rubric-Based Reinforcement Learningarxiv-2605.12474 Sparse Blocked context onlyMay 12, 2026
- LychSim: A Controllable and Interactive Simulation Framework for Vision Researcharxiv-2605.12449 Sparse Blocked context onlyMay 12, 2026
- Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamicsarxiv-2605.12178 Sparse Blocked context onlyMay 12, 2026
- MoCam: Unified Novel View Synthesis via Structured Denoising Dynamicsarxiv-2605.12119 Sparse Blocked context onlyMay 12, 2026
- World Action Models: The Next Frontier in Embodied AIarxiv-2605.12090 Curated Related Blocked context onlyMay 12, 2026
- Entropy Polarity in Reinforcement Fine-Tuning: Direction, Asymmetry, and Controlarxiv-2605.11775 Sparse Blocked context onlyMay 12, 2026
- ShapeCodeBench: A Renewable Benchmark for Perception-to-Program Reconstruction of Synthetic Shape Scenesarxiv-2605.11680 Sparse Blocked context onlyMay 12, 2026
- The DAWN of World-Action Interactive Modelsarxiv-2605.11550 Sparse Blocked context onlyMay 12, 2026
- Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Modelsarxiv-2605.11459 Sparse Blocked context onlyMay 12, 2026
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoningarxiv-2605.11458 Sparse Blocked context onlyMay 12, 2026
- PresentAgent-2: Towards Generalist Multimodal Presentation Agentsarxiv-2605.11363 Sparse Blocked context onlyMay 12, 2026
- ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reductionarxiv-2605.11212 Sparse Blocked context onlyMay 11, 2026
- DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devicesarxiv-2605.10933 Sparse Blocked context onlyMay 11, 2026
- CapVector: Learning Transferable Capability Vectors in Parametric Space for Vision-Language-Action Modelsarxiv-2605.10903 Sparse Blocked context onlyMay 11, 2026
- RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewardsarxiv-2605.10899 Sparse Blocked context onlyMay 11, 2026
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Whyarxiv-2605.10889 Sparse Blocked context onlyMay 11, 2026
- BEACON: A Multimodal Dataset for Learning Behavioral Fingerprints from Gameplay Dataarxiv-2605.10867 Sparse Blocked context onlyMay 11, 2026
- From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-Worldarxiv-2605.10834 Sparse Blocked context onlyMay 11, 2026
- Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agentsarxiv-2605.10832 Sparse Blocked context onlyMay 11, 2026
- NanoResearch: Co-Evolving Skills, Memory, and Policy for Personalized Research Automationarxiv-2605.10813 Sparse Blocked context onlyMay 11, 2026
- GridProbe: Posterior-Probing for Adaptive Test-Time Compute in Long-Video VLMsarxiv-2605.10762 Sparse Blocked context onlyMay 11, 2026
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Imagearxiv-2605.10616 Sparse Blocked context onlyMay 11, 2026
- Follow the Mean: Reference-Guided Flow Matchingarxiv-2605.10302 Sparse Blocked context onlyMay 11, 2026
- Task-Aware Calibration: Provably Optimal Decoding in LLMsarxiv-2605.10202 Sparse Blocked context onlyMay 11, 2026
- When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviewsarxiv-2605.10171 Sparse Blocked context onlyMay 11, 2026
- Continual Harness: Online Adaptation for Self-Improving Foundation Agentsarxiv-2605.09998 Sparse Blocked context onlyMay 11, 2026
- PREPING: Building Agent Memory without Tasksarxiv-2605.13880 Sparse Blocked context onlyMay 11, 2026
- G-Zero: Self-Play for Open-Ended Generation from Zero Dataarxiv-2605.09959 Sparse Blocked context onlyMay 11, 2026
- Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Siliconarxiv-2605.09708 Sparse Blocked context onlyMay 10, 2026
- Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Evictionarxiv-2605.09649 Sparse Blocked context onlyMay 10, 2026
- Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Modelsarxiv-2605.09630 Sparse Blocked context onlyMay 10, 2026
- TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systemsarxiv-2605.09539 Sparse Blocked context onlyMay 10, 2026
- LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Modelsarxiv-2605.11011 Sparse Blocked context onlyMay 10, 2026
- Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairsarxiv-2605.09433 Sparse Blocked context onlyMay 10, 2026
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Modelsarxiv-2605.09241 Sparse Blocked context onlyMay 10, 2026
- FORTIS: Benchmarking Over-Privilege in Agent Skillsarxiv-2605.09163 Sparse Blocked context onlyMay 9, 2026
- Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMsarxiv-2605.09063 Sparse Blocked context onlyMay 9, 2026
- ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usagearxiv-2605.16363 Sparse Blocked context onlyMay 9, 2026
- CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Modelsarxiv-2605.08735 Curated Related Blocked context onlyMay 9, 2026
- Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimizationarxiv-2605.08659 Sparse Blocked context onlyMay 9, 2026
- FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestrationarxiv-2605.08520 Sparse Blocked context onlyMay 8, 2026
- Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challengearxiv-2605.08518 Sparse Blocked context onlyMay 8, 2026
- MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planningarxiv-2605.08512 Sparse Blocked context onlyMay 8, 2026
- jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Compositionarxiv-2605.08384 Sparse Blocked context onlyMay 8, 2026
- Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteriaarxiv-2605.08354 Sparse Blocked context onlyMay 8, 2026
- Normalizing Trajectory Modelsarxiv-2605.08078 Sparse Blocked context onlyMay 8, 2026
- Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibrationarxiv-2605.08077 Sparse Blocked context onlyMay 8, 2026
- The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agentsarxiv-2605.08060 Sparse Blocked context onlyMay 8, 2026
- Accurate and Efficient Statistical Testing for Word Semantic Breadtharxiv-2605.08048 Sparse Blocked context onlyMay 8, 2026
- Uncertainty-Aware Structured Data Extraction from Full CMR Reports via Distilled LLMsarxiv-2605.08045 Sparse Blocked context onlyMay 8, 2026
- SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generationarxiv-2605.08043 Sparse Blocked context onlyMay 8, 2026
- Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claimsarxiv-2605.08012 Sparse Blocked context onlyMay 8, 2026
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewersarxiv-2605.07905 Sparse Blocked context onlyMay 8, 2026
- SCENE: Recognizing Social Norms and Sanctioning in Group Chatsarxiv-2605.07823 Sparse Blocked context onlyMay 8, 2026
- Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoningarxiv-2605.07804 Sparse Blocked context onlyMay 8, 2026
- SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillationarxiv-2605.07711 Sparse Blocked context onlyMay 8, 2026
- DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domainarxiv-2605.07699 Sparse Blocked context onlyMay 8, 2026
- MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditingarxiv-2605.07646 Sparse Blocked context onlyMay 8, 2026
- Learning to Communicate Locally for Large-Scale Multi-Agent Pathfindingarxiv-2605.07637 Sparse Blocked context onlyMay 8, 2026
- Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agentsarxiv-2605.07630 Sparse Blocked context onlyMay 8, 2026
- Implicit Preference Alignment for Human Image Animationarxiv-2605.07545 Sparse Blocked context onlyMay 8, 2026
- From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMsarxiv-2605.07268 Sparse Blocked context onlyMay 8, 2026
- When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Modelsarxiv-2605.07260 Sparse Blocked context onlyMay 8, 2026
- SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Draftingarxiv-2605.07243 Sparse Blocked context onlyMay 8, 2026
- HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agentsarxiv-2605.07177 Sparse Blocked context onlyMay 8, 2026
- CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalizationarxiv-2605.07162 Sparse Blocked context onlyMay 8, 2026
- Learning Visual Feature-Based World Models via Residual Latent Actionarxiv-2605.07079 Sparse Blocked context onlyMay 8, 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment Videoarxiv-2605.06658 Sparse Blocked context onlyMay 7, 2026
- AI Co-Mathematician: Accelerating Mathematicians with Agentic AIarxiv-2605.06651 Sparse Blocked context onlyMay 7, 2026
- PianoCoRe: Combined and Refined Piano MIDI Datasetarxiv-2605.06627 Sparse Blocked context onlyMay 7, 2026
- SkillOS: Learning Skill Curation for Self-Evolving Agentsarxiv-2605.06614 Sparse Blocked context onlyMay 7, 2026
- GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMsarxiv-2605.06477 Sparse Blocked context onlyMay 7, 2026
- MiA-Signature: Approximating Global Activation for Long-Context Understandingarxiv-2605.06416 Sparse Blocked context onlyMay 7, 2026
- Continuous-Time Distribution Matching for Few-Step Diffusion Distillationarxiv-2605.06376 Sparse Blocked context onlyMay 7, 2026
- Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascadesarxiv-2605.06350 Sparse Blocked context onlyMay 7, 2026
- Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoningarxiv-2605.06326 Sparse Blocked context onlyMay 7, 2026
- OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Modelsarxiv-2605.06188 Sparse Blocked context onlyMay 7, 2026
- On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflowsarxiv-2605.06110 Sparse Blocked context onlyMay 7, 2026
- Uncovering Entity Identity Confusion in Multimodal Knowledge Editingarxiv-2605.06096 Sparse Blocked context onlyMay 7, 2026
- PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialoguearxiv-2605.06007 Sparse Blocked context onlyMay 7, 2026
- VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understandingarxiv-2605.05848 Sparse Blocked context onlyMay 7, 2026
- Retrieval from Within: An Intrinsic Capability of Attention-Based Modelsarxiv-2605.05806 Sparse Blocked context onlyMay 7, 2026
- Steering Visual Generation in Unified Multimodal Models with Understanding Supervisionarxiv-2605.05781 Sparse Blocked context onlyMay 7, 2026
- Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMsarxiv-2606.27378 Sparse Blocked context onlyMay 7, 2026
- Belief Memory: Agent Memory Under Partial Observabilityarxiv-2605.05583 Sparse Blocked context onlyMay 7, 2026
- Who Prices Cognitive Labor in the Age of Agents? Compute-Anchored Wagesarxiv-2605.05558 Sparse Blocked context onlyMay 7, 2026
- ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesisarxiv-2605.05485 Sparse Blocked context onlyMay 6, 2026
- PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual Worldarxiv-2605.05163 Sparse Blocked context onlyMay 6, 2026
- Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiersarxiv-2605.04984 Sparse Blocked context onlyMay 6, 2026
- A Foundation Model for Zero-Shot Logical Rule Inductionarxiv-2605.04916 Sparse Blocked context onlyMay 6, 2026
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agentsarxiv-2605.04808 Sparse Blocked context onlyMay 6, 2026
- CHE-TKG: Collaborative Historical Evidence and Evolutionary Dynamics Learning for Temporal Knowledge Graph Reasoningarxiv-2605.04652 Sparse Blocked context onlyMay 6, 2026
- SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agenciesarxiv-2605.04637 Sparse Blocked context onlyMay 6, 2026
- RemoteZero: Geospatial Reasoning with Zero Human Annotationsarxiv-2605.04451 Sparse Blocked context onlyMay 6, 2026
- The Scaling Properties of Implicit Deductive Reasoning in Transformersarxiv-2605.04330 Sparse Blocked context onlyMay 5, 2026
- Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extractionarxiv-2605.04221 Sparse Blocked context onlyMay 5, 2026
- Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluationarxiv-2605.04135 Sparse Blocked context onlyMay 5, 2026
- SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessmentarxiv-2605.04012 Sparse Blocked context onlyMay 5, 2026
- A Benchmark for Interactive World Models with a Unified Action Generation Frameworkarxiv-2605.03941 Sparse Blocked context onlyMay 5, 2026
- The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Modelsarxiv-2605.03936 Sparse Blocked context onlyMay 5, 2026
- Reproducing Complex Set-Compositional Information Retrievalarxiv-2605.03824 Sparse Blocked context onlyMay 5, 2026
- PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examinationarxiv-2605.03571 Sparse Blocked context onlyMay 5, 2026
- SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agentsarxiv-2605.03353 Sparse Blocked context onlyMay 5, 2026
- RAG over Thinking Traces Can Improve Reasoning Tasksarxiv-2605.03344 Sparse Blocked context onlyMay 5, 2026
- ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptomsarxiv-2605.03212 Sparse Blocked context onlyMay 4, 2026
- ARIS: Autonomous Research via Adversarial Multi-Agent Collaborationarxiv-2605.03042 Sparse Blocked context onlyMay 4, 2026
- MolmoAct2: Action Reasoning Models for Real-world Deploymentarxiv-2605.02881 Sparse Blocked context onlyMay 4, 2026
- HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harnessarxiv-2605.02396 Sparse Blocked context onlyMay 4, 2026
- Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decodingarxiv-2605.02290 Sparse Blocked context onlyMay 4, 2026
- PhysicianBench: Evaluating LLM Agents in Real-World EHR Environmentsarxiv-2605.02240 Sparse Blocked context onlyMay 4, 2026