300 canonical paper links on this archive page.
- Text-to-Image Models Need Less from Text Encoders Than You Thinkarxiv-2606.03715 Sparse Blocked context onlyJun 2, 2026
- When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Modelsarxiv-2606.03712 Sparse Blocked context onlyJun 2, 2026
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoningarxiv-2606.03603 Sparse Blocked context onlyJun 2, 2026
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matchingarxiv-2606.03577 Sparse Blocked context onlyJun 2, 2026
- ThoughtFold: Folding Reasoning Chains via Introspective Preference Learningarxiv-2606.03503 Sparse Blocked context onlyJun 2, 2026
- KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasksarxiv-2606.03458 Direct Blocked context onlyJun 2, 2026
- Large Language Models Are Overconfident in Their Own Responsesarxiv-2606.03437 Sparse Blocked context onlyJun 2, 2026
- MemTrain: Self-Supervised Context Memory Trainingarxiv-2606.03197 Sparse Blocked context onlyJun 2, 2026
- EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learningarxiv-2606.03108 Sparse Blocked context onlyJun 2, 2026
- Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scalingarxiv-2606.03102 Sparse Blocked context onlyJun 2, 2026
- The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMsarxiv-2606.03092 Sparse Blocked context onlyJun 2, 2026
- Self-Distilled Policy Gradientarxiv-2606.04036 Direct Blocked context onlyJun 2, 2026
- AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verificationarxiv-2606.03031 Sparse Blocked context onlyJun 2, 2026
- The Road Ahead in Autonomous Driving: The KITScenes Multimodal Datasetarxiv-2606.02956 Direct Blocked context onlyJun 1, 2026
- Do Transformers Need Three Projections? Systematic Study of QKV Variantsarxiv-2606.04032 Sparse Blocked context onlyJun 1, 2026
- Cosmos 3: Omnimodal World Models for Physical AIarxiv-2606.02800 Direct Blocked context onlyJun 1, 2026
- AURA: Action-Gated Memory for Robot Policies at Constant VRAMarxiv-2606.02775 Sparse Blocked context onlyJun 1, 2026
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modelingarxiv-2606.02578 Sparse Blocked context onlyJun 1, 2026
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillationarxiv-2606.02684 Sparse Blocked context onlyJun 1, 2026
- AdaCodec: A Predictive Visual Code for Video MLLMsarxiv-2606.02569 Sparse Blocked context onlyJun 1, 2026
- VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimizationarxiv-2606.02564 Sparse Blocked context onlyJun 1, 2026
- SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignmentarxiv-2606.02530 Sparse Blocked context onlyJun 1, 2026
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulationarxiv-2606.02470 Sparse Blocked context onlyJun 1, 2026
- K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contextsarxiv-2606.02404 Sparse Blocked context onlyJun 1, 2026
- A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RLarxiv-2606.02398 Sparse Blocked context onlyJun 1, 2026
- Policy and World Modeling Co-Training for Language Agentsarxiv-2606.02388 Sparse Blocked context onlyJun 1, 2026
- Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnessesarxiv-2606.02373 Direct Blocked context onlyJun 1, 2026
- RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Modelsarxiv-2606.02277 Sparse Blocked context onlyJun 1, 2026
- Geometric Latent Reasoning Induces Shorter Generations in LLMsarxiv-2606.02248 Sparse Blocked context onlyJun 1, 2026
- OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agentsarxiv-2606.02031 Sparse Blocked context onlyJun 1, 2026
- WALL-WM: Carving World Action Modeling at the Event Jointsarxiv-2606.01955 Sparse Blocked context onlyJun 1, 2026
- Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agentsarxiv-2606.01886 Curated Related Blocked context onlyJun 1, 2026
- LayerRoute: Input-Conditioned Adaptive Layer Skipping via LoRA Fine-Tuning for Agentic Language Modelsarxiv-2606.01838 Sparse Blocked context onlyJun 1, 2026
- Decentralized Instruction Tuning: Conflict-Aware Splitting and Weight Mergingarxiv-2606.01717 Sparse Blocked context onlyJun 1, 2026
- RCEM: Robust Conversational Search EMbedder in Distributional Shiftarxiv-2606.01697 Sparse Blocked context onlyJun 1, 2026
- Off-the-Shelf LLMs as Process Scorers: Training-Free Alternative to PRMs for Mathematical Reasoningarxiv-2606.01682 Sparse Blocked context onlyJun 1, 2026
- DOT-MoE: Differentiable Optimal Transport for MoEficationarxiv-2606.01666 Sparse Blocked context onlyJun 1, 2026
- ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagreearxiv-2606.01494 Sparse Blocked context onlyMay 31, 2026
- Agent Skills Should Go Beyond Text: The Case for Visual Skillsarxiv-2606.01414 Sparse Blocked context onlyMay 31, 2026
- LongAttnComp: Cross-Family Context Compression for Long-Context Reasoningarxiv-2606.01336 Sparse Blocked context onlyMay 31, 2026
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspacesarxiv-2606.01317 Sparse Blocked context onlyMay 31, 2026
- SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectoriesarxiv-2606.01311 Sparse Blocked context onlyMay 31, 2026
- BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolutionarxiv-2606.01286 Sparse Blocked context onlyMay 31, 2026
- Trust Region On-Policy Distillationarxiv-2606.01249 Sparse Blocked context onlyMay 31, 2026
- Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?arxiv-2606.01247 Sparse Blocked context onlyMay 31, 2026
- BraveGuard: From Open-World Threats to Safer Computer-Use Agentsarxiv-2606.01166 Curated Related Blocked context onlyMay 31, 2026
- $τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulationarxiv-2606.01027 Sparse Blocked context onlyMay 31, 2026
- FVSpec: Real-World Property-Based Tests as Lean Challengesarxiv-2606.01008 Sparse Blocked context onlyMay 31, 2026
- RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenesarxiv-2606.00828 Sparse Blocked context onlyMay 30, 2026
- SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memoryarxiv-2606.00825 Sparse Blocked context onlyMay 30, 2026
- Moxia: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoningarxiv-2606.00671 Sparse Blocked context onlyMay 30, 2026
- FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Searcharxiv-2606.00660 Sparse Blocked context onlyMay 30, 2026
- SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answeringarxiv-2606.00593 Sparse Blocked context onlyMay 30, 2026
- Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedbackarxiv-2606.00590 Sparse Blocked context onlyMay 30, 2026
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performancearxiv-2606.00467 Sparse Blocked context onlyMay 30, 2026
- SDR: Set-Distance Rewards for Radiology Report Generationarxiv-2606.00440 Curated Related Blocked context onlyMay 30, 2026
- Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidancearxiv-2606.00305 Sparse Blocked context onlyMay 29, 2026
- StressDream: Steering Video World Models for Robust Policy Evaluation and Improvementarxiv-2606.00267 Sparse Blocked context onlyMay 29, 2026
- MindZero: Learning Online Mental Reasoning With Zero Annotationsarxiv-2606.00240 Sparse Blocked context onlyMay 29, 2026
- Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Modelsarxiv-2605.31603 Sparse Blocked context onlyMay 29, 2026
- Linear Scaling Video VLMs for Long Video Understandingarxiv-2605.31598 Sparse Blocked context onlyMay 29, 2026
- LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewardsarxiv-2605.31584 Sparse Blocked context onlyMay 29, 2026
- BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengaliarxiv-2605.31483 Sparse Blocked context onlyMay 29, 2026
- PithTrain: A Compact and Agent-Native MoE Training Systemarxiv-2605.31463 Sparse Blocked context onlyMay 29, 2026
- Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliancearxiv-2608.12323 Sparse Blocked context onlyMay 29, 2026
- Mellum2 Technical Reportarxiv-2605.31268 Sparse Blocked context onlyMay 29, 2026
- COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillationarxiv-2605.31264 Direct Blocked context onlyMay 29, 2026
- iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learningarxiv-2605.31096 Sparse Blocked context onlyMay 29, 2026
- AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answeringarxiv-2605.31062 Sparse Blocked context onlyMay 29, 2026
- Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombinationarxiv-2605.31058 Sparse Blocked context onlyMay 29, 2026
- From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoorsarxiv-2605.31042 Sparse Blocked context onlyMay 29, 2026
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformerarxiv-2605.30940 Sparse Blocked context onlyMay 29, 2026
- MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraftarxiv-2605.30931 Sparse Blocked context onlyMay 29, 2026
- The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvementarxiv-2605.30888 Sparse Blocked context onlyMay 29, 2026
- dMoE: dLLMs with Learnable Block Expertsarxiv-2605.30876 Sparse Blocked context onlyMay 29, 2026
- Distilling LLM Feedback for Lean Theorem Provingarxiv-2605.30861 Sparse Blocked context onlyMay 29, 2026
- Hide-and-Seek in Trajectories: Discovering Failure Signals for VLA Runtime Monitoringarxiv-2605.30834 Sparse Blocked context onlyMay 29, 2026
- Function2Scene: 3D Indoor Scene Layout from Functional Specificationsarxiv-2605.30819 Sparse Blocked context onlyMay 29, 2026
- MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understandingarxiv-2605.30794 Sparse Blocked context onlyMay 29, 2026
- Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agentsarxiv-2605.30723 Sparse Blocked context onlyMay 29, 2026
- A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problemsarxiv-2606.00138 Sparse Blocked context onlyMay 28, 2026
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agentsarxiv-2605.30621 Sparse Blocked context onlyMay 28, 2026
- Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimesarxiv-2605.30581 Sparse Blocked context onlyMay 28, 2026
- Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decodearxiv-2605.30571 Sparse Blocked context onlyMay 28, 2026
- SchGen: PCB Schematic Generation with Semantic-Grounded Code Representationsarxiv-2605.30345 Curated Related Blocked context onlyMay 28, 2026
- Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detectionarxiv-2605.30344 Sparse Blocked context onlyMay 28, 2026
- Unlocking the Working Memory of Large Language Models for Latent Reasoningarxiv-2605.30343 Sparse Blocked context onlyMay 28, 2026
- SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?arxiv-2605.30329 Sparse Blocked context onlyMay 28, 2026
- Reasoning with Sampling: Cutting at Decision Pointsarxiv-2605.30327 Sparse Blocked context onlyMay 28, 2026
- Exploring Autonomous Agentic Data Engineering for Model Specializationarxiv-2605.30407 Sparse Blocked context onlyMay 28, 2026
- MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settingsarxiv-2605.30295 Sparse Blocked context onlyMay 28, 2026
- Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selectionarxiv-2605.30274 Sparse Blocked context onlyMay 28, 2026
- LLUMI: Improving LLM Writing Assistance for Mental Health Support with Online Community Feedbackarxiv-2605.30273 Sparse Blocked context onlyMay 28, 2026
- PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactionsarxiv-2605.30268 Sparse Blocked context onlyMay 28, 2026
- minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Modelsarxiv-2605.30263 Sparse Blocked context onlyMay 28, 2026
- How LoRA Remembers? A Parametric Memory Law for LLM Finetuningarxiv-2605.30260 Sparse Blocked context onlyMay 28, 2026
- Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learningarxiv-2605.30257 Sparse Blocked context onlyMay 28, 2026
- Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Modelsarxiv-2605.30251 Sparse Blocked context onlyMay 28, 2026
- GenClaw: Code-Driven Agentic Image Generationarxiv-2605.30248 Direct Blocked context onlyMay 28, 2026
- Knowing What to Solve Before How: Preplan Empowered LLM Mathematical Reasoningarxiv-2605.30245 Sparse Blocked context onlyMay 28, 2026
- Do Language Models Track Entities Across State Changes?arxiv-2605.30233 Sparse Blocked context onlyMay 28, 2026
- How's it going? Reinforcement learning in language models recruits a functional welfare axisarxiv-2605.30232 Sparse Blocked context onlyMay 28, 2026
- Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoningarxiv-2605.30231 Sparse Blocked context onlyMay 28, 2026
- GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in Germanarxiv-2605.30214 Sparse Blocked context onlyMay 28, 2026
- Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agentsarxiv-2605.30159 Sparse Blocked context onlyMay 28, 2026
- SEAL: Can Saturated Benchmarks Be Revived by LLM-as-a-Meta-Judge?arxiv-2605.30104 Sparse Blocked context onlyMay 28, 2026
- When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systemsarxiv-2605.30102 Sparse Blocked context onlyMay 28, 2026
- HEART-Bench: Do LLM Agents Exhibit Human-like Psychology?arxiv-2605.30058 Sparse Blocked context onlyMay 28, 2026
- Token Inflation: How Dishonest Providers Can Overcharge for Large Language Model Usagearxiv-2605.30040 Sparse Blocked context onlyMay 28, 2026
- Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learningarxiv-2605.30039 Sparse Blocked context onlyMay 28, 2026
- Teaching Values to Machines: Simulating Human-Like Behavior in LLMsarxiv-2605.30036 Sparse Blocked context onlyMay 28, 2026
- Give it Space! Explicit Disentangling of Positional and Semantic Representations in Encodersarxiv-2605.30022 Sparse Blocked context onlyMay 28, 2026
- Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMsarxiv-2605.30021 Sparse Blocked context onlyMay 28, 2026
- Latent Performance Profiling of Large Language Modelsarxiv-2605.30018 Sparse Blocked context onlyMay 28, 2026
- VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policiesarxiv-2605.30011 Sparse Blocked context onlyMay 28, 2026
- EarlyTom: Early Token Compression Completes Fast Video Understandingarxiv-2605.30010 Sparse Blocked context onlyMay 28, 2026
- LaRA: Layer-wise Representation Analysis for Detecting Data Contamination in RL Post-Trainingarxiv-2605.29888 Sparse Blocked context onlyMay 28, 2026
- CRITIC-R1: Learning Structured Critics for Retrieval-Augmented Generationarxiv-2605.29886 Sparse Blocked context onlyMay 28, 2026
- Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generationarxiv-2605.29861 Sparse Blocked context onlyMay 28, 2026
- ESPO: Early-Stopping Proximal Policy Optimizationarxiv-2605.29860 Sparse Blocked context onlyMay 28, 2026
- EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generationarxiv-2605.29847 Sparse Blocked context onlyMay 28, 2026
- Towards Localized and Disentangled Knowledge Editing for Multimodal Large Language Modelsarxiv-2605.29826 Sparse Blocked context onlyMay 28, 2026
- SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Searcharxiv-2605.29796 Sparse Blocked context onlyMay 28, 2026
- ActTraitBench: Quantifying the Knowledge-Decision Gap in Large Language Models via Human-Grounded Behavioral Validationarxiv-2605.29791 Sparse Blocked context onlyMay 28, 2026
- Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategiesarxiv-2605.29712 Sparse Blocked context onlyMay 28, 2026
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decodingarxiv-2605.29707 Sparse Blocked context onlyMay 28, 2026
- Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systemsarxiv-2605.29676 Sparse Blocked context onlyMay 28, 2026
- Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answeringarxiv-2605.29648 Sparse Blocked context onlyMay 28, 2026
- Training Deliberative Monitors for Black-Box Scheming Detectionarxiv-2605.29601 Sparse Blocked context onlyMay 28, 2026
- Brain-IT-VQA: From Brain Signals to Answersarxiv-2605.29588 Sparse Blocked context onlyMay 28, 2026
- LiteCoder-Terminal: Scaling Long-Horizon Terminal Environments for Learning Language Agentsarxiv-2605.29559 Sparse Blocked context onlyMay 28, 2026
- From Blind Guess to Informed Judgment: Teaching LLMs to Evaluate Materials by Building Knowledge-Augmented Preference Signalsarxiv-2605.29555 Sparse Blocked context onlyMay 28, 2026
- PhoneWorld: Scaling Phone-Use Agent Environmentsarxiv-2605.29486 Sparse Blocked context onlyMay 28, 2026
- Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memoryarxiv-2607.00017 Sparse Blocked context onlyMay 28, 2026
- Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluationarxiv-2605.29430 Sparse Blocked context onlyMay 28, 2026
- GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Modelsarxiv-2605.29398 Sparse Blocked context onlyMay 28, 2026
- WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interactionarxiv-2605.29341 Sparse Blocked context onlyMay 28, 2026
- GrepSeek: Training Search Agents for Direct Corpus Interactionarxiv-2605.29307 Curated Related Blocked context onlyMay 28, 2026
- Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Tracesarxiv-2605.29288 Sparse Blocked context onlyMay 28, 2026
- Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agentsarxiv-2605.29224 Sparse Blocked context onlyMay 28, 2026
- ReasonOps: Operator Segmentation for LLM Reasoning Tracesarxiv-2605.29192 Sparse Blocked context onlyMay 28, 2026
- Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalizationarxiv-2605.28969 Sparse Blocked context onlyMay 27, 2026
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Playersarxiv-2605.28816 Curated Related Blocked context onlyMay 27, 2026
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibrationarxiv-2605.28805 Sparse Blocked context onlyMay 27, 2026
- Rethinking Memory as Continuously Evolving Connectivityarxiv-2605.28773 Sparse Blocked context onlyMay 27, 2026
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systemsarxiv-2605.28732 Sparse Blocked context onlyMay 27, 2026
- LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?arxiv-2605.28721 Direct Blocked context onlyMay 27, 2026
- The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolicarxiv-2605.28700 Sparse Blocked context onlyMay 27, 2026
- AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentationarxiv-2605.28655 Direct Blocked context onlyMay 27, 2026
- GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselectionarxiv-2605.28534 Sparse Blocked context onlyMay 27, 2026
- Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learningarxiv-2605.28424 Sparse Blocked context onlyMay 27, 2026
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixesarxiv-2605.28421 Sparse Blocked context onlyMay 27, 2026
- HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMsarxiv-2605.28398 Sparse Blocked context onlyMay 27, 2026
- Pruning and Distilling Mixture-of-Experts into Dense Language Modelsarxiv-2605.28207 Sparse Blocked context onlyMay 27, 2026
- Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibrationarxiv-2605.28184 Sparse Blocked context onlyMay 27, 2026
- When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Modelsarxiv-2605.28181 Sparse Blocked context onlyMay 27, 2026
- OR-Space: A Full-Lifecycle Workspace Benchmark for Industrial Optimization Agentsarxiv-2605.28158 Sparse Blocked context onlyMay 27, 2026
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimizationarxiv-2605.28109 Sparse Blocked context onlyMay 27, 2026
- AsyncTool: Evaluating the Asynchronous Function Calling Capability under Multi-Task Scenariosarxiv-2605.27995 Sparse Blocked context onlyMay 27, 2026
- Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representationsarxiv-2605.27958 Sparse Blocked context onlyMay 27, 2026
- AI Research Agents Narrow Scientific Explorationarxiv-2605.27905 Sparse Blocked context onlyMay 27, 2026
- VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wildarxiv-2605.27882 Sparse Blocked context onlyMay 27, 2026
- Revealing Algorithmic Deductive Circuits for Logical Reasoningarxiv-2605.27824 Sparse Blocked context onlyMay 27, 2026
- Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systemsarxiv-2605.27766 Sparse Blocked context onlyMay 26, 2026
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluationarxiv-2605.27366 Sparse Blocked context onlyMay 26, 2026
- MobileMoE: Scaling On-Device Mixture of Expertsarxiv-2605.27358 Sparse Blocked context onlyMay 26, 2026
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biasesarxiv-2605.27355 Sparse Blocked context onlyMay 26, 2026
- SIA: Self Improving AI with Harness & Weight Updatesarxiv-2605.27276 Sparse Blocked context onlyMay 26, 2026
- Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)arxiv-2605.27268 Sparse Blocked context onlyMay 26, 2026
- Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systemsarxiv-2605.27492 Sparse Blocked context onlyMay 26, 2026
- GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulationarxiv-2605.27491 Sparse Blocked context onlyMay 26, 2026
- MRT: Masked Region Transformer for Layered Image Generation and Editing at Scalearxiv-2605.27235 Sparse Blocked context onlyMay 26, 2026
- Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environmentsarxiv-2605.27209 Sparse Blocked context onlyMay 26, 2026
- VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactionsarxiv-2605.27141 Sparse Blocked context onlyMay 26, 2026
- QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agentsarxiv-2605.27068 Sparse Blocked context onlyMay 26, 2026
- Share More, Search Less: Collaborative Parallel Thinking for Efficient Test-Time Scalingarxiv-2605.27030 Sparse Blocked context onlyMay 26, 2026
- Efficient Agentic Reinforcement Learning with On-Policy Intrinsic Knowledge Boundary Enhancementarxiv-2605.26952 Sparse Blocked context onlyMay 26, 2026
- Balancing Fidelity and Diversity in Diffusion Models via Symmetric Attention Decomposition: Hopfield Perspectivearxiv-2605.27476 Sparse Blocked context onlyMay 26, 2026
- AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systemsarxiv-2605.27466 Sparse Blocked context onlyMay 26, 2026
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion Modelsarxiv-2605.26632 Sparse Blocked context onlyMay 26, 2026
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligencearxiv-2605.26494 Sparse Blocked context onlyMay 26, 2026
- OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistantsarxiv-2605.26485 Sparse Blocked context onlyMay 26, 2026
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Researcharxiv-2605.26114 Direct Blocked context onlyMay 25, 2026
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AIarxiv-2605.26112 Sparse Blocked context onlyMay 25, 2026
- Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generationarxiv-2605.26111 Sparse Blocked context onlyMay 25, 2026
- Reinforcing Few-step Generators via Reward-Tilted Distribution Matchingarxiv-2605.26108 Sparse Blocked context onlyMay 25, 2026
- InstructSAM: Segment Any Instance with Any Instructionsarxiv-2605.26102 Sparse Blocked context onlyMay 25, 2026
- Language Models Need Sleeparxiv-2605.26099 Sparse Blocked context onlyMay 25, 2026
- Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital Worldarxiv-2605.26086 Sparse Blocked context onlyMay 25, 2026
- When Gradients Collide: Failure Modes of Multi-Objective Prompt Optimization for LLM Judgesarxiv-2605.26046 Sparse Blocked context onlyMay 25, 2026
- LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligencearxiv-2605.25979 Sparse Blocked context onlyMay 25, 2026
- Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agentsarxiv-2605.25971 Sparse Blocked context onlyMay 25, 2026
- Triplet-Block Diffusion RWKVarxiv-2605.25969 Sparse Blocked context onlyMay 25, 2026
- AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptionsarxiv-2605.25707 Sparse Blocked context onlyMay 25, 2026
- CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agentsarxiv-2605.25624 Sparse Blocked context onlyMay 25, 2026
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learningarxiv-2605.25604 Sparse Blocked context onlyMay 25, 2026
- Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-horizon Agentsarxiv-2605.25535 Sparse Blocked context onlyMay 25, 2026
- Not only where, But when: Temporal Scheduling for RLVRarxiv-2605.25381 Sparse Blocked context onlyMay 25, 2026
- Injecting Image Guidance into Text-Conditioned Diffusion Models at Inferencearxiv-2605.25191 Sparse Blocked context onlyMay 24, 2026
- Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Modelsarxiv-2605.25189 Curated Related Blocked context onlyMay 24, 2026
- STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Mediaarxiv-2605.25162 Sparse Blocked context onlyMay 24, 2026
- SimuWoB: Simulating Real-World Mobile Apps for Fast and Faithful GUI Agent Benchmarkingarxiv-2605.25160 Sparse Blocked context onlyMay 24, 2026
- Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimizationarxiv-2605.24960 Sparse Blocked context onlyMay 24, 2026
- NITP: Next Implicit Token Prediction for LLM Pre-trainingarxiv-2605.24956 Sparse Blocked context onlyMay 24, 2026
- Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuningarxiv-2605.24844 Sparse Blocked context onlyMay 24, 2026
- Macaron-A2UI: A Model for Generative UI in Personal Agentsarxiv-2605.24830 Sparse Blocked context onlyMay 24, 2026
- CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLMarxiv-2605.24786 Sparse Blocked context onlyMay 24, 2026
- Mix-MoE: Improving Multilingual Machine Translation of Large Language Models through Mixed MoEsarxiv-2605.24681 Sparse Blocked context onlyMay 23, 2026
- VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translationarxiv-2605.24675 Sparse Blocked context onlyMay 23, 2026
- Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systemsarxiv-2606.00090 Sparse Blocked context onlyMay 23, 2026
- ECHO: Terminal Agents Learn World Models for Freearxiv-2605.24517 Sparse Blocked context onlyMay 23, 2026
- SAM: State-Adaptive Memory for Long-Horizon Reasoning Agentarxiv-2605.24468 Sparse Blocked context onlyMay 23, 2026
- SEAL: Synergistic Co-Evolution of Agents and Learning Environmentsarxiv-2605.24426 Sparse Blocked context onlyMay 23, 2026
- Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflowsarxiv-2605.24219 Sparse Blocked context onlyMay 22, 2026
- When Does Multi-Agent RL Improve LLM Workflows? Workflow, Scale, and Policy-Sharing Tradeoffsarxiv-2605.24202 Sparse Blocked context onlyMay 22, 2026
- SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skillsarxiv-2605.24117 Sparse Blocked context onlyMay 22, 2026
- Geo-Align: Video Generation Alignment via Metric Geometry Rewardarxiv-2605.23903 Sparse Blocked context onlyMay 22, 2026
- PhotoFlow: Agentic 3D Virtual Photography Missionsarxiv-2605.23771 Sparse Blocked context onlyMay 22, 2026
- OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agentsarxiv-2605.23657 Sparse Blocked context onlyMay 22, 2026
- CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Testarxiv-2605.23491 Sparse Blocked context onlyMay 22, 2026
- StepAudio 2.5 Technical Reportarxiv-2605.23463 Sparse Blocked context onlyMay 22, 2026
- SSDAU: Structured Semantic Data Augmentation for Joint Entity and Relation Extractionarxiv-2605.23440 Sparse Blocked context onlyMay 22, 2026
- Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusionarxiv-2605.23346 Sparse Blocked context onlyMay 22, 2026
- SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Modelsarxiv-2605.23345 Sparse Blocked context onlyMay 22, 2026
- EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generationarxiv-2605.23271 Sparse Blocked context onlyMay 22, 2026
- FastKernels: Benchmarking GPU Kernel Generation in Productionarxiv-2605.23215 Sparse Blocked context onlyMay 22, 2026
- Vector Policy Optimization: Training for Diversity Improves Test-Time Searcharxiv-2605.22817 Sparse Blocked context onlyMay 21, 2026
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attentionarxiv-2605.22791 Sparse Blocked context onlyMay 21, 2026
- Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Modelsarxiv-2605.22732 Sparse Blocked context onlyMay 21, 2026
- WorldKV: Efficient World Memory with World Retrieval and Compressionarxiv-2605.22718 Sparse Blocked context onlyMay 21, 2026
- Swift Sampling: Selecting Temporal Surprises via Taylor Seriesarxiv-2605.22678 Sparse Blocked context onlyMay 21, 2026
- Self-Policy Distillation via Capability-Selective Subspace Projectionarxiv-2605.22675 Sparse Blocked context onlyMay 21, 2026
- SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformersarxiv-2605.22668 Sparse Blocked context onlyMay 21, 2026
- Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corporaarxiv-2605.22660 Sparse Blocked context onlyMay 21, 2026
- Boiling the Frog: A Multi-Turn Benchmark for Agentic Safetyarxiv-2605.22643 Sparse Blocked context onlyMay 21, 2026
- Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learningarxiv-2605.22642 Sparse Blocked context onlyMay 21, 2026
- Two is better than one: A Collapse-free Multi-Reward RLIF Training Frameworkarxiv-2605.22620 Sparse Blocked context onlyMay 21, 2026
- Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agentsarxiv-2605.22608 Sparse Blocked context onlyMay 21, 2026
- Beyond Temperature: Hyperfitting as a Late-Stage Geometric Expansionarxiv-2605.22579 Sparse Blocked context onlyMay 21, 2026
- VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesisarxiv-2605.22570 Sparse Blocked context onlyMay 21, 2026
- LANG: Reinforcement Learning for Multilingual Reasoning with Language-Adaptive Hint Guidancearxiv-2605.22567 Sparse Blocked context onlyMay 21, 2026
- SynAE: A Framework for Measuring the Quality of Synthetic Data for Tool-Calling Agent Evaluationsarxiv-2605.22564 Sparse Blocked context onlyMay 21, 2026
- FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learningarxiv-2605.22552 Sparse Blocked context onlyMay 21, 2026
- SpaceDG: Benchmarking Spatial Intelligence under Visual Degradationarxiv-2605.22536 Sparse Blocked context onlyMay 21, 2026
- Polite on the Surface, Wrong in Practice: A Curated Dataset for Fixing Honorific Failures in Multilingual Bangla Generationarxiv-2605.22487 Sparse Blocked context onlyMay 21, 2026
- DeferMem: Query-Time Evidence Distillation via Reinforcement Learning for Long-Term Memory QAarxiv-2605.22411 Sparse Blocked context onlyMay 21, 2026
- Unified Data Selection for LLM Reasoningarxiv-2605.22389 Sparse Blocked context onlyMay 21, 2026
- Bernini: Latent Semantic Planning for Video Diffusionarxiv-2605.22344 Direct Blocked context onlyMay 21, 2026
- Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RLarxiv-2605.22217 Sparse Blocked context onlyMay 21, 2026
- Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensemblesarxiv-2605.22177 Sparse Blocked context onlyMay 21, 2026
- Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agentsarxiv-2605.22148 Sparse Blocked context onlyMay 21, 2026
- Psy-Chronicle:A Structured Pipeline for Synthesizing Long-Horizon Campus Psychological Counseling Dialoguesarxiv-2605.22140 Sparse Blocked context onlyMay 21, 2026
- Efficient Agentic Reasoning Through Self-Regulated Simulative Planningarxiv-2605.22138 Sparse Blocked context onlyMay 21, 2026
- Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistencyarxiv-2605.22137 Sparse Blocked context onlyMay 21, 2026
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?arxiv-2605.22109 Sparse Blocked context onlyMay 21, 2026
- From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoningarxiv-2605.22074 Sparse Blocked context onlyMay 21, 2026
- Faithful-MR1: Faithful Multimodal Reasoning via Anchoring and Reinforcing Visual Attentionarxiv-2605.22072 Sparse Blocked context onlyMay 21, 2026
- FlyRoute: Self-Evolving Agent Profiling via Data Flywheel for Adaptive Task Routingarxiv-2605.22057 Sparse Blocked context onlyMay 21, 2026
- LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoningarxiv-2605.22012 Sparse Blocked context onlyMay 21, 2026
- RiT: Vanilla Diffusion Transformers Suffice in Representation Spacearxiv-2605.21981 Sparse Blocked context onlyMay 21, 2026
- The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncationarxiv-2605.21856 Sparse Blocked context onlyMay 21, 2026
- ACC: Compiling Agent Trajectories for Long-Context Trainingarxiv-2605.21850 Sparse Blocked context onlyMay 21, 2026
- Reflective Prompt Tuning through Language Model Function-Callingarxiv-2605.21781 Sparse Blocked context onlyMay 20, 2026
- How Far Will They Go? Red-Teaming Online Influence with Large Language Modelsarxiv-2605.22880 Sparse Blocked context onlyMay 20, 2026
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assemblyarxiv-2605.21625 Sparse Blocked context onlyMay 20, 2026
- Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuningarxiv-2605.21487 Sparse Blocked context onlyMay 20, 2026
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectoriesarxiv-2605.21468 Sparse Blocked context onlyMay 20, 2026
- DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewardsarxiv-2605.21467 Sparse Blocked context onlyMay 20, 2026
- Mem-$π$: Adaptive Memory through Learning When and What to Generatearxiv-2605.21463 Sparse Blocked context onlyMay 20, 2026
- OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generationarxiv-2605.21343 Sparse Blocked context onlyMay 20, 2026
- LamPO: A Lambda Style Policy Optimization for Reasoning Language Modelsarxiv-2605.21235 Sparse Blocked context onlyMay 20, 2026
- GradeLegal: Automated Grading for German Legal Casesarxiv-2605.21076 Sparse Blocked context onlyMay 20, 2026
- Q-ARVD: Quantizing Autoregressive Video Diffusion Modelsarxiv-2605.21072 Sparse Blocked context onlyMay 20, 2026
- Rethinking Cross-Layer Information Routing in Diffusion Transformersarxiv-2605.20708 Sparse Blocked context onlyMay 20, 2026
- IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Toolsarxiv-2605.20682 Sparse Blocked context onlyMay 20, 2026
- Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelinesarxiv-2605.20630 Sparse Blocked context onlyMay 20, 2026
- HRM-Text: Efficient Pretraining Beyond Scalingarxiv-2605.20613 Sparse Blocked context onlyMay 20, 2026
- MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generationarxiv-2605.20183 Sparse Blocked context onlyMay 19, 2026
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Modelsarxiv-2605.20177 Sparse Blocked context onlyMay 19, 2026
- ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoningarxiv-2605.20176 Sparse Blocked context onlyMay 19, 2026
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMsarxiv-2605.20315 Sparse Blocked context onlyMay 19, 2026
- Rethinking Visual Attribution for Chest X-ray Reasoning in Large Vision Language Modelsarxiv-2605.20158 Sparse Blocked context onlyMay 19, 2026
- PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Datasetarxiv-2605.20147 Sparse Blocked context onlyMay 19, 2026
- ThoughtTrace: Understanding User Thoughts in Real-World LLM Interactionsarxiv-2605.20087 Sparse Blocked context onlyMay 19, 2026
- CopT: Contrastive On-Policy Thinking with Continuous Spaces for General and Agentic Reasoningarxiv-2605.20075 Sparse Blocked context onlyMay 19, 2026
- Stage-adaptive Token Selection for Efficient Omni-modal LLMsarxiv-2605.20035 Sparse Blocked context onlyMay 19, 2026
- CogOmniControl: Reasoning-Driven Controllable Video Generation via Creative Intent Cognitionarxiv-2605.19995 Sparse Blocked context onlyMay 19, 2026
- Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memoryarxiv-2605.19952 Sparse Blocked context onlyMay 19, 2026
- PEEK: Context Map as an Orientation Cache for Long-Context LLM Agentsarxiv-2605.19932 Sparse Blocked context onlyMay 19, 2026
- Stitched Value Model for Diffusion Alignmentarxiv-2605.19804 Sparse Blocked context onlyMay 19, 2026
- OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyondarxiv-2605.19660 Direct Blocked context onlyMay 19, 2026
- optimize_anything: A Universal API for Optimizing any Text Parameterarxiv-2605.19633 Sparse Blocked context onlyMay 19, 2026
- LLMEval-Logic: A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardeningarxiv-2605.19597 Sparse Blocked context onlyMay 19, 2026
- ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparisonarxiv-2605.20278 Sparse Blocked context onlyMay 19, 2026
- MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimizationarxiv-2605.19330 Sparse Blocked context onlyMay 19, 2026
- Rethinking Muon Beyond Pretraining: Spectral Failures and High-Pass Remedies for VLA and RLVRarxiv-2605.19282 Sparse Blocked context onlyMay 19, 2026
- DataPrep-Bench: Benchmarking LLMs as Training Data Preparatorsarxiv-2607.20465 Sparse Blocked context onlyMay 19, 2026
- A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlookarxiv-2605.20266 Sparse Blocked context onlyMay 18, 2026
- Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videosarxiv-2605.18984 Sparse Blocked context onlyMay 18, 2026
- Aurora: Unified Video Editing with a Tool-Using Agentarxiv-2605.18748 Sparse Blocked context onlyMay 18, 2026