300 canonical paper links on this archive page.
- Code as Agent Harnessarxiv-2605.18747 Curated Related Blocked context onlyMay 18, 2026
- ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Looparxiv-2605.18746 Sparse Blocked context onlyMay 18, 2026
- Actionable World Representationarxiv-2605.18743 Sparse Blocked context onlyMay 18, 2026
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generationarxiv-2605.18739 Curated Related Blocked context onlyMay 18, 2026
- EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RLarxiv-2605.18703 Sparse Blocked context onlyMay 18, 2026
- Lance: Unified Multimodal Modeling by Multi-Task Synergyarxiv-2605.18678 Direct Blocked context onlyMay 18, 2026
- MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agentsarxiv-2605.18652 Sparse Blocked context onlyMay 18, 2026
- Language-Switching Triggers Take a Latent Detour Through Language Modelsarxiv-2605.18646 Sparse Blocked context onlyMay 18, 2026
- Post-Trained MoE Can Skip Half Experts via Self-Distillationarxiv-2605.18643 Sparse Blocked context onlyMay 18, 2026
- SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Sciencearxiv-2605.18630 Sparse Blocked context onlyMay 18, 2026
- Incantation: Natural Language as the Action Interface for Multi-Entity Video World Modelsarxiv-2605.18601 Direct Blocked context onlyMay 18, 2026
- OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understandingarxiv-2605.18577 Sparse Blocked context onlyMay 18, 2026
- It Takes Two: Complementary Self-Distillation for Contextual Integrity in LLMsarxiv-2605.20258 Sparse Blocked context onlyMay 18, 2026
- SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolutionarxiv-2605.18401 Sparse Blocked context onlyMay 18, 2026
- StableVLA: Towards Robust Vision-Language-Action Models without Extra Dataarxiv-2605.18287 Sparse Blocked context onlyMay 18, 2026
- Context Memorization for Efficient Long Context Generationarxiv-2605.18226 Sparse Blocked context onlyMay 18, 2026
- SIREM: Speech-Informed MRI Reconstruction with Learned Samplingarxiv-2605.18221 Sparse Blocked context onlyMay 18, 2026
- KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inferencearxiv-2605.18071 Sparse Blocked context onlyMay 18, 2026
- See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understandingarxiv-2605.18018 Sparse Blocked context onlyMay 18, 2026
- AtlasVA: Self-Evolving Visual Skill Memory for Teacher-Free VLM Agentsarxiv-2605.17933 Sparse Blocked context onlyMay 18, 2026
- Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDAarxiv-2605.17932 Sparse Blocked context onlyMay 18, 2026
- Evaluating Cognitive Age Alignment in Interactive AI Agentsarxiv-2605.17894 Sparse Blocked context onlyMay 18, 2026
- HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agentsarxiv-2605.17873 Sparse Blocked context onlyMay 18, 2026
- SNLP: Layer-Parallel Inference via Structured Newton Correctionsarxiv-2605.17842 Sparse Blocked context onlyMay 18, 2026
- Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Searcharxiv-2605.20244 Sparse Blocked context onlyMay 18, 2026
- Interactive Evaluation Requires a Design Sciencearxiv-2605.17829 Sparse Blocked context onlyMay 18, 2026
- Internalizing Tool Knowledge in Small Language Models via QLoRA Fine-Tuningarxiv-2605.17774 Sparse Blocked context onlyMay 18, 2026
- Entropy-Gradient Inversion: Moving Toward Internal Mechanism of Large Reasoning Modelsarxiv-2605.17770 Sparse Blocked context onlyMay 18, 2026
- OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantizationarxiv-2605.17757 Curated Related Blocked context onlyMay 18, 2026
- Agent Bazaar: Enabling Economic Alignment in Multi-Agent Marketplacesarxiv-2605.17698 Sparse Blocked context onlyMay 17, 2026
- Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?arxiv-2605.17694 Sparse Blocked context onlyMay 17, 2026
- Validate Your Authority: Benchmarking LLMs on Multi-Label Precedent Treatment Classificationarxiv-2605.17691 Sparse Blocked context onlyMay 17, 2026
- AI Agents May Always Fall for Prompt Injectionsarxiv-2605.17634 Sparse Blocked context onlyMay 17, 2026
- AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignmentarxiv-2605.17602 Sparse Blocked context onlyMay 17, 2026
- How Off-Policy Can GRPO Be? Mu-GRPO for Efficient LLM Reinforcement Learningarxiv-2605.17570 Sparse Blocked context onlyMay 17, 2026
- Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIsarxiv-2605.17558 Sparse Blocked context onlyMay 17, 2026
- RAG-based EEG-to-Text Translation Using Deep Learning and LLMsarxiv-2605.17503 Sparse Blocked context onlyMay 17, 2026
- VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systemsarxiv-2605.17467 Sparse Blocked context onlyMay 17, 2026
- FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsingarxiv-2605.17447 Sparse Blocked context onlyMay 17, 2026
- MemRepair: Hierarchical Memory for Agentic Repository-Level Vulnerability Repairarxiv-2605.17444 Sparse Blocked context onlyMay 17, 2026
- BELIEF: Structured Evidence Modeling and Uncertainty-Aware Fusion for Biomedical Question Answeringarxiv-2605.17435 Sparse Blocked context onlyMay 17, 2026
- QQJ: Quantifying Qualitative Judgment for Scalable and Human-Aligned Evaluation of Generative AIarxiv-2605.17382 Sparse Blocked context onlyMay 17, 2026
- Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interactionarxiv-2605.17360 Sparse Blocked context onlyMay 17, 2026
- Learning Transferable Topology Priors for Multi-Agent LLM Collaboration Across Domainsarxiv-2605.17359 Sparse Blocked context onlyMay 17, 2026
- AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answeringarxiv-2605.17352 Sparse Blocked context onlyMay 17, 2026
- Taming "Zombie'' Agents: A Markov State-Aware Framework for Resilient Multi-Agent Evolutionarxiv-2605.17348 Sparse Blocked context onlyMay 17, 2026
- Transitivity Meets Cyclicity: Explicit Preference Decomposition for Dynamic Large Language Model Alignmentarxiv-2605.17342 Sparse Blocked context onlyMay 17, 2026
- CyberCorrect: A Cybernetic Framework for Closed-Loop Self-Correction in Large Language Modelsarxiv-2605.17305 Sparse Blocked context onlyMay 17, 2026
- DISA: Offline Importance Sampling for Distribution-Matching LLM-RLarxiv-2605.17295 Sparse Blocked context onlyMay 17, 2026
- OProver: A Unified Framework for Agentic Formal Theorem Provingarxiv-2605.17283 Curated Related Blocked context onlyMay 17, 2026
- LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMsarxiv-2605.17260 Sparse Blocked context onlyMay 17, 2026
- From Runnable to Shippable: Multi-Agent Test-Driven Development for Generating Full-Stack Web Applications from Requirementsarxiv-2605.17242 Sparse Blocked context onlyMay 17, 2026
- Artificial Intolerance: Stigmatizing Language in Clinical Documentation Skews Large Language Model Decision-Makingarxiv-2605.17228 Sparse Blocked context onlyMay 17, 2026
- OpenJarvis: Personal AI, On Personal Devicesarxiv-2605.17172 Direct Blocked context onlyMay 16, 2026
- The Point of No Return: Counterfactual Localization of Deceptive Commitment in Language-Model Reasoningarxiv-2605.17113 Sparse Blocked context onlyMay 16, 2026
- Capturing LLM Capabilities via Evidence-Calibrated Query Clusteringarxiv-2605.17110 Sparse Blocked context onlyMay 16, 2026
- HEED: Density-Weighted Residual Alignment for Hybrid Vision-Language Model Distillationarxiv-2605.17093 Sparse Blocked context onlyMay 16, 2026
- RAGA: Reading-And-Graph-building-Agent for Autonomous Knowledge Graph Construction and Retrieval-Augmented Generationarxiv-2605.17072 Sparse Blocked context onlyMay 16, 2026
- D$^2$Evo: Dual Difficulty-Aware Self-Evolution for Data-Efficient Reinforcement Learningarxiv-2605.17037 Sparse Blocked context onlyMay 16, 2026
- Why Do Reasoning Models Lose Coverage? The Role of Data and Forks in the Roadarxiv-2605.17026 Sparse Blocked context onlyMay 16, 2026
- HalluScore: Large Language Model Hallucination Question Answering Benchmarkarxiv-2605.17007 Sparse Blocked context onlyMay 16, 2026
- Evaluation Drift in LLM Personality Induction: Are We Moving the Goalpost?arxiv-2605.16996 Sparse Blocked context onlyMay 16, 2026
- MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexingarxiv-2605.23986 Sparse Blocked context onlyMay 16, 2026
- Roll Out and Roll Back: Diffusion LLMs are Their Own Efficiency Teachersarxiv-2605.16941 Sparse Blocked context onlyMay 16, 2026
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Stepsarxiv-2605.16928 Sparse Blocked context onlyMay 16, 2026
- DriveSafe: A Framework for Risk Detection and Safety Suggestions in Driving Scenariosarxiv-2605.16892 Sparse Blocked context onlyMay 16, 2026
- PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarksarxiv-2605.16881 Sparse Blocked context onlyMay 16, 2026
- MixSD: Mixed Contextual Self-Distillation for Knowledge Injectionarxiv-2605.16865 Sparse Blocked context onlyMay 16, 2026
- Thinking with Patterns: Breaking the Perceptual Bottleneck in Visual Planning via Pattern Inductionarxiv-2605.16848 Sparse Blocked context onlyMay 16, 2026
- RTI-Bench: A Structured Dataset for Indian Right-to-Information Decision Analysisarxiv-2605.16843 Direct Blocked context onlyMay 16, 2026
- Decoupling KL and Trajectories: A Unified Perspective for SFT, DAgger, Offline RL, and OPD in LLM Distillationarxiv-2605.16826 Sparse Blocked context onlyMay 16, 2026
- Confidence Geometry Reveals Trace-Level Correctness in Large Language Model Reasoningarxiv-2605.16824 Sparse Blocked context onlyMay 16, 2026
- AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agentsarxiv-2605.16819 Sparse Blocked context onlyMay 16, 2026
- FIM-LoRA: Task-Informative Rank Allocation for LoRA via Calibration-Time Gradient-Variance Estimationarxiv-2605.16800 Sparse Blocked context onlyMay 16, 2026
- TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Compositionarxiv-2605.16790 Sparse Blocked context onlyMay 16, 2026
- The Unlearnability Phenomenon in RLVR for Language Modelsarxiv-2605.16787 Sparse Blocked context onlyMay 16, 2026
- Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Freearxiv-2605.16767 Sparse Blocked context onlyMay 16, 2026
- EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformersarxiv-2605.16745 Sparse Blocked context onlyMay 16, 2026
- CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?arxiv-2605.16679 Direct Blocked context onlyMay 15, 2026
- Look Before You Leap: Autonomous Exploration for LLM Agentsarxiv-2605.16143 Sparse Blocked context onlyMay 15, 2026
- PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Controlarxiv-2605.15963 Sparse Blocked context onlyMay 15, 2026
- Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP modelsarxiv-2605.15961 Sparse Blocked context onlyMay 15, 2026
- Unlocking Dense Metric Depth Estimation in VLMsarxiv-2605.15876 Sparse Blocked context onlyMay 15, 2026
- Agentic Discovery of Neural Architectures: AIRA-Compose and AIRA-Designarxiv-2605.15871 Sparse Blocked context onlyMay 15, 2026
- GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactionsarxiv-2605.15764 Sparse Blocked context onlyMay 15, 2026
- DimMem: Dimensional Structuring for Efficient Long-Term Agent Memoryarxiv-2605.15759 Sparse Blocked context onlyMay 15, 2026
- Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVRarxiv-2605.15726 Sparse Blocked context onlyMay 15, 2026
- Rule2DRC: Benchmarking LLM Agents for DRC Script Synthesis with Execution-Guided Test Generationarxiv-2605.15669 Sparse Blocked context onlyMay 15, 2026
- Calibrating LLMs with Semantic-level Rewardarxiv-2605.15588 Sparse Blocked context onlyMay 15, 2026
- Measuring Maximum Activations in Open Large Language Modelsarxiv-2605.15572 Sparse Blocked context onlyMay 15, 2026
- AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMsarxiv-2605.15565 Direct Blocked context onlyMay 15, 2026
- RoPE Distinguishes Neither Positions Nor Tokens in Long Contexts, Provablyarxiv-2605.15514 Sparse Blocked context onlyMay 15, 2026
- STS: Efficient Sparse Attention with Speculative Token Sparsityarxiv-2605.15508 Sparse Blocked context onlyMay 15, 2026
- FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Modelsarxiv-2605.15482 Sparse Blocked context onlyMay 14, 2026
- Video Models Can Reason with Verifiable Rewardsarxiv-2605.15458 Sparse Blocked context onlyMay 14, 2026
- Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolutionarxiv-2605.15301 Sparse Blocked context onlyMay 14, 2026
- PhysBrain 1.0 Technical Reportarxiv-2605.15298 Sparse Blocked context onlyMay 14, 2026
- Aligning Latent Geometry for Spherical Flow Matching in Image Generationarxiv-2605.15193 Sparse Blocked context onlyMay 14, 2026
- RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPOarxiv-2605.15190 Curated Related Blocked context onlyMay 14, 2026
- From Plans to Pixels: Learning to Plan and Orchestrate for Open-Ended Image Editingarxiv-2605.15181 Sparse Blocked context onlyMay 14, 2026
- SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformerarxiv-2605.15178 Direct Blocked context onlyMay 14, 2026
- ReactiveGWM: Steering NPC in Reactive Game World Modelsarxiv-2605.15256 Sparse Blocked context onlyMay 14, 2026
- Self-Distilled Agentic Reinforcement Learningarxiv-2605.15155 Direct Blocked context onlyMay 14, 2026
- Causal Forcing++: Scalable Few-Step Autoregressive Diffusion Distillation for Real-Time Interactive Video Generationarxiv-2605.15141 Direct Blocked context onlyMay 14, 2026
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Modelsarxiv-2605.15055 Sparse Blocked context onlyMay 14, 2026
- Orchard: An Open-Source Agentic Modeling Frameworkarxiv-2605.15040 Direct Blocked context onlyMay 14, 2026
- GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decodingarxiv-2605.15250 Sparse Blocked context onlyMay 14, 2026
- Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowingarxiv-2605.14978 Sparse Blocked context onlyMay 14, 2026
- Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QAarxiv-2605.14928 Sparse Blocked context onlyMay 14, 2026
- MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Modelsarxiv-2605.14906 Sparse Blocked context onlyMay 14, 2026
- Unlocking Complex Visual Generation via Closed-Loop Verified Reasoningarxiv-2605.14876 Sparse Blocked context onlyMay 14, 2026
- Holistic Evaluation and Failure Diagnosis of AI Agentsarxiv-2605.14865 Sparse Blocked context onlyMay 14, 2026
- Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysisarxiv-2605.14842 Sparse Blocked context onlyMay 14, 2026
- Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretrainingarxiv-2605.14747 Sparse Blocked context onlyMay 14, 2026
- IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulationarxiv-2605.14712 Sparse Blocked context onlyMay 14, 2026
- $π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflowsarxiv-2605.14678 Direct Blocked context onlyMay 14, 2026
- Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewardsarxiv-2605.14539 Sparse Blocked context onlyMay 14, 2026
- BEAM: Binary Expert Activation Masking for Dynamic Routing in MoEarxiv-2605.14438 Sparse Blocked context onlyMay 14, 2026
- Learning to Build the Environment: Self-Evolving Reasoning RL via Verifiable Environment Synthesisarxiv-2605.14392 Sparse Blocked context onlyMay 14, 2026
- Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoningarxiv-2605.14386 Sparse Blocked context onlyMay 14, 2026
- MetaAgent-X : Breaking the Ceiling of Automatic Multi-Agent Systems via End-to-End Reinforcement Learningarxiv-2605.14212 Curated Related Blocked context onlyMay 14, 2026
- Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoningarxiv-2605.14040 Sparse Blocked context onlyMay 13, 2026
- Model-Adaptive Tool Necessity Reveals the Knowing-Doing Gap in LLM Tool Usearxiv-2605.14038 Curated Related Blocked context onlyMay 13, 2026
- Mistletoe: Stealthy Acceleration-Collapse Attacks on Speculative Decodingarxiv-2605.14005 Sparse Blocked context onlyMay 13, 2026
- VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Usearxiv-2605.13989 Sparse Blocked context onlyMay 13, 2026
- Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Contextarxiv-2605.13831 Sparse Blocked context onlyMay 13, 2026
- EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agentsarxiv-2605.13941 Direct Blocked context onlyMay 13, 2026
- AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillationarxiv-2605.13724 Direct Blocked context onlyMay 13, 2026
- FlowCompile: An Optimizing Compiler for Structured LLM Workflowsarxiv-2605.13647 Sparse Blocked context onlyMay 13, 2026
- Multilingual OCR-Aware Fine-Tuning and Prompt-Guided Chain-of-Thought Reasoning for Multimodal Large Language Modelsarxiv-2605.16409 Sparse Blocked context onlyMay 13, 2026
- Qwen-Image-VAE-2.0 Technical Reportarxiv-2605.13565 Sparse Blocked context onlyMay 13, 2026
- RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitationarxiv-2605.13542 Sparse Blocked context onlyMay 13, 2026
- Many-Shot CoT-ICL: Making In-Context Learning Truly Learnarxiv-2605.13511 Sparse Blocked context onlyMay 13, 2026
- PersonalAI 2.0: Enhancing knowledge graph traversal/retrieval with planning mechanism for Personalized LLM Agentsarxiv-2605.13481 Sparse Blocked context onlyMay 13, 2026
- CogniFold: Always-On Proactive Memory via Cognitive Foldingarxiv-2605.13438 Direct Blocked context onlyMay 13, 2026
- Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Rankingarxiv-2605.18852 Sparse Blocked context onlyMay 13, 2026
- Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Callingarxiv-2605.13360 Sparse Blocked context onlyMay 13, 2026
- PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama Worldarxiv-2605.13169 Sparse Blocked context onlyMay 13, 2026
- Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modelingarxiv-2605.13062 Curated Related Blocked context onlyMay 13, 2026
- Context Training with Active Information Seekingarxiv-2605.13050 Sparse Blocked context onlyMay 13, 2026
- MAP: A Map-then-Act Paradigm for Long-Horizon Interactive Agent Reasoningarxiv-2605.13037 Sparse Blocked context onlyMay 13, 2026
- When Vision Speaks for Soundarxiv-2605.16403 Sparse Blocked context onlyMay 13, 2026
- Useful Memories Become Faulty When Continuously Updated by LLMsarxiv-2605.12978 Sparse Blocked context onlyMay 13, 2026
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluationarxiv-2605.12925 Sparse Blocked context onlyMay 13, 2026
- Revisiting DAgger in the Era of LLM-Agentsarxiv-2605.12913 Sparse Blocked context onlyMay 13, 2026
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligencearxiv-2605.12882 Sparse Blocked context onlyMay 13, 2026
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusionarxiv-2605.12825 Direct Blocked context onlyMay 12, 2026
- Training Large Language Models to Predict Clinical Eventsarxiv-2605.12817 Sparse Blocked context onlyMay 12, 2026
- Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffoldsarxiv-2605.18827 Sparse Blocked context onlyMay 12, 2026
- Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?arxiv-2605.12684 Sparse Blocked context onlyMay 12, 2026
- Covering Human Action Space for Computer Use: Data Synthesis and Benchmarkarxiv-2605.12501 Sparse Blocked context onlyMay 12, 2026
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecturearxiv-2605.12500 Sparse Blocked context onlyMay 12, 2026
- From Web to Pixels: Bringing Agentic Search into Visual Perceptionarxiv-2605.12497 Sparse Blocked context onlyMay 12, 2026
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Rewardarxiv-2605.12495 Sparse Blocked context onlyMay 12, 2026
- TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Trackingarxiv-2605.12587 Sparse Blocked context onlyMay 12, 2026
- Learning, Fast and Slow: Towards LLMs That Adapt Continuallyarxiv-2605.12484 Sparse Blocked context onlyMay 12, 2026
- ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agentsarxiv-2605.12481 Sparse Blocked context onlyMay 12, 2026
- Solve the Loop: Attractor Models for Language and Reasoningarxiv-2605.12466 Sparse Blocked context onlyMay 12, 2026
- Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputsarxiv-2605.12460 Sparse Blocked context onlyMay 12, 2026
- TextSeal: A Localized LLM Watermark for Provenance & Distillation Protectionarxiv-2605.12456 Sparse Blocked context onlyMay 12, 2026
- ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Mergingarxiv-2605.12419 Sparse Blocked context onlyMay 12, 2026
- Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modelingarxiv-2605.12411 Sparse Blocked context onlyMay 12, 2026
- $δ$-mem: Efficient Online Memory for Large Language Modelsarxiv-2605.12357 Sparse Blocked context onlyMay 12, 2026
- Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generationarxiv-2605.12305 Sparse Blocked context onlyMay 12, 2026
- Targeted Neuron Modulation via Contrastive Pair Searcharxiv-2605.12290 Sparse Blocked context onlyMay 12, 2026
- TokenRatio: Principled Token-Level Preference Optimization via Ratio Matchingarxiv-2605.12288 Sparse Blocked context onlyMay 12, 2026
- Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correctionarxiv-2605.12070 Sparse Blocked context onlyMay 12, 2026
- Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluationarxiv-2605.12034 Sparse Blocked context onlyMay 12, 2026
- Learning Agentic Policy from Action Guidancearxiv-2605.12004 Sparse Blocked context onlyMay 12, 2026
- On-Policy Self-Evolution via Failure Trajectories for Agentic Safety Alignmentarxiv-2605.11882 Sparse Blocked context onlyMay 12, 2026
- Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Modelsarxiv-2605.11854 Sparse Blocked context onlyMay 12, 2026
- Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillationarxiv-2605.11739 Sparse Blocked context onlyMay 12, 2026
- Position: LLM Inference Should Be Evaluated as Energy-to-Token Productionarxiv-2605.11733 Sparse Blocked context onlyMay 12, 2026
- AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensivearxiv-2605.11518 Sparse Blocked context onlyMay 12, 2026
- Agent-BRACE: Decoupling Beliefs from Actions in Long-Horizon Tasks via Verbalized State Uncertaintyarxiv-2605.11436 Sparse Blocked context onlyMay 12, 2026
- UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoningarxiv-2605.11400 Sparse Blocked context onlyMay 12, 2026
- An Empirical Study of Automating Agent Evaluationarxiv-2605.11378 Sparse Blocked context onlyMay 12, 2026
- The Many Faces of On-Policy Distillation: Pitfalls, Mechanisms, and Fixesarxiv-2605.11182 Sparse Blocked context onlyMay 11, 2026
- EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scalesarxiv-2605.11136 Sparse Blocked context onlyMay 11, 2026
- Dynamic Skill Lifecycle Management for Agentic Reinforcement Learningarxiv-2605.10923 Sparse Blocked context onlyMay 11, 2026
- RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmarkarxiv-2605.10921 Sparse Blocked context onlyMay 11, 2026
- Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Tracearxiv-2605.10913 Direct Blocked context onlyMay 11, 2026
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluationarxiv-2605.10912 Sparse Blocked context onlyMay 11, 2026
- Rethinking Agentic Search with Pi-Serini: Is Lexical Retrieval Sufficient?arxiv-2605.10848 Sparse Blocked context onlyMay 11, 2026
- Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVRarxiv-2605.10781 Sparse Blocked context onlyMay 11, 2026
- Auditing Multimodal LLM Raters: Central Tendency Bias in Clinical Ordinal Scoringarxiv-2605.16386 Sparse Blocked context onlyMay 11, 2026
- Qwen-Image-2.0 Technical Reportarxiv-2605.10730 Sparse Blocked context onlyMay 11, 2026
- Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoningarxiv-2605.16385 Sparse Blocked context onlyMay 11, 2026
- DeepRefine: Agent-Compiled Knowledge Refinement via Reinforcement Learningarxiv-2605.10488 Sparse Blocked context onlyMay 11, 2026
- SlimSpec: Low-Rank Draft LM-Head for Accelerated Speculative Decodingarxiv-2605.10453 Sparse Blocked context onlyMay 11, 2026
- WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictorsarxiv-2605.10434 Sparse Blocked context onlyMay 11, 2026
- SleepWalk: A Three-Tier Benchmark for Stress-Testing Instruction-Guided Vision-Language Navigationarxiv-2605.10376 Sparse Blocked context onlyMay 11, 2026
- Agent-ValueBench: A Comprehensive Benchmark for Evaluating Agent Valuesarxiv-2605.10365 Sparse Blocked context onlyMay 11, 2026
- TMAS: Scaling Test-Time Compute via Multi-Agent Synergyarxiv-2605.10344 Sparse Blocked context onlyMay 11, 2026
- PaperFit: Vision-in-the-Loop Typesetting Optimization for Scientific Documentsarxiv-2605.10341 Direct Blocked context onlyMay 11, 2026
- MemReread: Enhancing Agentic Long-Context Reasoning via Memory-Guided Rereadingarxiv-2605.10268 Sparse Blocked context onlyMay 11, 2026
- IndustryBench: Probing the Industrial Knowledge Boundaries of LLMsarxiv-2605.10267 Sparse Blocked context onlyMay 11, 2026
- Unsupervised Process Reward Modelsarxiv-2605.10158 Sparse Blocked context onlyMay 11, 2026
- Personalizing LLMs with Binary Feedback: A Preference-Corrected Optimization Frameworkarxiv-2605.10043 Sparse Blocked context onlyMay 11, 2026
- PlantMarkerBench: A Multi-Species Benchmark for Evidence-Grounded Plant Marker Reasoningarxiv-2605.10032 Sparse Blocked context onlyMay 11, 2026
- Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalizationarxiv-2605.09996 Sparse Blocked context onlyMay 11, 2026
- Merlin: Deterministic Byte-Exact Deduplication for Lossless Context Optimization in Large Language Model Inferencearxiv-2605.09990 Sparse Blocked context onlyMay 11, 2026
- HAGE: Harnessing Agentic Memory via RL-Driven Weighted Graph Evolutionarxiv-2605.09942 Sparse Blocked context onlyMay 11, 2026
- Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perceptionarxiv-2605.09936 Direct Blocked context onlyMay 11, 2026
- FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuningarxiv-2605.09932 Sparse Blocked context onlyMay 11, 2026
- Key-Value Meansarxiv-2605.09877 Direct Blocked context onlyMay 11, 2026
- The Metacognitive Probe: Five Behavioural Calibration Diagnostics for LLMsarxiv-2605.09844 Sparse Blocked context onlyMay 11, 2026
- Dystruct: Dynamically Structured Diffusion Language Model Decoding via Bayesian Inferencearxiv-2605.09820 Sparse Blocked context onlyMay 10, 2026
- LEAD: Length-Efficient Adaptive and Dynamic Reasoning for Large Language Modelsarxiv-2605.09806 Sparse Blocked context onlyMay 10, 2026
- EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descentarxiv-2605.09777 Sparse Blocked context onlyMay 10, 2026
- Learning Multi-Indicator Weights for Data Selection: A Joint Task-Model Adaptation Framework with Efficient Proxiesarxiv-2605.09665 Sparse Blocked context onlyMay 10, 2026
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMsarxiv-2605.09635 Sparse Blocked context onlyMay 10, 2026
- Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocolsarxiv-2605.09618 Sparse Blocked context onlyMay 10, 2026
- Edit-Based Refinement for Parallel Masked Diffusion Language Modelsarxiv-2605.09603 Sparse Blocked context onlyMay 10, 2026
- Crosslingual On-Policy Self-Distillation for Multilingual Reasoningarxiv-2605.09548 Sparse Blocked context onlyMay 10, 2026
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agentsarxiv-2605.09530 Curated Related Blocked context onlyMay 10, 2026
- Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoningarxiv-2605.09490 Sparse Blocked context onlyMay 10, 2026
- SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learningarxiv-2605.09423 Sparse Blocked context onlyMay 10, 2026
- DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verificationarxiv-2605.09269 Sparse Blocked context onlyMay 10, 2026
- SeePhys Pro: Diagnosing Modality Transfer and Blind-Training Effects in Multimodal RLVR for Physics Reasoningarxiv-2605.09266 Sparse Blocked context onlyMay 10, 2026
- Reinforcing Multimodal Reasoning Against Visual Degradationarxiv-2605.09262 Sparse Blocked context onlyMay 10, 2026
- LLM Agents Already Know When to Call Tools -- Even Without Reasoningarxiv-2605.09252 Sparse Blocked context onlyMay 10, 2026
- RigidFormer: Learning Rigid Dynamics using Transformersarxiv-2605.09196 Sparse Blocked context onlyMay 9, 2026
- MCP-Cosmos: World Model-Augmented Agents for Complex Task Execution in MCP Environmentsarxiv-2605.09131 Sparse Blocked context onlyMay 9, 2026
- LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?arxiv-2605.08985 Sparse Blocked context onlyMay 9, 2026
- Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimizationarxiv-2605.08978 Direct Blocked context onlyMay 9, 2026
- SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-trainingarxiv-2605.08738 Sparse Blocked context onlyMay 9, 2026
- The Extrapolation Cliff in On-Policy Distillation of Near-Deterministic Structured Outputsarxiv-2605.08737 Sparse Blocked context onlyMay 9, 2026
- AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systemsarxiv-2605.08715 Sparse Blocked context onlyMay 9, 2026
- RewardHarness: Self-Evolving Agentic Post-Trainingarxiv-2605.08703 Sparse Blocked context onlyMay 9, 2026
- MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AIarxiv-2605.08678 Sparse Blocked context onlyMay 9, 2026
- PAAC: Privacy-Aware Agentic Device-Cloud Collaborationarxiv-2605.08646 Sparse Blocked context onlyMay 9, 2026
- Large Language Models over Networks: Collaborative Intelligence under Resource Constraintsarxiv-2605.08626 Sparse Blocked context onlyMay 9, 2026
- Source or It Didn't Happen: A Multi-Agent Framework for Citation Hallucination Detectionarxiv-2605.08583 Sparse Blocked context onlyMay 9, 2026
- A Single Neuron Is Sufficient to Bypass Safety Alignment in Large Language Modelsarxiv-2605.08513 Sparse Blocked context onlyMay 8, 2026
- A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Modelsarxiv-2605.08504 Sparse Blocked context onlyMay 8, 2026
- Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Modelsarxiv-2605.08472 Sparse Blocked context onlyMay 8, 2026
- LLMs Improving LLMs: Agentic Discovery for Test-Time Scalingarxiv-2605.08083 Sparse Blocked context onlyMay 8, 2026
- Flow-OPD: On-Policy Distillation for Flow Matching Modelsarxiv-2605.08063 Sparse Blocked context onlyMay 8, 2026
- Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoningarxiv-2605.08061 Sparse Blocked context onlyMay 8, 2026
- CA-SQL: Complexity-Aware Inference Time Reasoning for Text-to-SQL via Exploration and Compute Budget Allocationarxiv-2605.08057 Sparse Blocked context onlyMay 8, 2026
- Fast Byte Latent Transformerarxiv-2605.08044 Sparse Blocked context onlyMay 8, 2026
- STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generationarxiv-2605.08029 Sparse Blocked context onlyMay 8, 2026
- Tool Calling is Linearly Readable and Steerable in Language Modelsarxiv-2605.07990 Sparse Blocked context onlyMay 8, 2026
- Ask Early, Ask Late, Ask Right: When Does Clarification Timing Matter for Long-Horizon Agents?arxiv-2605.07937 Sparse Blocked context onlyMay 8, 2026
- Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillationarxiv-2605.07924 Sparse Blocked context onlyMay 8, 2026
- Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancementarxiv-2605.07883 Sparse Blocked context onlyMay 8, 2026
- KL for a KL: On-Policy Distillation with Control Variate Baselinearxiv-2605.07865 Sparse Blocked context onlyMay 8, 2026
- Measuring and Mitigating the Distributional Gap Between Real and Simulated User Behaviorsarxiv-2605.07847 Sparse Blocked context onlyMay 8, 2026
- Anisotropic Modality Alignarxiv-2605.07825 Sparse Blocked context onlyMay 8, 2026
- CktFormalizer: Autoformalization of Natural Language into Circuit Representationsarxiv-2605.07782 Sparse Blocked context onlyMay 8, 2026
- Tracing Uncertainty in Language Model "Reasoning"arxiv-2605.07776 Sparse Blocked context onlyMay 8, 2026
- TextLDM: Language Modeling with Continuous Latent Diffusionarxiv-2605.07748 Sparse Blocked context onlyMay 8, 2026
- Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMsarxiv-2605.07731 Sparse Blocked context onlyMay 8, 2026
- SOD: Step-wise On-policy Distillation for Small Language Model Agentsarxiv-2605.07725 Sparse Blocked context onlyMay 8, 2026
- Memory-Efficient Looped Transformer: Decoupling Compute from Memory in Looped Language Modelsarxiv-2605.07721 Sparse Blocked context onlyMay 8, 2026
- Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoningarxiv-2605.07660 Sparse Blocked context onlyMay 8, 2026
- Reliable Chain-of-Thought via Prefix Consistencyarxiv-2605.07654 Sparse Blocked context onlyMay 8, 2026
- Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptationarxiv-2605.07647 Sparse Blocked context onlyMay 8, 2026
- Intent-Driven Semantic ID Generation for Grounded Conversational News Recommendationarxiv-2605.07613 Sparse Blocked context onlyMay 8, 2026
- Mathematical Reasoning via Intervention-Based Time-Series Causal Discovery Using LLMs as Concept Mastery Simulatorsarxiv-2605.07600 Sparse Blocked context onlyMay 8, 2026
- Your Language Model is Its Own Critic: Reinforcement Learning with Value Estimation from Actor's Internal Statesarxiv-2605.07579 Sparse Blocked context onlyMay 8, 2026
- Tracing the Arrow of Time: Diagnosing Temporal Information Flow in Video-LLMsarxiv-2605.07568 Sparse Blocked context onlyMay 8, 2026
- ExpThink: Experience-Guided Reinforcement Learning for Adaptive Chain-of-Thought Compressionarxiv-2605.07501 Sparse Blocked context onlyMay 8, 2026
- Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidancearxiv-2605.07461 Sparse Blocked context onlyMay 8, 2026
- Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMsarxiv-2605.07447 Sparse Blocked context onlyMay 8, 2026
- ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referringarxiv-2605.07415 Sparse Blocked context onlyMay 8, 2026
- Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifactsarxiv-2605.07395 Sparse Blocked context onlyMay 8, 2026
- BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioningarxiv-2605.07394 Sparse Blocked context onlyMay 8, 2026
- Gradient-Based LoRA Rank Allocation Under GRPO: An Empirical Studyarxiv-2605.07366 Sparse Blocked context onlyMay 8, 2026
- MISA: Mixture of Indexer Sparse Attention for Long-Context LLM Inferencearxiv-2605.07363 Sparse Blocked context onlyMay 8, 2026
- LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verificationarxiv-2605.07315 Sparse Blocked context onlyMay 8, 2026
- PaT: Planning-after-Trial for Efficient Test-Time Code Generationarxiv-2605.07248 Sparse Blocked context onlyMay 8, 2026
- Teaching Language Models to Think in Codearxiv-2605.07237 Sparse Blocked context onlyMay 8, 2026
- DiffRetriever: Parallel Representative Tokens for Retrieval with Diffusion Language Modelsarxiv-2605.07210 Sparse Blocked context onlyMay 8, 2026
- Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMsarxiv-2605.07153 Sparse Blocked context onlyMay 8, 2026
- Structural Rationale Distillation via Reasoning Space Compressionarxiv-2605.07139 Sparse Blocked context onlyMay 8, 2026
- Region4Web: Rethinking Observation Space Granularity for Web Agentsarxiv-2605.07134 Sparse Blocked context onlyMay 8, 2026
- The Position Curse: LLMs Struggle to Locate the Last Few Items in a Listarxiv-2605.07127 Sparse Blocked context onlyMay 8, 2026
- ModelLens: Finding the Best for Your Task from Myriads of Modelsarxiv-2605.07075 Sparse Blocked context onlyMay 8, 2026
- PACEvolve++: Improving Test-time Learning for Evolutionary Search Agentsarxiv-2605.07039 Sparse Blocked context onlyMay 7, 2026
- A$^2$RD: Agentic Autoregressive Diffusion for Long Video Consistencyarxiv-2605.06924 Sparse Blocked context onlyMay 7, 2026
- Conformal Agent Error Attributionarxiv-2605.06788 Sparse Blocked context onlyMay 7, 2026
- UniPool: A Globally Shared Expert Pool for Mixture-of-Expertsarxiv-2605.06665 Sparse Blocked context onlyMay 7, 2026
- EMO: Pretraining Mixture of Experts for Emergent Modularityarxiv-2605.06663 Sparse Blocked context onlyMay 7, 2026
- Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradientsarxiv-2605.06650 Sparse Blocked context onlyMay 7, 2026
- StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstractionarxiv-2605.06642 Sparse Blocked context onlyMay 7, 2026
- Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Studyarxiv-2605.06643 Sparse Blocked context onlyMay 7, 2026
- Can RL Teach Long-Horizon Reasoning to LLMs? Expressiveness Is Keyarxiv-2605.06638 Sparse Blocked context onlyMay 7, 2026
- Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agentsarxiv-2605.06635 Sparse Blocked context onlyMay 7, 2026
- UniSD: Towards a Unified Self-Distillation Framework for Large Language Modelsarxiv-2605.06597 Direct Blocked context onlyMay 7, 2026
- Long Context Pre-Training with Lighthouse Attentionarxiv-2605.06554 Sparse Blocked context onlyMay 7, 2026
- Continuous Latent Diffusion Language Modelarxiv-2605.06548 Sparse Blocked context onlyMay 7, 2026
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?arxiv-2605.06527 Sparse Blocked context onlyMay 7, 2026
- PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitorsarxiv-2605.06455 Sparse Blocked context onlyMay 7, 2026
- HumanNet: Scaling Human-centric Video Learning to One Million Hoursarxiv-2605.06747 Direct Blocked context onlyMay 7, 2026
- Don't Lose Focus: Activation Steering via Key-Orthogonal Projectionsarxiv-2605.06342 Sparse Blocked context onlyMay 7, 2026
- MANTRA: Synthesizing SMT-Validated Compliance Benchmarks for Tool-Using LLM Agentsarxiv-2605.06334 Sparse Blocked context onlyMay 7, 2026
- Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneityarxiv-2605.06327 Sparse Blocked context onlyMay 7, 2026
- Improving the Efficiency of Language Agent Teams with Adaptive Task Graphsarxiv-2605.06320 Sparse Blocked context onlyMay 7, 2026