300 canonical paper links on this archive page.
- GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluationarxiv-2606.22737 Sparse Blocked context onlyJun 22, 2026
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Doarxiv-2606.22565 Sparse Blocked context onlyJun 21, 2026
- VADAOrchestra: Neurosymbolic Orchestration of Adaptive Reasoning Workflowsarxiv-2606.22485 Sparse Blocked context onlyJun 21, 2026
- Knowledge-Graph Grounding Helps LLMs Only for Out-of-Training Knowledge: A Controlled Study on Clinical Question Answeringarxiv-2606.22419 Sparse Blocked context onlyJun 21, 2026
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystemsarxiv-2606.22388 Sparse Blocked context onlyJun 21, 2026
- Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agentsarxiv-2606.22207 Sparse Blocked context onlyJun 20, 2026
- BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Languagearxiv-2606.22138 Sparse Blocked context onlyJun 20, 2026
- Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decodingarxiv-2606.21906 Sparse Blocked context onlyJun 20, 2026
- CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasksarxiv-2606.21777 Sparse Blocked context onlyJun 19, 2026
- EvoEmbedding: Evolvable Representations for Long-Context Retrieval and Agentic Memoryarxiv-2606.21649 Sparse Blocked context onlyJun 19, 2026
- How Transparent is DiffusionGemma?arxiv-2606.20560 Sparse Blocked context onlyJun 18, 2026
- UNIEGO: Proxies as Mediators for Unified Egocentric Video Representation Learningarxiv-2606.20559 Sparse Blocked context onlyJun 18, 2026
- LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agentsarxiv-2606.20529 Sparse Blocked context onlyJun 18, 2026
- StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMsarxiv-2606.20527 Sparse Blocked context onlyJun 18, 2026
- HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretrainingarxiv-2606.20521 Sparse Blocked context onlyJun 18, 2026
- Multi-LCB: Extending LiveCodeBench to Multiple Programming Languagesarxiv-2606.20517 Sparse Blocked context onlyJun 18, 2026
- FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Miningarxiv-2606.20506 Sparse Blocked context onlyJun 18, 2026
- Beyond Global Replanning: Hierarchical Recovery for Cross-Device Agent Systemsarxiv-2606.20487 Sparse Blocked context onlyJun 18, 2026
- Your Mouse and Eyes Secretly Leak Your Preference: LLM Alignment using Implicit Feedback from Usersarxiv-2606.20482 Sparse Blocked context onlyJun 18, 2026
- Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiologyarxiv-2606.20477 Sparse Blocked context onlyJun 18, 2026
- FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flowsarxiv-2606.20404 Sparse Blocked context onlyJun 18, 2026
- PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedbackarxiv-2606.20287 Sparse Blocked context onlyJun 18, 2026
- Actionable Activation Directions for Detecting and Mitigating Emergent Misalignment Across Language Model Familiesarxiv-2606.20225 Sparse Blocked context onlyJun 18, 2026
- MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimizationarxiv-2606.20164 Sparse Blocked context onlyJun 18, 2026
- When Does Streaming Tool Use Help? Characterizing Tool-Intent Stabilization in Streaming Retrieval-Augmented Generationarxiv-2606.20113 Sparse Blocked context onlyJun 18, 2026
- HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridizationarxiv-2606.20097 Sparse Blocked context onlyJun 18, 2026
- Holo-World: Unified Camera, Object and Weather Control for Video World Modelarxiv-2606.20083 Sparse Blocked context onlyJun 18, 2026
- What Makes Effective Supervision in Latent Chain-of-Thought: An Information-Theoretic Analysisarxiv-2606.20075 Sparse Blocked context onlyJun 18, 2026
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learningarxiv-2606.20002 Sparse Blocked context onlyJun 18, 2026
- Multi-Agent Transactive Memoryarxiv-2606.19911 Sparse Blocked context onlyJun 18, 2026
- Large Language Models Do Not Always Need Readable Languagearxiv-2606.19857 Sparse Blocked context onlyJun 18, 2026
- Prompt, Plan, Extract: Zero-Shot Agentic LLMs Workflows for Lung Pathology Extraction from Clinical Narrativesarxiv-2606.19852 Sparse Blocked context onlyJun 18, 2026
- AtomMem: Building Simple and Effective Memory System for LLM Agents via Atomic Factsarxiv-2606.19847 Sparse Blocked context onlyJun 18, 2026
- CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Modelsarxiv-2606.19788 Sparse Blocked context onlyJun 18, 2026
- Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Modelsarxiv-2606.19750 Sparse Blocked context onlyJun 18, 2026
- NRITYAM: Language Models Meet Art and Heritage of Dancearxiv-2606.19727 Sparse Blocked context onlyJun 18, 2026
- NEST: Narrative Event Structures in Time for Long Video Understandingarxiv-2606.19706 Sparse Blocked context onlyJun 18, 2026
- Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agentsarxiv-2606.19704 Sparse Blocked context onlyJun 18, 2026
- Efficiently Representing Algorithms With Chain-of-Thought Transformersarxiv-2606.19697 Sparse Blocked context onlyJun 18, 2026
- Uncertainty Decomposition for Clarification Seeking in LLM Agentsarxiv-2606.19559 Curated Related Blocked context onlyJun 17, 2026
- PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Modelsarxiv-2606.19534 Curated Related Blocked context onlyJun 17, 2026
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?arxiv-2606.19531 Sparse Blocked context onlyJun 17, 2026
- Diffusion Language Models: An Experimental Analysisarxiv-2606.19475 Sparse Blocked context onlyJun 17, 2026
- Native Active Perception as Reasoning for Omni-Modal Understandingarxiv-2606.19341 Direct Blocked context onlyJun 17, 2026
- Learning User Simulators with Turing Rewardsarxiv-2606.19336 Sparse Blocked context onlyJun 17, 2026
- Playful Agentic Robot Learningarxiv-2606.19419 Sparse Blocked context onlyJun 17, 2026
- Rethinking Reward Supervision: Rubric-Conditioned Self-Distillationarxiv-2606.19327 Sparse Blocked context onlyJun 17, 2026
- Enhancing Decision-Making with Large Language Models through Multi-Agent Fictitious Playarxiv-2606.19308 Sparse Blocked context onlyJun 17, 2026
- Trade-offs in Medical LLM Adaptation: An Empirical Study in French QAarxiv-2606.19266 Curated Related Blocked context onlyJun 17, 2026
- Structured Inference with Large Language Gibbsarxiv-2606.19264 Sparse Blocked context onlyJun 17, 2026
- DreamReasoner-8B: Block-Size Curriculum Learning for Diffusion Reasoning Modelsarxiv-2606.19257 Sparse Blocked context onlyJun 17, 2026
- STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stabilityarxiv-2606.19236 Sparse Blocked context onlyJun 17, 2026
- RECOM: A Validity Discrimination Tradeoff in Automatic Metrics for Open Ended Reddit Question Answeringarxiv-2606.19218 Sparse Blocked context onlyJun 17, 2026
- Written by AI, Managed by AI: Semantic Space Control and Index Sickness Elimination Across 391 Consecutive Sessionsarxiv-2606.19121 Sparse Blocked context onlyJun 17, 2026
- Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillationarxiv-2606.19120 Sparse Blocked context onlyJun 17, 2026
- Sumi: Open Uniform Diffusion Language Model from Scratcharxiv-2606.19005 Sparse Blocked context onlyJun 17, 2026
- G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignmentarxiv-2606.18989 Sparse Blocked context onlyJun 17, 2026
- Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answeringarxiv-2606.18986 Sparse Blocked context onlyJun 17, 2026
- EfficientRollout: System-Aware Self-Speculative Decoding for RL Rolloutsarxiv-2606.18967 Sparse Blocked context onlyJun 17, 2026
- Thermodynamic Signatures of Reasoning: Free-Energy and Spectral-Form-Factor Diagnostics for Hallucination Detection in Large Language Modelsarxiv-2606.19404 Sparse Blocked context onlyJun 17, 2026
- GraphPO: Graph-based Policy Optimization for Reasoning Modelsarxiv-2606.18954 Sparse Blocked context onlyJun 17, 2026
- Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agentsarxiv-2606.18947 Sparse Blocked context onlyJun 17, 2026
- REVES: REvision and VErification--Augmented Training for Test-Time Scalingarxiv-2606.18910 Sparse Blocked context onlyJun 17, 2026
- SAGE: Stochastic Prompt Optimization via Agent-Guided Explorationarxiv-2606.18902 Sparse Blocked context onlyJun 17, 2026
- Learning Robust Pair Confidence for Multimodal Emotion-Cause Pair Extractionarxiv-2606.18893 Sparse Blocked context onlyJun 17, 2026
- Efficient Financial Language Understanding via Distillation with Synthetic Dataarxiv-2606.18875 Sparse Blocked context onlyJun 17, 2026
- ScholarSum: Student-Teacher Abstractive Summarization via Knowledge Graph Reasoning and Reflective Refinementarxiv-2606.18850 Sparse Blocked context onlyJun 17, 2026
- Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillationarxiv-2606.18844 Sparse Blocked context onlyJun 17, 2026
- Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learningarxiv-2606.18831 Sparse Blocked context onlyJun 17, 2026
- GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agentsarxiv-2606.18829 Sparse Blocked context onlyJun 17, 2026
- HandwritingAgent: Language-Driven Handwriting Synthesis in Scalable Vector Spacearxiv-2606.18788 Sparse Blocked context onlyJun 17, 2026
- RedactionBencharxiv-2606.18782 Sparse Blocked context onlyJun 17, 2026
- Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregationarxiv-2606.18781 Sparse Blocked context onlyJun 17, 2026
- Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkisharxiv-2606.18717 Direct Blocked context onlyJun 17, 2026
- Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmarkarxiv-2606.18648 Sparse Blocked context onlyJun 17, 2026
- MolmoMotion: Forecasting Point Trajectories in 3D with Language Instructionarxiv-2606.18558 Sparse Blocked context onlyJun 17, 2026
- Towards Scalable Customization and Deployment of Multi-Agent Systems for Enterprise Applicationsarxiv-2606.18502 Sparse Blocked context onlyJun 16, 2026
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Draftingarxiv-2606.18394 Sparse Blocked context onlyJun 16, 2026
- PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulationarxiv-2606.18375 Sparse Blocked context onlyJun 16, 2026
- Guava: An Effective and Universal Harness for Embodied Manipulationarxiv-2606.18363 Sparse Blocked context onlyJun 16, 2026
- Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unificationarxiv-2606.18249 Sparse Blocked context onlyJun 16, 2026
- Variable-Width Transformersarxiv-2606.18246 Sparse Blocked context onlyJun 16, 2026
- Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradientsarxiv-2606.18216 Sparse Blocked context onlyJun 16, 2026
- Looped World Modelsarxiv-2606.18208 Sparse Blocked context onlyJun 16, 2026
- RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skillsarxiv-2606.18203 Sparse Blocked context onlyJun 16, 2026
- Learning from the Self-future: On-policy Self-distillation for dLLMsarxiv-2606.18195 Sparse Blocked context onlyJun 16, 2026
- Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation Systemarxiv-2606.18112 Sparse Blocked context onlyJun 16, 2026
- Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Groundingarxiv-2606.18101 Sparse Blocked context onlyJun 16, 2026
- LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scalingarxiv-2606.18023 Sparse Blocked context onlyJun 16, 2026
- ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressionsarxiv-2606.17905 Sparse Blocked context onlyJun 16, 2026
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Modelsarxiv-2606.17846 Sparse Blocked context onlyJun 16, 2026
- EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intentarxiv-2606.17698 Sparse Blocked context onlyJun 16, 2026
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoningarxiv-2606.17682 Sparse Blocked context onlyJun 16, 2026
- Reinforcing Dual-Path Reasoning in Spatial Vision Language Modelsarxiv-2606.17539 Sparse Blocked context onlyJun 16, 2026
- GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planningarxiv-2606.17480 Sparse Blocked context onlyJun 16, 2026
- ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretrainingarxiv-2606.17200 Sparse Blocked context onlyJun 15, 2026
- Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systemsarxiv-2606.17182 Sparse Blocked context onlyJun 15, 2026
- Context-Aware RL for Agentic and Multimodal LLMsarxiv-2606.17053 Sparse Blocked context onlyJun 15, 2026
- Geometric Action Model for Robot Policy Learningarxiv-2606.17046 Sparse Blocked context onlyJun 15, 2026
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generationarxiv-2606.17030 Sparse Blocked context onlyJun 15, 2026
- ExpRL: Exploratory RL for LLM Mid-Trainingarxiv-2606.17024 Sparse Blocked context onlyJun 15, 2026
- TokenPilot: Cache-Efficient Context Management for LLM Agentsarxiv-2606.17016 Sparse Blocked context onlyJun 15, 2026
- DreamX-World 1.0: A General-Purpose Interactive World Modelarxiv-2606.16993 Direct Blocked context onlyJun 15, 2026
- Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciencesarxiv-2606.16905 Sparse Blocked context onlyJun 15, 2026
- Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Textsarxiv-2606.16845 Sparse Blocked context onlyJun 15, 2026
- No Resource, No Benchmarks, No Problem? Evaluating and Improving LLMs for Code Generation in No-Resource Languagesarxiv-2606.16827 Sparse Blocked context onlyJun 15, 2026
- How Much Can We Trust LLM Search Agents? Measuring Endorsement Vulnerability to Web Content Manipulationarxiv-2606.16821 Sparse Blocked context onlyJun 15, 2026
- Text-Vision Co-Instructed Image Editingarxiv-2606.16767 Sparse Blocked context onlyJun 15, 2026
- Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Modelsarxiv-2606.16700 Sparse Blocked context onlyJun 15, 2026
- Kairos: A Native World Model Stack for Physical AIarxiv-2606.16533 Direct Blocked context onlyJun 15, 2026
- How Post-Training Shapes Biological Reasoning Modelsarxiv-2606.16517 Sparse Blocked context onlyJun 15, 2026
- daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimizationarxiv-2606.16497 Sparse Blocked context onlyJun 15, 2026
- LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teachingarxiv-2606.16428 Sparse Blocked context onlyJun 15, 2026
- RL-Index: Reinforcement Learning for Retrieval Index Reasoningarxiv-2606.16316 Sparse Blocked context onlyJun 15, 2026
- VisualClaw: A Real-Time, Personalized Agent for the Physical Worldarxiv-2606.16295 Sparse Blocked context onlyJun 15, 2026
- Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Modelsarxiv-2606.16281 Sparse Blocked context onlyJun 15, 2026
- UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformerarxiv-2606.16255 Sparse Blocked context onlyJun 15, 2026
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Modelsarxiv-2606.16140 Sparse Blocked context onlyJun 15, 2026
- Thinking with Visual Groundingarxiv-2606.16122 Sparse Blocked context onlyJun 15, 2026
- Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMsarxiv-2606.16011 Sparse Blocked context onlyJun 14, 2026
- Beyond NL2Code: A Structured Survey of Multimodal Code Intelligencearxiv-2606.15932 Sparse Blocked context onlyJun 14, 2026
- SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasksarxiv-2606.15872 Sparse Blocked context onlyJun 14, 2026
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policiesarxiv-2606.15768 Sparse Blocked context onlyJun 14, 2026
- Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracksarxiv-2606.15534 Sparse Blocked context onlyJun 14, 2026
- Selective Synergistic Learning for Video Object-Centric Learningarxiv-2606.15527 Sparse Blocked context onlyJun 14, 2026
- Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalitiesarxiv-2606.15514 Sparse Blocked context onlyJun 13, 2026
- Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plusarxiv-2606.15345 Sparse Blocked context onlyJun 13, 2026
- Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoningarxiv-2606.15231 Sparse Blocked context onlyJun 13, 2026
- Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddingsarxiv-2606.15134 Sparse Blocked context onlyJun 13, 2026
- Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scalearxiv-2606.15079 Sparse Blocked context onlyJun 13, 2026
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoningarxiv-2606.15007 Sparse Blocked context onlyJun 12, 2026
- OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chainsarxiv-2606.14702 Sparse Blocked context onlyJun 12, 2026
- RepFusion: Leveraging Multimodal Priors for Denoising in Representation Spacearxiv-2606.14700 Sparse Blocked context onlyJun 12, 2026
- ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoningarxiv-2606.14697 Sparse Blocked context onlyJun 12, 2026
- From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AIarxiv-2606.14502 Sparse Blocked context onlyJun 12, 2026
- Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stackarxiv-2606.14409 Direct Blocked context onlyJun 12, 2026
- IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Productsarxiv-2606.14383 Sparse Blocked context onlyJun 12, 2026
- AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Propertiesarxiv-2606.14240 Sparse Blocked context onlyJun 12, 2026
- Implicit Reasoning for Large Language Model-based Generative Recommendationarxiv-2606.14142 Sparse Blocked context onlyJun 12, 2026
- FastContext: Training Efficient Repository Explorer for Coding Agentsarxiv-2606.14066 Curated Related Blocked context onlyJun 12, 2026
- LLM Agents Can See Code Repositoriesarxiv-2606.14061 Sparse Blocked context onlyJun 12, 2026
- Avatar V: Scaling Video-Reference Avatar Video Generationarxiv-2606.13872 Curated Related Blocked context onlyJun 11, 2026
- Aligning Quantum Operators with Large Language Modelsarxiv-2606.13811 Sparse Blocked context onlyJun 11, 2026
- EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environmentsarxiv-2606.13681 Sparse Blocked context onlyJun 11, 2026
- InterleaveThinker: Reinforcing Agentic Interleaved Generationarxiv-2606.13679 Sparse Blocked context onlyJun 11, 2026
- See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agentsarxiv-2606.13594 Sparse Blocked context onlyJun 11, 2026
- ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languagesarxiv-2606.13572 Sparse Blocked context onlyJun 11, 2026
- OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Dataarxiv-2606.13432 Sparse Blocked context onlyJun 11, 2026
- MiniMax Sparse Attentionarxiv-2606.13392 Direct Blocked context onlyJun 11, 2026
- HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizersarxiv-2606.13289 Sparse Blocked context onlyJun 11, 2026
- SICI: A Semantic-Pragmatic Complexity Index Reveals Regime Shifts in LLM Stance Detectionarxiv-2606.13189 Sparse Blocked context onlyJun 11, 2026
- EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledgearxiv-2606.13120 Sparse Blocked context onlyJun 11, 2026
- HarmProfile: Characterizing Harmful Distributions in Frontier LLMsarxiv-2608.14577 Sparse Blocked context onlyJun 11, 2026
- HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harnessarxiv-2606.12882 Sparse Blocked context onlyJun 11, 2026
- DailyReport: An Open-ended Benchmark for Evaluating Search Agents on Daily Search Tasksarxiv-2606.12871 Sparse Blocked context onlyJun 11, 2026
- Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisationarxiv-2606.12594 Sparse Blocked context onlyJun 10, 2026
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Modelsarxiv-2606.12412 Sparse Blocked context onlyJun 10, 2026
- World Pilot: Steering Vision-Language-Action Models with World-Action Priorsarxiv-2606.12403 Sparse Blocked context onlyJun 10, 2026
- Redesign Mixture-of-Experts Routers with Manifold Power Iterationarxiv-2606.12397 Sparse Blocked context onlyJun 10, 2026
- APPO: Agentic Procedural Policy Optimizationarxiv-2606.12384 Sparse Blocked context onlyJun 10, 2026
- Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalizationarxiv-2606.12373 Sparse Blocked context onlyJun 10, 2026
- Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Samplingarxiv-2606.12370 Sparse Blocked context onlyJun 10, 2026
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policiesarxiv-2606.12366 Sparse Blocked context onlyJun 10, 2026
- On Subquadratic Architectures: From Applications to Principlesarxiv-2606.12364 Sparse Blocked context onlyJun 10, 2026
- Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasksarxiv-2606.12344 Direct Blocked context onlyJun 10, 2026
- Measuring Epistemic Resilience of LLMs Under Misleading Medical Contextarxiv-2606.12291 Sparse Blocked context onlyJun 10, 2026
- VIA-SD: Verification via Intra-Model Routing for Speculative Decodingarxiv-2606.12243 Sparse Blocked context onlyJun 10, 2026
- InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoningarxiv-2606.12195 Sparse Blocked context onlyJun 10, 2026
- Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Applicationarxiv-2606.12191 Sparse Blocked context onlyJun 10, 2026
- World Model Self-Distillation: Training World Models to Solve General Tasksarxiv-2606.12072 Sparse Blocked context onlyJun 10, 2026
- Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skillsarxiv-2606.11897 Sparse Blocked context onlyJun 10, 2026
- Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Trainingarxiv-2606.11854 Sparse Blocked context onlyJun 10, 2026
- Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Codearxiv-2606.11817 Sparse Blocked context onlyJun 10, 2026
- Orchestra-o1: Omnimodal Agent Orchestrationarxiv-2606.13707 Sparse Blocked context onlyJun 10, 2026
- When is Your LLM Steerable?arxiv-2606.11599 Sparse Blocked context onlyJun 10, 2026
- Building Social World Models with Large Language Modelsarxiv-2606.11482 Sparse Blocked context onlyJun 9, 2026
- Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Modelsarxiv-2606.11409 Sparse Blocked context onlyJun 9, 2026
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Modelsarxiv-2606.11324 Sparse Blocked context onlyJun 9, 2026
- ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representationsarxiv-2606.11188 Sparse Blocked context onlyJun 9, 2026
- Next Forcing: Causal World Modeling with Multi-Chunk Predictionarxiv-2606.11187 Sparse Blocked context onlyJun 9, 2026
- EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agentsarxiv-2606.11182 Sparse Blocked context onlyJun 9, 2026
- Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronizationarxiv-2606.11180 Sparse Blocked context onlyJun 9, 2026
- The Role of Feedback Alignment in Self-Distillationarxiv-2606.11173 Sparse Blocked context onlyJun 9, 2026
- P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoningarxiv-2606.11152 Sparse Blocked context onlyJun 9, 2026
- UniPET: a universal network for high-quality PET image denoising across varied dose reduction factorsarxiv-2606.11131 Sparse Blocked context onlyJun 9, 2026
- TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learningarxiv-2606.11119 Sparse Blocked context onlyJun 9, 2026
- Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix Itarxiv-2606.11052 Sparse Blocked context onlyJun 9, 2026
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fieldsarxiv-2606.11042 Sparse Blocked context onlyJun 9, 2026
- Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learningarxiv-2606.10968 Sparse Blocked context onlyJun 9, 2026
- Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolutionarxiv-2606.10917 Sparse Blocked context onlyJun 9, 2026
- SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioningarxiv-2606.10804 Direct Blocked context onlyJun 9, 2026
- N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimizationarxiv-2606.10768 Sparse Blocked context onlyJun 9, 2026
- The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignmentarxiv-2606.10747 Sparse Blocked context onlyJun 9, 2026
- Decentralized Multi-Agent Systems with Shared Contextarxiv-2606.10662 Sparse Blocked context onlyJun 9, 2026
- Kwai Keye-VL-2.0 Technical Reportarxiv-2606.10651 Sparse Blocked context onlyJun 9, 2026
- Dynamic Linear Attentionarxiv-2606.10650 Sparse Blocked context onlyJun 9, 2026
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMsarxiv-2606.10646 Sparse Blocked context onlyJun 9, 2026
- One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QAarxiv-2606.10572 Sparse Blocked context onlyJun 9, 2026
- ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatoricsarxiv-2606.10479 Sparse Blocked context onlyJun 9, 2026
- WebChallenger: A Reliable and Efficient Generalist Web Agentarxiv-2606.10423 Sparse Blocked context onlyJun 9, 2026
- BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contextsarxiv-2606.10061 Sparse Blocked context onlyJun 8, 2026
- OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamicsarxiv-2606.09826 Sparse Blocked context onlyJun 8, 2026
- AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routingarxiv-2606.09811 Sparse Blocked context onlyJun 8, 2026
- iOSWorld: A Benchmark for Personally Intelligent Phone Agentsarxiv-2606.09764 Sparse Blocked context onlyJun 8, 2026
- SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Researcharxiv-2606.09730 Sparse Blocked context onlyJun 8, 2026
- SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasksarxiv-2606.09669 Sparse Blocked context onlyJun 8, 2026
- Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Textarxiv-2606.09585 Sparse Blocked context onlyJun 8, 2026
- SwiftVR: Real-Time One-Step Generative Video Restorationarxiv-2606.09516 Sparse Blocked context onlyJun 8, 2026
- Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Shortarxiv-2606.09380 Sparse Blocked context onlyJun 8, 2026
- SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Samplingarxiv-2606.09304 Sparse Blocked context onlyJun 8, 2026
- Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturationarxiv-2606.09131 Sparse Blocked context onlyJun 8, 2026
- Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gatingarxiv-2606.09068 Sparse Blocked context onlyJun 8, 2026
- Bridging the Agent-World Gap: Text World Models for LLM-based Agentsarxiv-2606.09032 Sparse Blocked context onlyJun 8, 2026
- TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMsarxiv-2606.09030 Sparse Blocked context onlyJun 8, 2026
- MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Trainingarxiv-2606.08788 Sparse Blocked context onlyJun 7, 2026
- WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesisarxiv-2606.08670 Sparse Blocked context onlyJun 7, 2026
- From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscapearxiv-2606.08625 Sparse Blocked context onlyJun 7, 2026
- OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioningarxiv-2606.08572 Sparse Blocked context onlyJun 7, 2026
- Trajectory-Refined Distillationarxiv-2606.08432 Sparse Blocked context onlyJun 7, 2026
- Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnessesarxiv-2606.08348 Direct Blocked context onlyJun 6, 2026
- Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?arxiv-2606.08063 Sparse Blocked context onlyJun 6, 2026
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoningarxiv-2606.08035 Sparse Blocked context onlyJun 6, 2026
- POISE: Position-Aware Undetectable Skill Injection on LLM Agentsarxiv-2606.07943 Sparse Blocked context onlyJun 6, 2026
- MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanismarxiv-2606.07512 Sparse Blocked context onlyJun 5, 2026
- PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streamsarxiv-2606.07454 Direct Blocked context onlyJun 5, 2026
- Watch, Remember, Reason: Human-View Video Understanding with MLLMsarxiv-2606.07433 Sparse Blocked context onlyJun 5, 2026
- VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulationarxiv-2606.07723 Sparse Blocked context onlyJun 5, 2026
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customizationarxiv-2606.07326 Sparse Blocked context onlyJun 5, 2026
- DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoningarxiv-2606.07299 Sparse Blocked context onlyJun 5, 2026
- MMAE: A Massive Multitask Audio Editing Benchmarkarxiv-2606.07229 Sparse Blocked context onlyJun 5, 2026
- Robotic Policy Adaptation via Weight-Space Meta-Learningarxiv-2606.07217 Sparse Blocked context onlyJun 5, 2026
- SuTRA : Structurally-Unified Tokenization with Root Awarenessarxiv-2608.18087 Sparse Blocked context onlyJun 5, 2026
- On the Geometry of On-Policy Distillationarxiv-2606.07082 Sparse Blocked context onlyJun 5, 2026
- Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seekingarxiv-2606.07689 Sparse Blocked context onlyJun 5, 2026
- Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priorsarxiv-2606.06891 Sparse Blocked context onlyJun 5, 2026
- Towards Retrieving Interaction Spaces for Agentic Searcharxiv-2606.06880 Sparse Blocked context onlyJun 5, 2026
- OpenSkill: Open-World Self-Evolution for LLM Agentsarxiv-2606.06741 Sparse Blocked context onlyJun 4, 2026
- A Geometric Account of Activation Steering through Angle-Norm Decompositionarxiv-2606.06735 Sparse Blocked context onlyJun 4, 2026
- Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillationarxiv-2606.06712 Sparse Blocked context onlyJun 4, 2026
- UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMsarxiv-2606.06622 Sparse Blocked context onlyJun 4, 2026
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMsarxiv-2606.06574 Sparse Blocked context onlyJun 4, 2026
- Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolutionarxiv-2606.06492 Sparse Blocked context onlyJun 4, 2026
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discoveryarxiv-2606.06473 Sparse Blocked context onlyJun 4, 2026
- Benchmark Everything Everywhere All at Oncearxiv-2606.06462 Sparse Blocked context onlyJun 4, 2026
- Latent Reasoning with Normalizing Flowsarxiv-2606.06447 Sparse Blocked context onlyJun 4, 2026
- Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussionsarxiv-2606.06443 Sparse Blocked context onlyJun 4, 2026
- Reinforcement Learning Elicits Contextual Learning of Unseen Language Translationarxiv-2606.06428 Sparse Blocked context onlyJun 4, 2026
- Unsupervised Skill Discovery for Agentic Data Analysisarxiv-2606.06416 Sparse Blocked context onlyJun 4, 2026
- Towards One-to-Many Temporal Groundingarxiv-2606.06294 Curated Related Blocked context onlyJun 4, 2026
- ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMsarxiv-2606.12451 Sparse Blocked context onlyJun 4, 2026
- AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understandingarxiv-2606.06155 Sparse Blocked context onlyJun 4, 2026
- Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedbackarxiv-2606.06113 Sparse Blocked context onlyJun 4, 2026
- IR3DE: A Linear Router for Large Language Modelsarxiv-2606.06098 Sparse Blocked context onlyJun 4, 2026
- LoomVideo: Unifying Multimodal Inputs into Video Generation and Editingarxiv-2606.06042 Curated Related Blocked context onlyJun 4, 2026
- Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillationarxiv-2606.05988 Sparse Blocked context onlyJun 4, 2026
- World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesisarxiv-2606.05979 Sparse Blocked context onlyJun 4, 2026
- LLM Explainability with Counterfactual Chains and Causal Graphsarxiv-2606.05972 Sparse Blocked context onlyJun 4, 2026
- Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rolloutsarxiv-2606.05922 Sparse Blocked context onlyJun 4, 2026
- Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Modelsarxiv-2606.05833 Sparse Blocked context onlyJun 4, 2026
- Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Predictionarxiv-2606.05769 Sparse Blocked context onlyJun 4, 2026
- SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agentsarxiv-2606.05761 Sparse Blocked context onlyJun 4, 2026
- Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latentsarxiv-2606.05753 Sparse Blocked context onlyJun 4, 2026
- Answer Presence Drives RAG Rewriting Gainsarxiv-2606.05633 Sparse Blocked context onlyJun 4, 2026
- AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraintsarxiv-2606.05622 Sparse Blocked context onlyJun 4, 2026
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmarkarxiv-2606.06538 Sparse Blocked context onlyJun 4, 2026
- AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agentsarxiv-2606.05557 Sparse Blocked context onlyJun 4, 2026
- ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?arxiv-2606.05553 Sparse Blocked context onlyJun 4, 2026
- BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understandingarxiv-2606.05515 Sparse Blocked context onlyJun 3, 2026
- Agents' Last Examarxiv-2606.05405 Direct Blocked context onlyJun 3, 2026
- What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systemsarxiv-2606.05304 Sparse Blocked context onlyJun 3, 2026
- Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Dataarxiv-2606.05122 Sparse Blocked context onlyJun 3, 2026
- Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Casesarxiv-2606.05112 Sparse Blocked context onlyJun 3, 2026
- VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understandingarxiv-2606.05259 Sparse Blocked context onlyJun 3, 2026
- Flash-WAM: Modality-Aware Distillation for World Action Modelsarxiv-2606.05254 Sparse Blocked context onlyJun 3, 2026
- Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Gamearxiv-2606.04978 Sparse Blocked context onlyJun 3, 2026
- Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learningarxiv-2606.04923 Sparse Blocked context onlyJun 3, 2026
- GRAIL: Gradient-Reweighted Advantages for Reinforcement Learning with Verifiable Rewardsarxiv-2606.04889 Sparse Blocked context onlyJun 3, 2026
- Optimizing the Cost-Quality Tradeoff of Agentic Theorem Provers in Leanarxiv-2606.04883 Sparse Blocked context onlyJun 3, 2026
- Rethinking Continual Experience Internalization for Self-Evolving LLM Agentsarxiv-2606.04703 Sparse Blocked context onlyJun 3, 2026
- Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generationarxiv-2606.04527 Sparse Blocked context onlyJun 3, 2026
- MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generationarxiv-2606.04513 Sparse Blocked context onlyJun 3, 2026
- SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inferencearxiv-2606.04511 Sparse Blocked context onlyJun 3, 2026
- SePO: Self-Evolving Prompt Agent for System Prompt Optimizationarxiv-2606.04465 Sparse Blocked context onlyJun 3, 2026
- Stateful Visual Encoders for Vision-Language Modelsarxiv-2606.04433 Sparse Blocked context onlyJun 3, 2026
- Online Skill Learning for Web Agents via State-Grounded Dynamic Retrievalarxiv-2606.04391 Sparse Blocked context onlyJun 3, 2026
- Video2LoRA: Parametric Video Internalization for Vision-Language Modelsarxiv-2606.04351 Sparse Blocked context onlyJun 3, 2026
- Can Generalist Agents Automate Data Curation?arxiv-2606.04261 Sparse Blocked context onlyJun 2, 2026
- Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectoryarxiv-2606.06523 Sparse Blocked context onlyJun 2, 2026
- Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skillarxiv-2606.03980 Sparse Blocked context onlyJun 2, 2026
- Language Models Need Sleep: Learning to Self-Modify and Consolidate Memoriesarxiv-2606.03979 Sparse Blocked context onlyJun 2, 2026
- Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoningarxiv-2606.03965 Sparse Blocked context onlyJun 2, 2026
- SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstructionarxiv-2606.03940 Sparse Blocked context onlyJun 2, 2026
- Value-Aware Stochastic KV Cache Eviction for Reasoning Modelsarxiv-2606.03928 Sparse Blocked context onlyJun 2, 2026
- Benchmarking Visual State Tracking in Multimodal Video Understandingarxiv-2606.03920 Sparse Blocked context onlyJun 2, 2026
- Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agentsarxiv-2606.03895 Sparse Blocked context onlyJun 2, 2026
- OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMsarxiv-2606.03890 Sparse Blocked context onlyJun 2, 2026
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Managementarxiv-2606.03841 Sparse Blocked context onlyJun 2, 2026
- Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?arxiv-2606.03782 Sparse Blocked context onlyJun 2, 2026
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Modelsarxiv-2606.03748 Curated Related Blocked context onlyJun 2, 2026
- Qwen-Image-Flash: Beyond Objective Designarxiv-2606.03746 Sparse Blocked context onlyJun 2, 2026