300 canonical paper links on this archive page.
- UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generationarxiv-2607.06306 Sparse Blocked context onlyJul 7, 2026
- AlayaWorld: Long-Horizon and Playable Video World Generationarxiv-2607.06291 Direct Blocked context onlyJul 7, 2026
- AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluationarxiv-2607.06624 Sparse Blocked context onlyJul 7, 2026
- PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languagesarxiv-2607.05992 Sparse Blocked context onlyJul 7, 2026
- TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Trainingarxiv-2607.05804 Sparse Blocked context onlyJul 7, 2026
- Image2Sim: Scaling Embodied Navigation via Generative Neural Simulatorarxiv-2607.05765 Sparse Blocked context onlyJul 7, 2026
- Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decodingarxiv-2607.05722 Sparse Blocked context onlyJul 7, 2026
- Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memoryarxiv-2607.05511 Curated Related Blocked context onlyJul 6, 2026
- LLM-as-a-Verifier: A General-Purpose Verification Frameworkarxiv-2607.05391 Direct Blocked context onlyJul 6, 2026
- MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcingarxiv-2607.05376 Sparse Blocked context onlyJul 6, 2026
- Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondentsarxiv-2608.14606 Sparse Blocked context onlyJul 6, 2026
- Unified Audio Intelligence Without Regressing on Text Intelligencearxiv-2607.05196 Sparse Blocked context onlyJul 6, 2026
- DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generationarxiv-2607.05147 Direct Blocked context onlyJul 6, 2026
- ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streamsarxiv-2607.09759 Sparse Blocked context onlyJul 6, 2026
- InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalizationarxiv-2607.04988 Sparse Blocked context onlyJul 6, 2026
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Betterarxiv-2607.04884 Sparse Blocked context onlyJul 6, 2026
- Multi-Turn On-Policy Distillation with Prefix Replayarxiv-2607.04763 Sparse Blocked context onlyJul 6, 2026
- Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Modelsarxiv-2607.04461 Sparse Blocked context onlyJul 5, 2026
- ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blogarxiv-2607.04438 Direct Blocked context onlyJul 5, 2026
- dOPSD: On-Policy Self-Distillation for Diffusion Language Modelsarxiv-2607.04428 Sparse Blocked context onlyJul 5, 2026
- UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learningarxiv-2607.04425 Sparse Blocked context onlyJul 5, 2026
- LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RLarxiv-2607.04412 Sparse Blocked context onlyJul 5, 2026
- Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenizationarxiv-2607.04064 Sparse Blocked context onlyJul 5, 2026
- CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generationarxiv-2607.03803 Sparse Blocked context onlyJul 4, 2026
- Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Modelsarxiv-2607.03751 Sparse Blocked context onlyJul 4, 2026
- Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Processarxiv-2607.03748 Sparse Blocked context onlyJul 4, 2026
- Attending to Multimodal Generation One Token at a Timearxiv-2607.03738 Sparse Blocked context onlyJul 4, 2026
- SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibearxiv-2607.03451 Sparse Blocked context onlyJul 3, 2026
- Bibby AI: An Editor-Native Agentic Platform for Academic Research, Writing, and Publishingarxiv-2607.05435 Sparse Blocked context onlyJul 3, 2026
- PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentationarxiv-2607.03068 Sparse Blocked context onlyJul 3, 2026
- Spectral Rewiring for Exploration, Purification, and Model Mergingarxiv-2607.03065 Sparse Blocked context onlyJul 3, 2026
- Hierarchical Sparse Attention Done Right: Toward Infinite Context Modelingarxiv-2607.02980 Sparse Blocked context onlyJul 3, 2026
- Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioningarxiv-2607.02963 Sparse Blocked context onlyJul 3, 2026
- WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memoryarxiv-2607.02517 Sparse Blocked context onlyJul 2, 2026
- From SRA to Self-Flow: Data Augmentation or Self-Supervision?arxiv-2607.02508 Sparse Blocked context onlyJul 2, 2026
- Reasoning LLM Improves Speaker Recognition in Long-form TV Dramasarxiv-2607.02504 Sparse Blocked context onlyJul 2, 2026
- Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learningarxiv-2607.02490 Sparse Blocked context onlyJul 2, 2026
- Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAsarxiv-2607.02466 Sparse Blocked context onlyJul 2, 2026
- Will Scaling Improve Social Simulation with LLMs?arxiv-2607.02464 Sparse Blocked context onlyJul 2, 2026
- GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluationarxiv-2607.02642 Sparse Blocked context onlyJul 2, 2026
- Know Your Source: A Public Knowledge Store for Media Background Checksarxiv-2607.02383 Sparse Blocked context onlyJul 2, 2026
- HULAT2 at MER-TRANS 2026: Governed Multi-Agent Simplification for Spanish Easy-to-Read Generationarxiv-2607.02381 Sparse Blocked context onlyJul 2, 2026
- SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplacesarxiv-2607.02345 Sparse Blocked context onlyJul 2, 2026
- CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoningarxiv-2607.02262 Sparse Blocked context onlyJul 2, 2026
- Bayesian Sparse Low-Rank Adaptation for Large Language Model Uncertainty Estimationarxiv-2607.02182 Sparse Blocked context onlyJul 2, 2026
- HaloGuard 1.0: An Open Weights Constitutional Classifier for Multilingual AI Safetyarxiv-2607.02079 Sparse Blocked context onlyJul 2, 2026
- EduArt: An educational-level benchmark for evaluating art history knowledge in large language modelsarxiv-2607.02007 Sparse Blocked context onlyJul 2, 2026
- Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editingarxiv-2607.01978 Sparse Blocked context onlyJul 2, 2026
- Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimizationarxiv-2607.01972 Sparse Blocked context onlyJul 2, 2026
- Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsingarxiv-2607.01964 Sparse Blocked context onlyJul 2, 2026
- Rank-Then-Act: Reward-Free Control from Frame-Order Progressarxiv-2607.01897 Sparse Blocked context onlyJul 2, 2026
- PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generationarxiv-2607.01883 Sparse Blocked context onlyJul 2, 2026
- Safety Targeted Embedding Exploit via Refinementarxiv-2607.01859 Sparse Blocked context onlyJul 2, 2026
- VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizonarxiv-2607.01804 Direct Blocked context onlyJul 2, 2026
- PARTREP: Learning What to Repeat for Decoder-only LLMsarxiv-2607.01792 Sparse Blocked context onlyJul 2, 2026
- Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Trainingarxiv-2607.01763 Sparse Blocked context onlyJul 2, 2026
- Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editingarxiv-2607.01690 Sparse Blocked context onlyJul 2, 2026
- AgenticDataBench: A Comprehensive Benchmark for Data Agentsarxiv-2607.01647 Sparse Blocked context onlyJul 2, 2026
- Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Samplingarxiv-2607.01642 Direct Blocked context onlyJul 2, 2026
- BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systemsarxiv-2607.01600 Sparse Blocked context onlyJul 2, 2026
- Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Modelarxiv-2607.01595 Sparse Blocked context onlyJul 2, 2026
- Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Frameworkarxiv-2607.01581 Sparse Blocked context onlyJul 2, 2026
- DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agentsarxiv-2607.01557 Sparse Blocked context onlyJul 2, 2026
- Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewritingarxiv-2607.01457 Sparse Blocked context onlyJul 1, 2026
- FaithMed: Training LLMs For Faithful Evidence-Based Medical Reasoningarxiv-2607.01440 Sparse Blocked context onlyJul 1, 2026
- Discrete Diffusion Language Models for Interactive Radiology Report Draftingarxiv-2607.01436 Sparse Blocked context onlyJul 1, 2026
- IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMsarxiv-2607.01431 Direct Blocked context onlyJul 1, 2026
- MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answeringarxiv-2607.01420 Sparse Blocked context onlyJul 1, 2026
- Multi-Objective Exploration and Preference Optimization via Mutual Informationarxiv-2607.01392 Sparse Blocked context onlyJul 1, 2026
- Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Trainingarxiv-2607.01232 Sparse Blocked context onlyJul 1, 2026
- Theoria: Rewrite-Acceptability Verification over Informal Reasoning Statesarxiv-2607.01223 Sparse Blocked context onlyJul 1, 2026
- Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillationarxiv-2607.01208 Sparse Blocked context onlyJul 1, 2026
- Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoningarxiv-2607.01191 Sparse Blocked context onlyJul 1, 2026
- QuasiMoTTo: Quasi-Monte Carlo Test-Time Scalingarxiv-2607.01179 Sparse Blocked context onlyJul 1, 2026
- Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguityarxiv-2607.01153 Sparse Blocked context onlyJul 1, 2026
- Autonomous Scientific Discovery via Iterative Meta-Reflectionarxiv-2607.01131 Sparse Blocked context onlyJul 1, 2026
- Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarkingarxiv-2607.01103 Sparse Blocked context onlyJul 1, 2026
- Message Passing Enables Efficient Reasoningarxiv-2607.01077 Sparse Blocked context onlyJul 1, 2026
- MemSyco-Bench: Benchmarking Sycophancy in Agent Memoryarxiv-2607.01071 Sparse Blocked context onlyJul 1, 2026
- Agentic generation of verifiable rules for deterministic, self-expanding reaction classificationarxiv-2607.01061 Sparse Blocked context onlyJul 1, 2026
- Understanding Large Language Modelsarxiv-2607.01006 Sparse Blocked context onlyJul 1, 2026
- Logit-Contribution Scoring Identifies Non-Literal Retrieval Headsarxiv-2607.01002 Sparse Blocked context onlyJul 1, 2026
- Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombinationarxiv-2607.00924 Sparse Blocked context onlyJul 1, 2026
- From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narrativesarxiv-2607.00918 Sparse Blocked context onlyJul 1, 2026
- Valdi: Value Diffusion World Modelsarxiv-2607.00917 Sparse Blocked context onlyJul 1, 2026
- ABot-M0.5: Unified Mobility-and-Manipulation World Action Modelarxiv-2607.00678 Direct Blocked context onlyJul 1, 2026
- Multi-Turn Agentic Scientific Literature Search via Workflow Inductionarxiv-2607.00597 Sparse Blocked context onlyJul 1, 2026
- A Task-State Representation for Long-Horizon Mobile GUI Agentsarxiv-2607.00502 Sparse Blocked context onlyJul 1, 2026
- MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photosarxiv-2607.00491 Sparse Blocked context onlyJul 1, 2026
- Efficient Multilingual Reasoning Transfer via Progressive Code-Switchingarxiv-2607.00485 Sparse Blocked context onlyJul 1, 2026
- Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinkingarxiv-2607.00482 Sparse Blocked context onlyJul 1, 2026
- Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learningarxiv-2607.00461 Sparse Blocked context onlyJul 1, 2026
- Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoisingarxiv-2607.00407 Sparse Blocked context onlyJul 1, 2026
- DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoningarxiv-2607.00341 Sparse Blocked context onlyJul 1, 2026
- Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worldsarxiv-2607.00276 Sparse Blocked context onlyJun 30, 2026
- Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexityarxiv-2607.00248 Sparse Blocked context onlyJun 30, 2026
- GEAR: Guided End-to-End AutoRegression for Image Synthesisarxiv-2606.32039 Sparse Blocked context onlyJun 30, 2026
- Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMsarxiv-2606.32032 Sparse Blocked context onlyJun 30, 2026
- When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errorsarxiv-2606.32029 Sparse Blocked context onlyJun 30, 2026
- Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachersarxiv-2606.32020 Sparse Blocked context onlyJun 30, 2026
- Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoningarxiv-2606.31825 Sparse Blocked context onlyJun 30, 2026
- MuSViT: A Foundation Vision Model for Sheet Music Representationarxiv-2606.31811 Sparse Blocked context onlyJun 30, 2026
- DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generationarxiv-2606.31537 Sparse Blocked context onlyJun 30, 2026
- FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agentsarxiv-2606.31522 Sparse Blocked context onlyJun 30, 2026
- Xiaomi-GUI-0 Technical Reportarxiv-2606.31410 Sparse Blocked context onlyJun 30, 2026
- Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teamingarxiv-2606.31227 Sparse Blocked context onlyJun 30, 2026
- HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agentsarxiv-2606.31179 Sparse Blocked context onlyJun 30, 2026
- ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industriesarxiv-2606.31163 Sparse Blocked context onlyJun 30, 2026
- SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inferencearxiv-2606.31145 Sparse Blocked context onlyJun 30, 2026
- AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generationarxiv-2607.00052 Sparse Blocked context onlyJun 30, 2026
- DOPD: Dual On-policy Distillationarxiv-2606.30626 Curated Related Blocked context onlyJun 29, 2026
- SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessionsarxiv-2606.30573 Sparse Blocked context onlyJun 29, 2026
- Orca: The World is in Your Mindarxiv-2606.30534 Direct Blocked context onlyJun 29, 2026
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Trainingarxiv-2606.30406 Sparse Blocked context onlyJun 29, 2026
- Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threatsarxiv-2606.30259 Sparse Blocked context onlyJun 29, 2026
- TACO: Tool-Augmented Credit Optimization for Agentic Tool Usearxiv-2606.30251 Sparse Blocked context onlyJun 29, 2026
- Grounding LLM Reasoning under Incomplete Graph Evidencearxiv-2606.30247 Sparse Blocked context onlyJun 29, 2026
- CaresAI at CT-DEB26: Detecting Dosing Errors In Clinical Trials Using Domain-Specific Transformer Embeddings and Classification Modelsarxiv-2606.30236 Sparse Blocked context onlyJun 29, 2026
- EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failuresarxiv-2606.30219 Sparse Blocked context onlyJun 29, 2026
- Before Thinking, Learn to Decide: Proactive Routing for Efficient Visual Reasoningarxiv-2606.30217 Sparse Blocked context onlyJun 29, 2026
- Forewarned is Forearmed: When Non-Sequential Embedding Turns Into an Anomaly Detectorarxiv-2606.30196 Sparse Blocked context onlyJun 29, 2026
- CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Grapharxiv-2606.30175 Sparse Blocked context onlyJun 29, 2026
- DNA Language Models: An Assessment of Pre-Training for Fine-Tuning Tasksarxiv-2606.30140 Sparse Blocked context onlyJun 29, 2026
- SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generationarxiv-2606.30124 Sparse Blocked context onlyJun 29, 2026
- Efficient Retrieval-Augmented Generation via Token Co-occurrence Graphsarxiv-2606.30093 Sparse Blocked context onlyJun 29, 2026
- Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogatesarxiv-2606.30085 Sparse Blocked context onlyJun 29, 2026
- One Forward Beats Two: InnerZoom for Accurate and Efficient GUI Groundingarxiv-2606.30084 Sparse Blocked context onlyJun 29, 2026
- MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMsarxiv-2606.30026 Sparse Blocked context onlyJun 29, 2026
- Parametric Skillsarxiv-2606.30015 Sparse Blocked context onlyJun 29, 2026
- LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboardarxiv-2606.30005 Sparse Blocked context onlyJun 29, 2026
- Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoningarxiv-2606.29985 Sparse Blocked context onlyJun 29, 2026
- DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillationarxiv-2606.29961 Sparse Blocked context onlyJun 29, 2026
- LatentRevise: Learning from Zero-Hit Reasoningarxiv-2606.29938 Sparse Blocked context onlyJun 29, 2026
- SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematicsarxiv-2606.29894 Sparse Blocked context onlyJun 29, 2026
- Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistencyarxiv-2606.29876 Sparse Blocked context onlyJun 29, 2026
- ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generationarxiv-2606.29869 Sparse Blocked context onlyJun 29, 2026
- KbSD: Knowledge Boundary aware Self-Distillation for Behavioral Calibration in Agentic Searcharxiv-2606.29863 Sparse Blocked context onlyJun 29, 2026
- Mandol: An Agglomerative Agent Memory System for Long-Term Conversationsarxiv-2606.29778 Direct Blocked context onlyJun 29, 2026
- LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agentsarxiv-2606.30697 Sparse Blocked context onlyJun 29, 2026
- SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attributionarxiv-2606.29713 Sparse Blocked context onlyJun 29, 2026
- Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compressionarxiv-2606.29712 Sparse Blocked context onlyJun 29, 2026
- GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshotsarxiv-2606.29705 Sparse Blocked context onlyJun 29, 2026
- Can MLLMs Critique Like Humans? Evaluating Open-Ended Aesthetic Reasoning in Multimodal Large Language Modelsarxiv-2606.29689 Sparse Blocked context onlyJun 29, 2026
- How LLMs See Creativity: Zero-Shot Scoring of Visual Creativity with Interpretable Reasoningarxiv-2606.29672 Sparse Blocked context onlyJun 29, 2026
- Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literaturearxiv-2606.29667 Sparse Blocked context onlyJun 29, 2026
- Hybrid Retriever Evolution for Multimodal Document Reasoning Agentsarxiv-2606.29648 Sparse Blocked context onlyJun 28, 2026
- How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classificationarxiv-2606.29605 Sparse Blocked context onlyJun 28, 2026
- SurrogateShield: Beyond Redaction for High-Utility, Privacy-Preserving LLM Interactionsarxiv-2606.29567 Sparse Blocked context onlyJun 28, 2026
- Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLMarxiv-2606.29563 Sparse Blocked context onlyJun 28, 2026
- AURORA: Asymmetry and Update-Induced Rotation for Robust Hallucination Detection in Large Language Modelsarxiv-2606.29545 Sparse Blocked context onlyJun 28, 2026
- OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasksarxiv-2606.29537 Direct Blocked context onlyJun 28, 2026
- The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learningarxiv-2606.29526 Sparse Blocked context onlyJun 28, 2026
- Learning Transferable Dynamics Priors from Action to World Modelingarxiv-2606.29501 Sparse Blocked context onlyJun 28, 2026
- Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropyarxiv-2606.29489 Sparse Blocked context onlyJun 28, 2026
- To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregationarxiv-2606.29481 Sparse Blocked context onlyJun 28, 2026
- Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extractionarxiv-2606.29445 Sparse Blocked context onlyJun 28, 2026
- Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defensearxiv-2606.29441 Sparse Blocked context onlyJun 28, 2026
- EntroRouter: Learning Efficient Model Routing via Entropy Regulationarxiv-2606.29424 Sparse Blocked context onlyJun 28, 2026
- TriageRA-CCF: Source-Side Clinical Confidence and Coverage Signals for Adaptive Rank Budgeting in Medical LLMsarxiv-2606.29375 Sparse Blocked context onlyJun 28, 2026
- Hierarchical Experimentalist Agentsarxiv-2606.29315 Sparse Blocked context onlyJun 28, 2026
- MirrorPPR: Exemplar-Based Portrait Photo Retouchingarxiv-2606.29308 Sparse Blocked context onlyJun 28, 2026
- Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learningarxiv-2606.29280 Sparse Blocked context onlyJun 28, 2026
- Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Factsarxiv-2606.29279 Sparse Blocked context onlyJun 28, 2026
- MIThinker: A Plug-and-Play Policy-Optimized Thinker For Motivational Interviewing Counselingarxiv-2606.29265 Sparse Blocked context onlyJun 28, 2026
- Travel-Oriented Reasoning Large Language Model via Domain-Specific Knowledge Graphsarxiv-2606.29254 Sparse Blocked context onlyJun 28, 2026
- Multi-Block Diffusion Language Modelsarxiv-2606.29215 Sparse Blocked context onlyJun 28, 2026
- Evidence-Informed LLM Beliefs for Continual Scientific Discoveryarxiv-2606.29182 Sparse Blocked context onlyJun 28, 2026
- DistilledGemma: Balanced Efficiency-Accuracy for Person-Place Relation Extraction from Multilingual Historical Articlesarxiv-2606.29130 Sparse Blocked context onlyJun 28, 2026
- Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasksarxiv-2606.29082 Sparse Blocked context onlyJun 27, 2026
- Low-cost concept-based localized explanations: How far can we get with training-free approaches?arxiv-2606.29069 Sparse Blocked context onlyJun 27, 2026
- The strength of clinical evidence is recoverable from language model representations but not from their stated gradesarxiv-2606.29034 Sparse Blocked context onlyJun 27, 2026
- BERTomelo: Your Portuguese Encoder Best Friendarxiv-2606.28999 Sparse Blocked context onlyJun 27, 2026
- Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Dataarxiv-2606.28963 Sparse Blocked context onlyJun 27, 2026
- EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Controlarxiv-2606.28938 Sparse Blocked context onlyJun 27, 2026
- PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMsarxiv-2606.28898 Sparse Blocked context onlyJun 27, 2026
- LAMP: Lean-based Agentic framework with MCP and Proof Repairarxiv-2606.28841 Sparse Blocked context onlyJun 27, 2026
- Labeling Training Data for Entity Matching Using Large Language Modelsarxiv-2606.28823 Sparse Blocked context onlyJun 27, 2026
- Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathiarxiv-2606.28796 Sparse Blocked context onlyJun 27, 2026
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?arxiv-2606.28733 Sparse Blocked context onlyJun 27, 2026
- Modular Cognitive Architecture Emerges in Large Language Modelsarxiv-2608.13567 Sparse Blocked context onlyJun 27, 2026
- DataComp-VLM: Improved Open Datasets for Vision-Language Modelsarxiv-2606.28551 Sparse Blocked context onlyJun 26, 2026
- Turn-Averaged SAEs for Feature Discovery and Long-Context Attributionarxiv-2606.28548 Sparse Blocked context onlyJun 26, 2026
- Developmental Trajectories of Situation Modeling and Mentalizing in Transformer Language Modelsarxiv-2606.28524 Sparse Blocked context onlyJun 26, 2026
- TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agentsarxiv-2606.28480 Sparse Blocked context onlyJun 26, 2026
- PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perceptionarxiv-2606.28322 Sparse Blocked context onlyJun 26, 2026
- Towards Automating Scientific Review with Google's Paper Assistant Toolarxiv-2606.28277 Sparse Blocked context onlyJun 26, 2026
- GBC: Gradient-Based Connections for Optimizing Multi-Agent Systemsarxiv-2606.28187 Sparse Blocked context onlyJun 26, 2026
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Predictionarxiv-2606.28186 Sparse Blocked context onlyJun 26, 2026
- Translation as a Bridging Action: Transferring Manipulation Skills from Humans to Robotsarxiv-2606.28133 Sparse Blocked context onlyJun 26, 2026
- PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulationarxiv-2606.28128 Sparse Blocked context onlyJun 26, 2026
- Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generationarxiv-2606.27978 Sparse Blocked context onlyJun 26, 2026
- ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answeringarxiv-2606.27974 Sparse Blocked context onlyJun 26, 2026
- Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoningarxiv-2606.27828 Sparse Blocked context onlyJun 26, 2026
- Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?arxiv-2606.27755 Sparse Blocked context onlyJun 26, 2026
- Dockerless: Environment-Free Program Verifier for Coding Agentsarxiv-2606.28436 Sparse Blocked context onlyJun 26, 2026
- ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrievalarxiv-2606.27708 Curated Related Blocked context onlyJun 26, 2026
- When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Searcharxiv-2606.27669 Sparse Blocked context onlyJun 26, 2026
- Building to the Test: Coding Agents Deliver What You Check, Not What You Requestedarxiv-2606.28430 Sparse Blocked context onlyJun 26, 2026
- Masked Language Flow Modelsarxiv-2606.27617 Sparse Blocked context onlyJun 26, 2026
- Qwen-Image-2.0-RL Technical Reportarxiv-2606.27608 Sparse Blocked context onlyJun 25, 2026
- Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agentsarxiv-2606.27595 Sparse Blocked context onlyJun 25, 2026
- EntMTP: Accelerating LLM Inference with Entropy Guided Multi Token Predictionarxiv-2606.27550 Sparse Blocked context onlyJun 25, 2026
- PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generationarxiv-2606.30673 Sparse Blocked context onlyJun 25, 2026
- PhysiFormer: Learning to Simulate Mechanics in World Spacearxiv-2606.27364 Sparse Blocked context onlyJun 25, 2026
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolutionarxiv-2606.27313 Sparse Blocked context onlyJun 25, 2026
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Modelsarxiv-2606.27288 Sparse Blocked context onlyJun 25, 2026
- EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecastingarxiv-2606.27277 Sparse Blocked context onlyJun 25, 2026
- CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attentionarxiv-2606.27229 Sparse Blocked context onlyJun 25, 2026
- LISA: Likelihood Score Alignment for Visual-condition Controllable Generationarxiv-2606.27192 Sparse Blocked context onlyJun 25, 2026
- Just how sure are you? Improving Verbalized Uncertainty Calibration in Medical VQAarxiv-2606.27023 Sparse Blocked context onlyJun 25, 2026
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generationarxiv-2606.26907 Curated Related Blocked context onlyJun 25, 2026
- Confidence-Aware Tool Orchestration for Robust Video Understandingarxiv-2606.26904 Sparse Blocked context onlyJun 25, 2026
- Information-Aware KV Cache Compression for Long Reasoningarxiv-2606.26875 Sparse Blocked context onlyJun 25, 2026
- Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placementarxiv-2606.27409 Sparse Blocked context onlyJun 25, 2026
- OPID: On-Policy Skill Distillation for Agentic Reinforcement Learningarxiv-2606.26790 Direct Blocked context onlyJun 25, 2026
- What the LLM Should Not Say: Boundary-Aware Context Grounding for A Seven-Channel EEG Agentarxiv-2606.26519 Sparse Blocked context onlyJun 25, 2026
- Epiphany-Aware KV Cache Eviction Without the Attention Matrixarxiv-2606.26472 Sparse Blocked context onlyJun 25, 2026
- MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generationarxiv-2606.26087 Sparse Blocked context onlyJun 24, 2026
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agentsarxiv-2606.26080 Sparse Blocked context onlyJun 24, 2026
- Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Modelsarxiv-2606.26079 Sparse Blocked context onlyJun 24, 2026
- How Robust is OCR-Reasoning? Evaluating OCR-Reasoning Robustness of Vision-Language Models under Visual Perturbationsarxiv-2606.26041 Sparse Blocked context onlyJun 24, 2026
- AI translation of literary texts is "fine", but readers still prefer human translationsarxiv-2606.26040 Sparse Blocked context onlyJun 24, 2026
- Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix Itarxiv-2606.26027 Sparse Blocked context onlyJun 24, 2026
- In-Context World Modeling for Robotic Controlarxiv-2606.26025 Sparse Blocked context onlyJun 24, 2026
- SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Modelsarxiv-2606.25990 Direct Blocked context onlyJun 24, 2026
- Weave of Formal Thoughtarxiv-2606.25987 Sparse Blocked context onlyJun 24, 2026
- Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Textsarxiv-2606.25935 Sparse Blocked context onlyJun 24, 2026
- SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignmentarxiv-2606.25821 Sparse Blocked context onlyJun 24, 2026
- Beyond Function Calling: Benchmarking Tool-Using Agents under Tool-Environment Unreliabilityarxiv-2606.25819 Sparse Blocked context onlyJun 24, 2026
- ShutterMuse: Capture-Time Photography Guidance with MLLMsarxiv-2606.25763 Curated Related Blocked context onlyJun 24, 2026
- Uncertainty Quantification for Computer-Use Agents: A Benchmark across Vision-Language Models and GUI Grounding Datasetsarxiv-2606.25760 Sparse Blocked context onlyJun 24, 2026
- OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learningarxiv-2606.25757 Sparse Blocked context onlyJun 24, 2026
- RAS: Measuring LLM Safety Through Refusal Alignmentarxiv-2606.25750 Sparse Blocked context onlyJun 24, 2026
- BitNet Text Embeddingsarxiv-2606.25674 Sparse Blocked context onlyJun 24, 2026
- MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correctionarxiv-2606.25651 Sparse Blocked context onlyJun 24, 2026
- Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agentsarxiv-2606.25632 Sparse Blocked context onlyJun 24, 2026
- Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraintsarxiv-2606.25605 Sparse Blocked context onlyJun 24, 2026
- Riazi-8B: An Urdu Large Language Model for Mathematical Reasoningarxiv-2606.25568 Sparse Blocked context onlyJun 24, 2026
- SFL-MTSC: Leveraging Semantic Frame-Level Multi-Task Self-Consistency for Robust Multi-Intent Spoken Language Understandingarxiv-2606.25552 Sparse Blocked context onlyJun 24, 2026
- Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoningarxiv-2606.25524 Sparse Blocked context onlyJun 24, 2026
- Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Modelsarxiv-2606.25473 Direct Blocked context onlyJun 24, 2026
- PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Modelsarxiv-2606.25442 Sparse Blocked context onlyJun 24, 2026
- A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Modelsarxiv-2606.25380 Sparse Blocked context onlyJun 24, 2026
- TheoremGraph: Bridging Formal and Informal Mathematicsarxiv-2606.25363 Direct Blocked context onlyJun 24, 2026
- Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routingarxiv-2606.25354 Sparse Blocked context onlyJun 24, 2026
- Hybrid-IR: Dual-Path Hybrid Retrieval with Iterative Reasoning for Complex Medical Question Answeringarxiv-2606.25338 Sparse Blocked context onlyJun 24, 2026
- V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoningarxiv-2606.25319 Sparse Blocked context onlyJun 24, 2026
- Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generationarxiv-2606.25306 Sparse Blocked context onlyJun 24, 2026
- Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reportsarxiv-2606.28393 Sparse Blocked context onlyJun 23, 2026
- ASAP: Agent-System Co-Design for Wall-Clock-Centered Auto HPO Research for ML Experimentsarxiv-2606.25207 Sparse Blocked context onlyJun 23, 2026
- To Isolate or to Score? Model-Adaptive Assessment for Cost-Efficient Multi-Agent RAGarxiv-2606.25191 Sparse Blocked context onlyJun 23, 2026
- The cognitive, affective, and behavioral expression of self-stigma among people who use drugs in online substance use communitiesarxiv-2606.25143 Sparse Blocked context onlyJun 23, 2026
- LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challengesarxiv-2606.25057 Sparse Blocked context onlyJun 23, 2026
- Do Thinking Tokens Help with Safety?arxiv-2606.25013 Sparse Blocked context onlyJun 23, 2026
- FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representationarxiv-2606.24874 Sparse Blocked context onlyJun 23, 2026
- OpenThoughts-Agent: Data Recipes for Agentic Modelsarxiv-2606.24855 Direct Blocked context onlyJun 23, 2026
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generationarxiv-2606.24849 Sparse Blocked context onlyJun 23, 2026
- SHERLOC: Structured Diagnostic Localization for Code Repair Agentsarxiv-2606.24820 Direct Blocked context onlyJun 23, 2026
- Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commercearxiv-2606.24783 Sparse Blocked context onlyJun 23, 2026
- Are We Ready For An Agent-Native Memory System?arxiv-2606.24775 Sparse Blocked context onlyJun 23, 2026
- Posterior Refinement: Fast Language Generation via Any-Order Flow Mapsarxiv-2606.24773 Sparse Blocked context onlyJun 23, 2026
- CANDLE: Character-level Arabic Noise Deduplication using Lightweight Encoderarxiv-2606.24758 Sparse Blocked context onlyJun 23, 2026
- Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelityarxiv-2606.24623 Sparse Blocked context onlyJun 23, 2026
- Qwen-AgentWorld: Language World Models for General Agentsarxiv-2606.24597 Direct Blocked context onlyJun 23, 2026
- AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferabilityarxiv-2606.24589 Sparse Blocked context onlyJun 23, 2026
- Cross-Lingual Exploration for Parametric Knowledgearxiv-2606.24579 Sparse Blocked context onlyJun 23, 2026
- Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoningarxiv-2606.24548 Sparse Blocked context onlyJun 23, 2026
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?arxiv-2606.24530 Sparse Blocked context onlyJun 23, 2026
- AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoningarxiv-2606.24526 Sparse Blocked context onlyJun 23, 2026
- A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trialarxiv-2606.24510 Sparse Blocked context onlyJun 23, 2026
- Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methodsarxiv-2606.24484 Sparse Blocked context onlyJun 23, 2026
- Why Do Accumulated Transformations Extrapolate?arxiv-2606.24975 Sparse Blocked context onlyJun 23, 2026
- An LLM-based Two-Stage Transformer Framework for Cross-Domain Bearing Fault Diagnosis with Limited Dataarxiv-2606.24459 Sparse Blocked context onlyJun 23, 2026
- Bayesian control for coding agentsarxiv-2606.24453 Sparse Blocked context onlyJun 23, 2026
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learningarxiv-2606.24428 Sparse Blocked context onlyJun 23, 2026
- PETRA: Transforming Web Text for Petroleum-Engineering Domain Adaptationarxiv-2606.24346 Sparse Blocked context onlyJun 23, 2026
- Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessmentarxiv-2606.24331 Sparse Blocked context onlyJun 23, 2026
- Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decodingarxiv-2606.24957 Sparse Blocked context onlyJun 23, 2026
- AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compressionarxiv-2606.24286 Sparse Blocked context onlyJun 23, 2026
- CALIBER: Calibrating Confidence Before and After Reasoning in Language Modelsarxiv-2606.24281 Sparse Blocked context onlyJun 23, 2026
- Pigeonholing: Bad prompts hurt models to collapse and make mistakesarxiv-2606.24267 Sparse Blocked context onlyJun 23, 2026
- SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalizationarxiv-2606.24259 Sparse Blocked context onlyJun 23, 2026
- AsyncOPD: How Stale Can On-Policy Distillation Be?arxiv-2606.24143 Sparse Blocked context onlyJun 23, 2026
- Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learningarxiv-2606.24133 Sparse Blocked context onlyJun 23, 2026
- CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compressionarxiv-2606.24083 Sparse Blocked context onlyJun 23, 2026
- Towards Spec Learning: Inference-Time Alignment from Preference Pairsarxiv-2606.24004 Sparse Blocked context onlyJun 22, 2026
- Critique of Agent Modelarxiv-2606.23991 Sparse Blocked context onlyJun 22, 2026
- Mind the Heads: Topological Representation Alignment for Multimodal LLMsarxiv-2606.23885 Sparse Blocked context onlyJun 22, 2026
- The Hitchhiker's Guide to Agentic AI: From Foundations to Systemsarxiv-2606.24937 Sparse Blocked context onlyJun 22, 2026
- Causal Discovery in the Era of Agentsarxiv-2606.23608 Sparse Blocked context onlyJun 22, 2026
- Tmax: A simple recipe for terminal agentsarxiv-2606.23321 Direct Blocked context onlyJun 22, 2026
- Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memoryarxiv-2606.23195 Sparse Blocked context onlyJun 22, 2026
- ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillationarxiv-2606.23104 Sparse Blocked context onlyJun 22, 2026
- Unlimited OCR Worksarxiv-2606.23050 Direct Blocked context onlyJun 22, 2026
- Training Open Models for Agentic Phone Usearxiv-2606.23049 Sparse Blocked context onlyJun 22, 2026
- Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agentsarxiv-2606.22953 Sparse Blocked context onlyJun 22, 2026
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agentsarxiv-2606.22883 Sparse Blocked context onlyJun 22, 2026
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoningarxiv-2606.22873 Sparse Blocked context onlyJun 22, 2026
- KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Rerankingarxiv-2606.22807 Sparse Blocked context onlyJun 22, 2026
- RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representationarxiv-2606.22749 Sparse Blocked context onlyJun 22, 2026