300 canonical paper links on this archive page.
- When Reviews Disagree: Fine-Grained Contradiction Analysis in Scientific Peer Reviewsarxiv-2605.10171 Sparse Blocked context onlyMay 11, 2026
- PREPING: Building Agent Memory without Tasksarxiv-2605.13880 Sparse Blocked context onlyMay 11, 2026
- G-Zero: Self-Play for Open-Ended Generation from Zero Dataarxiv-2605.09959 Sparse Blocked context onlyMay 11, 2026
- Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Siliconarxiv-2605.09708 Sparse Blocked context onlyMay 10, 2026
- Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Evictionarxiv-2605.09649 Sparse Blocked context onlyMay 10, 2026
- Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Modelsarxiv-2605.09630 Sparse Blocked context onlyMay 10, 2026
- TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systemsarxiv-2605.09539 Sparse Blocked context onlyMay 10, 2026
- LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Modelsarxiv-2605.11011 Sparse Blocked context onlyMay 10, 2026
- Offline Preference Optimization for Rectified Flow with Noise-Tracked Pairsarxiv-2605.09433 Sparse Blocked context onlyMay 10, 2026
- Sub-JEPA: Subspace Gaussian Regularization for Stable End-to-End World Modelsarxiv-2605.09241 Sparse Blocked context onlyMay 10, 2026
- FORTIS: Benchmarking Over-Privilege in Agent Skillsarxiv-2605.09163 Sparse Blocked context onlyMay 9, 2026
- Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMsarxiv-2605.09063 Sparse Blocked context onlyMay 9, 2026
- ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usagearxiv-2605.16363 Sparse Blocked context onlyMay 9, 2026
- CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Modelsarxiv-2605.08735 Curated Related Blocked context onlyMay 9, 2026
- Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimizationarxiv-2605.08659 Sparse Blocked context onlyMay 9, 2026
- FlashEvolve: Accelerating Agent Self-Evolution with Asynchronous Stage Orchestrationarxiv-2605.08520 Sparse Blocked context onlyMay 8, 2026
- Results and Retrospective Analysis of the CODS 2025 AssetOpsBench Challengearxiv-2605.08518 Sparse Blocked context onlyMay 8, 2026
- MoMo: Conditioned Contrastive Representation Learning for Preference-Modulated Planningarxiv-2605.08512 Sparse Blocked context onlyMay 8, 2026
- jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Compositionarxiv-2605.08384 Sparse Blocked context onlyMay 8, 2026
- Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteriaarxiv-2605.08354 Sparse Blocked context onlyMay 8, 2026
- Normalizing Trajectory Modelsarxiv-2605.08078 Sparse Blocked context onlyMay 8, 2026
- Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibrationarxiv-2605.08077 Sparse Blocked context onlyMay 8, 2026
- The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agentsarxiv-2605.08060 Sparse Blocked context onlyMay 8, 2026
- Accurate and Efficient Statistical Testing for Word Semantic Breadtharxiv-2605.08048 Sparse Blocked context onlyMay 8, 2026
- Uncertainty-Aware Structured Data Extraction from Full CMR Reports via Distilled LLMsarxiv-2605.08045 Sparse Blocked context onlyMay 8, 2026
- SCOPE: Structured Decomposition and Conditional Skill Orchestration for Complex Image Generationarxiv-2605.08043 Sparse Blocked context onlyMay 8, 2026
- Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claimsarxiv-2605.08012 Sparse Blocked context onlyMay 8, 2026
- CoCoReviewBench: A Completeness- and Correctness-Oriented Benchmark for AI Reviewersarxiv-2605.07905 Sparse Blocked context onlyMay 8, 2026
- SCENE: Recognizing Social Norms and Sanctioning in Group Chatsarxiv-2605.07823 Sparse Blocked context onlyMay 8, 2026
- Prune-OPD: Efficient and Reliable On-Policy Distillation for Long-Horizon Reasoningarxiv-2605.07804 Sparse Blocked context onlyMay 8, 2026
- SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillationarxiv-2605.07711 Sparse Blocked context onlyMay 8, 2026
- DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domainarxiv-2605.07699 Sparse Blocked context onlyMay 8, 2026
- MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditingarxiv-2605.07646 Sparse Blocked context onlyMay 8, 2026
- Learning to Communicate Locally for Large-Scale Multi-Agent Pathfindingarxiv-2605.07637 Sparse Blocked context onlyMay 8, 2026
- Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agentsarxiv-2605.07630 Sparse Blocked context onlyMay 8, 2026
- Implicit Preference Alignment for Human Image Animationarxiv-2605.07545 Sparse Blocked context onlyMay 8, 2026
- From 0-Order Selection to 2-Order Judgment: Combinatorial Hardening Exposes Compositional Failures in Frontier LLMsarxiv-2605.07268 Sparse Blocked context onlyMay 8, 2026
- When Are Experts Misrouted? Counterfactual Routing Analysis in Mixture-of-Experts Language Modelsarxiv-2605.07260 Sparse Blocked context onlyMay 8, 2026
- SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Draftingarxiv-2605.07243 Sparse Blocked context onlyMay 8, 2026
- HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agentsarxiv-2605.07177 Sparse Blocked context onlyMay 8, 2026
- CLIPer: Tailoring Diverse User Preference via Classifier-Guided Inference-Time Personalizationarxiv-2605.07162 Sparse Blocked context onlyMay 8, 2026
- Learning Visual Feature-Based World Models via Residual Latent Actionarxiv-2605.07079 Sparse Blocked context onlyMay 8, 2026
- Relit-LiVE: Relight Video by Jointly Learning Environment Videoarxiv-2605.06658 Sparse Blocked context onlyMay 7, 2026
- AI Co-Mathematician: Accelerating Mathematicians with Agentic AIarxiv-2605.06651 Sparse Blocked context onlyMay 7, 2026
- PianoCoRe: Combined and Refined Piano MIDI Datasetarxiv-2605.06627 Sparse Blocked context onlyMay 7, 2026
- SkillOS: Learning Skill Curation for Self-Evolving Agentsarxiv-2605.06614 Sparse Blocked context onlyMay 7, 2026
- GeoStack: A Framework for Quasi-Abelian Knowledge Composition in VLMsarxiv-2605.06477 Sparse Blocked context onlyMay 7, 2026
- MiA-Signature: Approximating Global Activation for Long-Context Understandingarxiv-2605.06416 Sparse Blocked context onlyMay 7, 2026
- Continuous-Time Distribution Matching for Few-Step Diffusion Distillationarxiv-2605.06376 Sparse Blocked context onlyMay 7, 2026
- Is Escalation Worth It? A Decision-Theoretic Characterization of LLM Cascadesarxiv-2605.06350 Sparse Blocked context onlyMay 7, 2026
- Teaching Thinking Models to Reason with Tools: A Full-Pipeline Recipe for Tool-Integrated Reasoningarxiv-2605.06326 Sparse Blocked context onlyMay 7, 2026
- OPSD Compresses What RLVR Teaches: A Post-RL Compaction Stage for Reasoning Modelsarxiv-2605.06188 Sparse Blocked context onlyMay 7, 2026
- On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflowsarxiv-2605.06110 Sparse Blocked context onlyMay 7, 2026
- Uncovering Entity Identity Confusion in Multimodal Knowledge Editingarxiv-2605.06096 Sparse Blocked context onlyMay 7, 2026
- PersonaKit (PK): A Plug-and-Play Platform for User Testing Diverse Roles in Full-Duplex Dialoguearxiv-2605.06007 Sparse Blocked context onlyMay 7, 2026
- VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understandingarxiv-2605.05848 Sparse Blocked context onlyMay 7, 2026
- Steering Visual Generation in Unified Multimodal Models with Understanding Supervisionarxiv-2605.05781 Sparse Blocked context onlyMay 7, 2026
- Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMsarxiv-2606.27378 Sparse Blocked context onlyMay 7, 2026
- Belief Memory: Agent Memory Under Partial Observabilityarxiv-2605.05583 Sparse Blocked context onlyMay 7, 2026
- Who Prices Cognitive Labor in the Age of Agents? Compute-Anchored Wagesarxiv-2605.05558 Sparse Blocked context onlyMay 7, 2026
- ReaComp: Compiling LLM Reasoning into Symbolic Solvers for Efficient Program Synthesisarxiv-2605.05485 Sparse Blocked context onlyMay 6, 2026
- PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual Worldarxiv-2605.05163 Sparse Blocked context onlyMay 6, 2026
- Self-Induced Outcome Potential: Turn-Level Credit Assignment for Agents without Verifiersarxiv-2605.04984 Sparse Blocked context onlyMay 6, 2026
- A Foundation Model for Zero-Shot Logical Rule Inductionarxiv-2605.04916 Sparse Blocked context onlyMay 6, 2026
- DecodingTrust-Agent Platform (DTap): A Controllable and Interactive Red-Teaming Platform for AI Agentsarxiv-2605.04808 Sparse Blocked context onlyMay 6, 2026
- CHE-TKG: Collaborative Historical Evidence and Evolutionary Dynamics Learning for Temporal Knowledge Graph Reasoningarxiv-2605.04652 Sparse Blocked context onlyMay 6, 2026
- SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agenciesarxiv-2605.04637 Sparse Blocked context onlyMay 6, 2026
- RemoteZero: Geospatial Reasoning with Zero Human Annotationsarxiv-2605.04451 Sparse Blocked context onlyMay 6, 2026
- The Scaling Properties of Implicit Deductive Reasoning in Transformersarxiv-2605.04330 Sparse Blocked context onlyMay 5, 2026
- Self-Prompting Small Language Models for Privacy-Sensitive Clinical Information Extractionarxiv-2605.04221 Sparse Blocked context onlyMay 5, 2026
- Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluationarxiv-2605.04135 Sparse Blocked context onlyMay 5, 2026
- SymptomAI: Towards a Conversational AI Agent for Everyday Symptom Assessmentarxiv-2605.04012 Sparse Blocked context onlyMay 5, 2026
- A Benchmark for Interactive World Models with a Unified Action Generation Frameworkarxiv-2605.03941 Sparse Blocked context onlyMay 5, 2026
- The Counterexample Game: Iterated Conceptual Analysis and Repair in Language Modelsarxiv-2605.03936 Sparse Blocked context onlyMay 5, 2026
- Reproducing Complex Set-Compositional Information Retrievalarxiv-2605.03824 Sparse Blocked context onlyMay 5, 2026
- PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examinationarxiv-2605.03571 Sparse Blocked context onlyMay 5, 2026
- SkCC: Portable and Secure Skill Compilation for Cross-Framework LLM Agentsarxiv-2605.03353 Sparse Blocked context onlyMay 5, 2026
- RAG over Thinking Traces Can Improve Reasoning Tasksarxiv-2605.03344 Sparse Blocked context onlyMay 5, 2026
- ADAPTS: Agentic Decomposition for Automated Protocol-agnostic Tracking of Symptomsarxiv-2605.03212 Sparse Blocked context onlyMay 4, 2026
- HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harnessarxiv-2605.02396 Sparse Blocked context onlyMay 4, 2026
- Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decodingarxiv-2605.02290 Sparse Blocked context onlyMay 4, 2026
- Generative Modeling with Orbit-Space Particle Flow Matchingarxiv-2605.02222 Sparse Blocked context onlyMay 4, 2026
- T$^2$PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learningarxiv-2605.02178 Sparse Blocked context onlyMay 4, 2026
- Video Generation with Predictive Latentsarxiv-2605.02134 Sparse Blocked context onlyMay 4, 2026
- Hallucinations Undermine Trust; Metacognition is a Way Forwardarxiv-2605.01428 Sparse Blocked context onlyMay 2, 2026
- TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videosarxiv-2605.01234 Sparse Blocked context onlyMay 2, 2026
- Agentic AI Systems Should Be Designed as Marginal Token Allocatorsarxiv-2605.01214 Sparse Blocked context onlyMay 2, 2026
- WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wildarxiv-2605.01018 Sparse Blocked context onlyMay 1, 2026
- LASE: Language-Adversarial Speaker Encoding for Indic Cross-Script Identity Preservationarxiv-2605.00777 Sparse Blocked context onlyMay 1, 2026
- Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoringarxiv-2605.00754 Sparse Blocked context onlyMay 1, 2026
- BlenderRAG: High-Fidelity 3D Object Generation via Retrieval-Augmented Code Synthesisarxiv-2605.00632 Sparse Blocked context onlyMay 1, 2026
- Learning while Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policiesarxiv-2605.00416 Sparse Blocked context onlyMay 1, 2026
- Towards Customized Multimodal Role-Playarxiv-2605.08129 Curated Related Blocked context onlyMay 1, 2026
- CGM-JEPA: Learning Consistent Continuous Glucose Monitor Representations via Predictive Self-Supervised Pretrainingarxiv-2605.00933 Sparse Blocked context onlyMay 1, 2026
- Online Self-Calibration Against Hallucination in Vision-Language Modelsarxiv-2605.00323 Sparse Blocked context onlyMay 1, 2026
- Code World Model Preparedness Reportarxiv-2605.00932 Sparse Blocked context onlyMay 1, 2026
- Linking spatial biology and clinical histology via Haikuarxiv-2605.00925 Sparse Blocked context onlyApr 30, 2026
- Representation Fréchet Loss for Visual Generationarxiv-2604.28190 Sparse Blocked context onlyApr 30, 2026
- Synthetic Computers at Scale for Long-Horizon Productivity Simulationarxiv-2604.28181 Sparse Blocked context onlyApr 30, 2026
- Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientistsarxiv-2604.28158 Sparse Blocked context onlyApr 30, 2026
- Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RLarxiv-2604.28123 Sparse Blocked context onlyApr 30, 2026
- World Model for Robot Learning: A Comprehensive Surveyarxiv-2605.00080 Sparse Blocked context onlyApr 30, 2026
- WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environmentsarxiv-2604.27776 Sparse Blocked context onlyApr 30, 2026
- Heterogeneous Scientific Foundation Model Collaborationarxiv-2604.27351 Sparse Blocked context onlyApr 30, 2026
- Step-level Optimization for Efficient Computer-use Agentsarxiv-2604.27151 Sparse Blocked context onlyApr 29, 2026
- Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversationsarxiv-2604.27093 Sparse Blocked context onlyApr 29, 2026
- Length Value Model: Scalable Value Pretraining for Token-Level Length Modelingarxiv-2604.27039 Sparse Blocked context onlyApr 29, 2026
- Hypencoder Revisited: Reproducibility and Analysis of Non-Linear Scoring for First-Stage Retrievalarxiv-2604.27037 Sparse Blocked context onlyApr 29, 2026
- Accelerating RL Post-Training Rollouts via System-Integrated Speculative Decodingarxiv-2604.26779 Sparse Blocked context onlyApr 29, 2026
- FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probingarxiv-2604.26186 Sparse Blocked context onlyApr 29, 2026
- Operating-Layer Controls for Onchain Language-Model Agents Under Real Capitalarxiv-2604.26091 Sparse Blocked context onlyApr 28, 2026
- DV-World: Benchmarking Data Visualization Agents in Real-World Scenariosarxiv-2604.25914 Sparse Blocked context onlyApr 28, 2026
- How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuumarxiv-2604.25907 Sparse Blocked context onlyApr 28, 2026
- Mutual Forcing: Dual-Mode Self-Evolution for Fast Autoregressive Audio-Video Character Generationarxiv-2604.25819 Sparse Blocked context onlyApr 28, 2026
- MAIC-UI: Making Interactive Courseware with Generative UIarxiv-2604.25806 Sparse Blocked context onlyApr 28, 2026
- KinDER: A Physical Reasoning Benchmark for Robot Learning and Planningarxiv-2604.25788 Sparse Blocked context onlyApr 28, 2026
- PSP: An Interpretable Per-Dimension Accent Benchmark for Indic Text-to-Speecharxiv-2604.25476 Sparse Blocked context onlyApr 28, 2026
- R$^3$-SQL: Ranking Reward and Resampling for Text-to-SQLarxiv-2604.25325 Sparse Blocked context onlyApr 28, 2026
- BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debatearxiv-2604.25203 Sparse Blocked context onlyApr 28, 2026
- IAM: Identity-Aware Human Motion and Shape Joint Generationarxiv-2604.25164 Sparse Blocked context onlyApr 28, 2026
- ViPO: Visual Preference Optimization at Scalearxiv-2604.24953 Sparse Blocked context onlyApr 27, 2026
- Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimizationarxiv-2604.24952 Sparse Blocked context onlyApr 27, 2026
- Co-Director: Agentic Generative Video Storytellingarxiv-2604.24842 Sparse Blocked context onlyApr 27, 2026
- AutoGUI-v2: A Comprehensive Multi-Modal GUI Functionality Understanding Benchmarkarxiv-2604.24441 Sparse Blocked context onlyApr 27, 2026
- Diffusion Templates: A Unified Plugin Framework for Controllable Diffusionarxiv-2604.24351 Sparse Blocked context onlyApr 27, 2026
- TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agentsarxiv-2604.24005 Sparse Blocked context onlyApr 27, 2026
- MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generationarxiv-2604.23789 Sparse Blocked context onlyApr 26, 2026
- ClawMark: A Living-World Benchmark for Multi-Turn, Multi-Day, Multimodal Coworker Agentsarxiv-2604.23781 Sparse Blocked context onlyApr 26, 2026
- Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanismsarxiv-2604.23775 Sparse Blocked context onlyApr 26, 2026
- AnalogRetriever: Learning Cross-Modal Representations for Analog Circuit Retrievalarxiv-2604.23195 Sparse Blocked context onlyApr 25, 2026
- ProEval: Proactive Failure Discovery and Efficient Performance Estimation for Generative AI Evaluationarxiv-2604.23099 Sparse Blocked context onlyApr 25, 2026
- Code Broker: A Multi-Agent System for Automated Code Quality Assessmentarxiv-2604.23088 Sparse Blocked context onlyApr 25, 2026
- Representational Harms in LLM-Generated Narratives Against Global Majority Nationalitiesarxiv-2604.22749 Sparse Blocked context onlyApr 24, 2026
- Agentic World Modeling: Foundations, Capabilities, Laws, and Beyondarxiv-2604.22748 Sparse Blocked context onlyApr 24, 2026
- Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settingsarxiv-2604.22662 Sparse Blocked context onlyApr 24, 2026
- From graphemic dependence to lexical structure: a Markovian perspective on Dante's Commediaarxiv-2604.22626 Sparse Blocked context onlyApr 24, 2026
- FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editingarxiv-2604.22586 Sparse Blocked context onlyApr 24, 2026
- QuantClaw: Precision Where It Matters for OpenClawarxiv-2604.22577 Curated Related Blocked context onlyApr 24, 2026
- Data-Free Contribution Estimation in Federated Learning using Gradient von Neumann Entropyarxiv-2604.22562 Sparse Blocked context onlyApr 24, 2026
- Cross-Stage Coherence in Hierarchical Driving VQA: Explicit Baselines and Learned Gated Context Projectorsarxiv-2604.22560 Sparse Blocked context onlyApr 24, 2026
- Using Embedding Models to Improve Probabilistic Race Predictionarxiv-2604.22555 Sparse Blocked context onlyApr 24, 2026
- ArmSSL: Adversarial Robust Black-Box Watermarking for Self-Supervised Learning Pre-trained Encodersarxiv-2604.22550 Curated Related Blocked context onlyApr 24, 2026
- Superminds Test: Actively Evaluating Collective Intelligence of Agent Society via Probing Agentsarxiv-2604.22452 Sparse Blocked context onlyApr 24, 2026
- From Skills to Talent: Organising Heterogeneous Agents as a Real-World Companyarxiv-2604.22446 Sparse Blocked context onlyApr 24, 2026
- CognitiveTwin: Robust Multi-Modal Digital Twins for Predicting Cognitive Decline in Alzheimer's Diseasearxiv-2604.22428 Sparse Blocked context onlyApr 24, 2026
- From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variablesarxiv-2604.22416 Sparse Blocked context onlyApr 24, 2026
- Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Setsarxiv-2604.22294 Sparse Blocked context onlyApr 24, 2026
- STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generationarxiv-2604.22282 Sparse Blocked context onlyApr 24, 2026
- Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QAarxiv-2604.22239 Sparse Blocked context onlyApr 24, 2026
- Preserve Support, Not Correspondence: Dynamic Routing for Offline Reinforcement Learningarxiv-2604.22229 Sparse Blocked context onlyApr 24, 2026
- A Co-Evolutionary Theory of Human-AI Coexistence: Mutualism, Governance, and Dynamics in Complex Societiesarxiv-2604.22227 Sparse Blocked context onlyApr 24, 2026
- An LLM-Driven Closed-Loop Autonomous Learning Framework for Robots Facing Uncovered Tasks in Open Environmentsarxiv-2604.22199 Sparse Blocked context onlyApr 24, 2026
- From Global to Local: Rethinking CLIP Feature Aggregation for Person Re-Identificationarxiv-2604.22190 Sparse Blocked context onlyApr 24, 2026
- Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systemsarxiv-2604.22154 Sparse Blocked context onlyApr 24, 2026
- SketchVLM: Vision language models can annotate images to explain thoughts and guide usersarxiv-2604.22875 Sparse Blocked context onlyApr 23, 2026
- Removing Sandbagging in LLMs by Training with Weak Supervisionarxiv-2604.22082 Sparse Blocked context onlyApr 23, 2026
- Sound Agentic Science Requires Adversarial Experimentsarxiv-2604.22080 Sparse Blocked context onlyApr 23, 2026
- LayerBoost: Layer-Aware Attention Reduction for Efficient LLMsarxiv-2604.22050 Sparse Blocked context onlyApr 23, 2026
- Source-Modality Monitoring in Vision-Language Modelsarxiv-2604.22038 Sparse Blocked context onlyApr 23, 2026
- Probing Visual Planning in Image Editing Modelsarxiv-2604.22868 Sparse Blocked context onlyApr 23, 2026
- Universal Transformers Need Memory: Depth-State Trade-offs in Adaptive Recursive Reasoningarxiv-2604.21999 Sparse Blocked context onlyApr 23, 2026
- Soft Anisotropic Diagrams for Differentiable Image Representationarxiv-2604.21984 Sparse Blocked context onlyApr 23, 2026
- Seeing Fast and Slow: Learning the Flow of Time in Videosarxiv-2604.21931 Sparse Blocked context onlyApr 23, 2026
- UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detectionarxiv-2604.21904 Sparse Blocked context onlyApr 23, 2026
- WorldMark: A Unified Benchmark Suite for Interactive Video World Modelsarxiv-2604.21686 Sparse Blocked context onlyApr 23, 2026
- Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Modelsarxiv-2604.21523 Sparse Blocked context onlyApr 23, 2026
- VLAA-GUI: Knowing When to Stop, Recover, and Search, A Modular Framework for GUI Automationarxiv-2604.21375 Sparse Blocked context onlyApr 23, 2026
- Cross-Entropy Is Load-Bearing: A Pre-Registered Scope Test of the K-Way Energy Probe on Bidirectional Predictive Codingarxiv-2604.21286 Sparse Blocked context onlyApr 23, 2026
- Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithmsarxiv-2604.21131 Sparse Blocked context onlyApr 22, 2026
- Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweightingarxiv-2604.21082 Sparse Blocked context onlyApr 22, 2026
- TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stoppingarxiv-2604.21057 Sparse Blocked context onlyApr 22, 2026
- The Last Harness You'll Ever Buildarxiv-2604.21003 Sparse Blocked context onlyApr 22, 2026
- DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitationarxiv-2604.20841 Sparse Blocked context onlyApr 22, 2026
- Working Memory Constraints Scaffold Learning in Transformers under Data Scarcityarxiv-2604.20789 Sparse Blocked context onlyApr 22, 2026
- SWE-chat: Coding Agent Interactions From Real Users in the Wildarxiv-2604.20779 Sparse Blocked context onlyApr 22, 2026
- Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agentsarxiv-2604.20572 Sparse Blocked context onlyApr 22, 2026
- MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptationarxiv-2604.20468 Sparse Blocked context onlyApr 22, 2026
- Decoding Text Spans for Efficient and Accurate Named-Entity Recognitionarxiv-2604.20447 Sparse Blocked context onlyApr 22, 2026
- MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skillsarxiv-2604.20441 Sparse Blocked context onlyApr 22, 2026
- Building a Precise Video Language with Human-AI Oversightarxiv-2604.21718 Sparse Blocked context onlyApr 22, 2026
- Cortex 2.0: Grounding World Models in Real-World Industrial Deploymentarxiv-2604.20246 Sparse Blocked context onlyApr 22, 2026
- UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modelingarxiv-2604.19734 Sparse Blocked context onlyApr 21, 2026
- Taming Actor-Observer Asymmetry in Agents via Dialectical Alignmentarxiv-2604.19548 Sparse Blocked context onlyApr 21, 2026
- ATTN-FIQA: Interpretable Attention-based Face Image Quality Assessment with Vision Transformersarxiv-2604.22841 Sparse Blocked context onlyApr 21, 2026
- Does Self-Consistency Improve the Recall of Encyclopedic Knowledge?arxiv-2604.19395 Sparse Blocked context onlyApr 21, 2026
- TEMPO: Scaling Test-time Training for Large Reasoning Modelsarxiv-2604.19295 Sparse Blocked context onlyApr 21, 2026
- ClawNet: Human-Symbiotic Agent Network for Cross-User Autonomous Cooperationarxiv-2604.19211 Sparse Blocked context onlyApr 21, 2026
- ClawEnvKit: Automatic Environment Generation for Claw-Like Agentsarxiv-2604.18543 Sparse Blocked context onlyApr 20, 2026
- OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanationarxiv-2604.18486 Sparse Blocked context onlyApr 20, 2026
- MARCO: Navigating the Unseen Space of Semantic Correspondencearxiv-2604.18267 Sparse Blocked context onlyApr 20, 2026
- Mitigating Multimodal Hallucination via Phase-wise Self-rewardarxiv-2604.17982 Sparse Blocked context onlyApr 20, 2026
- Latent Preference Modeling for Cross-Session Personalized Tool Callingarxiv-2604.17886 Sparse Blocked context onlyApr 20, 2026
- GraSP: Graph-Structured Skill Compositions for LLM Agentsarxiv-2604.17870 Sparse Blocked context onlyApr 20, 2026
- Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systemsarxiv-2604.17487 Sparse Blocked context onlyApr 19, 2026
- UniMesh: Unifying 3D Mesh Understanding and Generationarxiv-2604.17472 Sparse Blocked context onlyApr 19, 2026
- SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agentsarxiv-2604.17308 Sparse Blocked context onlyApr 19, 2026
- The Continuity Layer: Why Intelligence Needs an Architecture for What It Carries Forwardarxiv-2604.17273 Sparse Blocked context onlyApr 19, 2026
- Bolzano: Case Studies in LLM-Assisted Mathematical Researcharxiv-2604.16989 Sparse Blocked context onlyApr 18, 2026
- The Cognitive Penalty: Ablating System 1 and System 2 Reasoning in Edge-Native SLMs for Decentralized Consensusarxiv-2604.16913 Sparse Blocked context onlyApr 18, 2026
- Crowded in B-Space: Calibrating Shared Directions for LoRA Mergingarxiv-2604.16826 Sparse Blocked context onlyApr 18, 2026
- Benign Fine-Tuning Breaks Safety Alignment in Audio LLMsarxiv-2604.16659 Sparse Blocked context onlyApr 17, 2026
- Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Modelsarxiv-2604.16593 Sparse Blocked context onlyApr 17, 2026
- Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMsarxiv-2604.16060 Sparse Blocked context onlyApr 17, 2026
- Elucidating the SNR-t Bias of Diffusion Probabilistic Modelsarxiv-2604.16044 Sparse Blocked context onlyApr 17, 2026
- Cut Your Losses! Learning to Prune Paths Early for Efficient Parallel Reasoningarxiv-2604.16029 Sparse Blocked context onlyApr 17, 2026
- Where does output diversity collapse in post-training?arxiv-2604.16027 Sparse Blocked context onlyApr 17, 2026
- VoxMind: An End-to-End Agentic Spoken Dialogue Systemarxiv-2604.15710 Sparse Blocked context onlyApr 17, 2026
- Stargazer: A Scalable Model-Fitting Benchmark Environment for AI Agents under Astrophysical Constraintsarxiv-2604.15664 Sparse Blocked context onlyApr 17, 2026
- Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utilityarxiv-2604.15579 Sparse Blocked context onlyApr 16, 2026
- (1D) Ordered Tokens Enable Efficient Test-Time Searcharxiv-2604.15453 Sparse Blocked context onlyApr 16, 2026
- LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectoriesarxiv-2604.15311 Sparse Blocked context onlyApr 16, 2026
- GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokensarxiv-2604.15284 Sparse Blocked context onlyApr 16, 2026
- Prism: Symbolic Superoptimization of Tensor Programsarxiv-2604.15272 Sparse Blocked context onlyApr 16, 2026
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmasarxiv-2604.15267 Sparse Blocked context onlyApr 16, 2026
- Agentic Microphysics: A Manifesto for Generative AI Safetyarxiv-2604.15236 Sparse Blocked context onlyApr 16, 2026
- Blue Data Intelligence Layer: Streaming Data and Agents for Multi-source Multi-modal Data-Centric Applicationsarxiv-2604.15233 Sparse Blocked context onlyApr 16, 2026
- RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomographyarxiv-2604.15231 Sparse Blocked context onlyApr 16, 2026
- Context Over Content: Exposing Evaluation Faking in Automated Judgesarxiv-2604.15224 Sparse Blocked context onlyApr 16, 2026
- AI-Assisted Requirements Engineering: An Empirical Evaluation Relative to Expert Judgmentarxiv-2604.15222 Sparse Blocked context onlyApr 16, 2026
- Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understandingarxiv-2604.15210 Sparse Blocked context onlyApr 16, 2026
- Benchmarking Classical Coverage Path Planning Heuristics on Irregular Hexagonal Grids for Maritime Coverage Scenariosarxiv-2604.15202 Sparse Blocked context onlyApr 16, 2026
- Meituan Merchant Business Diagnosis via Policy-Guided Dual-Process User Simulationarxiv-2604.15190 Sparse Blocked context onlyApr 16, 2026
- PRL-Bench: A Comprehensive Benchmark Evaluating LLMs' Capabilities in Frontier Physics Researcharxiv-2604.15411 Sparse Blocked context onlyApr 16, 2026
- VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Modelsarxiv-2604.15188 Sparse Blocked context onlyApr 16, 2026
- Agent-Aided Design for Dynamic CAD Modelsarxiv-2604.15184 Sparse Blocked context onlyApr 16, 2026
- LLMs Gaming Verifiers: RLVR can Lead to Reward Hackingarxiv-2604.15149 Sparse Blocked context onlyApr 16, 2026
- SRMU: Relevance-Gated Updates for Streaming Hyperdimensional Memoriesarxiv-2604.15121 Sparse Blocked context onlyApr 16, 2026
- HyperSpace: A Generalized Framework for Spatial Encoding in Hyperdimensional Representationsarxiv-2604.15113 Sparse Blocked context onlyApr 16, 2026
- Where are the Humans? A Scoping Review of Fairness in Multi-agent AI Systemsarxiv-2604.15078 Sparse Blocked context onlyApr 16, 2026
- CoGrid & the Multi-User Gymnasium: A Framework for Multi-Agent Experimentationarxiv-2604.15044 Sparse Blocked context onlyApr 16, 2026
- From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bencharxiv-2604.15037 Sparse Blocked context onlyApr 16, 2026
- Autogenesis: A Self-Evolving Agent Protocolarxiv-2604.15034 Sparse Blocked context onlyApr 16, 2026
- COEVO: Co-Evolutionary Framework for Joint Functional Correctness and PPA Optimization in LLM-Based RTL Generationarxiv-2604.15001 Sparse Blocked context onlyApr 16, 2026
- CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Predictionarxiv-2604.14651 Sparse Blocked context onlyApr 16, 2026
- Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Modelsarxiv-2604.14629 Sparse Blocked context onlyApr 16, 2026
- DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselinesarxiv-2604.14314 Sparse Blocked context onlyApr 15, 2026
- SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environmentsarxiv-2604.14144 Sparse Blocked context onlyApr 15, 2026
- C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferencesarxiv-2604.13618 Sparse Blocked context onlyApr 15, 2026
- Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understandingarxiv-2604.13313 Sparse Blocked context onlyApr 14, 2026
- Unleashing Implicit Rewards: Prefix-Value Learning for Distribution-Level Optimizationarxiv-2604.13197 Sparse Blocked context onlyApr 14, 2026
- Exploration and Exploitation Errors Are Measurable for Language Model Agentsarxiv-2604.13151 Sparse Blocked context onlyApr 14, 2026
- Learning Versatile Humanoid Manipulation with Touch Dreamingarxiv-2604.13015 Sparse Blocked context onlyApr 14, 2026
- GlotOCR Bench: OCR Models Still Struggle Beyond a Handful of Unicode Scriptsarxiv-2604.12978 Sparse Blocked context onlyApr 14, 2026
- VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenizationarxiv-2604.12887 Sparse Blocked context onlyApr 14, 2026
- Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literaturearxiv-2604.12243 Sparse Blocked context onlyApr 14, 2026
- Artificial Intelligence Index Report 2026arxiv-2606.15708 Sparse Blocked context onlyApr 14, 2026
- Domain-Specific Latent Representations Improve the Fidelity of Diffusion-Based Medical Image Super-Resolutionarxiv-2604.12152 Sparse Blocked context onlyApr 14, 2026
- Beyond Perception Errors: Semantic Fixation in Large Vision-Language Modelsarxiv-2604.12119 Sparse Blocked context onlyApr 13, 2026
- TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignmentarxiv-2604.12012 Sparse Blocked context onlyApr 13, 2026
- Narrative-Driven Paper-to-Slide Generation via ArcDeckarxiv-2604.11969 Sparse Blocked context onlyApr 13, 2026
- HDR Video Generation via Latent Alignment with Logarithmic Encodingarxiv-2604.11788 Sparse Blocked context onlyApr 13, 2026
- Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasksarxiv-2604.11753 Sparse Blocked context onlyApr 13, 2026
- LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignmentarxiv-2604.11689 Sparse Blocked context onlyApr 13, 2026
- RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Timearxiv-2604.11626 Sparse Blocked context onlyApr 13, 2026
- Seeing Through Touch: Tactile-Driven Visual Localization of Material Regionsarxiv-2604.11579 Sparse Blocked context onlyApr 13, 2026
- SemaClaw: A Step Towards General-Purpose Personal AI Agents through Harness Engineeringarxiv-2604.11548 Sparse Blocked context onlyApr 13, 2026
- Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understandingarxiv-2604.11177 Sparse Blocked context onlyApr 13, 2026
- ActorMind: Emulating Human Actor Reasoning for Speech Role-Playingarxiv-2604.11103 Sparse Blocked context onlyApr 13, 2026
- Sema Code: Decoupling AI Coding Agents into Programmable, Embeddable Infrastructurearxiv-2604.11045 Sparse Blocked context onlyApr 13, 2026
- You Only Judge Once: Multi-response Reward Modeling in a Single Forward Passarxiv-2604.10966 Sparse Blocked context onlyApr 13, 2026
- PokeRL: Reinforcement Learning for Pokemon Redarxiv-2604.10812 Sparse Blocked context onlyApr 12, 2026
- Reason Only When Needed: Efficient Generative Reward Modeling via Model-Internal Uncertaintyarxiv-2604.10072 Sparse Blocked context onlyApr 11, 2026
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Modelsarxiv-2604.08545 Sparse Blocked context onlyApr 9, 2026
- Meta-learning In-Context Enables Training-Free Cross Subject Brain Decodingarxiv-2604.08537 Sparse Blocked context onlyApr 9, 2026
- PSI: Shared State as the Missing Layer for Coherent AI-Generated Instruments in Personal AI Agentsarxiv-2604.08529 Sparse Blocked context onlyApr 9, 2026
- 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decodingarxiv-2604.08645 Sparse Blocked context onlyApr 9, 2026
- What They Saw, Not Just Where They Looked: Semantic Scanpath Similarity via VLMs and NLP metricarxiv-2604.08494 Sparse Blocked context onlyApr 9, 2026
- TTVS: Boosting Self-Exploring Reinforcement Learning via Test-time Variational Synthesisarxiv-2604.08468 Sparse Blocked context onlyApr 9, 2026
- CrashSight: A Phase-Aware, Infrastructure-Centric Video Benchmark for Traffic Crash Scene Understanding and Reasoningarxiv-2604.08457 Sparse Blocked context onlyApr 9, 2026
- On-board Telemetry Monitoring in Autonomous Satellites: Challenges and Opportunitiesarxiv-2604.08424 Sparse Blocked context onlyApr 9, 2026
- Exploring Temporal Representation in Neural Processes for Multimodal Action Predictionarxiv-2604.08418 Sparse Blocked context onlyApr 9, 2026
- TASU2: Controllable CTC Simulation for Alignment and Low-Resource Adaptation of Speech LLMsarxiv-2604.08384 Sparse Blocked context onlyApr 9, 2026
- Scalable Neural Decoders for Practical Fault-Tolerant Quantum Computationarxiv-2604.08358 Sparse Blocked context onlyApr 9, 2026
- Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimizationarxiv-2604.08324 Sparse Blocked context onlyApr 9, 2026
- Activation Steering for Aligned Open-ended Generation without Sacrificing Coherencearxiv-2604.08169 Sparse Blocked context onlyApr 9, 2026
- Multimodal Reasoning with LLM for Encrypted Traffic Interpretation: A Benchmarkarxiv-2604.08140 Sparse Blocked context onlyApr 9, 2026
- SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinkingarxiv-2604.07922 Sparse Blocked context onlyApr 9, 2026
- An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooksarxiv-2604.07883 Sparse Blocked context onlyApr 9, 2026
- MemReader: From Passive to Active Extraction for Long-Term Agent Memoryarxiv-2604.07877 Sparse Blocked context onlyApr 9, 2026
- Symbiotic-MoE: Unlocking the Synergy between Generation and Understandingarxiv-2604.07753 Sparse Blocked context onlyApr 9, 2026
- DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verificationarxiv-2604.07622 Sparse Blocked context onlyApr 8, 2026
- Reasoning Graphs: Deterministic Agent Accuracy through Evidence-Centric Chain-of-Thought Feedbackarxiv-2604.07595 Sparse Blocked context onlyApr 8, 2026
- From Ground Truth to Measurement: A Statistical Framework for Human Labelingarxiv-2604.07591 Sparse Blocked context onlyApr 8, 2026
- Lexical Tone is Hard to Quantize: Probing Discrete Speech Units in Mandarin and Yorùbáarxiv-2604.07467 Sparse Blocked context onlyApr 8, 2026
- OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligencearxiv-2604.07296 Sparse Blocked context onlyApr 8, 2026
- LaScA: Language-Conditioned Scalable Modelling of Affective Dynamicsarxiv-2604.07193 Sparse Blocked context onlyApr 8, 2026
- Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answeringarxiv-2604.07116 Sparse Blocked context onlyApr 8, 2026
- Cognitive Loop of Thought: Reversible Hierarchical Markov Chain for Efficient Mathematical Reasoningarxiv-2604.06805 Sparse Blocked context onlyApr 8, 2026
- When Is Thinking Enough? Early Exit via Sufficiency Assessment for Efficient Reasoningarxiv-2604.06787 Sparse Blocked context onlyApr 8, 2026
- Multi-Faceted Self-Consistent Preference Alignment for Query Rewriting in Conversational Searcharxiv-2604.06771 Sparse Blocked context onlyApr 8, 2026
- Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Trackingarxiv-2606.15673 Sparse Blocked context onlyApr 8, 2026
- ATANT: An Evaluation Framework for AI Continuityarxiv-2604.06710 Sparse Blocked context onlyApr 8, 2026
- A Parameter-Efficient Transfer Learning Approach through Multitask Prompt Distillation and Decomposition for Clinical NLParxiv-2604.06650 Sparse Blocked context onlyApr 8, 2026
- Feedback Adaptation for Retrieval-Augmented Generationarxiv-2604.06647 Sparse Blocked context onlyApr 8, 2026
- The Detection--Extraction Gap: Models Know the Answer Before They Can Say Itarxiv-2604.06613 Sparse Blocked context onlyApr 8, 2026
- Does a Global Perspective Help Prune Sparse MoEs Elegantly?arxiv-2604.06542 Sparse Blocked context onlyApr 8, 2026
- Transformer See, Transformer Do: Copying as an Intermediate Step in Learning Analogical Reasoningarxiv-2604.06501 Sparse Blocked context onlyApr 7, 2026
- When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don'tarxiv-2604.06422 Sparse Blocked context onlyApr 7, 2026
- Say Something Else: Rethinking Contextual Privacy as Information Sufficiencyarxiv-2604.06409 Sparse Blocked context onlyApr 7, 2026
- The Illusion of Superposition? A Principled Analysis of Latent Thinking in Language Modelsarxiv-2604.06374 Sparse Blocked context onlyApr 7, 2026