300 canonical paper links on this archive page.
- SoK: Agentic Retrieval-Augmented Generation (RAG): Taxonomy, Architectures, Evaluation, and Research Directionsarxiv-2603.07379 Sparse Blocked context onlyMar 7, 2026
- The Third Ambition: Artificial Intelligence and the Science of Human Behaviorarxiv-2603.07329 Sparse Blocked context onlyMar 7, 2026
- Lying to Win: Assessing LLM Deception through Human-AI Games and Parallel-World Probingarxiv-2603.07202 Sparse Blocked context onlyMar 7, 2026
- Governance Architecture for Autonomous Agent Systems: Threats, Framework, and Engineering Practicearxiv-2603.07191 Sparse Blocked context onlyMar 7, 2026
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledgearxiv-2603.07131 Sparse Blocked context onlyMar 7, 2026
- Enhancing Consistency of Werewolf AI through Dialogue Summarization and Persona Informationarxiv-2603.07111 Sparse Blocked context onlyMar 7, 2026
- Entropy-Aware On-Policy Distillation of Language Modelsarxiv-2603.07079 Sparse Blocked context onlyMar 7, 2026
- Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervisionarxiv-2603.07025 Sparse Blocked context onlyMar 7, 2026
- AutoChecklist: Composable Pipelines for Checklist Generation and Scoring with LLM-as-a-Judgearxiv-2603.07019 Sparse Blocked context onlyMar 7, 2026
- Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Modelsarxiv-2603.07017 Sparse Blocked context onlyMar 7, 2026
- Chart-RL: Generalized Chart Comprehension via Reinforcement Learning with Verifiable Rewardsarxiv-2603.06958 Sparse Blocked context onlyMar 7, 2026
- Reforming the Mechanism: Editing Reasoning Patterns in LLMs with Circuit Reshapingarxiv-2603.06923 Sparse Blocked context onlyMar 6, 2026
- A Dynamic Self-Evolving Extraction Systemarxiv-2603.06915 Sparse Blocked context onlyMar 6, 2026
- LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Modelsarxiv-2603.06874 Sparse Blocked context onlyMar 6, 2026
- "Dark Triad" Model Organisms of Misalignment: Narrow Fine-Tuning Mirrors Human Antisocial Behaviorarxiv-2603.06816 Sparse Blocked context onlyMar 6, 2026
- Beyond Rows to Reasoning: Agentic Retrieval for Multimodal Spreadsheet Understanding and Editingarxiv-2603.06503 Sparse Blocked context onlyMar 6, 2026
- COLD-Steer: Steering Large Language Models via In-Context One-step Learning Dynamicsarxiv-2603.06495 Sparse Blocked context onlyMar 6, 2026
- Abductive Reasoning with Syllogistic Forms in Large Language Modelsarxiv-2603.06428 Sparse Blocked context onlyMar 6, 2026
- From Prompting to Preference Optimization: A Comparative Study of LLM-based Automated Essay Scoringarxiv-2603.06424 Sparse Blocked context onlyMar 6, 2026
- Evaluation of Deontic Conditional Reasoning in Large Language Models: The Case of Wason's Selection Taskarxiv-2603.06416 Sparse Blocked context onlyMar 6, 2026
- Qworld: Question-Specific Evaluation Criteria for LLMsarxiv-2603.23522 Sparse Blocked context onlyMar 6, 2026
- Mind the Gap: Pitfalls of LLM Alignment with Asian Public Opinionarxiv-2603.06264 Sparse Blocked context onlyMar 6, 2026
- SPOT: Span-level Pause-of-Thought for Efficient and Interpretable Latent Reasoning in Large Language Modelsarxiv-2603.06222 Sparse Blocked context onlyMar 6, 2026
- LIT-RAGBench: Benchmarking Generator Capabilities of Large Language Models in Retrieval-Augmented Generationarxiv-2603.06198 Sparse Blocked context onlyMar 6, 2026
- Wisdom of the AI Crowd (AI-CROWD) for Ground Truth Approximation in Content Analysis: A Research Protocol & Validation Using Eleven Large Language Modelsarxiv-2603.06197 Sparse Blocked context onlyMar 6, 2026
- Diffusion Language Models Are Natively Length-Awarearxiv-2603.06123 Sparse Blocked context onlyMar 6, 2026
- Making Implicit Premises Explicit in Logical Understanding of Enthymemesarxiv-2603.06114 Sparse Blocked context onlyMar 6, 2026
- Experiences Build Characters: The Linguistic Origins and Functional Impact of LLM Personalityarxiv-2603.06088 Sparse Blocked context onlyMar 6, 2026
- ViewFusion: Structured Spatial Thinking Chains for Multi-View Reasoningarxiv-2603.06024 Sparse Blocked context onlyMar 6, 2026
- MASFactory: A Graph-centric Framework for Orchestrating LLM-Based Multi-Agent Systems with Vibe Graphingarxiv-2603.06007 Direct Blocked context onlyMar 6, 2026
- Implicit Style Conditioning: A Structured Style-Rewrite Framework for Low-Resource Character Modelingarxiv-2603.05933 Sparse Blocked context onlyMar 6, 2026
- X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMsarxiv-2603.24596 Sparse Blocked context onlyMar 6, 2026
- Learning Next Action Predictors from Human-Computer Interactionarxiv-2603.05923 Sparse Blocked context onlyMar 6, 2026
- VerChol -- Grammar-First Tokenization for Agglutinative Languagesarxiv-2603.05883 Sparse Blocked context onlyMar 6, 2026
- ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learningarxiv-2603.05863 Sparse Blocked context onlyMar 6, 2026
- Knowing without Acting: The Disentangled Geometry of Safety Mechanisms in Large Language Modelsarxiv-2603.05773 Sparse Blocked context onlyMar 6, 2026
- Depth Charge: Jailbreak Large Language Models from Deep Safety Attention Headsarxiv-2603.05772 Sparse Blocked context onlyMar 6, 2026
- Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thoughtarxiv-2603.05488 Curated Related Blocked context onlyMar 5, 2026
- Leveraging LLM Parametric Knowledge for Fact Checking without Retrievalarxiv-2603.05471 Sparse Blocked context onlyMar 5, 2026
- Distributed Partial Information Puzzles: Examining Common Ground Construction Under Epistemic Asymmetryarxiv-2603.05450 Sparse Blocked context onlyMar 5, 2026
- DiSCTT: Consensus-Guided Self-Curriculum for Efficient Test-Time Adaptation in Reasoningarxiv-2603.05357 Sparse Blocked context onlyMar 5, 2026
- Building Effective AI Coding Agents for the Terminal: Scaffolding, Harness, Context Engineering, and Lessons Learnedarxiv-2603.05344 Sparse Blocked context onlyMar 5, 2026
- WavSLM: Single-Stream Speech Language Modeling via WavLM Distillationarxiv-2603.05299 Sparse Blocked context onlyMar 5, 2026
- SarcasmMiner: A Dual-Track Post-Training Framework for Robust Audio-Visual Sarcasm Reasoningarxiv-2603.05275 Sparse Blocked context onlyMar 5, 2026
- Balancing Coverage and Draft Latency in Vocabulary Trimming for Faster Speculative Decodingarxiv-2603.05210 Sparse Blocked context onlyMar 5, 2026
- Distilling Formal Logic into Neural Spaces: A Kernel Alignment Approach for Signal Temporal Logicarxiv-2603.05198 Sparse Blocked context onlyMar 5, 2026
- C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoningarxiv-2603.05167 Sparse Blocked context onlyMar 5, 2026
- LBM: Hierarchical Large Auto-Bidding Model via Reasoning and Actingarxiv-2603.05134 Sparse Blocked context onlyMar 5, 2026
- NeuronMoE: Neuron-Guided Mixture-of-Experts for Efficient Multilingual LLM Extensionarxiv-2603.05046 Sparse Blocked context onlyMar 5, 2026
- Survive at All Costs: Exploring LLM's Risky Behaviors under Survival Pressurearxiv-2603.05028 Sparse Blocked context onlyMar 5, 2026
- S5-SHB Agent: Society 5.0 enabled Multi-model Agentic Blockchain Framework for Smart Homearxiv-2603.05027 Sparse Blocked context onlyMar 5, 2026
- ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contextsarxiv-2603.04992 Direct Blocked context onlyMar 5, 2026
- When Weak LLMs Speak with Confidence, Preference Alignment Gets Strongerarxiv-2603.04968 Sparse Blocked context onlyMar 5, 2026
- VisionPangu: A Compact and Fine-Grained Multimodal Assistant with 1.7B Parametersarxiv-2603.04957 Sparse Blocked context onlyMar 5, 2026
- Retrieval-Augmented Generation with Covariate Time Seriesarxiv-2603.04951 Sparse Blocked context onlyMar 5, 2026
- AILS-NTUA at SemEval-2026 Task 10: Agentic LLMs for Psycholinguistic Marker Extraction and Conspiracy Endorsement Detectionarxiv-2603.04921 Sparse Blocked context onlyMar 5, 2026
- Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Modelsarxiv-2603.04893 Sparse Blocked context onlyMar 5, 2026
- Beyond the Context Window: A Cost-Performance Analysis of Fact-Based Memory vs. Long-Context LLMs for Persistent Agentsarxiv-2603.04814 Sparse Blocked context onlyMar 5, 2026
- TSEmbed: Unlocking Task Scaling in Universal Multimodal Embeddingsarxiv-2603.04772 Sparse Blocked context onlyMar 5, 2026
- Stacked from One: Multi-Scale Self-Injection for Context Window Extensionarxiv-2603.04759 Sparse Blocked context onlyMar 5, 2026
- DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrievalarxiv-2603.04743 Sparse Blocked context onlyMar 5, 2026
- IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluationarxiv-2603.04738 Sparse Blocked context onlyMar 5, 2026
- Optimizing Language Models for Crosslingual Knowledge Consistencyarxiv-2603.04678 Sparse Blocked context onlyMar 4, 2026
- Using Vision + Language Models to Predict Item Difficultyarxiv-2603.04670 Sparse Blocked context onlyMar 4, 2026
- Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Developmentarxiv-2603.04601 Sparse Blocked context onlyMar 4, 2026
- TaxonRL: Reinforcement Learning with Intermediate Rewards for Interpretable Fine-Grained Visual Reasoningarxiv-2603.04380 Sparse Blocked context onlyMar 4, 2026
- World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddingsarxiv-2603.04317 Sparse Blocked context onlyMar 4, 2026
- Memex(RL): Scaling Long-Horizon LLM Agents via Indexed Experience Memoryarxiv-2603.04257 Sparse Blocked context onlyMar 4, 2026
- Retrieval or Representation? Reassessing Benchmark Gaps in Multilingual and Visually Rich RAGarxiv-2603.04238 Sparse Blocked context onlyMar 4, 2026
- When Do Language Models Endorse Limitations on Human Rights Principles?arxiv-2603.04217 Sparse Blocked context onlyMar 4, 2026
- Bielik-Q2-Sharp: A Comparative Study of Extreme 2-bit Quantization Methods for a Polish 11B Language Modelarxiv-2603.04162 Sparse Blocked context onlyMar 4, 2026
- Traces of Social Competence in Large Language Modelsarxiv-2603.04161 Sparse Blocked context onlyMar 4, 2026
- BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoningarxiv-2603.04124 Sparse Blocked context onlyMar 4, 2026
- Monitoring Emergent Reward Hacking During Generation via Internal Activationsarxiv-2603.04069 Sparse Blocked context onlyMar 4, 2026
- Who Judges the Judge? Evaluating LLM-as-a-Judge for French Medical open-ended QAarxiv-2603.04033 Sparse Blocked context onlyMar 4, 2026
- Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effectsarxiv-2603.03915 Sparse Blocked context onlyMar 4, 2026
- From Threat Intelligence to Firewall Rules: Semantic Relations in Hybrid AI Agent and Expert System Architecturesarxiv-2603.03911 Sparse Blocked context onlyMar 4, 2026
- Semantic Bridging Domains: Pseudo-Source as Test-Time Connectorarxiv-2603.03844 Sparse Blocked context onlyMar 4, 2026
- SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integrationarxiv-2603.03823 Sparse Blocked context onlyMar 4, 2026
- MOOSE-Star: Unlocking Tractable Training for Scientific Discovery by Breaking the Complexity Barrierarxiv-2603.03756 Sparse Blocked context onlyMar 4, 2026
- Confidence-Calibrated Small-Large Language Model Collaboration for Cost-Efficient Reasoningarxiv-2603.03752 Sparse Blocked context onlyMar 4, 2026
- CONCUR: Benchmarking LLMs for Concurrent Code Generationarxiv-2603.03683 Sparse Blocked context onlyMar 4, 2026
- Self-Sovereign Agentarxiv-2604.08551 Sparse Blocked context onlyMar 4, 2026
- MIND: Unified Inquiry and Diagnosis RL with Criteria Grounded Clinical Supports for Psychiatric Consultationarxiv-2603.03677 Sparse Blocked context onlyMar 4, 2026
- Can LLM Aid in Solving Constraints with Inductive Definitions?arxiv-2603.03668 Sparse Blocked context onlyMar 4, 2026
- ByteFlow: Language Modeling through Adaptive Byte Compression without a Tokenizerarxiv-2603.03583 Sparse Blocked context onlyMar 3, 2026
- Build, Judge, Optimize: A Blueprint for Continuous Improvement of Multi-Agent Consumer Assistantsarxiv-2603.03565 Sparse Blocked context onlyMar 3, 2026
- Tucano 2 Cool: Better Open Source LLMs for Portuguesearxiv-2603.03543 Sparse Blocked context onlyMar 3, 2026
- Farther the Shift, Sparser the Representation: Analyzing OOD Mechanisms in LLMsarxiv-2603.03415 Sparse Blocked context onlyMar 3, 2026
- Code2Math: Can Your Code Agent Effectively Evolve Math Problems Through Exploration?arxiv-2603.03202 Sparse Blocked context onlyMar 3, 2026
- MoD-DPO: Towards Mitigating Cross-modal Hallucinations in Omni LLMs using Modality Decoupled Preference Optimizationarxiv-2603.03192 Sparse Blocked context onlyMar 3, 2026
- Agentic AI-based Coverage Closure for Formal Verificationarxiv-2603.03147 Sparse Blocked context onlyMar 3, 2026
- TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Modelsarxiv-2603.03081 Sparse Blocked context onlyMar 3, 2026
- PrivMedChat: End-to-End Differentially Private RLHF for Medical Dialogue Systemsarxiv-2603.03054 Sparse Blocked context onlyMar 3, 2026
- MaBERT:A Padding Safe Interleaved Transformer Mamba Hybrid Encoder for Efficient Extended Context Masked Language Modelingarxiv-2603.03001 Sparse Blocked context onlyMar 3, 2026
- Contextualized Privacy Defense for LLM Agentsarxiv-2603.02983 Sparse Blocked context onlyMar 3, 2026
- Learning to Generate and Extract: A Multi-Agent Collaboration Framework For Zero-shot Document-level Event Arguments Extractionarxiv-2603.02909 Sparse Blocked context onlyMar 3, 2026
- Eval4Sim: An Evaluation Framework for Persona Simulationarxiv-2603.02876 Sparse Blocked context onlyMar 3, 2026
- A Browser-based Open Source Assistant for Multimodal Content Verificationarxiv-2603.02842 Sparse Blocked context onlyMar 3, 2026
- OCR or Not? Rethinking Document Information Extraction in the MLLMs Era with Real-World Large-Scale Datasetsarxiv-2603.02789 Sparse Blocked context onlyMar 3, 2026
- Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarksarxiv-2603.04459 Curated Related Blocked context onlyMar 3, 2026
- Sensory-Aware Sequential Recommendation via Review-Distilled Representationsarxiv-2603.02709 Sparse Blocked context onlyMar 3, 2026
- Graph-GRPO: Stabilizing Multi-Agent Topology Learning via Group Relative Policy Optimizationarxiv-2603.02701 Sparse Blocked context onlyMar 3, 2026
- ITLC at SemEval-2026 Task 11: Normalization and Deterministic Parsing for Formal Reasoning in LLMsarxiv-2603.02676 Sparse Blocked context onlyMar 3, 2026
- Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approachesarxiv-2603.02655 Sparse Blocked context onlyMar 3, 2026
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learningarxiv-2603.02637 Sparse Blocked context onlyMar 3, 2026
- Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Modelsarxiv-2603.02631 Sparse Blocked context onlyMar 3, 2026
- Think, But Don't Overthink: Reproducing Recursive Language Modelsarxiv-2603.02615 Sparse Blocked context onlyMar 3, 2026
- MUSE: A Run-Centric Platform for Multimodal Unified Safety Evaluation of Large Language Modelsarxiv-2603.02482 Sparse Blocked context onlyMar 3, 2026
- Detecting AI-Generated Essays in Writing Assessment: Responsible Use and Generalizability Across LLMsarxiv-2603.02353 Sparse Blocked context onlyMar 2, 2026
- Characterizing Memorization in Diffusion Language Models: Generalized Extraction and Sampling Effectsarxiv-2603.02333 Sparse Blocked context onlyMar 2, 2026
- Reasoning Core: A Scalable Procedural Data Generation Suite for Symbolic Pre-training and Post-Trainingarxiv-2603.02208 Sparse Blocked context onlyMar 2, 2026
- LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulationsarxiv-2603.02128 Sparse Blocked context onlyMar 2, 2026
- Recursive Think-Answer Process for LLMs and VLMsarxiv-2603.02099 Sparse Blocked context onlyMar 2, 2026
- ClinConsensus: A Consensus-Based Benchmark for Evaluating Chinese Medical LLMs across Difficulty Levelsarxiv-2603.02097 Sparse Blocked context onlyMar 2, 2026
- Learning from Synthetic Data Improves Multi-hop Reasoningarxiv-2603.02091 Sparse Blocked context onlyMar 2, 2026
- GenDB: The Next Generation of Query Processing -- Synthesized, Not Engineeredarxiv-2603.02081 Sparse Blocked context onlyMar 2, 2026
- Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planningarxiv-2603.02070 Sparse Blocked context onlyMar 2, 2026
- EstLLM: Enhancing Estonian Capabilities in Multilingual LLMs via Continued Pretraining and Post-Trainingarxiv-2603.02041 Sparse Blocked context onlyMar 2, 2026
- Learning to Read Where to Look: Disease-Aware Vision-Language Pretraining for 3D CTarxiv-2603.02026 Sparse Blocked context onlyMar 2, 2026
- MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoningarxiv-2603.02024 Sparse Blocked context onlyMar 2, 2026
- AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversationsarxiv-2603.01966 Sparse Blocked context onlyMar 2, 2026
- Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experimentarxiv-2603.01950 Sparse Blocked context onlyMar 2, 2026
- Efficient RLVR Training via Weighted Mutual Information Data Selectionarxiv-2603.01907 Sparse Blocked context onlyMar 2, 2026
- KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Modelsarxiv-2603.01875 Sparse Blocked context onlyMar 2, 2026
- Let the Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answeringarxiv-2603.01853 Sparse Blocked context onlyMar 2, 2026
- ALTER: Asymmetric LoRA for Token-Entropy-Guided Unlearning of LLMsarxiv-2603.01792 Sparse Blocked context onlyMar 2, 2026
- FreeAct: Freeing Activations for LLM Quantizationarxiv-2603.01776 Sparse Blocked context onlyMar 2, 2026
- TopoCurate:Modeling Interaction Topology for Tool-Use Agent Trainingarxiv-2603.01714 Sparse Blocked context onlyMar 2, 2026
- Surgical Post-Training: Cutting Errors, Keeping Knowledgearxiv-2603.01683 Sparse Blocked context onlyMar 2, 2026
- Graph-of-Mark: Promote Spatial Reasoning in Multimodal Language Models with Graph-Based Visual Promptingarxiv-2603.06663 Sparse Blocked context onlyMar 2, 2026
- LexChronos: An Agentic Framework for Structured Event Timeline Extraction in Indian Jurisprudencearxiv-2603.01651 Sparse Blocked context onlyMar 2, 2026
- Learning to Draft: Adaptive Speculative Decoding with Reinforcement Learningarxiv-2603.01639 Sparse Blocked context onlyMar 2, 2026
- ToolRLA: Multiplicative Reward Decomposition for Tool-Integrated Agentsarxiv-2603.01620 Sparse Blocked context onlyMar 2, 2026
- Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMsarxiv-2603.01502 Sparse Blocked context onlyMar 2, 2026
- ProtRLSearch: A Multi-Round Multimodal Protein Search Agent with Large Language Models Trained via Reinforcement Learningarxiv-2603.01464 Sparse Blocked context onlyMar 2, 2026
- From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agentsarxiv-2603.01455 Sparse Blocked context onlyMar 2, 2026
- Enhancing Persona Following at Decoding Time via Dynamic Importance Estimation for Role-Playing Agentsarxiv-2603.01438 Sparse Blocked context onlyMar 2, 2026
- Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamicsarxiv-2603.01426 Sparse Blocked context onlyMar 2, 2026
- LaSER: Internalizing Explicit Reasoning into Latent Space for Dense Retrievalarxiv-2603.01425 Curated Related Blocked context onlyMar 2, 2026
- Toward Graph-Tokenizing Large Language Models with Reconstructive Graph Instruction Tuningarxiv-2603.01385 Sparse Blocked context onlyMar 2, 2026
- PanCanBench: A Comprehensive Benchmark for Evaluating Large Language Models in Pancreatic Oncologyarxiv-2603.01343 Sparse Blocked context onlyMar 2, 2026
- MetaState: Persistent Working Memory for Discrete Diffusion Language Modelsarxiv-2603.01331 Sparse Blocked context onlyMar 2, 2026
- Learn Hard Problems During RL with Reference Guided Fine-tuningarxiv-2603.01223 Sparse Blocked context onlyMar 1, 2026
- Reasoning Boosts Opinion Alignment in LLMsarxiv-2603.01214 Sparse Blocked context onlyMar 1, 2026
- GroupGPT: A Token-efficient and Privacy-preserving Agentic Framework for Multi-User Chat Assistantarxiv-2603.01059 Sparse Blocked context onlyMar 1, 2026
- MC-Search: Evaluating and Enhancing Multimodal Agentic Search with Structured Long Reasoning Chainsarxiv-2603.00873 Sparse Blocked context onlyMar 1, 2026
- PPC-MT: Parallel Point Cloud Completion with Mamba-Transformer Hybrid Architecturearxiv-2603.00870 Sparse Blocked context onlyMar 1, 2026
- MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicinearxiv-2603.00842 Sparse Blocked context onlyMar 1, 2026
- Constitutional Black-Box Monitoring for Scheming in LLM Agentsarxiv-2603.00829 Sparse Blocked context onlyFeb 28, 2026
- Qwen3-Coder-Next Technical Reportarxiv-2603.00729 Sparse Blocked context onlyFeb 28, 2026
- RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Modelsarxiv-2603.00724 Sparse Blocked context onlyFeb 28, 2026
- RAVEL: Reasoning Agents for Validating and Evaluating LLM Text Synthesisarxiv-2603.00686 Sparse Blocked context onlyFeb 28, 2026
- BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languagesarxiv-2603.00634 Sparse Blocked context onlyFeb 28, 2026
- TraceSIR: A Multi-Agent Framework for Structured Analysis and Reporting of Agentic Execution Tracesarxiv-2603.00623 Sparse Blocked context onlyFeb 28, 2026
- From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generationarxiv-2603.00612 Sparse Blocked context onlyFeb 28, 2026
- LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Modelsarxiv-2603.00592 Sparse Blocked context onlyFeb 28, 2026
- Super Research: Answering Highly Complex Questions with Large Language Models through Super Deep and Super Wide Researcharxiv-2603.00582 Sparse Blocked context onlyFeb 28, 2026
- Draft-Thinking: Learning Efficient Reasoning in Long Chain-of-Thought LLMsarxiv-2603.00578 Curated Related Blocked context onlyFeb 28, 2026
- CoMoL: Efficient Mixture of LoRA Experts via Dynamic Core Space Mergingarxiv-2603.00573 Sparse Blocked context onlyFeb 28, 2026
- RTLocating: Intent-aware RTL Localization for Hardware Design Iterationarxiv-2603.00434 Sparse Blocked context onlyFeb 28, 2026
- LLM-Bootstrapped Targeted Finding Guidance for Factual MLLM-based Medical Report Generationarxiv-2603.00426 Sparse Blocked context onlyFeb 28, 2026
- Distribution-Aware Companding Quantization of Large Language Modelsarxiv-2603.00364 Sparse Blocked context onlyFeb 27, 2026
- Transformers Remember First, Forget Last: Dual-Process Interference in LLMsarxiv-2603.00270 Sparse Blocked context onlyFeb 27, 2026
- Do LLMs Benefit From Their Own Words?arxiv-2602.24287 Sparse Blocked context onlyFeb 27, 2026
- Controllable Reasoning Models Are Private Thinkersarxiv-2602.24210 Sparse Blocked context onlyFeb 27, 2026
- Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volumearxiv-2602.24195 Sparse Blocked context onlyFeb 27, 2026
- Task-Centric Acceleration of Small-Language Modelsarxiv-2602.24174 Sparse Blocked context onlyFeb 27, 2026
- ArgLLM-App: An Interactive System for Argumentative Reasoning with Large Language Modelsarxiv-2602.24172 Sparse Blocked context onlyFeb 27, 2026
- Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verificationarxiv-2602.24111 Sparse Blocked context onlyFeb 27, 2026
- Recycling Failures: Salvaging Exploration in RLVR via Fine-Grained Off-Policy Guidancearxiv-2602.24110 Sparse Blocked context onlyFeb 27, 2026
- Preference Packing: Efficient Preference Optimization for Large Language Modelsarxiv-2602.24082 Sparse Blocked context onlyFeb 27, 2026
- A Novel Hierarchical Multi-Agent System for Payments Using LLMsarxiv-2602.24068 Sparse Blocked context onlyFeb 27, 2026
- Task Complexity Matters: An Empirical Study of Reasoning in LLMs for Sentiment Analysisarxiv-2602.24060 Sparse Blocked context onlyFeb 27, 2026
- LK Losses: Direct Acceptance Rate Optimization for Speculative Decodingarxiv-2602.23881 Sparse Blocked context onlyFeb 27, 2026
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scalearxiv-2602.23866 Sparse Blocked context onlyFeb 27, 2026
- UTPTrack: Towards Simple and Unified Token Pruning for Visual Trackingarxiv-2602.23734 Sparse Blocked context onlyFeb 27, 2026
- HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exitarxiv-2602.23699 Sparse Blocked context onlyFeb 27, 2026
- TRIZ-RAGNER: A Retrieval-Augmented Large Language Model for TRIZ-Aware Named Entity Recognition in Patent-Based Contradiction Miningarxiv-2602.23656 Sparse Blocked context onlyFeb 27, 2026
- Multi-Agent Causal Reasoning for Suicide Ideation Detection Through Online Conversationsarxiv-2602.23577 Sparse Blocked context onlyFeb 27, 2026
- IDP Accelerator: Agentic Document Intelligence from Extraction to Compliance Validationarxiv-2602.23481 Sparse Blocked context onlyFeb 26, 2026
- FHIRPath-QA: Executable Question Answering over FHIR Electronic Health Recordsarxiv-2602.23479 Sparse Blocked context onlyFeb 26, 2026
- CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Eraarxiv-2602.23452 Sparse Blocked context onlyFeb 26, 2026
- Truncated Step-Level Sampling with Process Rewards for Retrieval-Augmented Reasoningarxiv-2602.23440 Sparse Blocked context onlyFeb 26, 2026
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transportarxiv-2602.23353 Sparse Blocked context onlyFeb 26, 2026
- Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoningarxiv-2602.23351 Sparse Blocked context onlyFeb 26, 2026
- Toward Expert Investment Teams:A Multi-Agent LLM System with Fine-Grained Trading Tasksarxiv-2602.23330 Sparse Blocked context onlyFeb 26, 2026
- A Mixture-of-Experts Model for Multimodal Emotion Recognition in Conversationsarxiv-2602.23300 Sparse Blocked context onlyFeb 26, 2026
- CXReasonAgent: Evidence-Grounded Diagnostic Reasoning Agent for Chest X-raysarxiv-2602.23276 Sparse Blocked context onlyFeb 26, 2026
- Discourse-Aware Dual-Track Streaming Response for Low-Latency Spoken Dialogue Systemsarxiv-2602.23266 Sparse Blocked context onlyFeb 26, 2026
- AgentDropoutV2: Optimizing Information Flow in Multi-Agent Systems via Test-Time Rectify-or-Reject Pruningarxiv-2602.23258 Sparse Blocked context onlyFeb 26, 2026
- Why Diffusion Language Models Struggle with Truly Parallel (Non-Autoregressive) Decoding?arxiv-2602.23225 Sparse Blocked context onlyFeb 26, 2026
- InnerQ: Hardware-aware Tuning-free Quantization of KV Cache for Large Language Modelsarxiv-2602.23200 Sparse Blocked context onlyFeb 26, 2026
- SC-Arena: A Natural Language Benchmark for Single-Cell Reasoning with Knowledge-Augmented Evaluationarxiv-2602.23199 Sparse Blocked context onlyFeb 26, 2026
- ESAA: Event Sourcing for Autonomous Agents in LLM-Based Software Engineeringarxiv-2602.23193 Sparse Blocked context onlyFeb 26, 2026
- Multi-Agent Large Language Model Based Emotional Detoxification Through Personalized Intensity Control for Consumer Protectionarxiv-2602.23123 Sparse Blocked context onlyFeb 26, 2026
- Automated Vulnerability Detection in Source Code Using Deep Representation Learningarxiv-2602.23121 Sparse Blocked context onlyFeb 26, 2026
- Assessing Deanonymization Risks with Stylometry-Assisted LLM Agentarxiv-2602.23079 Sparse Blocked context onlyFeb 26, 2026
- CiteLLM: An Agentic Platform for Trustworthy Scientific Reference Discoveryarxiv-2602.23075 Sparse Blocked context onlyFeb 26, 2026
- Residual Koopman Spectral Profiling for Predicting and Preventing Transformer Training Instabilityarxiv-2602.22988 Sparse Blocked context onlyFeb 26, 2026
- SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopyarxiv-2602.22971 Sparse Blocked context onlyFeb 26, 2026
- MM-NeuroOnco: A Multimodal Benchmark and Instruction Dataset for MRI-Based Brain Tumor Diagnosisarxiv-2602.22955 Sparse Blocked context onlyFeb 26, 2026
- pMoE: Prompting Diverse Experts Together Wins More in Visual Adaptationarxiv-2602.22938 Sparse Blocked context onlyFeb 26, 2026
- OmniGAIA: Towards Native Omni-Modal AI Agentsarxiv-2602.22897 Sparse Blocked context onlyFeb 26, 2026
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasksarxiv-2602.22817 Sparse Blocked context onlyFeb 26, 2026
- MiroFlow: Towards High-Performance and Robust Open-Source Agent Framework for General Deep Research Tasksarxiv-2602.22808 Direct Blocked context onlyFeb 26, 2026
- Towards Better RL Training Data Utilization via Second-Order Rolloutarxiv-2602.22765 Sparse Blocked context onlyFeb 26, 2026
- Know What You Know: Metacognitive Entropy Calibration for Verifiable RL Reasoningarxiv-2602.22751 Sparse Blocked context onlyFeb 26, 2026
- AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentationarxiv-2602.22740 Sparse Blocked context onlyFeb 26, 2026
- AgentSentry: Mitigating Indirect Prompt Injection in LLM Agents via Temporal Causal Diagnostics and Context Purificationarxiv-2602.22724 Sparse Blocked context onlyFeb 26, 2026
- RLHFless: Serverless Computing for Efficient RLHFarxiv-2602.22718 Sparse Blocked context onlyFeb 26, 2026
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsarxiv-2602.22716 Sparse Blocked context onlyFeb 26, 2026
- IMMACULATE: A Practical LLM Auditing Framework via Verifiable Computationarxiv-2602.22700 Sparse Blocked context onlyFeb 26, 2026
- Reinforcing Real-world Service Agents: Balancing Utility and Cost in Task-oriented Dialoguearxiv-2602.22697 Sparse Blocked context onlyFeb 26, 2026
- SUPERGLASSES: Benchmarking Vision Language Models as Intelligent Agents for AI Smart Glassesarxiv-2602.22683 Sparse Blocked context onlyFeb 26, 2026
- Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Acceleratorsarxiv-2602.22647 Direct Blocked context onlyFeb 26, 2026
- Search-P1: Path-Centric Reward Shaping for Stable and Efficient Agentic RAG Trainingarxiv-2602.22576 Sparse Blocked context onlyFeb 26, 2026
- Stable Adaptive Thinking via Advantage Shaping and Length-Aware Gradient Regulationarxiv-2602.22556 Sparse Blocked context onlyFeb 26, 2026
- RAIN-Merging: A Gradient-Free Method to Enhance Instruction Following in Large Reasoning Models with Preserved Thinking Formatarxiv-2602.22538 Sparse Blocked context onlyFeb 26, 2026
- Mind the Gap in Cultural Alignment: Task-Aware Culture Management for Large Language Modelsarxiv-2602.22475 Sparse Blocked context onlyFeb 25, 2026
- Bridging Latent Reasoning and Target-Language Generation via Retrieval-Transition Headsarxiv-2602.22453 Sparse Blocked context onlyFeb 25, 2026
- Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Promptsarxiv-2602.22359 Sparse Blocked context onlyFeb 25, 2026
- Decoder-based Sense Knowledge Distillationarxiv-2602.22351 Sparse Blocked context onlyFeb 25, 2026
- Improving Parametric Knowledge Access in Reasoning Language Modelsarxiv-2602.22193 Sparse Blocked context onlyFeb 25, 2026
- GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RLarxiv-2602.22190 Sparse Blocked context onlyFeb 25, 2026
- Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Adsarxiv-2602.22299 Sparse Blocked context onlyFeb 25, 2026
- DySCO: Dynamic Attention-Scaling Decoding for Long-Context LMsarxiv-2602.22175 Sparse Blocked context onlyFeb 25, 2026
- Dynamic Personality Adaptation in Large Language Models via State Machinesarxiv-2602.22157 Sparse Blocked context onlyFeb 25, 2026
- IndicIFEval: A Benchmark for Verifiable Instruction-Following Evaluation in 14 Indic Languagesarxiv-2602.22125 Sparse Blocked context onlyFeb 25, 2026
- SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agentsarxiv-2602.22124 Sparse Blocked context onlyFeb 25, 2026
- Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inferencearxiv-2602.22090 Sparse Blocked context onlyFeb 25, 2026
- Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Modelsarxiv-2602.22072 Sparse Blocked context onlyFeb 25, 2026
- RADAR: Reasoning as Discrimination with Aligned Representations for LLM-based Knowledge Graph Reasoningarxiv-2602.21951 Sparse Blocked context onlyFeb 25, 2026
- MEDSYN: Benchmarking Multi-EviDence SYNthesis in Complex Clinical Cases for Multimodal Large Language Modelsarxiv-2602.21950 Sparse Blocked context onlyFeb 25, 2026
- Large Language Models are Algorithmically Blindarxiv-2602.21947 Sparse Blocked context onlyFeb 25, 2026
- FinReasoning: A Hierarchical Benchmark for Reliable Financial Research Reportingarxiv-2603.19254 Sparse Blocked context onlyFeb 25, 2026
- ExpLang: Improved Exploration and Exploitation in LLM Reasoning with On-Policy Thinking Language Selectionarxiv-2602.21887 Sparse Blocked context onlyFeb 25, 2026
- DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAsarxiv-2602.21864 Sparse Blocked context onlyFeb 25, 2026
- ProactiveMobile: A Comprehensive Benchmark for Boosting Proactive Intelligence on Mobile Devicesarxiv-2602.21858 Sparse Blocked context onlyFeb 25, 2026
- Distill and Align Decomposition for Enhanced Claim Verificationarxiv-2602.21857 Sparse Blocked context onlyFeb 25, 2026
- FewMMBench: A Benchmark for Multimodal Few-Shot Learningarxiv-2602.21854 Direct Blocked context onlyFeb 25, 2026
- DocDjinn: Controllable Synthetic Document Generation with VLMs and Handwriting Diffusionarxiv-2602.21824 Sparse Blocked context onlyFeb 25, 2026
- Prompt Architecture Determines Reasoning Quality: A Variable Isolation Study on the Car Wash Problemarxiv-2602.21814 Sparse Blocked context onlyFeb 25, 2026
- D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Modelsarxiv-2602.21786 Curated Related Blocked context onlyFeb 25, 2026
- Explore-on-Graph: Incentivizing Autonomous Exploration of Large Language Models on Knowledge Graphs with Path-refined Reward Modelingarxiv-2602.21728 Sparse Blocked context onlyFeb 25, 2026
- Two-Stage Active Distribution Network Voltage Control via LLM-RL Collaboration: A Hybrid Knowledge-Data-Driven Approacharxiv-2602.21715 Sparse Blocked context onlyFeb 25, 2026
- Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Modelsarxiv-2602.21704 Sparse Blocked context onlyFeb 25, 2026
- Hierarchical LLM-Based Multi-Agent Framework with Prompt Optimization for Multi-Robot Task Planningarxiv-2602.21670 Sparse Blocked context onlyFeb 25, 2026
- DWA-KD: Dual-Space Weighting and Time-Warped Alignment for Cross-Tokenizer Knowledge Distillationarxiv-2602.21669 Sparse Blocked context onlyFeb 25, 2026
- Sparsity Induction for Accurate Post-Training Pruning of Large Language Modelsarxiv-2602.21652 Sparse Blocked context onlyFeb 25, 2026
- Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusionarxiv-2602.21646 Sparse Blocked context onlyFeb 25, 2026
- RuCL: Stratified Rubric-Based Curriculum Learning for Multimodal Large Language Model Reasoningarxiv-2602.21628 Sparse Blocked context onlyFeb 25, 2026
- When More Is Less: A Systematic Analysis of Spatial and Commonsense Information for Visual Spatial Reasoningarxiv-2602.21619 Sparse Blocked context onlyFeb 25, 2026
- Virtual Biopsy for Intracranial Tumors Diagnosis on MRIarxiv-2602.21613 Sparse Blocked context onlyFeb 25, 2026
- Structurally Aligned Subtask-Level Memory for Software Engineering Agentsarxiv-2602.21611 Sparse Blocked context onlyFeb 25, 2026
- ARLArena: A Unified Framework for Stable Agentic Reinforcement Learningarxiv-2602.21534 Sparse Blocked context onlyFeb 25, 2026
- One Brain, Omni Modalities: Towards Unified Non-Invasive Brain Decoding with Large Language Modelsarxiv-2602.21522 Sparse Blocked context onlyFeb 25, 2026
- Beyond Refusal: Probing the Limits of Agentic Self-Correction for Semantic Sensitive Informationarxiv-2602.21496 Sparse Blocked context onlyFeb 25, 2026
- Both Ends Count! Just How Good are LLM Agents at "Text-to-Big SQL"?arxiv-2602.21480 Sparse Blocked context onlyFeb 25, 2026
- VecGlypher: Unified Vector Glyph Generation with Language Modelsarxiv-2602.21461 Sparse Blocked context onlyFeb 25, 2026
- Revisiting Text Ranking in Deep Researcharxiv-2602.21456 Sparse Blocked context onlyFeb 25, 2026
- MINAR: Mechanistic Interpretability for Neural Algorithmic Reasoningarxiv-2602.21442 Sparse Blocked context onlyFeb 24, 2026
- Causal Decoding for Hallucination-Resistant Multimodal Large Language Modelsarxiv-2602.21441 Sparse Blocked context onlyFeb 24, 2026
- Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learningarxiv-2602.21420 Sparse Blocked context onlyFeb 24, 2026
- Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibrationarxiv-2602.21368 Sparse Blocked context onlyFeb 24, 2026
- A Hierarchical Multi-Agent System for Autonomous Discovery in Geoscientific Data Archivesarxiv-2602.21351 Sparse Blocked context onlyFeb 24, 2026
- ToolMATH: A Math Tool Benchmark for Realistic Long-Horizon Multi-Tool Reasoningarxiv-2602.21265 Sparse Blocked context onlyFeb 24, 2026
- Under the Influence: Quantifying Persuasion and Vigilance in Large Language Modelsarxiv-2602.21262 Sparse Blocked context onlyFeb 24, 2026
- MERRY: Semantically Decoupled Evaluation of Multimodal Emotional and Role Consistencies of Role-Playing Agentsarxiv-2602.21941 Sparse Blocked context onlyFeb 24, 2026
- Actor-Curator: Co-adaptive Curriculum Learning via Policy-Improvement Bandits for RL Post-Trainingarxiv-2602.20532 Sparse Blocked context onlyFeb 24, 2026
- E-MMKGR: A Unified Multimodal Knowledge Graph Framework for E-commerce Applicationsarxiv-2602.20877 Sparse Blocked context onlyFeb 24, 2026
- On Data Engineering for Scaling LLM Terminal Capabilitiesarxiv-2602.21193 Direct Blocked context onlyFeb 24, 2026
- Case-Aware LLM-as-a-Judge Evaluation for Enterprise-Scale RAG Systemsarxiv-2602.20379 Sparse Blocked context onlyFeb 23, 2026
- MedCLIPSeg: Probabilistic Vision-Language Adaptation for Data-Efficient and Generalizable Medical Image Segmentationarxiv-2602.20423 Curated Related Blocked context onlyFeb 23, 2026
- MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systemsarxiv-2602.19843 Sparse Blocked context onlyFeb 23, 2026
- RegionRoute: Regional Style Transfer with Diffusion Modelarxiv-2602.19254 Curated Related Blocked context onlyFeb 22, 2026
- K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Modelarxiv-2602.19128 Curated Related Blocked context onlyFeb 22, 2026
- Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understandingarxiv-2602.18702 Sparse Blocked context onlyFeb 21, 2026
- TAG: Thinking with Action Unit Grounding for Facial Expression Recognitionarxiv-2602.18763 Sparse Blocked context onlyFeb 21, 2026
- Decoding as Optimisation on the Probability Simplex: From Top-K to Top-P (Nucleus) to Best-of-K Samplersarxiv-2602.18292 Sparse Blocked context onlyFeb 20, 2026
- MIRA: Memory-Integrated Reinforcement Learning Agent with Limited LLM Guidancearxiv-2602.17930 Sparse Blocked context onlyFeb 20, 2026
- Stable Asynchrony: Variance-Controlled Off-Policy RL for LLMsarxiv-2602.17616 Sparse Blocked context onlyFeb 19, 2026
- Fine-Grained Uncertainty Quantification for Long-Form Language Model Outputs: A Comparative Studyarxiv-2602.17431 Sparse Blocked context onlyFeb 19, 2026
- Arcee Trinity Large Technical Reportarxiv-2602.17004 Curated Related Blocked context onlyFeb 19, 2026
- Sketch2Feedback: Grammar-in-the-Loop Framework for Rubric-Aligned Feedback on Student STEM Diagramsarxiv-2602.18520 Sparse Blocked context onlyFeb 19, 2026
- MeGU: Machine-Guided Unlearning with Target Feature Disentanglementarxiv-2602.17088 Sparse Blocked context onlyFeb 19, 2026
- Trojan Horses in Recruiting: A Red-Teaming Case Study on Indirect Prompt Injection in Standard vs. Reasoning Modelsarxiv-2602.18514 Sparse Blocked context onlyFeb 19, 2026
- GeneZip: Region-Aware Compression for Long Context DNA Modelingarxiv-2602.17739 Curated Related Blocked context onlyFeb 19, 2026
- Retrospective In-Context Learning for Temporal Credit Assignment with Large Language Modelsarxiv-2602.17497 Sparse Blocked context onlyFeb 19, 2026
- AdaptOrch: Task-Adaptive Multi-Agent Orchestration in the Era of LLM Performance Convergencearxiv-2602.16873 Sparse Blocked context onlyFeb 18, 2026
- Toward Scalable Verifiable Reward: Proxy State-Based Evaluation for Multi-turn Tool-Calling LLM Agentsarxiv-2602.16246 Sparse Blocked context onlyFeb 18, 2026
- DeepVision-103K: A Visually Diverse, Broad-Coverage, and Verifiable Mathematical Dataset for Multimodal Reasoningarxiv-2602.16742 Direct Blocked context onlyFeb 18, 2026
- Team of Thoughts: Efficient Test-time Scaling of Agentic Systems through Orchestrated Tool Callingarxiv-2602.16485 Sparse Blocked context onlyFeb 18, 2026
- One-step Language Modeling via Continuous Denoisingarxiv-2602.16813 Sparse Blocked context onlyFeb 18, 2026
- Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agentsarxiv-2602.16943 Sparse Blocked context onlyFeb 18, 2026
- VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detectionarxiv-2602.16681 Sparse Blocked context onlyFeb 18, 2026
- Updating Parametric Knowledge with Context Distillation Retains Post-Training Capabilitiesarxiv-2602.16093 Sparse Blocked context onlyFeb 17, 2026
- Evidence-Grounded Subspecialty Reasoning: Evaluating a Curated Clinical Intelligence Layer on the 2025 Endocrinology Board-Style Examinationarxiv-2602.16050 Sparse Blocked context onlyFeb 17, 2026
- MAEB: Massive Audio Embedding Benchmarkarxiv-2602.16008 Sparse Blocked context onlyFeb 17, 2026
- DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splittingarxiv-2602.15958 Sparse Blocked context onlyFeb 17, 2026