300 canonical paper links on this archive page.
- DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splittingarxiv-2602.15958 Sparse Blocked context onlyFeb 17, 2026
- Avey-Barxiv-2602.15814 Sparse Blocked context onlyFeb 17, 2026
- ViTaB-A: Evaluating Multimodal Large Language Models on Visual Table Attributionarxiv-2602.15769 Sparse Blocked context onlyFeb 17, 2026
- ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Modelsarxiv-2602.15758 Sparse Blocked context onlyFeb 17, 2026
- Recursive Concept Evolution for Compositional Reasoning in Large Language Modelsarxiv-2602.15725 Sparse Blocked context onlyFeb 17, 2026
- A Content-Based Framework for Cybersecurity Refusal Decisions in Large Language Modelsarxiv-2602.15689 Sparse Blocked context onlyFeb 17, 2026
- STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokensarxiv-2602.15620 Sparse Blocked context onlyFeb 17, 2026
- In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generationsarxiv-2602.15456 Sparse Blocked context onlyFeb 17, 2026
- TAROT: Test-driven and Capability-adaptive Curriculum Reinforcement Fine-tuning for Code Generation with Large Language Modelsarxiv-2602.15449 Sparse Blocked context onlyFeb 17, 2026
- SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognitionarxiv-2603.10005 Sparse Blocked context onlyFeb 17, 2026
- World-Model-Augmented Web Agents with Action Correctionarxiv-2602.15384 Sparse Blocked context onlyFeb 17, 2026
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systemsarxiv-2602.15382 Sparse Blocked context onlyFeb 17, 2026
- Orchestration-Free Customer Service Automation: A Privacy-Preserving and Flowchart-Guided Frameworkarxiv-2602.15377 Sparse Blocked context onlyFeb 17, 2026
- Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memoryarxiv-2602.15313 Sparse Blocked context onlyFeb 17, 2026
- OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interactionarxiv-2602.15197 Sparse Blocked context onlyFeb 16, 2026
- Weight space Detection of Backdoors in LoRA Adaptersarxiv-2602.15195 Sparse Blocked context onlyFeb 16, 2026
- AIC CTU@AVerImaTeC: dual-retriever RAG for image-text fact checkingarxiv-2602.15190 Sparse Blocked context onlyFeb 16, 2026
- Seeing to Generalize: How Visual Data Corrects Binding Shortcutsarxiv-2602.15183 Sparse Blocked context onlyFeb 16, 2026
- Protecting Language Models Against Unauthorized Distillation through Trace Rewritingarxiv-2602.15143 Sparse Blocked context onlyFeb 16, 2026
- Learning User Interests via Reasoning and Distillation for Cross-Domain News Recommendationarxiv-2602.15005 Sparse Blocked context onlyFeb 16, 2026
- Counterfactual Fairness Evaluation of LLM-Based Contact Center Agent Quality Assurance Systemarxiv-2602.14970 Sparse Blocked context onlyFeb 16, 2026
- BFS-PO: Best-First Search for Large Reasoning Modelsarxiv-2602.14917 Sparse Blocked context onlyFeb 16, 2026
- Physical Commonsense Reasoning for Lower-Resourced Languages and Dialects: a Study on Basquearxiv-2602.14812 Sparse Blocked context onlyFeb 16, 2026
- Overthinking Loops in Agents: A Structural Risk via MCP Toolsarxiv-2602.14798 Sparse Blocked context onlyFeb 16, 2026
- Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignmentarxiv-2602.14777 Sparse Blocked context onlyFeb 16, 2026
- Unlocking Reasoning Capability on Machine Translation in Large Language Modelsarxiv-2602.14763 Sparse Blocked context onlyFeb 16, 2026
- Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformersarxiv-2602.14760 Sparse Blocked context onlyFeb 16, 2026
- Rethinking the Role of LLMs in Time Series Forecastingarxiv-2602.14744 Sparse Blocked context onlyFeb 16, 2026
- Evolutionary System Prompt Learning for Reinforcement Learning in LLMsarxiv-2602.14697 Sparse Blocked context onlyFeb 16, 2026
- Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacksarxiv-2602.14689 Sparse Blocked context onlyFeb 16, 2026
- Crowdsourcing Piedmontese to Test LLMs on Non-Standard Orthographyarxiv-2602.14675 Sparse Blocked context onlyFeb 16, 2026
- GradMAP: Faster Layer Pruning with Gradient Metric and Projection Compensationarxiv-2602.14649 Sparse Blocked context onlyFeb 16, 2026
- Beyond the Prompt in Large Language Models: Comprehension, In-Context Learning, and Chain-of-Thoughtarxiv-2603.10000 Sparse Blocked context onlyFeb 16, 2026
- Alignment Adapter to Improve the Performance of Compressed Deep Learning Modelsarxiv-2602.14635 Sparse Blocked context onlyFeb 16, 2026
- The Wikidata Query Logs Datasetarxiv-2602.14594 Sparse Blocked context onlyFeb 16, 2026
- MATEO: A Multimodal Benchmark for Temporal Reasoning and Planning in LVLMsarxiv-2602.14589 Sparse Blocked context onlyFeb 16, 2026
- Explainable Token-level Noise Filtering for LLM Fine-tuning Datasetsarxiv-2602.14536 Sparse Blocked context onlyFeb 16, 2026
- Beyond Translation: Evaluating Mathematical Reasoning Capabilities of LLMs in Sinhala and Tamilarxiv-2602.14517 Sparse Blocked context onlyFeb 16, 2026
- Query as Anchor: Scenario-Adaptive User Representation via Large Language Modelarxiv-2602.14492 Sparse Blocked context onlyFeb 16, 2026
- BETA-Labeling for Multilingual Dataset Construction in Low-Resource IRarxiv-2602.14488 Sparse Blocked context onlyFeb 16, 2026
- TikArt: Stabilizing Aperture-Guided Fine-Grained Visual Reasoning with Reinforcement Learningarxiv-2602.14482 Sparse Blocked context onlyFeb 16, 2026
- Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report v1.5arxiv-2602.14457 Sparse Blocked context onlyFeb 16, 2026
- Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learningarxiv-2602.14451 Sparse Blocked context onlyFeb 16, 2026
- LLM-Guided Knowledge Distillation for Temporal Knowledge Graph Reasoningarxiv-2602.14428 Sparse Blocked context onlyFeb 16, 2026
- WavePhaseNet: A DFT-Based Method for Constructing Semantic Conceptual Hierarchy Structures (SCHS)arxiv-2602.14419 Sparse Blocked context onlyFeb 16, 2026
- Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detectionarxiv-2602.14408 Sparse Blocked context onlyFeb 16, 2026
- Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Modelsarxiv-2602.14386 Sparse Blocked context onlyFeb 16, 2026
- InnoEval: On Research Idea Evaluation as a Knowledge-Grounded, Multi-Perspective Reasoning Problemarxiv-2602.14367 Sparse Blocked context onlyFeb 16, 2026
- MCPShield: A Security Cognition Layer for Adaptive Trust Calibration in Model Context Protocol Agentsarxiv-2602.14281 Sparse Blocked context onlyFeb 15, 2026
- Whom to Query for What: Adaptive Group Elicitation via Multi-Turn LLM Interactionsarxiv-2602.14279 Sparse Blocked context onlyFeb 15, 2026
- STATe-of-Thoughts: Structured Action Templates for Tree-of-Thoughtsarxiv-2602.14265 Direct Blocked context onlyFeb 15, 2026
- Detecting LLM Hallucinations via Embedding Cluster Geometry: A Three-Type Taxonomy with Measurable Signaturesarxiv-2602.14259 Sparse Blocked context onlyFeb 15, 2026
- REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agentsarxiv-2602.14234 Sparse Blocked context onlyFeb 15, 2026
- The Interspeech 2026 Audio Reasoning Challenge: Evaluating Reasoning Process Quality for Audio Reasoning Models and Agentsarxiv-2602.14224 Sparse Blocked context onlyFeb 15, 2026
- Knowing When Not to Answer: Abstention-Aware Scientific Reasoningarxiv-2602.14189 Sparse Blocked context onlyFeb 15, 2026
- UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Modelarxiv-2602.14178 Sparse Blocked context onlyFeb 15, 2026
- Investigation for Relative Voice Impression Estimationarxiv-2602.14172 Sparse Blocked context onlyFeb 15, 2026
- Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answeringarxiv-2602.14162 Sparse Blocked context onlyFeb 15, 2026
- A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processingarxiv-2602.14158 Sparse Blocked context onlyFeb 15, 2026
- Algebraic Quantum Intelligence: A New Framework for Reproducible Machine Creativityarxiv-2602.14130 Sparse Blocked context onlyFeb 15, 2026
- CCiV: A Benchmark for Structure, Rhythm and Quality in LLM-Generated Chinese \textit{Ci} Poetryarxiv-2602.14081 Sparse Blocked context onlyFeb 15, 2026
- Annotation-Efficient Vision-Language Model Adaptation to the Polish Language Using the LLaVA Frameworkarxiv-2602.14073 Sparse Blocked context onlyFeb 15, 2026
- Open Rubric System: Scaling Reinforcement Learning with Pairwise Adaptive Rubricarxiv-2602.14069 Sparse Blocked context onlyFeb 15, 2026
- Context Shapes LLMs Retrieval-Augmented Fact-Checking Effectivenessarxiv-2602.14044 Sparse Blocked context onlyFeb 15, 2026
- GRRM: Group Relative Reward Modeling for Machine Translationarxiv-2602.14028 Sparse Blocked context onlyFeb 15, 2026
- The Sufficiency-Conciseness Trade-off in LLM Self-Explanation from an Information Bottleneck Perspectivearxiv-2602.14002 Sparse Blocked context onlyFeb 15, 2026
- Chain-of-Thought Reasoning with Large Language Models for Clinical Alzheimer's Disease Assessment and Diagnosisarxiv-2602.13979 Sparse Blocked context onlyFeb 15, 2026
- MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Marsarxiv-2602.13961 Sparse Blocked context onlyFeb 15, 2026
- From Pixels to Policies: Reinforcing Spatial Reasoning in Language Models for Content-Aware Layout Designarxiv-2602.13912 Sparse Blocked context onlyFeb 14, 2026
- Evaluating Prompt Engineering Techniques for RAG in Small Language Models: A Multi-Hop QA Approacharxiv-2602.13890 Sparse Blocked context onlyFeb 14, 2026
- Bridging the Multilingual Safety Divide: Efficient, Culturally-Aware Alignment for Global South Languagesarxiv-2602.13867 Sparse Blocked context onlyFeb 14, 2026
- Tutoring Large Language Models to be Domain-adaptive, Precise, and Safearxiv-2602.13860 Sparse Blocked context onlyFeb 14, 2026
- Do Mixed-Vendor Multi-Agent LLMs Improve Clinical Diagnosis?arxiv-2603.04421 Sparse Blocked context onlyFeb 14, 2026
- Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mindarxiv-2602.13832 Sparse Blocked context onlyFeb 14, 2026
- Elo-Evolve: A Co-evolutionary Framework for Language Model Alignmentarxiv-2602.13575 Sparse Blocked context onlyFeb 14, 2026
- Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learningarxiv-2602.13562 Sparse Blocked context onlyFeb 14, 2026
- CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modelingarxiv-2602.13191 Sparse Blocked context onlyFeb 13, 2026
- BrowseComp-$V^3$: A Visual, Vertical, and Verifiable Benchmark for Multimodal Browsing Agentsarxiv-2602.12876 Sparse Blocked context onlyFeb 13, 2026
- MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMsarxiv-2602.12705 Sparse Blocked context onlyFeb 13, 2026
- To Mix or To Merge: Toward Multi-Domain Reinforcement Learning for Large Language Modelsarxiv-2602.12566 Sparse Blocked context onlyFeb 13, 2026
- Sparse Autoencoders are Capable LLM Jailbreak Mitigatorsarxiv-2602.12418 Sparse Blocked context onlyFeb 12, 2026
- propella-1: Multi-Property Document Annotation for LLM Data Curation at Scalearxiv-2602.12414 Sparse Blocked context onlyFeb 12, 2026
- On-Policy Context Distillation for Language Modelsarxiv-2602.12275 Sparse Blocked context onlyFeb 12, 2026
- Think like a Scientist: Physics-guided LLM Agent for Equation Discoveryarxiv-2602.12259 Sparse Blocked context onlyFeb 12, 2026
- Query-focused and Memory-aware Reranker for Long Context Processingarxiv-2602.12192 Sparse Blocked context onlyFeb 12, 2026
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolationarxiv-2602.12125 Direct Blocked context onlyFeb 12, 2026
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penaltyarxiv-2602.12113 Sparse Blocked context onlyFeb 12, 2026
- Tiny Recursive Reasoning with Mamba-2 Attention Hybridarxiv-2602.12078 Sparse Blocked context onlyFeb 12, 2026
- AdaptEvolve: Improving Efficiency of Evolutionary AI Agents through Adaptive Model Selectionarxiv-2602.11931 Sparse Blocked context onlyFeb 12, 2026
- Zooming without Zooming: Region-to-Image Distillation for Fine-Grained Multimodal Perceptionarxiv-2602.11858 Sparse Blocked context onlyFeb 12, 2026
- Predicting LLM Output Length via Entropy-Guided Representationsarxiv-2602.11812 Sparse Blocked context onlyFeb 12, 2026
- TSR: Trajectory-Search Rollouts for Multi-Turn RL of LLM Agentsarxiv-2602.11767 Sparse Blocked context onlyFeb 12, 2026
- MiniCPM-SALA: Hybridizing Sparse and Linear Attention for Efficient Long-Context Modelingarxiv-2602.11761 Sparse Blocked context onlyFeb 12, 2026
- Native Reasoning Models: Training Language Models to Reason on Unverifiable Dataarxiv-2602.11549 Sparse Blocked context onlyFeb 12, 2026
- Multimodal Fact-Level Attribution for Verifiable Reasoningarxiv-2602.11509 Sparse Blocked context onlyFeb 12, 2026
- When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitrationarxiv-2602.11488 Sparse Blocked context onlyFeb 12, 2026
- Voxtral Realtimearxiv-2602.11298 Curated Related Blocked context onlyFeb 11, 2026
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuningarxiv-2602.11149 Sparse Blocked context onlyFeb 11, 2026
- Embedding Inversion via Conditional Masked Diffusion Language Modelsarxiv-2602.11047 Sparse Blocked context onlyFeb 11, 2026
- Learning Page Order in Shuffled WOO Releasesarxiv-2602.11040 Sparse Blocked context onlyFeb 11, 2026
- LoRA-Squeeze: Simple and Effective Post-Tuning and In-Tuning Compression of LoRA Modulesarxiv-2602.10993 Sparse Blocked context onlyFeb 11, 2026
- Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language Modelsarxiv-2602.10953 Sparse Blocked context onlyFeb 11, 2026
- Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluationarxiv-2602.11224 Sparse Blocked context onlyFeb 11, 2026
- Understand Then Memory: A Cognitive Gist-Driven RAG Framework with Global Semantic Diffusionarxiv-2602.15895 Sparse Blocked context onlyFeb 11, 2026
- To Think or Not To Think, That is The Question for Large Reasoning Models in Theory of Mind Tasksarxiv-2602.10625 Sparse Blocked context onlyFeb 11, 2026
- Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parametersarxiv-2602.10604 Sparse Blocked context onlyFeb 11, 2026
- Neuro-Symbolic Synergy for Interactive World Modelingarxiv-2602.10480 Sparse Blocked context onlyFeb 11, 2026
- TestExplora: Benchmarking LLMs for Proactive Bug Discovery via Repository-Level Test Generationarxiv-2602.10471 Sparse Blocked context onlyFeb 11, 2026
- When Tables Go Crazy: Evaluating Multimodal Models on French Financial Documentsarxiv-2602.10384 Sparse Blocked context onlyFeb 11, 2026
- Versor: A Geometric Sequence Architecturearxiv-2602.10195 Sparse Blocked context onlyFeb 10, 2026
- UI-Venus-1.5 Technical Reportarxiv-2602.09082 Direct Blocked context onlyFeb 9, 2026
- Beyond Transcripts: A Renewed Perspective on Audio Chapteringarxiv-2602.08979 Sparse Blocked context onlyFeb 9, 2026
- PBLean: Pseudo-Boolean Proof Certificates for Lean 4arxiv-2602.08692 Sparse Blocked context onlyFeb 9, 2026
- ValueFlow: Measuring the Propagation of Value Perturbations in Multi-Agent LLM Systemsarxiv-2602.08567 Sparse Blocked context onlyFeb 9, 2026
- Automating Computational Reproducibility in Social Science: Comparing Prompt-Based and Agent-Based Approachesarxiv-2602.08561 Sparse Blocked context onlyFeb 9, 2026
- Breaking the Factorization Barrier in Diffusion Language Modelsarxiv-2603.00045 Sparse Blocked context onlyFeb 9, 2026
- Transforming Science Learning Materials in the Era of Artificial Intelligencearxiv-2602.18470 Sparse Blocked context onlyFeb 8, 2026
- AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineeringarxiv-2602.07906 Sparse Blocked context onlyFeb 8, 2026
- Rethinking the Value of Agent-Generated Tests for LLM-Based Software Engineering Agentsarxiv-2602.07900 Sparse Blocked context onlyFeb 8, 2026
- Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projectionarxiv-2602.07892 Sparse Blocked context onlyFeb 8, 2026
- The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decodingarxiv-2603.03305 Sparse Blocked context onlyFeb 8, 2026
- AngelSlim: A more accessible, comprehensive, and efficient toolkit for large model compressionarxiv-2602.21233 Direct Blocked context onlyFeb 7, 2026
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?arxiv-2602.18464 Sparse Blocked context onlyFeb 6, 2026
- PACIFIC: Can LLMs Discern the Traits Influencing Your Preferences? Evaluating Personality-Driven Preference Alignment in LLMsarxiv-2602.07181 Sparse Blocked context onlyFeb 6, 2026
- On Randomness in Agentic Evalsarxiv-2602.07150 Sparse Blocked context onlyFeb 6, 2026
- SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacksarxiv-2602.06854 Sparse Blocked context onlyFeb 6, 2026
- "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skillsarxiv-2602.06547 Sparse Blocked context onlyFeb 6, 2026
- LogicSkills: A Structured Benchmark for Formal Reasoning in Large Language Modelsarxiv-2602.06533 Sparse Blocked context onlyFeb 6, 2026
- From Conflict to Consensus: Boosting Medical Reasoning via Multi-Round Agentic RAGarxiv-2603.03292 Sparse Blocked context onlyFeb 6, 2026
- Stopping Computation for Converged Tokens in Masked Diffusion-LM Decodingarxiv-2602.06412 Sparse Blocked context onlyFeb 6, 2026
- RoPE-LIME: RoPE-Space Locality + Sparse-K Sampling for Efficient LLM Attributionarxiv-2602.06275 Sparse Blocked context onlyFeb 6, 2026
- Self-Improving World Modelling with Latent Actionsarxiv-2602.06130 Sparse Blocked context onlyFeb 5, 2026
- Multi-Token Prediction via Self-Distillationarxiv-2602.06019 Sparse Blocked context onlyFeb 5, 2026
- OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scalearxiv-2602.05711 Sparse Blocked context onlyFeb 5, 2026
- Rewards as Labels: Revisiting RLVR from a Classification Perspectivearxiv-2602.05630 Sparse Blocked context onlyFeb 5, 2026
- EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimizationarxiv-2602.05165 Sparse Blocked context onlyFeb 5, 2026
- GreekMMLU: A Native-Sourced Multitask Benchmark for Evaluating Language Models in Greekarxiv-2602.05150 Sparse Blocked context onlyFeb 5, 2026
- CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generationarxiv-2602.04856 Sparse Blocked context onlyFeb 4, 2026
- When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?arxiv-2602.04755 Sparse Blocked context onlyFeb 4, 2026
- WideSeek-R1: Exploring Width Scaling for Broad Information Seeking via Multi-Agent Reinforcement Learningarxiv-2602.04634 Direct Blocked context onlyFeb 4, 2026
- Contextual Drag: How Errors in the Context Affect LLM Reasoningarxiv-2602.04288 Sparse Blocked context onlyFeb 4, 2026
- Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summariesarxiv-2603.04413 Sparse Blocked context onlyFeb 3, 2026
- SpatiaLab: Can Vision-Language Models Perform Spatial Reasoning in the Wild?arxiv-2602.03916 Sparse Blocked context onlyFeb 3, 2026
- Fix the Structural Bottleneck: Context Compression via Explicit Information Transmissionarxiv-2602.03784 Sparse Blocked context onlyFeb 3, 2026
- OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answeringarxiv-2602.03707 Sparse Blocked context onlyFeb 3, 2026
- SEAD: Self-Evolving Agent for Multi-Turn Service Dialoguearxiv-2602.03548 Sparse Blocked context onlyFeb 3, 2026
- SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Trainingarxiv-2602.03411 Sparse Blocked context onlyFeb 3, 2026
- Towards Distillation-Resistant Large Language Models: An Information-Theoretic Perspectivearxiv-2602.03396 Sparse Blocked context onlyFeb 3, 2026
- POP: Prefill-Only Pruning for Efficient Large Model Inferencearxiv-2602.03295 Sparse Blocked context onlyFeb 3, 2026
- Accordion-Thinking: Self-Regulated Step Summaries for Efficient and Readable LLM Reasoningarxiv-2602.03249 Sparse Blocked context onlyFeb 3, 2026
- FASA: Frequency-aware Sparse Attentionarxiv-2602.03152 Sparse Blocked context onlyFeb 3, 2026
- ChemPro: A Progressive Chemistry Benchmark for Large Language Modelsarxiv-2602.03108 Sparse Blocked context onlyFeb 3, 2026
- LatentMem: Customizing Latent Memory for Multi-Agent Systemsarxiv-2602.03036 Sparse Blocked context onlyFeb 3, 2026
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling Modelsarxiv-2602.03022 Sparse Blocked context onlyFeb 3, 2026
- Dynamic Mixed-Precision Routing for Efficient Multi-step LLM Interactionarxiv-2602.02711 Sparse Blocked context onlyFeb 2, 2026
- Proof-RM: A Scalable and Generalizable Reward Model for Math Proofarxiv-2602.02377 Sparse Blocked context onlyFeb 2, 2026
- Vision-DeepResearch Benchmark: Rethinking Visual and Textual Search for Multimodal Large Language Modelsarxiv-2602.02185 Direct Blocked context onlyFeb 2, 2026
- DCoPilot: Generative AI-Empowered Policy Adaptation for Dynamic Data Center Operationsarxiv-2602.02137 Sparse Blocked context onlyFeb 2, 2026
- Out of the Memory Barrier: A Highly Memory Efficient Training System for LLMs with Million-Token Contextsarxiv-2602.02108 Sparse Blocked context onlyFeb 2, 2026
- COMI: Coarse-to-fine Context Compression via Marginal Information Gainarxiv-2602.01719 Sparse Blocked context onlyFeb 2, 2026
- Mechanistic Indicators of Steering Effectiveness in Large Language Modelsarxiv-2602.01716 Sparse Blocked context onlyFeb 2, 2026
- Restoring Exploration after Post-Training: Latent Exploration Decoding for Large Reasoning Modelsarxiv-2602.01698 Sparse Blocked context onlyFeb 2, 2026
- InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMsarxiv-2602.01554 Sparse Blocked context onlyFeb 2, 2026
- ConPress: Learning Efficient Reasoning from Multi-Question Contextual Pressurearxiv-2602.01472 Sparse Blocked context onlyFeb 1, 2026
- CRAFT: Calibrated Reasoning with Answer-Faithful Traces via Reinforcement Learning for Multi-Hop Question Answeringarxiv-2602.01348 Sparse Blocked context onlyFeb 1, 2026
- What If We Allocate Test-Time Compute Adaptively?arxiv-2602.01070 Sparse Blocked context onlyFeb 1, 2026
- Beyond Static Instruction: A Multi-agent AI Framework for Adaptive Augmented Reality Robot Trainingarxiv-2603.00016 Sparse Blocked context onlyJan 31, 2026
- Can Small Language Models Handle Context-Summarized Multi-Turn Customer-Service QA? A Synthetic Data-Driven Comparative Evaluationarxiv-2602.00665 Sparse Blocked context onlyJan 31, 2026
- Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMsarxiv-2602.00564 Sparse Blocked context onlyJan 31, 2026
- PAIR-Former: Budgeted Relational Multi-Instance Learning for Functional miRNA Target Predictionarxiv-2602.00465 Sparse Blocked context onlyJan 31, 2026
- PaperBanana: Automating Academic Illustration for AI Scientistsarxiv-2601.23265 Direct Blocked context onlyJan 30, 2026
- Mem-T: Densifying Rewards for Long-Horizon Memory Agentsarxiv-2601.23014 Sparse Blocked context onlyJan 30, 2026
- TSPO: Breaking the Double Homogenization Dilemma in Multi-turn Search Policy Optimizationarxiv-2601.22776 Sparse Blocked context onlyJan 30, 2026
- KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Modelsarxiv-2602.11184 Sparse Blocked context onlyJan 30, 2026
- Time-Annealed Perturbation Sampling: Diverse Generation for Diffusion Language Modelsarxiv-2601.22629 Sparse Blocked context onlyJan 30, 2026
- From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agentsarxiv-2601.22607 Sparse Blocked context onlyJan 30, 2026
- Mock Worlds, Real Skills: Building Small Agentic Language Models with Synthetic Tasks, Simulated Environments, and Rubric-Based Rewardsarxiv-2601.22511 Sparse Blocked context onlyJan 30, 2026
- LLM Compression by Block Removal with Constrained Binary Optimizationarxiv-2602.00161 Sparse Blocked context onlyJan 29, 2026
- RedSage: A Cybersecurity Generalist LLMarxiv-2601.22159 Sparse Blocked context onlyJan 29, 2026
- Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirersarxiv-2601.22139 Sparse Blocked context onlyJan 29, 2026
- WebArbiter: A Principle-Guided Reasoning Process Reward Model for Web Agentsarxiv-2601.21872 Sparse Blocked context onlyJan 29, 2026
- MoHETS: Long-term Time Series Forecasting with Mixture-of-Heterogeneous-Expertsarxiv-2601.21866 Sparse Blocked context onlyJan 29, 2026
- Embodied Task Planning via Graph-Informed Action Generation with Large Language Modelarxiv-2601.21841 Sparse Blocked context onlyJan 29, 2026
- Indic-TunedLens: Interpreting Multilingual Models in Indian Languagesarxiv-2602.15038 Direct Blocked context onlyJan 29, 2026
- FIT to Forget: Robust Continual Unlearning for Large Language Modelsarxiv-2601.21682 Sparse Blocked context onlyJan 29, 2026
- Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluationarxiv-2601.21464 Sparse Blocked context onlyJan 29, 2026
- CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embeddingarxiv-2601.21262 Sparse Blocked context onlyJan 29, 2026
- MuVaC: A Variational Causal Framework for Multimodal Sarcasm Understanding in Dialoguesarxiv-2601.20451 Sparse Blocked context onlyJan 28, 2026
- Text-only adaptation in LLM-based ASR through text denoisingarxiv-2601.20900 Curated Related Blocked context onlyJan 28, 2026
- HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-bencharxiv-2601.20255 Sparse Blocked context onlyJan 28, 2026
- Structural Anchor Pruning: Training-Free Multi-Vector Compression for Visual Document Retrievalarxiv-2601.20107 Sparse Blocked context onlyJan 27, 2026
- Understanding LLM Failures: A Multi-Tape Turing Machine Analysis of Systematic Errors in Language Model Reasoningarxiv-2602.15868 Sparse Blocked context onlyJan 27, 2026
- LLMs versus the Halting Problem: Revisiting Program Termination Predictionarxiv-2601.18987 Sparse Blocked context onlyJan 26, 2026
- Flatter Tokens are More Valuable for Speculative Draft Model Trainingarxiv-2601.18902 Sparse Blocked context onlyJan 26, 2026
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Modelsarxiv-2601.18734 Sparse Blocked context onlyJan 26, 2026
- \textsc{NaVIDA}: Vision-Language Navigation with Inverse Dynamics Augmentationarxiv-2601.18188 Sparse Blocked context onlyJan 26, 2026
- Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agentsarxiv-2601.18077 Sparse Blocked context onlyJan 26, 2026
- EFT-CoT: A Multi-Agent Chain-of-Thought Framework for Emotion-Focused Therapyarxiv-2601.17842 Sparse Blocked context onlyJan 25, 2026
- Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Modelsarxiv-2601.17387 Sparse Blocked context onlyJan 24, 2026
- Decoupling Strategy and Execution in Task-Focused Dialogue via Goal-Oriented Preference Optimizationarxiv-2602.15854 Sparse Blocked context onlyJan 24, 2026
- EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutionsarxiv-2602.00095 Sparse Blocked context onlyJan 23, 2026
- IntelliAsk: Learning to Ask High-Quality Research Questions via RLVRarxiv-2602.15849 Sparse Blocked context onlyJan 23, 2026
- The Mouth is Not the Brain: Bridging Energy-Based World Models and Language Generationarxiv-2601.17094 Sparse Blocked context onlyJan 23, 2026
- Computer Environments Elicit General Agentic Intelligence in LLMsarxiv-2601.16206 Sparse Blocked context onlyJan 22, 2026
- ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Modelsarxiv-2601.15812 Sparse Blocked context onlyJan 22, 2026
- YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Modelsarxiv-2601.15588 Sparse Blocked context onlyJan 22, 2026
- Multi-Persona Thinking for Bias Mitigation in Large Language Modelsarxiv-2601.15488 Sparse Blocked context onlyJan 21, 2026
- The Effect of Scripts and Formats on LLM Numeracyarxiv-2601.15251 Sparse Blocked context onlyJan 21, 2026
- The Flexibility Trap: Why Arbitrary Order Limits Reasoning Potential in Diffusion Language Modelsarxiv-2601.15165 Sparse Blocked context onlyJan 21, 2026
- Knowledge Graphs are Implicit Reward Models: Path-Derived Signals Enable Compositional Reasoningarxiv-2601.15160 Sparse Blocked context onlyJan 21, 2026
- Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Modelsarxiv-2601.14758 Sparse Blocked context onlyJan 21, 2026
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understandingarxiv-2601.14724 Sparse Blocked context onlyJan 21, 2026
- Forest-Chat: Adapting Vision-Language Agents for Interactive Forest Change Analysisarxiv-2601.14637 Sparse Blocked context onlyJan 21, 2026
- Report for NSF Workshop on AI for Electronic Design Automationarxiv-2601.14541 Sparse Blocked context onlyJan 20, 2026
- VisTIRA: Closing the Image-Text Modality Gap in Visual Math Reasoning via Structured Tool Integrationarxiv-2601.14440 Sparse Blocked context onlyJan 20, 2026
- LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systemsarxiv-2601.14053 Sparse Blocked context onlyJan 20, 2026
- Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Modelsarxiv-2601.14004 Sparse Blocked context onlyJan 20, 2026
- Agentic SPARQL: Evaluating SPARQL-MCP-powered Intelligent Agents on the Federated KGQA Benchmarkarxiv-2603.06582 Sparse Blocked context onlyJan 20, 2026
- Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoringarxiv-2601.13879 Sparse Blocked context onlyJan 20, 2026
- Yuan3.0 Ultra: A Trillion-Parameter Enterprise-Oriented MoE LLMarxiv-2601.14327 Sparse Blocked context onlyJan 20, 2026
- Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Searcharxiv-2601.13719 Sparse Blocked context onlyJan 20, 2026
- Agentic Conversational Search with Contextualized Reasoning via Reinforcement Learningarxiv-2601.13115 Sparse Blocked context onlyJan 19, 2026
- ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generationarxiv-2601.12983 Sparse Blocked context onlyJan 19, 2026
- The Bitter Lesson of Diffusion Language Models for Agentic Workflows: A Comprehensive Reality Checkarxiv-2601.12979 Sparse Blocked context onlyJan 19, 2026
- Replayable Financial Agents: A Determinism-Faithfulness Assurance Harness for Tool-Using LLM Agentsarxiv-2601.15322 Sparse Blocked context onlyJan 17, 2026
- Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protectionarxiv-2601.12033 Sparse Blocked context onlyJan 17, 2026
- PEARL: Self-Evolving Assistant for Time Management with Reinforcement Learningarxiv-2601.11957 Sparse Blocked context onlyJan 17, 2026
- LSTM-MAS: A Long Short-Term Memory Inspired Multi-Agent System for Long-Context Understandingarxiv-2601.11913 Sparse Blocked context onlyJan 17, 2026
- Neural Chain-of-Thought Search: Searching the Optimal Reasoning Path to Enhance Large Language Modelsarxiv-2601.11340 Sparse Blocked context onlyJan 16, 2026
- Language of Thought Shapes Output Diversity in Large Language Modelsarxiv-2601.11227 Sparse Blocked context onlyJan 16, 2026
- T*: Progressive Block Scaling for Masked Diffusion Language Models Through Trajectory Aware Reinforcement Learningarxiv-2601.11214 Sparse Blocked context onlyJan 16, 2026
- TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speecharxiv-2601.11178 Sparse Blocked context onlyJan 16, 2026
- Generating metamers of human scene understandingarxiv-2601.11675 Sparse Blocked context onlyJan 16, 2026
- AJAR: Adaptive Jailbreak Architecture for Red-teamingarxiv-2601.10971 Sparse Blocked context onlyJan 16, 2026
- Beyond Max Tokens: Stealthy Resource Amplification via Tool Calling Chains in LLM Agentsarxiv-2601.10955 Sparse Blocked context onlyJan 16, 2026
- Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Modelsarxiv-2601.10679 Sparse Blocked context onlyJan 15, 2026
- Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Groundingarxiv-2601.10611 Sparse Blocked context onlyJan 15, 2026
- Learning Latency-Aware Orchestration for Multi-Agent Systemsarxiv-2601.10560 Sparse Blocked context onlyJan 15, 2026
- Toward Ultra-Long-Horizon Agentic Science: Cognitive Accumulation for Machine Learning Engineeringarxiv-2601.10402 Sparse Blocked context onlyJan 15, 2026
- TRIM: Hybrid Inference via Targeted Stepwise Routing in Multi-Step Reasoning Tasksarxiv-2601.10245 Sparse Blocked context onlyJan 15, 2026
- HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive Patternsarxiv-2601.10198 Sparse Blocked context onlyJan 15, 2026
- Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planningarxiv-2601.09708 Sparse Blocked context onlyJan 14, 2026
- SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metricsarxiv-2601.09487 Sparse Blocked context onlyJan 14, 2026
- GIFT: Reconciling Post-Training Objectives via Finite-Temperature Gibbs Initializationarxiv-2601.09233 Sparse Blocked context onlyJan 14, 2026
- FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figuresarxiv-2601.08026 Sparse Blocked context onlyJan 12, 2026
- VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understandingarxiv-2601.07986 Sparse Blocked context onlyJan 12, 2026
- Multilingual, Multimodal Pipeline for Creating Authentic and Structured Fact-Checked Claim Datasetarxiv-2601.07985 Sparse Blocked context onlyJan 12, 2026
- Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inferencearxiv-2601.07667 Sparse Blocked context onlyJan 12, 2026
- Thinking Before Constraining: A Unified Decoding Framework for Large Language Modelsarxiv-2601.07525 Sparse Blocked context onlyJan 12, 2026
- Two Pathways to Truthfulness: On the Intrinsic Encoding of LLM Hallucinationsarxiv-2601.07422 Sparse Blocked context onlyJan 12, 2026
- Reward Modeling from Natural Language Human Feedbackarxiv-2601.07349 Sparse Blocked context onlyJan 12, 2026
- Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluationarxiv-2601.07338 Sparse Blocked context onlyJan 12, 2026
- VLM-CAD: VLM-Optimized Collaborative Agent Design Workflow for Analog Circuit Sizingarxiv-2601.07315 Sparse Blocked context onlyJan 12, 2026
- NRR-Phi: Text-to-State Mapping for Ambiguity Preservation in LLM Inferencearxiv-2601.19933 Sparse Blocked context onlyJan 12, 2026
- CascadeMind at SemEval-2026 Task 4: A Hybrid Neuro-Symbolic Cascade for Narrative Similarityarxiv-2601.19931 Sparse Blocked context onlyJan 12, 2026
- EVM-QuestBench: An Execution-Grounded Benchmark for Natural-Language Transaction Code Generationarxiv-2601.06565 Sparse Blocked context onlyJan 10, 2026
- LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Modelsarxiv-2601.06550 Sparse Blocked context onlyJan 10, 2026
- Pantagruel: Unified Self-Supervised Encoders for French Text and Speecharxiv-2601.05911 Sparse Blocked context onlyJan 9, 2026
- Classroom AI: Large Language Models as Grade-Specific Teachersarxiv-2601.06225 Sparse Blocked context onlyJan 9, 2026
- HEART: A Unified Benchmark for Assessing Humans and LLMs in Emotional Support Dialoguearxiv-2601.19922 Sparse Blocked context onlyJan 9, 2026
- Over-Searching in Search-Augmented Large Language Modelsarxiv-2601.05503 Sparse Blocked context onlyJan 9, 2026
- Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Modelsarxiv-2601.05366 Sparse Blocked context onlyJan 8, 2026
- A Two-Stage Multitask Vision-Language Framework for Explainable Crop Disease Visual Question Answeringarxiv-2601.05143 Direct Blocked context onlyJan 8, 2026
- Token-Level LLM Collaboration via FusionRoutearxiv-2601.05106 Sparse Blocked context onlyJan 8, 2026
- DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Modelsarxiv-2601.04823 Sparse Blocked context onlyJan 8, 2026
- LAMB: LLM-based Audio Captioning with Modality Gap Bridging via Cauchy-Schwarz Divergencearxiv-2601.04658 Sparse Blocked context onlyJan 8, 2026
- Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systemsarxiv-2604.04936 Sparse Blocked context onlyJan 8, 2026
- Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Biasarxiv-2601.06194 Sparse Blocked context onlyJan 8, 2026
- CircuitLM: A Multi-Agent LLM-Aided Design Framework for Generating Circuit Schematics from Natural Language Promptsarxiv-2601.04505 Sparse Blocked context onlyJan 8, 2026
- Vision-Language Agents for Interactive Forest Change Analysisarxiv-2601.04497 Sparse Blocked context onlyJan 8, 2026
- PILOT: Planning via Internalized Latent Optimization Trajectories for Large Language Modelsarxiv-2601.19917 Sparse Blocked context onlyJan 7, 2026
- Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agentsarxiv-2601.03785 Sparse Blocked context onlyJan 7, 2026
- EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering for Enhanced Alignment and Reasoningarxiv-2601.03471 Sparse Blocked context onlyJan 6, 2026
- Prompting Underestimates LLM Capability for Time Series Classificationarxiv-2601.03464 Sparse Blocked context onlyJan 6, 2026
- AnatomiX, an Anatomy-Aware Grounded Multimodal Large Language Model for Chest X-Ray Interpretationarxiv-2601.03191 Sparse Blocked context onlyJan 6, 2026
- One Sample to Rule Them All: Extreme Data Efficiency in Multidiscipline Reasoning with Reinforcement Learningarxiv-2601.03111 Sparse Blocked context onlyJan 6, 2026
- Enhancing Moral Diagnosis and Correction in Large Language Modelsarxiv-2601.03079 Sparse Blocked context onlyJan 6, 2026
- SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answeringarxiv-2601.03014 Sparse Blocked context onlyJan 6, 2026
- Towards Faithful Reasoning in Comics for Small MLLMsarxiv-2601.02991 Sparse Blocked context onlyJan 6, 2026
- Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoningarxiv-2601.02970 Sparse Blocked context onlyJan 6, 2026
- Enhancing Multilingual RAG Systems with Debiased Language Preference-Guided Query Fusionarxiv-2601.02956 Sparse Blocked context onlyJan 6, 2026
- Beyond the Black Box: A Survey on the Theory and Mechanism of Large Language Modelsarxiv-2601.02907 Sparse Blocked context onlyJan 6, 2026
- HAL: Inducing Human-likeness in LLMs with Alignmentarxiv-2601.02813 Sparse Blocked context onlyJan 6, 2026
- Stratified Hazard Sampling: Minimal-Variance Event Scheduling for CTMC/DTMC Discrete Diffusion and Flow Modelsarxiv-2601.02799 Sparse Blocked context onlyJan 6, 2026
- SYNAPSE: Empowering LLM Agents with Episodic-Semantic Memory via Spreading Activationarxiv-2601.02744 Sparse Blocked context onlyJan 6, 2026
- When Do Tools and Planning Help Large Language Models Think? A Cost- and Latency-Aware Benchmarkarxiv-2601.02663 Sparse Blocked context onlyJan 6, 2026
- ModeX: Evaluator-Free Best-of-N Selection for Open-Ended Generationarxiv-2601.02535 Sparse Blocked context onlyJan 5, 2026
- Agentic Retoucher for Text-To-Image Generationarxiv-2601.02046 Sparse Blocked context onlyJan 5, 2026
- CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solvingarxiv-2601.01874 Sparse Blocked context onlyJan 5, 2026
- WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metricsarxiv-2601.02430 Sparse Blocked context onlyJan 5, 2026
- JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Modelsarxiv-2601.01627 Sparse Blocked context onlyJan 4, 2026
- Speculative Decoding: Performance or Illusion?arxiv-2601.11580 Sparse Blocked context onlyDec 31, 2025
- The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMsarxiv-2601.00097 Sparse Blocked context onlyDec 31, 2025
- RAIR: A Rule-Aware Benchmark Uniting Challenging Long-Tail and Visual Salience Subset for E-commerce Relevance Assessmentarxiv-2512.24943 Sparse Blocked context onlyDec 31, 2025
- Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speecharxiv-2512.24517 Sparse Blocked context onlyDec 30, 2025
- Multi-Agent LLMs for Generating Research Limitationsarxiv-2601.11578 Sparse Blocked context onlyDec 30, 2025
- Activation Steering for Masked Diffusion Language Modelsarxiv-2512.24143 Sparse Blocked context onlyDec 30, 2025
- Diversity or Precision? A Deep Dive into Next Token Predictionarxiv-2512.22955 Sparse Blocked context onlyDec 28, 2025
- Beg to Differ: Understanding Reasoning-Answer Misalignment Across Languagesarxiv-2512.22712 Sparse Blocked context onlyDec 27, 2025
- Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2arxiv-2512.22671 Sparse Blocked context onlyDec 27, 2025