300 canonical paper links on this archive page.
- GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNAarxiv-2510.23868 Sparse Blocked context onlyOct 27, 2025
- Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perceptionarxiv-2510.23853 Sparse Blocked context onlyOct 27, 2025
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Languagearxiv-2510.23828 Sparse Blocked context onlyOct 27, 2025
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?arxiv-2510.23587 Sparse Blocked context onlyOct 27, 2025
- RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translationarxiv-2510.23571 Sparse Blocked context onlyOct 27, 2025
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligencearxiv-2510.23538 Sparse Blocked context onlyOct 27, 2025
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Modelsarxiv-2510.22758 Sparse Blocked context onlyOct 26, 2025
- VisJudge-Bench: Aesthetics and Quality Assessment of Visualizationsarxiv-2510.22373 Curated Related Blocked context onlyOct 25, 2025
- WAON: Large-Scale Japanese Image-Text Pair Dataset for Improving Model Performance on Japanese Cultural Tasksarxiv-2510.22276 Sparse Blocked context onlyOct 25, 2025
- VisCoder2: Building Multi-Language Visualization Coding Agentsarxiv-2510.23642 Sparse Blocked context onlyOct 24, 2025
- PARL: Prompt-based Agents for Reinforcement Learningarxiv-2510.21306 Sparse Blocked context onlyOct 24, 2025
- AgentBound: Securing Execution Boundaries of AI Agentsarxiv-2510.21236 Sparse Blocked context onlyOct 24, 2025
- Estonian Native Large Language Model Benchmarkarxiv-2510.21193 Sparse Blocked context onlyOct 24, 2025
- Support-Contra Asymmetry in LLM Explanationsarxiv-2510.21884 Sparse Blocked context onlyOct 23, 2025
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculationarxiv-2510.20812 Sparse Blocked context onlyOct 23, 2025
- RELOOP: Recursive Retrieval with Multi-Hop Reasoner and Planners for Heterogeneous QAarxiv-2510.20505 Sparse Blocked context onlyOct 23, 2025
- Robust Preference Alignment via Directional Neighborhood Consensusarxiv-2510.20498 Sparse Blocked context onlyOct 23, 2025
- CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativityarxiv-2510.20091 Sparse Blocked context onlyOct 23, 2025
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoningarxiv-2510.19807 Sparse Blocked context onlyOct 22, 2025
- ToolDreamer: Instilling LLM Reasoning Into Tool Retrieversarxiv-2510.19791 Sparse Blocked context onlyOct 22, 2025
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inferencearxiv-2510.19669 Sparse Blocked context onlyOct 22, 2025
- A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklistarxiv-2510.19139 Sparse Blocked context onlyOct 22, 2025
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsarxiv-2510.18876 Sparse Blocked context onlyOct 21, 2025
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoningarxiv-2510.18814 Sparse Blocked context onlyOct 21, 2025
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judgearxiv-2510.18196 Sparse Blocked context onlyOct 21, 2025
- Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Modelsarxiv-2510.18077 Sparse Blocked context onlyOct 20, 2025
- SPACeR: Self-Play Anchoring with Centralized Reference Modelsarxiv-2510.18060 Sparse Blocked context onlyOct 20, 2025
- Annotation-Efficient Universal Honesty Alignmentarxiv-2510.17509 Sparse Blocked context onlyOct 20, 2025
- StreamingThinker: Large Language Models Can Think While Readingarxiv-2510.17238 Sparse Blocked context onlyOct 20, 2025
- Soft-Masked Diffusion Language Modelsarxiv-2510.17206 Sparse Blocked context onlyOct 20, 2025
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Modelsarxiv-2510.16917 Sparse Blocked context onlyOct 19, 2025
- Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Modelsarxiv-2510.16727 Sparse Blocked context onlyOct 19, 2025
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimizationarxiv-2510.16635 Sparse Blocked context onlyOct 18, 2025
- Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safetyarxiv-2510.16492 Sparse Blocked context onlyOct 18, 2025
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attributionarxiv-2510.16439 Sparse Blocked context onlyOct 18, 2025
- SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detectionarxiv-2510.16219 Sparse Blocked context onlyOct 17, 2025
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstractionarxiv-2510.15863 Sparse Blocked context onlyOct 17, 2025
- HypoSpace: Evaluating LLM Creativity as Set-Valued Hypothesis Generators under Underdeterminationarxiv-2510.15614 Sparse Blocked context onlyOct 17, 2025
- SAG-Agent: Enabling Long-Horizon Reasoning in Strategy Games via Dynamic Knowledge Graphsarxiv-2510.15259 Sparse Blocked context onlyOct 17, 2025
- GUIrilla: A Scalable Framework for Automated Desktop UI Explorationarxiv-2510.16051 Sparse Blocked context onlyOct 16, 2025
- Composition-Grounded Data Synthesis for Visual Reasoningarxiv-2510.15040 Sparse Blocked context onlyOct 16, 2025
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agentsarxiv-2510.14967 Sparse Blocked context onlyOct 16, 2025
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summariesarxiv-2510.14751 Sparse Blocked context onlyOct 16, 2025
- E2Edev: Benchmarking Large Language Models in End-to-End Software Development Taskarxiv-2510.14509 Sparse Blocked context onlyOct 16, 2025
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Followingarxiv-2510.14420 Sparse Blocked context onlyOct 16, 2025
- PluriHopRAG: Exhaustive, Recall-Sensitive QA Through Corpus-Specific Document Structure Learningarxiv-2510.14377 Sparse Blocked context onlyOct 16, 2025
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimizationarxiv-2510.14150 Direct Blocked context onlyOct 15, 2025
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completionarxiv-2510.13702 Sparse Blocked context onlyOct 15, 2025
- Closing the Gap Between Text and Speech Understanding in LLMsarxiv-2510.13632 Sparse Blocked context onlyOct 15, 2025
- MemoTime: Memory-Augmented Temporal Knowledge Graph Enhanced Large Language Model Reasoningarxiv-2510.13614 Sparse Blocked context onlyOct 15, 2025
- Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discoursearxiv-2510.13417 Sparse Blocked context onlyOct 15, 2025
- Mismatch Aware Guidance for Robust Emotion Control in Auto-Regressive TTS Modelsarxiv-2510.13293 Sparse Blocked context onlyOct 15, 2025
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanismarxiv-2510.13170 Sparse Blocked context onlyOct 15, 2025
- On the Reasoning Abilities of Masked Diffusion Language Modelsarxiv-2510.13117 Sparse Blocked context onlyOct 15, 2025
- Schema for In-Context Learningarxiv-2510.13905 Sparse Blocked context onlyOct 14, 2025
- Reveal-to-Revise: Explainable Bias-Aware Generative Modeling with Multimodal Attentionarxiv-2510.12957 Sparse Blocked context onlyOct 14, 2025
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differencesarxiv-2510.13900 Sparse Blocked context onlyOct 14, 2025
- Toward LLM-Supported Automated Assessment of Critical Thinking Subskillsarxiv-2510.12915 Sparse Blocked context onlyOct 14, 2025
- Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generationarxiv-2510.12460 Sparse Blocked context onlyOct 14, 2025
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agentsarxiv-2510.15994 Sparse Blocked context onlyOct 14, 2025
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editingarxiv-2510.12121 Sparse Blocked context onlyOct 14, 2025
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulationarxiv-2510.13884 Sparse Blocked context onlyOct 14, 2025
- R-WoM: Retrieval-augmented World Model For Computer-use Agentsarxiv-2510.11892 Sparse Blocked context onlyOct 13, 2025
- StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Modelsarxiv-2510.11618 Sparse Blocked context onlyOct 13, 2025
- Unlocking the Potential of Diffusion Language Models through Template Infillingarxiv-2510.13870 Sparse Blocked context onlyOct 13, 2025
- ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Modelsarxiv-2510.13860 Sparse Blocked context onlyOct 13, 2025
- DropVLA: An Action-Level Backdoor Attack on Vision--Language--Action Modelsarxiv-2510.10932 Sparse Blocked context onlyOct 13, 2025
- DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Modelsarxiv-2510.10846 Sparse Blocked context onlyOct 12, 2025
- Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Promptingarxiv-2510.10528 Sparse Blocked context onlyOct 12, 2025
- FML-bench: Benchmarking Machine Learning Agents for Scientific Researcharxiv-2510.10472 Sparse Blocked context onlyOct 12, 2025
- EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editingarxiv-2510.13851 Sparse Blocked context onlyOct 11, 2025
- Language steering in latent space to mitigate unintended code-switchingarxiv-2510.13849 Sparse Blocked context onlyOct 11, 2025
- You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMsarxiv-2510.10223 Sparse Blocked context onlyOct 11, 2025
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMsarxiv-2510.09885 Sparse Blocked context onlyOct 10, 2025
- GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Dataarxiv-2510.09580 Sparse Blocked context onlyOct 10, 2025
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answeringarxiv-2510.09351 Sparse Blocked context onlyOct 10, 2025
- ATLAS: Adaptive Trading with LLM AgentS Through Dynamic Prompt Optimization and Multi-Agent Coordinationarxiv-2510.15949 Sparse Blocked context onlyOct 10, 2025
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Expertsarxiv-2510.09278 Sparse Blocked context onlyOct 10, 2025
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Modelsarxiv-2510.09259 Sparse Blocked context onlyOct 10, 2025
- Clear Roads, Clear Vision: Advancements in Multi-Weather Restoration for Smart Transportationarxiv-2510.09228 Sparse Blocked context onlyOct 10, 2025
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsarxiv-2510.09201 Sparse Blocked context onlyOct 10, 2025
- FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMsarxiv-2510.08886 Direct Blocked context onlyOct 10, 2025
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Codingarxiv-2510.08804 Sparse Blocked context onlyOct 9, 2025
- How Reliable is Language Model Micro-Benchmarking?arxiv-2510.08730 Sparse Blocked context onlyOct 9, 2025
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresightarxiv-2510.08713 Direct Blocked context onlyOct 9, 2025
- DeepPrune: Parallel Scaling without Inter-trace Redundancyarxiv-2510.08483 Sparse Blocked context onlyOct 9, 2025
- If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Modelsarxiv-2510.08388 Sparse Blocked context onlyOct 9, 2025
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?arxiv-2510.07962 Direct Blocked context onlyOct 9, 2025
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recallarxiv-2510.07896 Sparse Blocked context onlyOct 9, 2025
- Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energyarxiv-2510.08646 Sparse Blocked context onlyOct 9, 2025
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewardsarxiv-2510.07774 Sparse Blocked context onlyOct 9, 2025
- EconCausal: A Context-Aware Causal Reasoning Benchmark for Large Language Models in Social Sciencearxiv-2510.07231 Sparse Blocked context onlyOct 8, 2025
- Search-R3: Unifying Reasoning and Embedding in Large Language Modelsarxiv-2510.07048 Sparse Blocked context onlyOct 8, 2025
- Native Hybrid Attention for Efficient Sequence Modelingarxiv-2510.07019 Sparse Blocked context onlyOct 8, 2025
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipelinearxiv-2510.06800 Sparse Blocked context onlyOct 8, 2025
- How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effectsarxiv-2510.06700 Sparse Blocked context onlyOct 8, 2025
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratcharxiv-2510.06670 Sparse Blocked context onlyOct 8, 2025
- Peeking inside the Black-Box: Reinforcement Learning for Explainable and Accurate Relation Extractionarxiv-2510.06198 Sparse Blocked context onlyOct 7, 2025
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Modelsarxiv-2510.06107 Sparse Blocked context onlyOct 7, 2025
- Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerabilityarxiv-2510.06084 Sparse Blocked context onlyOct 7, 2025
- Prompt reinforcing for long-term planning of large language modelsarxiv-2510.05921 Sparse Blocked context onlyOct 7, 2025
- Early Multimodal Prediction of Cross-Lingual Meme Virality on Reddit: A Time-Window Analysisarxiv-2510.05761 Sparse Blocked context onlyOct 7, 2025
- Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficultyarxiv-2510.05534 Sparse Blocked context onlyOct 7, 2025
- Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEsarxiv-2510.05278 Sparse Blocked context onlyOct 6, 2025
- Slm-mux: Orchestrating small language models for reasoningarxiv-2510.05077 Sparse Blocked context onlyOct 6, 2025
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materialsarxiv-2510.04704 Sparse Blocked context onlyOct 6, 2025
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAarxiv-2510.04682 Sparse Blocked context onlyOct 6, 2025
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Modelsarxiv-2510.04618 Direct Blocked context onlyOct 6, 2025
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoningarxiv-2510.04573 Sparse Blocked context onlyOct 6, 2025
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluationarxiv-2510.04265 Sparse Blocked context onlyOct 5, 2025
- AlphaApollo: A System for Deep Agentic Reasoningarxiv-2510.06261 Sparse Blocked context onlyOct 5, 2025
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environmentsarxiv-2510.04142 Sparse Blocked context onlyOct 5, 2025
- PoLi-RL: A Point-to-List Reinforcement Learning Framework for Conditional Semantic Textual Similarityarxiv-2510.04080 Sparse Blocked context onlyOct 5, 2025
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoningarxiv-2510.04072 Sparse Blocked context onlyOct 5, 2025
- What Scales in Cross-Entropy Scaling Law?arxiv-2510.04067 Sparse Blocked context onlyOct 5, 2025
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Modelsarxiv-2510.03721 Sparse Blocked context onlyOct 4, 2025
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learningarxiv-2510.03669 Sparse Blocked context onlyOct 4, 2025
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operationsarxiv-2510.03666 Sparse Blocked context onlyOct 4, 2025
- AgentHub: A Registry for Discoverable, Verifiable, and Reproducible AI Agentsarxiv-2510.03495 Sparse Blocked context onlyOct 3, 2025
- Attention-Aligned Reasoning for Large Language Modelsarxiv-2510.03223 Sparse Blocked context onlyOct 3, 2025
- Unraveling Syntax: How Language Models Learn Context-Free Grammarsarxiv-2510.02524 Sparse Blocked context onlyOct 2, 2025
- ExGRPO: Learning to Reason from Experiencearxiv-2510.02245 Direct Blocked context onlyOct 2, 2025
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?arxiv-2510.02209 Sparse Blocked context onlyOct 2, 2025
- VL-KnG: Persistent Spatiotemporal Knowledge Graphs from Egocentric Video for Embodied Scene Understandingarxiv-2510.01483 Sparse Blocked context onlyOct 1, 2025
- CurES: From Gradient Analysis to Efficient Curriculum Learning for Reasoning LLMsarxiv-2510.01037 Sparse Blocked context onlyOct 1, 2025
- Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scalingarxiv-2510.01025 Sparse Blocked context onlyOct 1, 2025
- Erase to Improve: Erasable Reinforcement Learning for Search-Augmented LLMsarxiv-2510.00861 Sparse Blocked context onlyOct 1, 2025
- ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMsarxiv-2510.00857 Sparse Blocked context onlyOct 1, 2025
- Stochastic Self-Organization in Multi-Agent Systemsarxiv-2510.00685 Sparse Blocked context onlyOct 1, 2025
- ReSeek: A Self-Correcting Framework for Search Agents with Instructive Rewardsarxiv-2510.00568 Sparse Blocked context onlyOct 1, 2025
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphsarxiv-2510.00507 Sparse Blocked context onlyOct 1, 2025
- Training Large Language Models To Reason In Parallel With Global Forking Tokensarxiv-2510.05132 Sparse Blocked context onlyOct 1, 2025
- PromptLoop: Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignmentarxiv-2510.00430 Sparse Blocked context onlyOct 1, 2025
- Towards Self-Evolving Benchmarks: Synthesizing Agent Trajectories via Test-Time Exploration under Validate-by-Reproduce Paradigmarxiv-2510.00415 Sparse Blocked context onlyOct 1, 2025
- PrefDisco: Benchmarking Proactive Personalized Reasoningarxiv-2510.00177 Sparse Blocked context onlySep 30, 2025
- MENLO: From Preferences to Proficiency -- Evaluating and Modeling Native-like Quality Across 47 Languagesarxiv-2509.26601 Direct Blocked context onlySep 30, 2025
- OffTopicEval: When Large Language Models Enter the Wrong Chat, Almost Always!arxiv-2509.26495 Sparse Blocked context onlySep 30, 2025
- Your Agent May Misevolve: Emergent Risks in Self-evolving LLM Agentsarxiv-2509.26354 Sparse Blocked context onlySep 30, 2025
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editingarxiv-2509.26346 Curated Related Blocked context onlySep 30, 2025
- Latent Thinking Optimization: Your Latent Reasoning Language Model Secretly Encodes Reward Signals in Its Latent Thoughtsarxiv-2509.26314 Sparse Blocked context onlySep 30, 2025
- SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIParxiv-2509.26036 Sparse Blocked context onlySep 30, 2025
- ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attackarxiv-2509.25843 Sparse Blocked context onlySep 30, 2025
- Overthinking Reduction with Decoupled Rewards and Curriculum Data Schedulingarxiv-2509.25827 Sparse Blocked context onlySep 30, 2025
- v-HUB: A Benchmark for Video Humor Understanding from Vision and Soundarxiv-2509.25773 Sparse Blocked context onlySep 30, 2025
- LD-MoLE: Learnable Dynamic Routing for Mixture of LoRA Expertsarxiv-2509.25684 Sparse Blocked context onlySep 30, 2025
- Generative Value Conflicts Reveal LLM Prioritiesarxiv-2509.25369 Sparse Blocked context onlySep 29, 2025
- Pretraining with hierarchical memories: separating long-tail and common knowledgearxiv-2510.02375 Sparse Blocked context onlySep 29, 2025
- Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agentsarxiv-2509.25302 Sparse Blocked context onlySep 29, 2025
- Towards Personalized Deep Research: Benchmarks and Evaluationsarxiv-2509.25106 Sparse Blocked context onlySep 29, 2025
- Ultra-Fast Language Generation via Discrete Diffusion Divergence Instructarxiv-2509.25035 Sparse Blocked context onlySep 29, 2025
- Agentic Exploration of Physics Modelsarxiv-2509.24978 Sparse Blocked context onlySep 29, 2025
- MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipesarxiv-2509.24945 Direct Blocked context onlySep 29, 2025
- Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMsarxiv-2509.24857 Sparse Blocked context onlySep 29, 2025
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Modelsarxiv-2509.24803 Direct Blocked context onlySep 29, 2025
- Inducing Dyslexia in Vision Language Modelsarxiv-2509.24597 Sparse Blocked context onlySep 29, 2025
- Speculative Verification: Exploiting Information Gain to Refine Speculative Decodingarxiv-2509.24328 Sparse Blocked context onlySep 29, 2025
- DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Modelsarxiv-2509.24296 Sparse Blocked context onlySep 29, 2025
- G-reasoner: Foundation Models for Unified Reasoning over Graph-structured Knowledgearxiv-2509.24276 Sparse Blocked context onlySep 29, 2025
- Pragmatic Inference for Moral Reasoning Acquisition: Generalization via Metapragmatic Linksarxiv-2509.24102 Sparse Blocked context onlySep 28, 2025
- Uncovering Grounding IDs: How External Cues Shape Multimodal Bindingarxiv-2509.24072 Sparse Blocked context onlySep 28, 2025
- SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Modelsarxiv-2509.23863 Sparse Blocked context onlySep 28, 2025
- From What to Why: A Multi-Agent System for Evidence-based Chemical Reaction Condition Reasoningarxiv-2509.23768 Sparse Blocked context onlySep 28, 2025
- Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factualityarxiv-2509.23765 Sparse Blocked context onlySep 28, 2025
- SafeSearch: Automated Red-Teaming of LLM-Based Search Agentsarxiv-2509.23694 Sparse Blocked context onlySep 28, 2025
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensionalarxiv-2509.23499 Sparse Blocked context onlySep 27, 2025
- Mapping Overlaps in Benchmarks through Perplexity in the Wildarxiv-2509.23488 Sparse Blocked context onlySep 27, 2025
- Your Models Have Thought Enough: Training Large Reasoning Models to Stop Overthinkingarxiv-2509.23392 Sparse Blocked context onlySep 27, 2025
- Learning to Reason in Structured In-context Environments with Reinforcement Learningarxiv-2509.23330 Sparse Blocked context onlySep 27, 2025
- p-less Sampling: A Robust Hyperparameter-Free Approach for LLM Decodingarxiv-2509.23234 Sparse Blocked context onlySep 27, 2025
- AutoEP: LLMs-Driven Automation of Hyperparameter Evolution for Metaheuristic Algorithmsarxiv-2509.23189 Sparse Blocked context onlySep 27, 2025
- RHYTHM: Reasoning with Hierarchical Temporal Tokenization for Human Mobilityarxiv-2509.23115 Sparse Blocked context onlySep 27, 2025
- General Exploratory Bonus for Optimistic Exploration in RLHFarxiv-2510.03269 Sparse Blocked context onlySep 27, 2025
- Look Back to Reason Forward: Revisitable Memory for Long-Context LLM Agentsarxiv-2509.23040 Sparse Blocked context onlySep 27, 2025
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddingsarxiv-2509.22925 Sparse Blocked context onlySep 26, 2025
- HEART: Emotionally-Driven Test-Time Scaling of Language Modelsarxiv-2509.22876 Sparse Blocked context onlySep 26, 2025
- Critique-Coder: Enhancing Coder Models by Critique Reinforcement Learningarxiv-2509.22824 Sparse Blocked context onlySep 26, 2025
- StateX: Enhancing RNN Recall via Post-training State Expansionarxiv-2509.22630 Sparse Blocked context onlySep 26, 2025
- GeoSketch: A Neural-Symbolic Approach to Geometric Multimodal Reasoning with Auxiliary Line Construction and Affine Transformationarxiv-2509.22460 Sparse Blocked context onlySep 26, 2025
- FeatBench: Towards More Realistic Evaluation of Feature-level Code Generationarxiv-2509.22237 Sparse Blocked context onlySep 26, 2025
- LogiPart: Local Large Language Models for Data Exploration at Scale with Logical Partitioningarxiv-2509.22211 Sparse Blocked context onlySep 26, 2025
- Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillationarxiv-2509.22193 Sparse Blocked context onlySep 26, 2025
- Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decodingarxiv-2509.22134 Sparse Blocked context onlySep 26, 2025
- SciTS: Scientific Time Series Understanding and Generation with LLMsarxiv-2510.03255 Sparse Blocked context onlySep 26, 2025
- MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inferencearxiv-2509.22750 Sparse Blocked context onlySep 26, 2025
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Modelsarxiv-2509.21991 Sparse Blocked context onlySep 26, 2025
- Multimodal Neural Operators for Real-Time Biomechanical Modelling of Traumatic Brain Injuryarxiv-2510.03248 Sparse Blocked context onlySep 26, 2025
- ProPerSim: Developing Proactive and Personalized AI Assistants through User-Assistant Simulationarxiv-2509.21730 Sparse Blocked context onlySep 26, 2025
- ReviewScore: Misinformed Peer Review Detection with Large Language Modelsarxiv-2509.21679 Sparse Blocked context onlySep 25, 2025
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devicesarxiv-2510.05109 Sparse Blocked context onlySep 25, 2025
- AutoClimDS: Climate Data Science Agentic AI -- A Knowledge Graph is All You Needarxiv-2509.21553 Sparse Blocked context onlySep 25, 2025
- RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewardsarxiv-2509.21319 Direct Blocked context onlySep 25, 2025
- UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic Languagesarxiv-2509.21294 Sparse Blocked context onlySep 25, 2025
- Tree Search for LLM Agent Reinforcement Learningarxiv-2509.21240 Direct Blocked context onlySep 25, 2025
- Sigma: Semantically Informative Pre-training for Skeleton-based Sign Language Understandingarxiv-2509.21223 Sparse Blocked context onlySep 25, 2025
- CLAUSE: Agentic Neuro-Symbolic Knowledge Graph Reasoning via Dynamic Learnable Context Engineeringarxiv-2509.21035 Sparse Blocked context onlySep 25, 2025
- Predicting LLM Reasoning Performance with Small Proxy Modelarxiv-2509.21013 Sparse Blocked context onlySep 25, 2025
- MARS: toward more efficient multi-agent collaboration for LLM reasoningarxiv-2509.20502 Sparse Blocked context onlySep 24, 2025
- EpidemIQs: Prompt-to-Paper LLM Agents for Epidemic Modeling and Analysisarxiv-2510.00024 Sparse Blocked context onlySep 24, 2025
- From Text to Talk: Audio-Language Model Needs Non-Autoregressive Joint Trainingarxiv-2509.20072 Sparse Blocked context onlySep 24, 2025
- SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Modelsarxiv-2509.19270 Direct Blocked context onlySep 23, 2025
- Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactionsarxiv-2509.18847 Sparse Blocked context onlySep 23, 2025
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inferencearxiv-2509.18487 Sparse Blocked context onlySep 23, 2025
- Through the Lens of Human-Human Collaboration: A Configurable Research Platform for Exploring Human-Agent Collaborationarxiv-2509.18008 Sparse Blocked context onlySep 22, 2025
- A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialoguearxiv-2509.17766 Sparse Blocked context onlySep 22, 2025
- Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translationarxiv-2509.17349 Sparse Blocked context onlySep 22, 2025
- LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimizationarxiv-2509.17183 Sparse Blocked context onlySep 21, 2025
- AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluationarxiv-2509.16952 Sparse Blocked context onlySep 21, 2025
- Can GRPO Boost Complex Multimodal Table Understanding?arxiv-2509.16889 Sparse Blocked context onlySep 21, 2025
- Distribution-Aligned Decoding for Efficient LLM Task Adaptationarxiv-2509.15888 Sparse Blocked context onlySep 19, 2025
- Beyond Words: Enhancing Desire, Emotion, and Sentiment Recognition with Non-Verbal Cuesarxiv-2509.15540 Sparse Blocked context onlySep 19, 2025
- ATTS: Asynchronous Test-Time Scaling via Conformal Predictionarxiv-2509.15148 Sparse Blocked context onlySep 18, 2025
- GeoResponder: Towards Building Geospatial LLMs for Time-Critical Disaster Responsearxiv-2509.19354 Sparse Blocked context onlySep 18, 2025
- Reasoning Efficiently Through Adaptive Chain-of-Thought Compression: A Self-Optimizing Frameworkarxiv-2509.14093 Sparse Blocked context onlySep 17, 2025
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Togglesarxiv-2509.13615 Sparse Blocked context onlySep 17, 2025
- ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarizationarxiv-2509.13313 Sparse Blocked context onlySep 16, 2025
- From Next Token Prediction to (STRIPS) World Modelsarxiv-2509.13389 Sparse Blocked context onlySep 16, 2025
- Learning to Optimize Multi-Objective Alignment Through Dynamic Reward Weightingarxiv-2509.11452 Sparse Blocked context onlySep 14, 2025
- No Answer Needed: Predicting LLM Answer Accuracy from Question-Only Linear Probesarxiv-2509.10625 Sparse Blocked context onlySep 12, 2025
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMsarxiv-2509.09677 Sparse Blocked context onlySep 11, 2025
- Evolution and compression in LLMs: On the emergence of human-aligned categorizationarxiv-2509.08093 Sparse Blocked context onlySep 9, 2025
- New Insights into Optimal Alignment of Acoustic and Linguistic Representations for Knowledge Transfer in ASRarxiv-2509.05609 Sparse Blocked context onlySep 6, 2025
- Post-training Large Language Models for Diverse High-Quality Responsesarxiv-2509.04784 Sparse Blocked context onlySep 5, 2025
- Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenariosarxiv-2509.04403 Direct Blocked context onlySep 4, 2025
- Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoningarxiv-2509.03345 Sparse Blocked context onlySep 3, 2025
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflectionarxiv-2509.03113 Sparse Blocked context onlySep 3, 2025
- Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVRarxiv-2509.02522 Sparse Blocked context onlySep 2, 2025
- Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generationarxiv-2509.02510 Sparse Blocked context onlySep 2, 2025
- LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialoguearxiv-2509.02292 Sparse Blocked context onlySep 2, 2025
- Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Modelsarxiv-2509.01350 Sparse Blocked context onlySep 1, 2025
- TempCore: Are Video QA Benchmarks Temporally Grounded? A Frame Selection Sensitivity Analysis and Benchmarkarxiv-2509.01167 Sparse Blocked context onlySep 1, 2025
- When Thinking Backfires: Mechanistic Insights Into Reasoning-Induced Misalignmentarxiv-2509.00544 Sparse Blocked context onlyAug 30, 2025
- PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chainsarxiv-2508.21787 Sparse Blocked context onlyAug 29, 2025
- On the Theoretical Limitations of Embedding-Based Retrievalarxiv-2508.21038 Direct Blocked context onlyAug 28, 2025
- EO-1: An Open Unified Embodied Foundation Model for General Robot Controlarxiv-2508.21112 Sparse Blocked context onlyAug 28, 2025
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problemsarxiv-2508.20373 Curated Related Blocked context onlyAug 28, 2025
- AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenariosarxiv-2508.19988 Sparse Blocked context onlyAug 27, 2025
- Diffusion Language Models Know the Answer Before Decodingarxiv-2508.19982 Sparse Blocked context onlyAug 27, 2025
- LaTeXTrans: Structured LaTeX Translation with Multi-Agent Coordinationarxiv-2508.18791 Direct Blocked context onlyAug 26, 2025
- VistaWise: Building Cost-Effective Agent with Cross-Modal Knowledge Graph for Minecraftarxiv-2508.18722 Sparse Blocked context onlyAug 26, 2025
- Latent Self-Consistency for Reliable Majority-Set Selection in Short- and Long-Answer Reasoningarxiv-2508.18395 Sparse Blocked context onlyAug 25, 2025
- How Quantization Shapes Bias in Large Language Modelsarxiv-2508.18088 Sparse Blocked context onlyAug 25, 2025
- MedRepBench: A Comprehensive Benchmark for Medical Report Interpretationarxiv-2508.16674 Direct Blocked context onlyAug 21, 2025
- AmbiSQL: Interactive Ambiguity Detection and Resolution for Text-to-SQLarxiv-2508.15276 Sparse Blocked context onlyAug 21, 2025
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMsarxiv-2508.14896 Sparse Blocked context onlyAug 20, 2025
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulationarxiv-2508.13998 Sparse Blocked context onlyAug 19, 2025
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimizationarxiv-2508.13993 Sparse Blocked context onlyAug 19, 2025
- The Collaboration Paradox: Why Generative AI Requires Both Strategic Intelligence and Operational Stability in Supply Chain Managementarxiv-2508.13942 Sparse Blocked context onlyAug 19, 2025
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Explorationarxiv-2508.13755 Sparse Blocked context onlyAug 19, 2025
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generationarxiv-2508.13587 Sparse Blocked context onlyAug 19, 2025
- TASER: Table Agents for Schema-guided Extraction and Recommendationarxiv-2508.13404 Sparse Blocked context onlyAug 18, 2025
- TaoSR1: The Thinking Model for E-commerce Relevance Searcharxiv-2508.12365 Sparse Blocked context onlyAug 17, 2025
- CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressuresarxiv-2508.11915 Sparse Blocked context onlyAug 16, 2025
- SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communicationarxiv-2508.11733 Sparse Blocked context onlyAug 15, 2025
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasksarxiv-2508.11360 Sparse Blocked context onlyAug 15, 2025
- Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rulesarxiv-2508.10971 Sparse Blocked context onlyAug 14, 2025
- Agentic Design Review Systemarxiv-2508.10745 Sparse Blocked context onlyAug 14, 2025
- GenOM: Ontology Matching with Description Generation and Large Language Modelarxiv-2508.10703 Sparse Blocked context onlyAug 14, 2025
- For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMsarxiv-2508.10180 Sparse Blocked context onlyAug 13, 2025
- From Context to Intent: Reasoning-Guided Function-Level Code Completionarxiv-2508.09537 Sparse Blocked context onlyAug 13, 2025
- PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoningarxiv-2508.09521 Sparse Blocked context onlyAug 13, 2025
- IAG: Input-aware Backdoor Attack on VLM-based Visual Groundingarxiv-2508.09456 Sparse Blocked context onlyAug 13, 2025
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokensarxiv-2508.09325 Sparse Blocked context onlyAug 12, 2025
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environmentsarxiv-2508.08791 Sparse Blocked context onlyAug 12, 2025
- DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrievalarxiv-2508.07995 Direct Blocked context onlyAug 11, 2025
- 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoningarxiv-2508.07667 Sparse Blocked context onlyAug 11, 2025
- Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimizationarxiv-2508.07629 Sparse Blocked context onlyAug 11, 2025
- SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detectionarxiv-2508.06803 Sparse Blocked context onlyAug 9, 2025
- Memp: Exploring Agent Procedural Memoryarxiv-2508.06433 Sparse Blocked context onlyAug 8, 2025
- UR$^2$: Unify RAG and Reasoning through Reinforcement Learningarxiv-2508.06165 Sparse Blocked context onlyAug 8, 2025
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generationarxiv-2508.06046 Sparse Blocked context onlyAug 8, 2025
- GroundAct: Can LLM Agents Ground Actions in Environmental States?arxiv-2508.05614 Sparse Blocked context onlyAug 7, 2025
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policyarxiv-2508.05592 Sparse Blocked context onlyAug 7, 2025
- Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoningarxiv-2508.05282 Sparse Blocked context onlyAug 7, 2025
- TURA: Tool-Augmented Unified Retrieval Agent for AI Searcharxiv-2508.04604 Sparse Blocked context onlyAug 6, 2025
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learningarxiv-2508.04581 Sparse Blocked context onlyAug 6, 2025
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Momentsarxiv-2508.04204 Sparse Blocked context onlyAug 6, 2025
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"arxiv-2508.04086 Sparse Blocked context onlyAug 6, 2025
- CoAct-1: Computer-using Multi-Agent System with Coding Actionsarxiv-2508.03923 Sparse Blocked context onlyAug 5, 2025
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMsarxiv-2508.02066 Sparse Blocked context onlyAug 4, 2025
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agentsarxiv-2508.01503 Sparse Blocked context onlyAug 2, 2025
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Dataarxiv-2508.01450 Sparse Blocked context onlyAug 2, 2025
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Trainingarxiv-2508.00414 Direct Blocked context onlyAug 1, 2025
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimizationarxiv-2508.00222 Sparse Blocked context onlyJul 31, 2025
- DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Spacearxiv-2607.25675 Sparse Blocked context onlyJul 28, 2026
- CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimizationarxiv-2607.25659 Sparse Blocked context onlyJul 28, 2026
- Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansionarxiv-2607.25572 Curated Related Blocked context onlyJul 28, 2026
- Where Quality Breaks in Compressed Short-Text Generation: Staged Bottleneck Localizationarxiv-2607.24176 Sparse Blocked context onlyJul 27, 2026
- Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsificationarxiv-2607.24027 Sparse Blocked context onlyJul 27, 2026
- Characterizing Warp Divergence from Pascal to Blackwellarxiv-2607.23402 Sparse Blocked context onlyJul 26, 2026
- IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languagesarxiv-2607.23242 Direct Blocked context onlyJul 25, 2026
- Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Modelsarxiv-2607.21936 Sparse Blocked context onlyJul 24, 2026
- Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-Onarxiv-2607.21694 Sparse Blocked context onlyJul 23, 2026
- SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generationarxiv-2607.21553 Sparse Blocked context onlyJul 23, 2026
- Self Gradient Forcing: Native Long Video Extrapolationarxiv-2607.20368 Sparse Blocked context onlyJul 22, 2026
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Modelarxiv-2607.20058 Sparse Blocked context onlyJul 22, 2026
- Delineate Anything v2: A Global Foundation Model for Field Delineationarxiv-2607.19069 Sparse Blocked context onlyJul 21, 2026
- FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicryarxiv-2607.18227 Sparse Blocked context onlyJul 20, 2026
- Three-Body Scattering for Generative Modelingarxiv-2607.18198 Sparse Blocked context onlyJul 20, 2026
- Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraintsarxiv-2607.18144 Sparse Blocked context onlyJul 20, 2026
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shiftarxiv-2607.17524 Sparse Blocked context onlyJul 20, 2026