300 canonical paper links on this archive page.
- Syntactic Framing Fragility: An Audit of Robustness in LLM Ethical Decisionsarxiv-2601.09724 Sparse Blocked context onlyDec 27, 2025
- ADMEDTAGGER: an annotation framework for distillation of expert knowledge for the Polish medical languagearxiv-2601.09722 Sparse Blocked context onlyDec 27, 2025
- Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifoldsarxiv-2512.22473 Sparse Blocked context onlyDec 27, 2025
- The Bayesian Geometry of Transformer Attentionarxiv-2512.22471 Sparse Blocked context onlyDec 27, 2025
- Monadic Context Engineeringarxiv-2512.22431 Sparse Blocked context onlyDec 27, 2025
- CricBench: A Multilingual Benchmark for Evaluating LLMs in Cricket Analyticsarxiv-2512.21877 Sparse Blocked context onlyDec 26, 2025
- Parallel Token Prediction for Language Modelsarxiv-2512.21323 Sparse Blocked context onlyDec 24, 2025
- Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanationarxiv-2512.21066 Sparse Blocked context onlyDec 24, 2025
- Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillationarxiv-2512.20908 Sparse Blocked context onlyDec 24, 2025
- MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMsarxiv-2512.20822 Sparse Blocked context onlyDec 23, 2025
- Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profilesarxiv-2512.20780 Sparse Blocked context onlyDec 23, 2025
- Generalization of RLVR Using Causal Reasoning as a Testbedarxiv-2512.20760 Sparse Blocked context onlyDec 23, 2025
- AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agentarxiv-2512.20745 Sparse Blocked context onlyDec 23, 2025
- Coherence in the brain unfolds across separable temporal regimesarxiv-2512.20481 Sparse Blocked context onlyDec 23, 2025
- Reason2Decide: Rationale-Driven Multi-Task Learningarxiv-2512.20074 Sparse Blocked context onlyDec 23, 2025
- Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?arxiv-2512.19980 Sparse Blocked context onlyDec 23, 2025
- Stop saying LLM: Large Discourse Models (LDM) and Artificial Discursive Agent (ADA)?arxiv-2512.19117 Sparse Blocked context onlyDec 22, 2025
- Training-Free Global Geometric Association for 4D LiDAR Panoptic Segmentationarxiv-2512.18991 Sparse Blocked context onlyDec 22, 2025
- From Word to World: Can Large Language Models be Implicit Text-based World Models?arxiv-2512.18832 Sparse Blocked context onlyDec 21, 2025
- Towards Efficient Agents: A Co-Design of Inference Architecture and Systemarxiv-2512.18337 Sparse Blocked context onlyDec 20, 2025
- MACE-Dance: Motion-Appearance Cascaded Experts for Music-Driven Dance Video Generationarxiv-2512.18181 Sparse Blocked context onlyDec 20, 2025
- DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generationarxiv-2512.17776 Sparse Blocked context onlyDec 19, 2025
- Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQLarxiv-2512.17053 Sparse Blocked context onlyDec 18, 2025
- Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learningarxiv-2512.16917 Sparse Blocked context onlyDec 18, 2025
- Refusal Steering: Fine-grained Control over LLM Refusal Behaviour for Sensitive Topicsarxiv-2512.16602 Sparse Blocked context onlyDec 18, 2025
- TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Modelsarxiv-2512.16523 Sparse Blocked context onlyDec 18, 2025
- Adaptation of Agentic AI: A Survey of Post-Training, Memory, and Skillsarxiv-2512.16301 Sparse Blocked context onlyDec 18, 2025
- Stepwise Think-Critique: A Unified Framework for Robust and Interpretable LLM Reasoningarxiv-2512.15662 Sparse Blocked context onlyDec 17, 2025
- The Moralization Corpus: Frame-Based Annotation and Analysis of Moralizing Speech Acts across Diverse Text Genresarxiv-2512.15248 Sparse Blocked context onlyDec 17, 2025
- MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Serversarxiv-2512.15163 Sparse Blocked context onlyDec 17, 2025
- Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learningarxiv-2512.15146 Sparse Blocked context onlyDec 17, 2025
- Cross-Tokenizer Likelihood Scoring Algorithms for Language Model Distillationarxiv-2512.14954 Sparse Blocked context onlyDec 16, 2025
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMsarxiv-2512.14698 Curated Related Blocked context onlyDec 16, 2025
- GRAFT: Grid-Aware Load Forecasting with Multi-Source Textual Alignment and Fusionarxiv-2512.14400 Sparse Blocked context onlyDec 16, 2025
- RePo: Language Models with Context Re-Positioningarxiv-2512.14391 Sparse Blocked context onlyDec 16, 2025
- Towards Interactive Intelligence for Digital Humansarxiv-2512.13674 Sparse Blocked context onlyDec 15, 2025
- ReFusion: A Diffusion Large Language Model with Parallel Autoregressive Decodingarxiv-2512.13586 Sparse Blocked context onlyDec 15, 2025
- NRR-Core: Non-Resolution Reasoning as a Computational Framework for Contextual Identity and Ambiguity Preservationarxiv-2512.13478 Sparse Blocked context onlyDec 15, 2025
- What Makes an Ideal Quote? Recommending "Unexpected yet Rational" Quotations via Noveltyarxiv-2602.22220 Sparse Blocked context onlyDec 15, 2025
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Trainingarxiv-2512.13043 Sparse Blocked context onlyDec 15, 2025
- Understanding Structured Financial Data with LLMs: A Case Study on Fraud Detectionarxiv-2512.13040 Sparse Blocked context onlyDec 15, 2025
- Information-Consistent Language Model Recommendations through Group Relative Policy Optimizationarxiv-2512.12858 Sparse Blocked context onlyDec 14, 2025
- Does Tone Change the Answer? Evaluating Prompt Politeness Effects on Modern LLMs: GPT, Gemini, and LLaMAarxiv-2512.12812 Sparse Blocked context onlyDec 14, 2025
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Spacearxiv-2512.12623 Sparse Blocked context onlyDec 14, 2025
- Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolutionarxiv-2512.10696 Sparse Blocked context onlyDec 11, 2025
- SWAA: Sliding Window Attention Adaptation for Efficient and Quality Preserving Long Context Processingarxiv-2512.10411 Sparse Blocked context onlyDec 11, 2025
- Parallel Decoder Transformer: Planner-Seeded Latent Coordination for Synchronized Parallel Decodingarxiv-2512.10054 Sparse Blocked context onlyDec 10, 2025
- Toward Closed-loop Molecular Discovery via Language Model, Property Alignment and Strategic Searcharxiv-2512.09566 Sparse Blocked context onlyDec 10, 2025
- Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languagesarxiv-2512.08777 Sparse Blocked context onlyDec 9, 2025
- Automatic Essay Scoring and Feedback Generation in Basque Language Learningarxiv-2512.08713 Sparse Blocked context onlyDec 9, 2025
- QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Modelsarxiv-2512.08646 Sparse Blocked context onlyDec 9, 2025
- Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoningarxiv-2512.08639 Sparse Blocked context onlyDec 9, 2025
- Training Language Models to Use Prolog as a Toolarxiv-2512.07407 Sparse Blocked context onlyDec 8, 2025
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Modelsarxiv-2512.07222 Sparse Blocked context onlyDec 8, 2025
- STaRR: Spatial-Temporal Token-Dynamics-Aware Responsive Remasking for Diffusion Language Modelsarxiv-2601.04205 Sparse Blocked context onlyDec 7, 2025
- ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wildarxiv-2512.06721 Sparse Blocked context onlyDec 7, 2025
- Think-While-Generating: On-the-Fly Reasoning for Personalized Long-Form Generationarxiv-2512.06690 Sparse Blocked context onlyDec 7, 2025
- Towards Small Language Models for Security Query Generation in SOC Workflowsarxiv-2512.06660 Sparse Blocked context onlyDec 7, 2025
- Conflict-Aware Fusion: Resolving Logic Inertia in Large Language Models via Structured Cognitive Priorsarxiv-2512.06393 Sparse Blocked context onlyDec 6, 2025
- ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answeringarxiv-2512.05430 Sparse Blocked context onlyDec 5, 2025
- LexGenius: An Expert-Level Benchmark for Large Language Models in Legal General Intelligencearxiv-2512.04578 Sparse Blocked context onlyDec 4, 2025
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisitionarxiv-2512.03794 Sparse Blocked context onlyDec 3, 2025
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMsarxiv-2512.03324 Sparse Blocked context onlyDec 3, 2025
- Is Vibe Coding Safe? Benchmarking Vulnerability of Agent-Generated Code in Real-World Tasksarxiv-2512.03262 Sparse Blocked context onlyDec 2, 2025
- From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?arxiv-2512.03005 Sparse Blocked context onlyDec 2, 2025
- Process-Centric Analysis of Agentic Software Systemsarxiv-2512.02393 Sparse Blocked context onlyDec 2, 2025
- OGD4All: A Framework for Accessible Interaction with Geospatial Open Government Data Based on Large Language Modelsarxiv-2602.00012 Sparse Blocked context onlyNov 30, 2025
- OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusionarxiv-2512.00234 Sparse Blocked context onlyNov 28, 2025
- ORCA: Open-ended Response Correctness Assessment for Audio Question Answeringarxiv-2512.09066 Sparse Blocked context onlyNov 28, 2025
- Dripper: Token-Efficient Main HTML Extraction with a Lightweight LMarxiv-2511.23119 Sparse Blocked context onlyNov 28, 2025
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answeringarxiv-2511.22715 Sparse Blocked context onlyNov 27, 2025
- DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQAarxiv-2511.22521 Sparse Blocked context onlyNov 27, 2025
- PAT: Accelerating LLM Decoding via Prefix-Aware Attention with Resource Efficient Multi-Tile Kernelarxiv-2511.22333 Sparse Blocked context onlyNov 27, 2025
- Hybrid Stackelberg Game and Diffusion-based Auction for Two-tier Agentic AI Task Offloading in Internet of Agentsarxiv-2511.22076 Sparse Blocked context onlyNov 27, 2025
- MedEyes: Learning Dynamic Visual Focus for Medical Progressive Diagnosisarxiv-2511.22018 Sparse Blocked context onlyNov 27, 2025
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Tweedie Discrete Diffusionarxiv-2511.21542 Sparse Blocked context onlyNov 26, 2025
- SPHINX: A Synthetic Environment for Visual Perception and Reasoningarxiv-2511.20814 Sparse Blocked context onlyNov 25, 2025
- Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detectionarxiv-2511.20162 Sparse Blocked context onlyNov 25, 2025
- Stabilizing Off-Policy Training for Long-Horizon LLM Agent via Turn-Level Importance Sampling and Clipping-Triggered Normalizationarxiv-2511.20718 Sparse Blocked context onlyNov 25, 2025
- ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Modelsarxiv-2512.07843 Sparse Blocked context onlyNov 24, 2025
- CDLM: Consistency Diffusion Language Models For Faster Samplingarxiv-2511.19269 Sparse Blocked context onlyNov 24, 2025
- SO-Bench: A Structural Output Evaluation of Multimodal LLMsarxiv-2511.21750 Sparse Blocked context onlyNov 23, 2025
- Scaling Implicit Fields via Hypernetwork-Driven Multiscale Coordinate Transformationsarxiv-2511.18387 Sparse Blocked context onlyNov 23, 2025
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularizationarxiv-2511.17938 Sparse Blocked context onlyNov 22, 2025
- A cross-species neural foundation model for end-to-end speech decodingarxiv-2511.21740 Sparse Blocked context onlyNov 21, 2025
- REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agentarxiv-2511.17442 Sparse Blocked context onlyNov 21, 2025
- Estonian WinoGrande Dataset: Comparative Analysis of LLM Performance on Human and Machine Translationarxiv-2511.17290 Sparse Blocked context onlyNov 21, 2025
- HUMORCHAIN: Theory-Guided Multi-Stage Reasoning for Interpretable Multimodal Humor Generationarxiv-2511.21732 Sparse Blocked context onlyNov 21, 2025
- Bridging Symbolic Control and Neural Reasoning in LLM Agents: Structured Cognitive Loop with a Governance Layerarxiv-2511.17673 Sparse Blocked context onlyNov 21, 2025
- Pharos-ESG: A Framework for Multimodal Parsing, Contextual Narration, and Hierarchical Labeling of ESG Reportarxiv-2511.16417 Sparse Blocked context onlyNov 20, 2025
- AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimizationarxiv-2511.15915 Sparse Blocked context onlyNov 19, 2025
- German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studiesarxiv-2511.21722 Sparse Blocked context onlyNov 19, 2025
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skippingarxiv-2511.15690 Sparse Blocked context onlyNov 19, 2025
- When to Think and When to Look: Uncertainty-Guided Lookbackarxiv-2511.15613 Sparse Blocked context onlyNov 19, 2025
- Masked Auto-Regressive Variational Acceleration: Fast Inference Makes Practical Reinforcement Learningarxiv-2511.15190 Sparse Blocked context onlyNov 19, 2025
- SVBRD-LLM: Self-Verifying Behavioral Rule Discovery for Autonomous Vehicle Identificationarxiv-2511.14977 Sparse Blocked context onlyNov 18, 2025
- Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Modelsarxiv-2511.14900 Sparse Blocked context onlyNov 18, 2025
- From Competition to Coordination: Market Making as a Scalable Framework for Safe and Aligned Multi-Agent LLM Systemsarxiv-2511.17621 Sparse Blocked context onlyNov 18, 2025
- Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distributionarxiv-2511.14275 Sparse Blocked context onlyNov 18, 2025
- PRISM: Prompt-Refined In-Context System Modelling for Financial Retrievalarxiv-2511.14130 Sparse Blocked context onlyNov 18, 2025
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restorationarxiv-2511.14099 Sparse Blocked context onlyNov 18, 2025
- Cost-Effective Communication: An Auction-based Method for Language Agent Interactionarxiv-2511.13193 Sparse Blocked context onlyNov 17, 2025
- From Passive to Persuasive: Steering Emotional Nuance in Human-AI Negotiationarxiv-2511.12832 Sparse Blocked context onlyNov 16, 2025
- Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMsarxiv-2511.12710 Sparse Blocked context onlyNov 16, 2025
- Co-Layout: LLM-driven Co-optimization for Interior Layoutarxiv-2511.12474 Sparse Blocked context onlyNov 16, 2025
- Mobile-Agent-RAG: Driving Smart Multi-Agent Coordination with Contextual Knowledge Empowerment for Long-Horizon Mobile Automationarxiv-2511.12254 Sparse Blocked context onlyNov 15, 2025
- EARL: Entropy-Aware RL Alignment of LLMs for Reliable RTL Code Generationarxiv-2511.12033 Sparse Blocked context onlyNov 15, 2025
- Context-Emotion Aware Therapeutic Dialogue Generation: A Multi-component Reinforcement Learning Approach to Language Models for Mental Health Supportarxiv-2511.11884 Sparse Blocked context onlyNov 14, 2025
- Conformal Constrained Policy Optimization for Cost-Effective LLM Agentsarxiv-2511.11828 Sparse Blocked context onlyNov 14, 2025
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scalingarxiv-2511.11793 Sparse Blocked context onlyNov 14, 2025
- LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Modelsarxiv-2511.11334 Sparse Blocked context onlyNov 14, 2025
- CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generationarxiv-2511.11104 Sparse Blocked context onlyNov 14, 2025
- Multimodal Peer Review Simulation with Actionable To-Do Recommendations for Community-Aware Manuscript Revisionsarxiv-2511.10902 Sparse Blocked context onlyNov 14, 2025
- From Efficiency to Adaptivity: A Deeper Look at Adaptive Reasoning in Large Language Modelsarxiv-2511.10788 Sparse Blocked context onlyNov 13, 2025
- Beyond Elicitation: Provision-based Prompt Optimization for Knowledge-Intensive Tasksarxiv-2511.10465 Sparse Blocked context onlyNov 13, 2025
- Quality Assurance of LLM-generated Code: Addressing Non-Functional Quality Characteristicsarxiv-2511.10271 Sparse Blocked context onlyNov 13, 2025
- RadHiera: Semantic Hierarchical Reinforcement Learning for Medical Report Generationarxiv-2511.10065 Sparse Blocked context onlyNov 13, 2025
- Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basquearxiv-2511.09396 Sparse Blocked context onlyNov 12, 2025
- POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translationarxiv-2511.09232 Sparse Blocked context onlyNov 12, 2025
- Human or LLM as Standardized Patients? A Comparative Study for Medical Educationarxiv-2511.14783 Sparse Blocked context onlyNov 12, 2025
- Towards Hyper-Efficient RAG Systems in VecDBs: Distributed Parallel Multi-Resolution Vector Searcharxiv-2511.16681 Sparse Blocked context onlyNov 12, 2025
- $π$-Attention: Periodic Sparse Transformers for Efficient Long-Context Modelingarxiv-2511.10696 Curated Related Blocked context onlyNov 12, 2025
- Alphacast: An Interaction-Driven Agentic Reasoning Framework for Cognition-Inspired Time Series Forecastingarxiv-2511.08947 Sparse Blocked context onlyNov 12, 2025
- Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agentsarxiv-2511.10687 Sparse Blocked context onlyNov 11, 2025
- AlphaResearch: Accelerating New Algorithm Discovery with Language Modelsarxiv-2511.08522 Sparse Blocked context onlyNov 11, 2025
- Automatic Paper Reviewing with Heterogeneous Graph Reasoning over LLM-Simulated Reviewer-Author Debatesarxiv-2511.08317 Sparse Blocked context onlyNov 11, 2025
- Multimodal LLMs Do Not Compose Skills Optimally Across Modalitiesarxiv-2511.08113 Sparse Blocked context onlyNov 11, 2025
- Beyond Fact Retrieval: Episodic Memory for RAG with Generative Semantic Workspacesarxiv-2511.07587 Sparse Blocked context onlyNov 10, 2025
- Q-RAG: Long Context Multi-step Retrieval via Value-based Embedder Trainingarxiv-2511.07328 Sparse Blocked context onlyNov 10, 2025
- Graph Representation-based Model Poisoning on the Heterogeneous Internet of Agentsarxiv-2511.07176 Sparse Blocked context onlyNov 10, 2025
- Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMsarxiv-2511.05933 Sparse Blocked context onlyNov 8, 2025
- IDALC: A Semi-Supervised Framework for Intent Detection and Active Learning based Correctionarxiv-2511.05921 Sparse Blocked context onlyNov 8, 2025
- Q$^2$: Quantization-Aware Gradient Balancing and Attention Alignment for Low-Bit Quantizationarxiv-2511.05898 Sparse Blocked context onlyNov 8, 2025
- OckBench: Measuring the Efficiency of LLM Reasoningarxiv-2511.05722 Sparse Blocked context onlyNov 7, 2025
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scalearxiv-2511.05705 Sparse Blocked context onlyNov 7, 2025
- Steering Language Models with Weight Arithmeticarxiv-2511.05408 Sparse Blocked context onlyNov 7, 2025
- Making Knowledge Accessible: Divergent Readability-Accuracy Strategies of Mistral and QWen in Biomedical Text Simplificationarxiv-2511.05080 Sparse Blocked context onlyNov 7, 2025
- AgentExpt: Automating AI Experiment Design with LLM-based Resource Retrieval Agentarxiv-2511.04921 Sparse Blocked context onlyNov 7, 2025
- Modeling Clinical Uncertainty in Radiology Reports: from Explicit Uncertainty Markers to Implicit Reasoning Pathwaysarxiv-2511.04506 Sparse Blocked context onlyNov 6, 2025
- STARS: Synchronous Token Alignment for Robust Supervision in Large Language Modelsarxiv-2511.03827 Sparse Blocked context onlyNov 5, 2025
- Grounded Misunderstandings in Asymmetric Dialogue: A Perspectivist Annotation Scheme for MapTaskarxiv-2511.03718 Sparse Blocked context onlyNov 5, 2025
- CareMedEval dataset: Evaluating Critical Appraisal and Reasoning in the Biomedical Fieldarxiv-2511.03441 Sparse Blocked context onlyNov 5, 2025
- EQ-Negotiator: Dynamic Emotional Personas Empower Small Language Models for Edge-Deployable Credit Negotiationarxiv-2511.03370 Sparse Blocked context onlyNov 5, 2025
- Silenced Biases: The Dark Side LLMs Learned to Refusearxiv-2511.03369 Sparse Blocked context onlyNov 5, 2025
- From Five Dimensions to Many: Large Language Models as Precise and Interpretable Psychological Profilersarxiv-2511.03235 Sparse Blocked context onlyNov 5, 2025
- Mina: A Multilingual LLM-Powered Legal Assistant Agent for Bangladesh for Empowering Access to Justicearxiv-2511.08605 Sparse Blocked context onlyNov 4, 2025
- CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agentsarxiv-2511.02734 Sparse Blocked context onlyNov 4, 2025
- PETra: A Multilingual Corpus of Pragmatic Explicitation in Translationarxiv-2511.02721 Sparse Blocked context onlyNov 4, 2025
- An Interdisciplinary and Cross-Task Review on Missing Data Imputationarxiv-2511.01196 Sparse Blocked context onlyNov 3, 2025
- Surfacing Subtle Stereotypes: A Multilingual, Debate-Oriented Evaluation of Modern LLMsarxiv-2511.01187 Sparse Blocked context onlyNov 3, 2025
- TSVer: A Benchmark for Fact Verification Against Time-Series Evidencearxiv-2511.01101 Sparse Blocked context onlyNov 2, 2025
- Prompt-R1: Collaborative Automatic Prompting Framework via End-to-end Reinforcement Learningarxiv-2511.01016 Sparse Blocked context onlyNov 2, 2025
- Self-Consistency Is Losing Its Edge: Diminishing Returns and Rising Costs in Modern LLMsarxiv-2511.00751 Sparse Blocked context onlyNov 2, 2025
- SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoningarxiv-2510.27568 Sparse Blocked context onlyOct 31, 2025
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chainsarxiv-2510.27419 Sparse Blocked context onlyOct 31, 2025
- Beyond a Million Tokens: Benchmarking and Enhancing Long-Term Memory in LLMsarxiv-2510.27246 Sparse Blocked context onlyOct 31, 2025
- Glia: A Human-Inspired AI for Automated Systems Design and Optimizationarxiv-2510.27176 Sparse Blocked context onlyOct 31, 2025
- Reasoning Up the Instruction Ladder for Controllable Language Modelsarxiv-2511.04694 Sparse Blocked context onlyOct 30, 2025
- Emu3.5: Native Multimodal Models are World Learnersarxiv-2510.26583 Direct Blocked context onlyOct 30, 2025
- Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Modelsarxiv-2510.26577 Sparse Blocked context onlyOct 30, 2025
- Co-Evolving Latent Action World Modelsarxiv-2510.26433 Sparse Blocked context onlyOct 30, 2025
- The Geometry of Dialogue: Graphing Language Models to Reveal Synergistic Teams for Multi-Agent Collaborationarxiv-2510.26352 Sparse Blocked context onlyOct 30, 2025
- Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Modelsarxiv-2510.26241 Sparse Blocked context onlyOct 30, 2025
- Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoningarxiv-2510.25992 Sparse Blocked context onlyOct 29, 2025
- RECAP: Reproducing Copyrighted Data from LLMs Training with an Agentic Pipelinearxiv-2510.25941 Sparse Blocked context onlyOct 29, 2025
- Through the Judge's Eyes: Inferred Thinking Traces Improve Reliability of LLM Ratersarxiv-2510.25860 Sparse Blocked context onlyOct 29, 2025
- TheraMind: A Strategic and Adaptive Agent for Longitudinal Psychological Counselingarxiv-2510.25758 Sparse Blocked context onlyOct 29, 2025
- The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Executionarxiv-2510.25726 Direct Blocked context onlyOct 29, 2025
- ProMediate: A Socio-cognitive framework for evaluating proactive agents in multi-party negotiationarxiv-2510.25224 Sparse Blocked context onlyOct 29, 2025
- Activation-Space Personality Steering: Hybrid Layer Selection for Stable Trait Control in LLMsarxiv-2511.03738 Sparse Blocked context onlyOct 29, 2025
- World Simulation with Video Foundation Models for Physical AIarxiv-2511.00062 Sparse Blocked context onlyOct 28, 2025
- Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgisharxiv-2510.24856 Sparse Blocked context onlyOct 28, 2025
- Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agentsarxiv-2510.24702 Sparse Blocked context onlyOct 28, 2025
- Tongyi DeepResearch Technical Reportarxiv-2510.24701 Direct Blocked context onlyOct 28, 2025
- OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learningarxiv-2510.24636 Sparse Blocked context onlyOct 28, 2025
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generationarxiv-2510.24821 Sparse Blocked context onlyOct 28, 2025
- Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewardsarxiv-2510.24302 Sparse Blocked context onlyOct 28, 2025
- MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotationsarxiv-2510.24178 Sparse Blocked context onlyOct 28, 2025
- GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNAarxiv-2510.23868 Sparse Blocked context onlyOct 27, 2025
- Your LLM Agents are Temporally Blind: The Misalignment Between Tool Use Decisions and Human Time Perceptionarxiv-2510.23853 Sparse Blocked context onlyOct 27, 2025
- Beyond Understanding: Evaluating the Pragmatic Gap in LLMs' Cultural Processing of Figurative Languagearxiv-2510.23828 Sparse Blocked context onlyOct 27, 2025
- A Survey of Data Agents: Emerging Paradigm or Overstated Hype?arxiv-2510.23587 Sparse Blocked context onlyOct 27, 2025
- RobotArena $\infty$: Scalable Robot Benchmarking via Real-to-Sim Translationarxiv-2510.23571 Sparse Blocked context onlyOct 27, 2025
- JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligencearxiv-2510.23538 Sparse Blocked context onlyOct 27, 2025
- EchoMind: An Interrelated Multi-level Benchmark for Evaluating Empathetic Speech Language Modelsarxiv-2510.22758 Sparse Blocked context onlyOct 26, 2025
- VisJudge-Bench: Aesthetics and Quality Assessment of Visualizationsarxiv-2510.22373 Curated Related Blocked context onlyOct 25, 2025
- CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learningarxiv-2510.22282 Sparse Blocked context onlyOct 25, 2025
- WAON: Large-Scale Japanese Image-Text Pair Dataset for Improving Model Performance on Japanese Cultural Tasksarxiv-2510.22276 Sparse Blocked context onlyOct 25, 2025
- VisCoder2: Building Multi-Language Visualization Coding Agentsarxiv-2510.23642 Sparse Blocked context onlyOct 24, 2025
- PARL: Prompt-based Agents for Reinforcement Learningarxiv-2510.21306 Sparse Blocked context onlyOct 24, 2025
- AgentBound: Securing Execution Boundaries of AI Agentsarxiv-2510.21236 Sparse Blocked context onlyOct 24, 2025
- Estonian Native Large Language Model Benchmarkarxiv-2510.21193 Sparse Blocked context onlyOct 24, 2025
- Support-Contra Asymmetry in LLM Explanationsarxiv-2510.21884 Sparse Blocked context onlyOct 23, 2025
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculationarxiv-2510.20812 Sparse Blocked context onlyOct 23, 2025
- RELOOP: Recursive Retrieval with Multi-Hop Reasoner and Planners for Heterogeneous QAarxiv-2510.20505 Sparse Blocked context onlyOct 23, 2025
- Robust Preference Alignment via Directional Neighborhood Consensusarxiv-2510.20498 Sparse Blocked context onlyOct 23, 2025
- CreativityPrism: A Holistic Evaluation Framework for Large Language Model Creativityarxiv-2510.20091 Sparse Blocked context onlyOct 23, 2025
- Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoningarxiv-2510.19807 Sparse Blocked context onlyOct 22, 2025
- ToolDreamer: Instilling LLM Reasoning Into Tool Retrieversarxiv-2510.19791 Sparse Blocked context onlyOct 22, 2025
- DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inferencearxiv-2510.19669 Sparse Blocked context onlyOct 22, 2025
- MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designsarxiv-2510.19366 Sparse Blocked context onlyOct 22, 2025
- A Multi-faceted Analysis of Cognitive Abilities: Evaluating Prompt Methods with Large Language Models on the CONSORT Checklistarxiv-2510.19139 Sparse Blocked context onlyOct 22, 2025
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMsarxiv-2510.18876 Sparse Blocked context onlyOct 21, 2025
- A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoningarxiv-2510.18814 Sparse Blocked context onlyOct 21, 2025
- Contrastive Decoding Mitigates Score Range Bias in LLM-as-a-Judgearxiv-2510.18196 Sparse Blocked context onlyOct 21, 2025
- Chain-of-Thought Reasoning Improves Context-Aware Translation with Large Language Modelsarxiv-2510.18077 Sparse Blocked context onlyOct 20, 2025
- SPACeR: Self-Play Anchoring with Centralized Reference Modelsarxiv-2510.18060 Sparse Blocked context onlyOct 20, 2025
- Annotation-Efficient Universal Honesty Alignmentarxiv-2510.17509 Sparse Blocked context onlyOct 20, 2025
- StreamingThinker: Large Language Models Can Think While Readingarxiv-2510.17238 Sparse Blocked context onlyOct 20, 2025
- Soft-Masked Diffusion Language Modelsarxiv-2510.17206 Sparse Blocked context onlyOct 20, 2025
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Modelsarxiv-2510.16917 Sparse Blocked context onlyOct 19, 2025
- Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Modelsarxiv-2510.16727 Sparse Blocked context onlyOct 19, 2025
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimizationarxiv-2510.16635 Sparse Blocked context onlyOct 18, 2025
- Check Yourself Before You Wreck Yourself: Selectively Quitting Improves LLM Agent Safetyarxiv-2510.16492 Sparse Blocked context onlyOct 18, 2025
- FrugalPrompt: Reducing Contextual Overhead in Large Language Models via Token Attributionarxiv-2510.16439 Sparse Blocked context onlyOct 18, 2025
- SentinelNet: Safeguarding Multi-Agent Collaboration Through Credit-Based Dynamic Threat Detectionarxiv-2510.16219 Sparse Blocked context onlyOct 17, 2025
- PolySkill: Learning Generalizable Skills Through Polymorphic Abstractionarxiv-2510.15863 Sparse Blocked context onlyOct 17, 2025
- HypoSpace: Evaluating LLM Creativity as Set-Valued Hypothesis Generators under Underdeterminationarxiv-2510.15614 Sparse Blocked context onlyOct 17, 2025
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecyclearxiv-2510.16079 Sparse Blocked context onlyOct 17, 2025
- SAG-Agent: Enabling Long-Horizon Reasoning in Strategy Games via Dynamic Knowledge Graphsarxiv-2510.15259 Sparse Blocked context onlyOct 17, 2025
- GUIrilla: A Scalable Framework for Automated Desktop UI Explorationarxiv-2510.16051 Sparse Blocked context onlyOct 16, 2025
- Composition-Grounded Data Synthesis for Visual Reasoningarxiv-2510.15040 Sparse Blocked context onlyOct 16, 2025
- Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agentsarxiv-2510.14967 Sparse Blocked context onlyOct 16, 2025
- Beyond Multi-Token Prediction: Pretraining LLMs with Future Summariesarxiv-2510.14751 Sparse Blocked context onlyOct 16, 2025
- E2Edev: Benchmarking Large Language Models in End-to-End Software Development Taskarxiv-2510.14509 Sparse Blocked context onlyOct 16, 2025
- Instructions are all you need: Self-supervised Reinforcement Learning for Instruction Followingarxiv-2510.14420 Sparse Blocked context onlyOct 16, 2025
- PluriHopRAG: Exhaustive, Recall-Sensitive QA Through Corpus-Specific Document Structure Learningarxiv-2510.14377 Sparse Blocked context onlyOct 16, 2025
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimizationarxiv-2510.14150 Direct Blocked context onlyOct 15, 2025
- MVCustom: Multi-View Customized Diffusion via Geometric Latent Rendering and Completionarxiv-2510.13702 Sparse Blocked context onlyOct 15, 2025
- Closing the Gap Between Text and Speech Understanding in LLMsarxiv-2510.13632 Sparse Blocked context onlyOct 15, 2025
- MemoTime: Memory-Augmented Temporal Knowledge Graph Enhanced Large Language Model Reasoningarxiv-2510.13614 Sparse Blocked context onlyOct 15, 2025
- Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discoursearxiv-2510.13417 Sparse Blocked context onlyOct 15, 2025
- Mismatch Aware Guidance for Robust Emotion Control in Auto-Regressive TTS Modelsarxiv-2510.13293 Sparse Blocked context onlyOct 15, 2025
- Putting on the Thinking Hats: A Survey on Chain of Thought Fine-tuning from the Perspective of Human Reasoning Mechanismarxiv-2510.13170 Sparse Blocked context onlyOct 15, 2025
- On the Reasoning Abilities of Masked Diffusion Language Modelsarxiv-2510.13117 Sparse Blocked context onlyOct 15, 2025
- Schema for In-Context Learningarxiv-2510.13905 Sparse Blocked context onlyOct 14, 2025
- Reveal-to-Revise: Explainable Bias-Aware Generative Modeling with Multimodal Attentionarxiv-2510.12957 Sparse Blocked context onlyOct 14, 2025
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differencesarxiv-2510.13900 Sparse Blocked context onlyOct 14, 2025
- Toward LLM-Supported Automated Assessment of Critical Thinking Subskillsarxiv-2510.12915 Sparse Blocked context onlyOct 14, 2025
- Beyond Black-Box Interventions: Latent Probing for Faithful Retrieval-Augmented Generationarxiv-2510.12460 Sparse Blocked context onlyOct 14, 2025
- MCP Security Bench (MSB): Benchmarking Attacks Against Model Context Protocol in LLM Agentsarxiv-2510.15994 Sparse Blocked context onlyOct 14, 2025
- Precise Attribute Intensity Control in Large Language Models via Targeted Representation Editingarxiv-2510.12121 Sparse Blocked context onlyOct 14, 2025
- Too Open for Opinion? Embracing Open-Endedness in Large Language Models for Social Simulationarxiv-2510.13884 Sparse Blocked context onlyOct 14, 2025
- R-WoM: Retrieval-augmented World Model For Computer-use Agentsarxiv-2510.11892 Sparse Blocked context onlyOct 13, 2025
- StoryBox: Collaborative Multi-Agent Simulation for Hybrid Bottom-Up Long-Form Story Generation Using Large Language Modelsarxiv-2510.11618 Sparse Blocked context onlyOct 13, 2025
- Unlocking the Potential of Diffusion Language Models through Template Infillingarxiv-2510.13870 Sparse Blocked context onlyOct 13, 2025
- ShishuLM : Achieving Optimal and Efficient Parameterization with Low Attention Transformer Modelsarxiv-2510.13860 Sparse Blocked context onlyOct 13, 2025
- DropVLA: An Action-Level Backdoor Attack on Vision--Language--Action Modelsarxiv-2510.10932 Sparse Blocked context onlyOct 13, 2025
- DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Modelsarxiv-2510.10846 Sparse Blocked context onlyOct 12, 2025
- Merlin's Whisper: Enabling Efficient Reasoning in Large Language Models via Black-box Persuasive Promptingarxiv-2510.10528 Sparse Blocked context onlyOct 12, 2025
- FML-bench: Benchmarking Machine Learning Agents for Scientific Researcharxiv-2510.10472 Sparse Blocked context onlyOct 12, 2025
- EvoEdit: Evolving Null-space Alignment for Robust and Efficient Knowledge Editingarxiv-2510.13851 Sparse Blocked context onlyOct 11, 2025
- Language steering in latent space to mitigate unintended code-switchingarxiv-2510.13849 Sparse Blocked context onlyOct 11, 2025
- You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMsarxiv-2510.10223 Sparse Blocked context onlyOct 11, 2025
- Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMsarxiv-2510.09885 Sparse Blocked context onlyOct 10, 2025
- GraphMERT: Efficient and Scalable Distillation of Reliable Knowledge Graphs from Unstructured Dataarxiv-2510.09580 Sparse Blocked context onlyOct 10, 2025
- ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answeringarxiv-2510.09351 Sparse Blocked context onlyOct 10, 2025
- ATLAS: Adaptive Trading with LLM AgentS Through Dynamic Prompt Optimization and Multi-Agent Coordinationarxiv-2510.15949 Sparse Blocked context onlyOct 10, 2025
- CLARity: Reasoning Consistency Alone Can Teach Reinforced Expertsarxiv-2510.09278 Sparse Blocked context onlyOct 10, 2025
- Detecting Data Contamination from Reinforcement Learning Post-training for Large Language Modelsarxiv-2510.09259 Sparse Blocked context onlyOct 10, 2025
- Clear Roads, Clear Vision: Advancements in Multi-Weather Restoration for Smart Transportationarxiv-2510.09228 Sparse Blocked context onlyOct 10, 2025
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMsarxiv-2510.09201 Sparse Blocked context onlyOct 10, 2025
- FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMsarxiv-2510.08886 Direct Blocked context onlyOct 10, 2025
- MOSAIC: Multi-agent Orchestration for Task-Intelligent Scientific Codingarxiv-2510.08804 Sparse Blocked context onlyOct 9, 2025
- How Reliable is Language Model Micro-Benchmarking?arxiv-2510.08730 Sparse Blocked context onlyOct 9, 2025
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresightarxiv-2510.08713 Direct Blocked context onlyOct 9, 2025
- DeepPrune: Parallel Scaling without Inter-trace Redundancyarxiv-2510.08483 Sparse Blocked context onlyOct 9, 2025
- If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Modelsarxiv-2510.08388 Sparse Blocked context onlyOct 9, 2025
- LightReasoner: Can Small Language Models Teach Large Language Models Reasoning?arxiv-2510.07962 Sparse Blocked context onlyOct 9, 2025
- ACE: Attribution-Controlled Knowledge Editing for Multi-hop Factual Recallarxiv-2510.07896 Sparse Blocked context onlyOct 9, 2025
- Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energyarxiv-2510.08646 Sparse Blocked context onlyOct 9, 2025
- Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewardsarxiv-2510.07774 Sparse Blocked context onlyOct 9, 2025
- EconCausal: A Context-Aware Causal Reasoning Benchmark for Large Language Models in Social Sciencearxiv-2510.07231 Sparse Blocked context onlyOct 8, 2025
- Search-R3: Unifying Reasoning and Embedding in Large Language Modelsarxiv-2510.07048 Sparse Blocked context onlyOct 8, 2025
- Native Hybrid Attention for Efficient Sequence Modelingarxiv-2510.07019 Sparse Blocked context onlyOct 8, 2025
- FURINA: A Fully Customizable Role-Playing Benchmark via Scalable Multi-Agent Collaboration Pipelinearxiv-2510.06800 Sparse Blocked context onlyOct 8, 2025
- How Language Models Conflate Logical Validity with Plausibility: A Representational Analysis of Content Effectsarxiv-2510.06700 Sparse Blocked context onlyOct 8, 2025
- PIKA: Expert-Level Synthetic Datasets for Post-Training Alignment from Scratcharxiv-2510.06670 Sparse Blocked context onlyOct 8, 2025
- Peeking inside the Black-Box: Reinforcement Learning for Explainable and Accurate Relation Extractionarxiv-2510.06198 Sparse Blocked context onlyOct 7, 2025
- Distributional Semantics Tracing: A Framework for Explaining Hallucinations in Large Language Modelsarxiv-2510.06107 Sparse Blocked context onlyOct 7, 2025
- Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerabilityarxiv-2510.06084 Sparse Blocked context onlyOct 7, 2025
- Prompt reinforcing for long-term planning of large language modelsarxiv-2510.05921 Sparse Blocked context onlyOct 7, 2025
- Early Multimodal Prediction of Cross-Lingual Meme Virality on Reddit: A Time-Window Analysisarxiv-2510.05761 Sparse Blocked context onlyOct 7, 2025
- Revisiting Self-Play Preference Optimization: On the Role of Prompt Difficultyarxiv-2510.05534 Sparse Blocked context onlyOct 7, 2025
- Decoding Partial Differential Equations: Cross-Modal Adaptation of Decoder-only Models to PDEsarxiv-2510.05278 Sparse Blocked context onlyOct 6, 2025
- Slm-mux: Orchestrating small language models for reasoningarxiv-2510.05077 Sparse Blocked context onlyOct 6, 2025
- AtomWorld: A Benchmark for Evaluating Spatial Reasoning in Large Language Models on Crystalline Materialsarxiv-2510.04704 Sparse Blocked context onlyOct 6, 2025
- TiTok: Transfer Token-level Knowledge via Contrastive Excess to Transplant LoRAarxiv-2510.04682 Sparse Blocked context onlyOct 6, 2025
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Modelsarxiv-2510.04618 Sparse Blocked context onlyOct 6, 2025
- LaDiR: Latent Diffusion Enhances LLMs for Text Reasoningarxiv-2510.04573 Sparse Blocked context onlyOct 6, 2025
- Don't Pass@k: A Bayesian Framework for Large Language Model Evaluationarxiv-2510.04265 Sparse Blocked context onlyOct 5, 2025
- AlphaApollo: A System for Deep Agentic Reasoningarxiv-2510.06261 Sparse Blocked context onlyOct 5, 2025
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Environmentsarxiv-2510.04142 Sparse Blocked context onlyOct 5, 2025
- PoLi-RL: A Point-to-List Reinforcement Learning Framework for Conditional Semantic Textual Similarityarxiv-2510.04080 Sparse Blocked context onlyOct 5, 2025
- Slow-Fast Policy Optimization: Reposition-Before-Update for LLM Reasoningarxiv-2510.04072 Sparse Blocked context onlyOct 5, 2025
- What Scales in Cross-Entropy Scaling Law?arxiv-2510.04067 Sparse Blocked context onlyOct 5, 2025
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Modelsarxiv-2510.03721 Sparse Blocked context onlyOct 4, 2025
- Token Hidden Reward: Steering Exploration-Exploitation in Group Relative Deep Reinforcement Learningarxiv-2510.03669 Sparse Blocked context onlyOct 4, 2025
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operationsarxiv-2510.03666 Sparse Blocked context onlyOct 4, 2025
- AgentHub: A Registry for Discoverable, Verifiable, and Reproducible AI Agentsarxiv-2510.03495 Sparse Blocked context onlyOct 3, 2025