300 canonical paper links on this archive page.
- Interpreto: An Explainability Library for Transformersarxiv-2512.09730 Direct Blocked context onlyDec 10, 2025
- ODMA: On-Demand Memory Allocation Strategy for LLM Serving on LPDDR-Class Acceleratorsarxiv-2512.09427 Sparse Blocked context onlyDec 10, 2025
- What Triggers my Model? Contrastive Explanations Inform Gender Choices by Translation Modelsarxiv-2512.08440 Sparse Blocked context onlyDec 9, 2025
- ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Accessarxiv-2512.08193 Sparse Blocked context onlyDec 9, 2025
- Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safetyarxiv-2512.06227 Sparse Blocked context onlyDec 6, 2025
- Enhancing Retrieval-Augmented Generation with Entity Linking for Educational Platformsarxiv-2512.05967 Sparse Blocked context onlyDec 5, 2025
- SignRoundV2: Toward Closing the Performance Gap in Extremely Low-Bit Post-Training Quantization for LLMsarxiv-2512.04746 Sparse Blocked context onlyDec 4, 2025
- Randomized Masked Finetuning: An Efficient Way to Mitigate Memorization of PIIs in LLMsarxiv-2512.03310 Sparse Blocked context onlyDec 2, 2025
- promptolution: A Unified, Modular Framework for Prompt Optimizationarxiv-2512.02840 Sparse Blocked context onlyDec 2, 2025
- PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Modelsarxiv-2512.02764 Sparse Blocked context onlyDec 2, 2025
- From Veracity to Diffusion: Adressing Operational Challenges in Moving From Fake-News Detection to Information Disordersarxiv-2512.02552 Sparse Blocked context onlyDec 2, 2025
- Diffusion Model in Latent Space for Medical Image Segmentation Taskarxiv-2512.01292 Sparse Blocked context onlyDec 1, 2025
- Epistemic Bias Injection: Biasing LLMs via Selective Context Retrievalarxiv-2512.00804 Sparse Blocked context onlyNov 30, 2025
- Writing in Symbiosis: Mapping Human Creative Agency in the AI Eraarxiv-2512.13697 Sparse Blocked context onlyNov 28, 2025
- When Models Fabricate Credentials: Measuring How Professional Identity Suppresses Honest Self-Representationarxiv-2511.21569 Sparse Blocked context onlyNov 26, 2025
- A Systematic Study of In-the-Wild Model Merging for Large Language Modelsarxiv-2511.21437 Sparse Blocked context onlyNov 26, 2025
- Steering Awareness: Models Can Be Trained to Detect Activation Steeringarxiv-2511.21399 Sparse Blocked context onlyNov 26, 2025
- PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmarkarxiv-2511.21285 Sparse Blocked context onlyNov 26, 2025
- RosettaSpeech: Zero-Shot Speech-to-Speech Translation without Parallel Speecharxiv-2511.20974 Sparse Blocked context onlyNov 26, 2025
- RefTr: Recurrent Refinement of Confluent Trajectories for 3D Vascular Tree Centerlinesarxiv-2511.20823 Sparse Blocked context onlyNov 25, 2025
- A Machine Learning Approach for Detection of Mental Health Conditions and Cyberbullying from Social Mediaarxiv-2511.20001 Sparse Blocked context onlyNov 25, 2025
- Open-weight genome language model safeguards: Assessing robustness via adversarial fine-tuningarxiv-2511.19299 Sparse Blocked context onlyNov 24, 2025
- Periodic Asynchrony: An On-Policy Approach for Accelerating LLM Reinforcement Learningarxiv-2511.18871 Sparse Blocked context onlyNov 24, 2025
- Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cachearxiv-2511.18811 Sparse Blocked context onlyNov 24, 2025
- Empathetic Cascading Networks: A Multi-Stage Prompting Technique for Reducing Social Biases in Large Language Modelsarxiv-2511.18696 Sparse Blocked context onlyNov 24, 2025
- Hierarchical Dual-Strategy Unlearning for Biomedical and Healthcare Intelligence Using Imperfect and Privacy-Sensitive Medical Dataarxiv-2511.19498 Sparse Blocked context onlyNov 23, 2025
- Point of Order: Action-Aware LLM Persona Modeling for Realistic Civic Simulationarxiv-2511.17813 Sparse Blocked context onlyNov 21, 2025
- MUCH: A Multilingual Claim Hallucination Benchmarkarxiv-2511.17081 Sparse Blocked context onlyNov 21, 2025
- ConCISE: A Reference-Free Conciseness Evaluation Metric for LLM-Generated Answersarxiv-2511.16846 Sparse Blocked context onlyNov 20, 2025
- Based on Data Balancing and Model Improvement for Multi-Label Sentiment Classification Performance Enhancementarxiv-2511.14073 Sparse Blocked context onlyNov 18, 2025
- MedPT: A Massive Medical Question Answering Dataset for Brazilian-Portuguese Speakersarxiv-2511.11878 Sparse Blocked context onlyNov 14, 2025
- MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Modelsarxiv-2511.10262 Sparse Blocked context onlyNov 13, 2025
- LexInstructEval: Lexical Instruction Following Evaluation for Large Language Modelsarxiv-2511.17561 Sparse Blocked context onlyNov 13, 2025
- TransactionGPTarxiv-2511.08939 Sparse Blocked context onlyNov 12, 2025
- iSeal: Encrypted Fingerprinting for Reliable LLM Ownership Verificationarxiv-2511.08905 Sparse Blocked context onlyNov 12, 2025
- Moral Susceptibility and Robustness under Persona Role-Play in Large Language Modelsarxiv-2511.08565 Sparse Blocked context onlyNov 11, 2025
- Benchmarking Educational LLMs with Analytics: A Case Study on Gender Bias in Feedbackarxiv-2511.08225 Sparse Blocked context onlyNov 11, 2025
- Quantizing Whisper-small: How design choices affect ASR performancearxiv-2511.08093 Sparse Blocked context onlyNov 11, 2025
- Information Capacity: Evaluating the Efficiency of Large Language Models via Text Compressionarxiv-2511.08066 Sparse Blocked context onlyNov 11, 2025
- State of the Art in Text Classification for South Slavic Languages: Fine-Tuning or Prompting?arxiv-2511.07989 Sparse Blocked context onlyNov 11, 2025
- Unified Work Embeddings: Contrastive Learning of a Bidirectional Multi-task Rankerarxiv-2511.07969 Sparse Blocked context onlyNov 11, 2025
- SPOT: An Annotated French Corpus and Benchmark for Detecting Critical Interventions in Online Conversationsarxiv-2511.07405 Sparse Blocked context onlyNov 10, 2025
- QUARK: Quantization-Enabled Circuit Sharing for Transformer Acceleration by Exploiting Common Patterns in Nonlinear Operationsarxiv-2511.06767 Sparse Blocked context onlyNov 10, 2025
- Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translationarxiv-2511.06680 Sparse Blocked context onlyNov 10, 2025
- You Had One Job: Per-Task Quantization Using LLMs' Hidden Representationsarxiv-2511.06516 Sparse Blocked context onlyNov 9, 2025
- HatePrototypes: Interpretable and Transferable Representations for Implicit and Explicit Hate Speech Detectionarxiv-2511.06391 Sparse Blocked context onlyNov 9, 2025
- Injecting Falsehoods: Adversarial Man-in-the-Middle Attacks Undermining Factual Recall in LLMsarxiv-2511.05919 Sparse Blocked context onlyNov 8, 2025
- REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgmentarxiv-2511.07458 Sparse Blocked context onlyNov 6, 2025
- Graph-Based Alternatives to LLMs for Human Simulationarxiv-2511.02135 Sparse Blocked context onlyNov 3, 2025
- COFAP: A Universal Framework for COFs Adsorption Prediction through Designed Multi-Modal Extraction and Cross-Modal Synergyarxiv-2511.01946 Sparse Blocked context onlyNov 3, 2025
- Eyes on Target: Gaze-Aware Object Detection in Egocentric Videoarxiv-2511.01237 Sparse Blocked context onlyNov 3, 2025
- Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steeringarxiv-2511.00617 Sparse Blocked context onlyNov 1, 2025
- When Distributions Shifts: Causal Generalization for Low-Resource Languagesarxiv-2510.27512 Sparse Blocked context onlyOct 31, 2025
- Analysing Environmental Efficiency in AI for X-Ray Diagnosisarxiv-2511.07436 Sparse Blocked context onlyOct 31, 2025
- Probability Distributions Computed by Autoregressive Transformersarxiv-2510.27118 Sparse Blocked context onlyOct 31, 2025
- VISTA: Verification In Sequential Turn-based Assessmentarxiv-2510.27052 Sparse Blocked context onlyOct 30, 2025
- Frame Semantic Patterns for Identifying Underreporting of Notifiable Events in Healthcare: The Case of Gender-Based Violencearxiv-2510.26969 Sparse Blocked context onlyOct 30, 2025
- Evontree: Ontology Rule-Guided Self-Evolution of Large Language Modelsarxiv-2510.26683 Sparse Blocked context onlyOct 30, 2025
- SynBullying: A Multi LLM Synthetic Conversational Dataset for Cyberbullying Detectionarxiv-2511.11599 Curated Related Blocked context onlyOct 30, 2025
- Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain Textsarxiv-2510.24541 Sparse Blocked context onlyOct 28, 2025
- LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Dataarxiv-2510.24434 Sparse Blocked context onlyOct 28, 2025
- Quantifying Systemic Vulnerability in the Foundation Model Industryarxiv-2510.23421 Sparse Blocked context onlyOct 27, 2025
- SwiftEmbed: Ultra-Fast Text Embeddings via Static Token Lookup for Real-Time Applicationsarxiv-2510.24793 Sparse Blocked context onlyOct 27, 2025
- Fast-MIA: Efficient and Scalable Membership Inference for LLMsarxiv-2510.23074 Sparse Blocked context onlyOct 27, 2025
- Understanding In-Context Learning Beyond Transformers: An Investigation of State Space and Hybrid Architecturesarxiv-2510.23006 Sparse Blocked context onlyOct 27, 2025
- Rule-Based Explanations for Retrieval-Augmented LLM Systemsarxiv-2510.22689 Sparse Blocked context onlyOct 26, 2025
- From Slides to Chatbots: Enhancing Large Language Models with University Course Materialsarxiv-2510.22272 Sparse Blocked context onlyOct 25, 2025
- Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metricsarxiv-2510.22028 Sparse Blocked context onlyOct 24, 2025
- Steering Evaluation-Aware Language Models to Act Like They Are Deployedarxiv-2510.20487 Sparse Blocked context onlyOct 23, 2025
- Evaluating Latent Knowledge of Public Tabular Datasets in Large Language Modelsarxiv-2510.20351 Sparse Blocked context onlyOct 23, 2025
- Citation Failure: Definition, Analysis and Efficient Mitigationarxiv-2510.20303 Sparse Blocked context onlyOct 23, 2025
- Difficulty-Controllable Multiple-Choice Question Generation Using Large Language Models and Direct Preference Optimizationarxiv-2510.19265 Sparse Blocked context onlyOct 22, 2025
- OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mappingarxiv-2510.18999 Sparse Blocked context onlyOct 21, 2025
- LightMem: Lightweight and Efficient Memory-Augmented Generationarxiv-2510.18866 Sparse Blocked context onlyOct 21, 2025
- CEFR-Annotated WordNet: LLM-Based Proficiency-Guided Semantic Database for Language Learningarxiv-2510.18466 Sparse Blocked context onlyOct 21, 2025
- CoGate-LSTM: Prototype-Guided Feature-Space Gating for Mitigating Gradient Dilution in Imbalanced Toxic Comment Classificationarxiv-2510.17018 Sparse Blocked context onlyOct 19, 2025
- In Generative AI We (Dis)Trust? Computational Analysis of Trust and Distrust in Reddit Discussionsarxiv-2510.16173 Sparse Blocked context onlyOct 17, 2025
- Language Models are Injective and Hence Invertiblearxiv-2510.15511 Sparse Blocked context onlyOct 17, 2025
- From Binary to Bilingual: How the National Weather Service is Using Artificial Intelligence to Develop a Comprehensive Translation Programarxiv-2510.14369 Sparse Blocked context onlyOct 16, 2025
- Understanding the Ability of LLMs to Handle Character-Level Perturbationarxiv-2510.14365 Sparse Blocked context onlyOct 16, 2025
- Assessing Web Search Credibility and Response Groundedness in Chat Assistantsarxiv-2510.13749 Sparse Blocked context onlyOct 15, 2025
- Embedding-Based Context-Aware Rerankerarxiv-2510.13329 Sparse Blocked context onlyOct 15, 2025
- LLM Prompt Duel Optimizer: Efficient Label-Free Prompt Optimizationarxiv-2510.13907 Direct Blocked context onlyOct 14, 2025
- When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detectionarxiv-2510.12476 Sparse Blocked context onlyOct 14, 2025
- Community size rather than grammatical complexity better predicts Large Language Model accuracy in a novel Wug Testarxiv-2510.12463 Sparse Blocked context onlyOct 14, 2025
- Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformationarxiv-2510.11423 Sparse Blocked context onlyOct 13, 2025
- CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysisarxiv-2510.11233 Sparse Blocked context onlyOct 13, 2025
- Detecting Hallucinations in Authentic LLM-Human Interactionsarxiv-2510.10539 Sparse Blocked context onlyOct 12, 2025
- SPG: Sandwiched Policy Gradient for Masked Diffusion Language Modelsarxiv-2510.09541 Sparse Blocked context onlyOct 10, 2025
- MaP: A Unified Framework for Reliable Evaluation of Pre-training Dynamicsarxiv-2510.09295 Sparse Blocked context onlyOct 10, 2025
- Inflated Excellence or True Performance? Rethinking Medical Diagnostic Benchmarks with Dynamic Evaluationarxiv-2510.09275 Sparse Blocked context onlyOct 10, 2025
- A Linguistics-Aware LLM Watermarking via Syntactic Predictabilityarxiv-2510.13829 Sparse Blocked context onlyOct 10, 2025
- CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in Chinaarxiv-2510.08986 Sparse Blocked context onlyOct 10, 2025
- Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Frameworkarxiv-2510.08663 Sparse Blocked context onlyOct 9, 2025
- Emotionally Charged, Logically Blurred: AI-driven Emotional Framing Impairs Human Fallacy Detectionarxiv-2510.09695 Sparse Blocked context onlyOct 9, 2025
- Neuron-Level Analysis of Cultural Understanding in Large Language Modelsarxiv-2510.08284 Sparse Blocked context onlyOct 9, 2025
- Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminabilityarxiv-2510.08138 Sparse Blocked context onlyOct 9, 2025
- Fewer Weights, More Problems: A Practical Attack on LLM Pruningarxiv-2510.07985 Sparse Blocked context onlyOct 9, 2025
- RCPU: Rotation-Constrained Error Compensation for Structured Pruning of Large Language Modelsarxiv-2510.07782 Sparse Blocked context onlyOct 9, 2025
- Biasless Language Models Learn Unnaturally: How LLMs Fail to Distinguish the Possible from the Impossiblearxiv-2510.07178 Sparse Blocked context onlyOct 8, 2025
- TRIM: Token-wise Attention-Derived Saliency for Data-Efficient Instruction Tuningarxiv-2510.07118 Sparse Blocked context onlyOct 8, 2025
- Exposing Citation Vulnerabilities in Generative Enginesarxiv-2510.06823 Sparse Blocked context onlyOct 8, 2025
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policiesarxiv-2510.05725 Sparse Blocked context onlyOct 7, 2025
- SECA: Semantically Equivalent and Coherent Attacks for Eliciting LLM Hallucinationsarxiv-2510.04398 Sparse Blocked context onlyOct 5, 2025
- Large Language Models Hallucination: A Comprehensive Surveyarxiv-2510.06265 Sparse Blocked context onlyOct 5, 2025
- Cache-to-Cache: Direct Semantic Communication Between Large Language Modelsarxiv-2510.03215 Sparse Blocked context onlyOct 3, 2025
- Tree-based Dialogue Reinforced Policy Optimization for Red-Teaming Attacksarxiv-2510.02286 Sparse Blocked context onlyOct 2, 2025
- AccurateRAG: A Framework for Building Accurate Retrieval-Augmented Question-Answering Applicationsarxiv-2510.02243 Sparse Blocked context onlyOct 2, 2025
- Style over Story: Measuring LLM Narrative Preferences via Structured Selectionarxiv-2510.02025 Sparse Blocked context onlyOct 2, 2025
- Energy-Regularized Sequential Model Editing on Hyperspheresarxiv-2510.01172 Sparse Blocked context onlyOct 1, 2025
- KnowledgeSmith: Uncovering Knowledge Updating in LLMs with Model Editing and Unlearningarxiv-2510.02392 Sparse Blocked context onlyOct 1, 2025
- BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responsesarxiv-2510.00232 Sparse Blocked context onlySep 30, 2025
- Bringing Emerging Architectures to Sequence Labeling in NLParxiv-2509.25918 Sparse Blocked context onlySep 30, 2025
- RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivityarxiv-2509.25897 Sparse Blocked context onlySep 30, 2025
- Calibrating Verbalized Confidence with Self-Generated Distractorsarxiv-2509.25532 Sparse Blocked context onlySep 29, 2025
- The Rise of AfricaNLP: A Survey of Contributions, Contributors, Community Impact, and Bibliometric Analysisarxiv-2509.25477 Sparse Blocked context onlySep 29, 2025
- Incentive-Aligned Multi-Source LLM Summariesarxiv-2509.25184 Sparse Blocked context onlySep 29, 2025
- EasySteer: A Unified Framework for High-Performance and Extensible LLM Steeringarxiv-2509.25175 Direct Blocked context onlySep 29, 2025
- Pretraining Large Language Models with NVFP4arxiv-2509.25149 Sparse Blocked context onlySep 29, 2025
- VSSFlow: Unifying Video-conditioned Sound and Speech Generation via Joint Learningarxiv-2509.24773 Sparse Blocked context onlySep 29, 2025
- ProxyAttn: Guided Sparse Attention via Representative Headsarxiv-2509.24745 Sparse Blocked context onlySep 29, 2025
- Building Benchmarks from the Ground Up: Community-Centered Evaluation of LLMs in Healthcare Chatbot Settingsarxiv-2509.24506 Sparse Blocked context onlySep 29, 2025
- SUIT: Knowledge Editing with Subspace-Aware Key-Value Mappingsarxiv-2509.24502 Sparse Blocked context onlySep 29, 2025
- HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessmentarxiv-2509.24384 Sparse Blocked context onlySep 29, 2025
- Model Merging Scaling Laws in Large Language Modelsarxiv-2509.24244 Sparse Blocked context onlySep 29, 2025
- M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generationarxiv-2509.23728 Sparse Blocked context onlySep 28, 2025
- Internal Planning in Language Models: Characterizing Horizon and Branch Awarenessarxiv-2509.25260 Sparse Blocked context onlySep 28, 2025
- Dual-Space Smoothness for Robust and Balanced LLM Unlearningarxiv-2509.23362 Sparse Blocked context onlySep 27, 2025
- PonderLM-2: Pretraining LLM with Latent Thoughts in Continuous Spacearxiv-2509.23184 Sparse Blocked context onlySep 27, 2025
- Test-Time Policy Adaptation for Enhanced Multi-Turn Interactions with LLMsarxiv-2509.23166 Sparse Blocked context onlySep 27, 2025
- Semantic Voting: A Self-Evaluation-Free Approach for Efficient LLM Self-Improvement on Unverifiable Open-ended Tasksarxiv-2509.23067 Sparse Blocked context onlySep 27, 2025
- AutoPK: Leveraging LLMs and a Hybrid Similarity Metric for Advanced Retrieval of Pharmacokinetic Data from Complex Tables and Documentsarxiv-2510.00039 Sparse Blocked context onlySep 26, 2025
- Death of the Novel(ty): Beyond n-Gram Novelty as a Metric for Textual Creativityarxiv-2509.22641 Sparse Blocked context onlySep 26, 2025
- Benefits and Pitfalls of Reinforcement Learning for Language Model Planning: A Theoretical Perspectivearxiv-2509.22613 Sparse Blocked context onlySep 26, 2025
- CoSpaDi: Compressing LLMs via Calibration-Guided Sparse Dictionary Learningarxiv-2509.22075 Sparse Blocked context onlySep 26, 2025
- Fine-tuning Done Right in Model Editingarxiv-2509.22072 Sparse Blocked context onlySep 26, 2025
- Quokka: Accelerating Program Verification with LLMs via Invariant Synthesisarxiv-2509.21629 Sparse Blocked context onlySep 25, 2025
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Trainingarxiv-2509.21500 Sparse Blocked context onlySep 25, 2025
- Frequency-Aware Model Parameter Explorer: A new attribution method for improving explainabilityarxiv-2510.03245 Sparse Blocked context onlySep 25, 2025
- PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraintsarxiv-2509.21057 Sparse Blocked context onlySep 25, 2025
- OLaPh: Optimal Language Phonemizerarxiv-2509.20086 Sparse Blocked context onlySep 24, 2025
- Redundancy-as-Masking: Formalizing the Artificial Age Score (AAS) to Model Memory Aging in Generative AIarxiv-2510.01242 Sparse Blocked context onlySep 24, 2025
- Diversity Boosts AI-Generated Text Detectionarxiv-2509.18880 Sparse Blocked context onlySep 23, 2025
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generationarxiv-2509.18600 Sparse Blocked context onlySep 23, 2025
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For Perplexityarxiv-2509.18577 Sparse Blocked context onlySep 23, 2025
- Similarity Field Theory: A Mathematical Framework for Intelligencearxiv-2509.18218 Sparse Blocked context onlySep 21, 2025
- Llama-Mimi: Exploring the Limits of Flattened Speech Language Modelingarxiv-2509.14882 Sparse Blocked context onlySep 18, 2025
- Neural-Quantum-States Impurity Solver for Quantum Embedding Problemsarxiv-2509.12431 Sparse Blocked context onlySep 15, 2025
- The AI Memory Gap: Users Misremember What They Created With AI or Withoutarxiv-2509.11851 Sparse Blocked context onlySep 15, 2025
- Incongruent Positivity: When Miscalibrated Positivity Undermines Online Supportive Conversationsarxiv-2509.10184 Sparse Blocked context onlySep 12, 2025
- DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matchingarxiv-2509.09631 Sparse Blocked context onlySep 11, 2025
- OTESGN: Optimal Transport-Enhanced Syntactic-Semantic Graph Networks for Aspect-Based Sentiment Analysisarxiv-2509.08612 Sparse Blocked context onlySep 10, 2025
- SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledgearxiv-2509.07968 Sparse Blocked context onlySep 9, 2025
- FHIR-RAG-MEDS: Integrating HL7 FHIR with Retrieval-Augmented Large Language Models for Enhanced Medical Decision Supportarxiv-2509.07706 Sparse Blocked context onlySep 9, 2025
- LLM Analysis of 150+ years of German Parliamentary Debates on Migration Reveals Shift from Post-War Solidarity to Anti-Solidarity in the Last Decadearxiv-2509.07274 Sparse Blocked context onlySep 8, 2025
- BinaryShield: Cross-Service Threat Intelligence in LLM Services using Privacy-Preserving Fingerprintsarxiv-2509.05608 Sparse Blocked context onlySep 6, 2025
- No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadataarxiv-2509.05425 Sparse Blocked context onlySep 5, 2025
- From Editor to Dense Geometry Estimatorarxiv-2509.04338 Sparse Blocked context onlySep 4, 2025
- MultiWikiQA: A Reading Comprehension Benchmark in 300+ Languagesarxiv-2509.04111 Sparse Blocked context onlySep 4, 2025
- From Construction to Injection: Edit-Based Fingerprints for Large Language Modelsarxiv-2509.03122 Sparse Blocked context onlySep 3, 2025
- Agri-Query: A Case Study on RAG vs. Long-Context LLMs for Cross-Lingual Technical Question Answeringarxiv-2508.18093 Sparse Blocked context onlyAug 25, 2025
- MedRepBench: A Comprehensive Benchmark for Medical Report Interpretationarxiv-2508.16674 Direct Blocked context onlyAug 21, 2025
- AmbiSQL: Interactive Ambiguity Detection and Resolution for Text-to-SQLarxiv-2508.15276 Sparse Blocked context onlyAug 21, 2025
- Quantization Meets dLLMs: A Systematic Study of Post-training Quantization for Diffusion LLMsarxiv-2508.14896 Sparse Blocked context onlyAug 20, 2025
- REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Modelsarxiv-2508.14390 Sparse Blocked context onlyAug 20, 2025
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulationarxiv-2508.13998 Sparse Blocked context onlyAug 19, 2025
- Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimizationarxiv-2508.13993 Sparse Blocked context onlyAug 19, 2025
- The Collaboration Paradox: Why Generative AI Requires Both Strategic Intelligence and Operational Stability in Supply Chain Managementarxiv-2508.13942 Sparse Blocked context onlyAug 19, 2025
- Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Explorationarxiv-2508.13755 Sparse Blocked context onlyAug 19, 2025
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generationarxiv-2508.13587 Sparse Blocked context onlyAug 19, 2025
- TASER: Table Agents for Schema-guided Extraction and Recommendationarxiv-2508.13404 Sparse Blocked context onlyAug 18, 2025
- TaoSR1: The Thinking Model for E-commerce Relevance Searcharxiv-2508.12365 Sparse Blocked context onlyAug 17, 2025
- CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressuresarxiv-2508.11915 Sparse Blocked context onlyAug 16, 2025
- SCOPE: A Generative Approach for LLM Prompt Compressionarxiv-2508.15813 Sparse Blocked context onlyAug 16, 2025
- SafeSieve: From Heuristics to Experience in Progressive Pruning for LLM-based Multi-Agent Communicationarxiv-2508.11733 Sparse Blocked context onlyAug 15, 2025
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasksarxiv-2508.11360 Sparse Blocked context onlyAug 15, 2025
- Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rulesarxiv-2508.10971 Sparse Blocked context onlyAug 14, 2025
- Agentic Design Review Systemarxiv-2508.10745 Sparse Blocked context onlyAug 14, 2025
- GenOM: Ontology Matching with Description Generation and Large Language Modelarxiv-2508.10703 Sparse Blocked context onlyAug 14, 2025
- For-Value: Efficient Forward-Only Data Valuation for finetuning LLMs and VLMsarxiv-2508.10180 Sparse Blocked context onlyAug 13, 2025
- From Context to Intent: Reasoning-Guided Function-Level Code Completionarxiv-2508.09537 Sparse Blocked context onlyAug 13, 2025
- PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoningarxiv-2508.09521 Sparse Blocked context onlyAug 13, 2025
- IAG: Input-aware Backdoor Attack on VLM-based Visual Groundingarxiv-2508.09456 Sparse Blocked context onlyAug 13, 2025
- SegDAC: Visual Generalization in Reinforcement Learning via Dynamic Object Tokensarxiv-2508.09325 Sparse Blocked context onlyAug 12, 2025
- Feedback-Driven Tool-Use Improvements in Large Language Models via Automated Build Environmentsarxiv-2508.08791 Sparse Blocked context onlyAug 12, 2025
- Can LLMs Detect Their Confabulations? Estimating Reliability in Uncertainty-Aware Language Modelsarxiv-2508.08139 Sparse Blocked context onlyAug 11, 2025
- DIVER: A Multi-Stage Approach for Reasoning-intensive Information Retrievalarxiv-2508.07995 Direct Blocked context onlyAug 11, 2025
- 1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoningarxiv-2508.07667 Sparse Blocked context onlyAug 11, 2025
- Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimizationarxiv-2508.07629 Sparse Blocked context onlyAug 11, 2025
- SEVADE: Self-Evolving Multi-Agent Analysis with Decoupled Evaluation for Hallucination-Resistant Irony Detectionarxiv-2508.06803 Sparse Blocked context onlyAug 9, 2025
- Memp: Exploring Agent Procedural Memoryarxiv-2508.06433 Sparse Blocked context onlyAug 8, 2025
- UR$^2$: Unify RAG and Reasoning through Reinforcement Learningarxiv-2508.06165 Sparse Blocked context onlyAug 8, 2025
- EvolvR: Self-Evolving Pairwise Reasoning for Story Evaluation to Enhance Generationarxiv-2508.06046 Sparse Blocked context onlyAug 8, 2025
- GroundAct: Can LLM Agents Ground Actions in Environmental States?arxiv-2508.05614 Sparse Blocked context onlyAug 7, 2025
- MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policyarxiv-2508.05592 Sparse Blocked context onlyAug 7, 2025
- Not All Errors Are Created Equal: ASCoT Addresses Late-Stage Fragility in Efficient LLM Reasoningarxiv-2508.05282 Sparse Blocked context onlyAug 7, 2025
- TURA: Tool-Augmented Unified Retrieval Agent for AI Searcharxiv-2508.04604 Sparse Blocked context onlyAug 6, 2025
- Share Your Attention: Transformer Weight Sharing via Matrix-based Dictionary Learningarxiv-2508.04581 Sparse Blocked context onlyAug 6, 2025
- ReasoningGuard: Safeguarding Large Reasoning Models with Inference-time Safety Aha Momentsarxiv-2508.04204 Sparse Blocked context onlyAug 6, 2025
- ToolGrad: Efficient Tool-use Dataset Generation with Textual "Gradients"arxiv-2508.04086 Sparse Blocked context onlyAug 6, 2025
- CoAct-1: Computer-using Multi-Agent System with Coding Actionsarxiv-2508.03923 Sparse Blocked context onlyAug 5, 2025
- MolReasoner: Toward Effective and Interpretable Reasoning for Molecular LLMsarxiv-2508.02066 Sparse Blocked context onlyAug 4, 2025
- A Theory of Adaptive Scaffolding for LLM-Based Pedagogical Agentsarxiv-2508.01503 Sparse Blocked context onlyAug 2, 2025
- Towards Efficient Medical Reasoning with Minimal Fine-Tuning Dataarxiv-2508.01450 Sparse Blocked context onlyAug 2, 2025
- Cognitive Kernel-Pro: A Framework for Deep Research Agents and Agent Foundation Models Trainingarxiv-2508.00414 Direct Blocked context onlyAug 1, 2025
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimizationarxiv-2508.00222 Sparse Blocked context onlyJul 31, 2025
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time Groundingarxiv-2507.22025 Sparse Blocked context onlyJul 29, 2025
- Soft Head Selection for Injecting ICL-Derived Task Embeddingsarxiv-2507.20906 Sparse Blocked context onlyJul 28, 2025
- MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphsarxiv-2507.20804 Sparse Blocked context onlyJul 28, 2025
- RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Examsarxiv-2507.19666 Curated Related Blocked context onlyJul 25, 2025
- Shop-R1: Rewarding LLMs to Simulate Human Behavior in Online Shopping via Reinforcement Learningarxiv-2507.17842 Sparse Blocked context onlyJul 23, 2025
- CASCADE: LLM-Powered JavaScript Deobfuscator at Googlearxiv-2507.17691 Curated Related Blocked context onlyJul 23, 2025
- SimLens for Early Exit in Large Language Models: Eliciting Accurate Latent Predictions with One More Tokenarxiv-2507.17618 Sparse Blocked context onlyJul 23, 2025
- Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challengearxiv-2507.17288 Sparse Blocked context onlyJul 23, 2025
- Decoding Translation-Related Functional Sequences in 5'UTRs Using Interpretable Deep Learning Modelsarxiv-2507.16801 Sparse Blocked context onlyJul 22, 2025
- SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reportingarxiv-2507.16145 Sparse Blocked context onlyJul 22, 2025
- Position: Reasoning After Perception Means Reasoning Without Visionarxiv-2507.16863 Sparse Blocked context onlyJul 21, 2025
- Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detectionarxiv-2507.16861 Sparse Blocked context onlyJul 21, 2025
- Learning to Extract Rational Evidence via Reinforcement Learning for Retrieval-Augmented Generationarxiv-2507.15586 Sparse Blocked context onlyJul 21, 2025
- WebShaper: Agentically Data Synthesizing via Information-Seeking Formalizationarxiv-2507.15061 Sparse Blocked context onlyJul 20, 2025
- MUR: Momentum Uncertainty guided Reasoning for Large Language Modelsarxiv-2507.14958 Sparse Blocked context onlyJul 20, 2025
- Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speecharxiv-2507.13563 Direct Blocked context onlyJul 17, 2025
- QuestA: Expanding Reasoning Capacity in LLMs via Question Augmentationarxiv-2507.13266 Sparse Blocked context onlyJul 17, 2025
- Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verificationarxiv-2507.11662 Sparse Blocked context onlyJul 15, 2025
- OPENXRD: A Comprehensive Benchmark Framework for LLM/MLLM XRD Question Answeringarxiv-2507.09155 Sparse Blocked context onlyJul 12, 2025
- ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMsarxiv-2507.10593 Direct Blocked context onlyJul 11, 2025
- Knowledge Fusion via Bidirectional Information Aggregationarxiv-2507.08704 Sparse Blocked context onlyJul 11, 2025
- FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answeringarxiv-2507.07634 Sparse Blocked context onlyJul 10, 2025
- SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMsarxiv-2507.07610 Sparse Blocked context onlyJul 10, 2025
- Reward Models Can Improve Themselves: Reward-Guided Adversarial Failure Mode Discovery for Robust Reward Modelingarxiv-2507.06419 Sparse Blocked context onlyJul 8, 2025
- Optimus: A Robust Defense Framework for Mitigating Toxicity while Fine-Tuning Conversational AIarxiv-2507.05660 Sparse Blocked context onlyJul 8, 2025
- The Generalization Ridge: Information Flow in Natural Language Generationarxiv-2507.05387 Sparse Blocked context onlyJul 7, 2025
- Agentic Vehicles for Human-Centered Mobilityarxiv-2507.04996 Sparse Blocked context onlyJul 7, 2025
- STRUCTSENSE: A Task-Agnostic Agentic Framework for Structured Information Extraction with Human-In-The-Loop Evaluation and Benchmarkingarxiv-2507.03674 Direct Blocked context onlyJul 4, 2025
- GDGB: A Benchmark for Generative Dynamic Text-Attributed Graph Learningarxiv-2507.03267 Sparse Blocked context onlyJul 4, 2025
- Multimodal Mathematical Reasoning with Diverse Solving Perspectivearxiv-2507.02804 Sparse Blocked context onlyJul 3, 2025
- Intrinsic Fingerprint of LLMs: Continue Training is NOT All You Need to Steal A Model!arxiv-2507.03014 Sparse Blocked context onlyJul 2, 2025
- LEDOM: Reverse Language Modelarxiv-2507.01335 Sparse Blocked context onlyJul 2, 2025
- TransLaw: A Large-Scale Dataset and Multi-Agent Benchmark Simulating Professional Translation of Hong Kong Case Lawarxiv-2507.00875 Sparse Blocked context onlyJul 1, 2025
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignmentarxiv-2507.02768 Direct Blocked context onlyJul 1, 2025
- WebSailor: Navigating Super-human Reasoning for Web Agentarxiv-2507.02592 Sparse Blocked context onlyJul 1, 2025
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential Privacyarxiv-2507.02974 Sparse Blocked context onlyJun 30, 2025
- SPIRAL: Self-Play on Zero-Sum Games Incentivizes Reasoning via Multi-Agent Multi-Turn Reinforcement Learningarxiv-2506.24119 Sparse Blocked context onlyJun 30, 2025
- Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaborationarxiv-2507.02935 Sparse Blocked context onlyJun 26, 2025
- Complexity-aware fine-tuningarxiv-2506.21220 Sparse Blocked context onlyJun 26, 2025
- Multi-lingual Functional Evaluation for Large Language Modelsarxiv-2506.20793 Sparse Blocked context onlyJun 25, 2025
- $π$-CoT: Prolog-Initialized Chain-of-Thought Prompting for Multi-Hop Question-Answeringarxiv-2506.20642 Sparse Blocked context onlyJun 25, 2025
- KnowRL: Exploring Knowledgeable Reinforcement Learning for Factualityarxiv-2506.19807 Direct Blocked context onlyJun 24, 2025
- Breaking Barriers: Do Reinforcement Post Training Gains Transfer To Unseen Domains?arxiv-2506.19733 Sparse Blocked context onlyJun 24, 2025
- NaviAgent: Graph-Driven Bilevel Planning for Scalable Tool Orchestrationarxiv-2506.19500 Sparse Blocked context onlyJun 24, 2025
- LongWriter-Zero: Mastering Ultra-Long Text Generation via Reinforcement Learningarxiv-2506.18841 Direct Blocked context onlyJun 23, 2025
- Parallel Continuous Chain-of-Thought with Jacobi Iterationarxiv-2506.18582 Sparse Blocked context onlyJun 23, 2025
- PDF Retrieval Augmented Question Answeringarxiv-2506.18027 Sparse Blocked context onlyJun 22, 2025
- LLM Probability Concentration: How Alignment Shrinks the Generative Horizonarxiv-2506.17871 Sparse Blocked context onlyJun 22, 2025
- PersonalAI: A Systematic Comparison of Knowledge Graph Storage and Retrieval Approaches for Personalized LLM agentsarxiv-2506.17001 Sparse Blocked context onlyJun 20, 2025
- A Scoping Review of Synthetic Data Generation by Language Models in Biomedical Research and Application: Data Utility and Quality Perspectivesarxiv-2506.16594 Sparse Blocked context onlyJun 19, 2025
- Revela: Dense Retriever Learning via Language Modelingarxiv-2506.16552 Sparse Blocked context onlyJun 19, 2025
- When Does Divide and Conquer Work for Long Context LLM? A Noise Decomposition Frameworkarxiv-2506.16411 Sparse Blocked context onlyJun 19, 2025
- OJBench: A Competition Level Code Benchmark For Large Language Modelsarxiv-2506.16395 Direct Blocked context onlyJun 19, 2025
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Modelsarxiv-2506.15681 Sparse Blocked context onlyJun 18, 2025
- SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modellingarxiv-2506.15498 Sparse Blocked context onlyJun 18, 2025
- DeVisE: Behavioral Testing of Medical Large Language Modelsarxiv-2506.15339 Sparse Blocked context onlyJun 18, 2025
- Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problemarxiv-2506.17299 Sparse Blocked context onlyJun 17, 2025
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loopsarxiv-2506.14493 Sparse Blocked context onlyJun 17, 2025
- RedTopic: Toward Topic-Diverse Red Teaming of Large Language Modelsarxiv-2507.00026 Sparse Blocked context onlyJun 17, 2025
- AgentSynth: Scalable Task Generation for Generalist Computer-Use Agentsarxiv-2506.14205 Sparse Blocked context onlyJun 17, 2025
- Instruction Following by Principled Boosting Attention of Large Language Modelsarxiv-2506.13734 Sparse Blocked context onlyJun 16, 2025
- Language Agents for Hypothesis-driven Clinical Decision Making with Reinforcement Learningarxiv-2506.13474 Sparse Blocked context onlyJun 16, 2025
- DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulationarxiv-2506.13285 Sparse Blocked context onlyJun 16, 2025
- MemeMind: A Large-Scale Multimodal Dataset with Chain-of-Thought Reasoning for Harmful Meme Detectionarxiv-2506.18919 Sparse Blocked context onlyJun 15, 2025
- $\texttt{SPECS}$: Faster Test-Time Scaling through Speculative Draftsarxiv-2506.15733 Sparse Blocked context onlyJun 15, 2025
- From Time Series Analysis to Question Answering: A Survey in the LLM Eraarxiv-2506.11512 Sparse Blocked context onlyJun 13, 2025
- Spurious Rewards: Rethinking Training Signals in RLVRarxiv-2506.10947 Sparse Blocked context onlyJun 12, 2025
- VINCIE: Unlocking In-context Image Editing from Videoarxiv-2506.10941 Sparse Blocked context onlyJun 12, 2025
- Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden Principlesarxiv-2506.10848 Sparse Blocked context onlyJun 12, 2025
- Query-Level Uncertainty in Large Language Modelsarxiv-2506.09669 Sparse Blocked context onlyJun 11, 2025
- Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithmsarxiv-2506.09457 Sparse Blocked context onlyJun 11, 2025
- Structure-Augmented Reasoning Generationarxiv-2506.08364 Sparse Blocked context onlyJun 10, 2025
- Not All Tokens Matter: Towards Efficient LLM Reasoning via Token Significance in Reinforcement Learningarxiv-2506.08125 Sparse Blocked context onlyJun 9, 2025
- AbstRaL: Augmenting LLMs' Reasoning by Reinforcing Abstract Thinkingarxiv-2506.07751 Sparse Blocked context onlyJun 9, 2025
- Learning What Reinforcement Learning Can't: Interleaved Online Fine-Tuning for Hardest Questionsarxiv-2506.07527 Sparse Blocked context onlyJun 9, 2025
- When Style Breaks Safety: Defending LLMs Against Superficial Style Alignmentarxiv-2506.07452 Sparse Blocked context onlyJun 9, 2025
- Exploring Effective Strategies for Building a User-Configured GPT for Coding Classroom Dialoguesarxiv-2506.07194 Sparse Blocked context onlyJun 8, 2025
- A dependently-typed calculus of event telicity and culminativityarxiv-2506.06968 Sparse Blocked context onlyJun 8, 2025
- BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic Reasoningarxiv-2506.06955 Sparse Blocked context onlyJun 8, 2025
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generationarxiv-2506.06251 Direct Blocked context onlyJun 6, 2025
- Can Theoretical Physics Research Benefit from Language Agents?arxiv-2506.06214 Sparse Blocked context onlyJun 6, 2025
- Comparative Analysis of Modern Machine Learning Models for Retail Sales Forecastingarxiv-2506.05941 Sparse Blocked context onlyJun 6, 2025
- Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Modelsarxiv-2506.05850 Sparse Blocked context onlyJun 6, 2025
- When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generationarxiv-2506.05690 Sparse Blocked context onlyJun 6, 2025
- OPeRA: A Dataset of Observation, Persona, Rationale, and Action for Evaluating LLMs on Human Online Shopping Behavior Simulationarxiv-2506.05606 Sparse Blocked context onlyJun 5, 2025
- MMTU: A Massive Multi-Task Table Understanding and Reasoning Benchmarkarxiv-2506.05587 Direct Blocked context onlyJun 5, 2025
- Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replayarxiv-2506.05316 Sparse Blocked context onlyJun 5, 2025
- Resisting Contextual Interference in RAG via Parametric-Knowledge Reinforcementarxiv-2506.05154 Sparse Blocked context onlyJun 5, 2025
- The NTNU System at the S&I Challenge 2025 SLA Open Trackarxiv-2506.05121 Sparse Blocked context onlyJun 5, 2025
- Toward Automated Robustness Evaluation of Mathematical Reasoningarxiv-2506.05038 Sparse Blocked context onlyJun 5, 2025
- Sensory-Motor Control with Large Language Models via Iterative Policy Refinementarxiv-2506.04867 Sparse Blocked context onlyJun 5, 2025
- Accelerated Test-Time Scaling with Model-Free Speculative Samplingarxiv-2506.04708 Sparse Blocked context onlyJun 5, 2025
- "Don't Do That!": Guiding Embodied Systems through Large Language Model-based Constraint Generationarxiv-2506.04500 Sparse Blocked context onlyJun 4, 2025
- High Accuracy, Less Talk (HALT): Reliable LLMs through Capability-Aligned Finetuningarxiv-2506.04051 Sparse Blocked context onlyJun 4, 2025