ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering

Daeyong Kwon, SeungHeon Doh, Juhan Nam · Dec 5, 2025 · Citations: 0

Abstract

Recent advances in large language models (LLMs) have transformed open-domain question answering, yet their effectiveness in music-related reasoning remains limited due to sparse music knowledge in pretraining data. While music information retrieval and computational musicology have explored structured and multimodal understanding, few resources support factual and contextual music question answering (MQA) grounded in artist metadata or historical context. We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA. Experiments show that RAG markedly improves factual accuracy; open-source models gain up to +56.8 percentage points (for example, Qwen3 8B improves from 35.0 to 91.8), approaching proprietary model performance. RAG-style fine-tuning further boosts both factual recall and contextual reasoning, improving results on both in-domain and out-of-domain benchmarks. MusWikiDB also yields approximately 6 percentage points higher accuracy and 40% faster retrieval than a general-purpose Wikipedia corpus. We release MusWikiDB and ArtistMus to advance research in music information retrieval and domain-specific question answering, establishing a foundation for retrieval-augmented reasoning in culturally rich domains such as music.

HFEPX Relevance Assessment

This paper appears adjacent to HFEPX scope (human-feedback/eval), but does not show strong direct protocol evidence in metadata/abstract.

Eval-Fit Score

5/100 • Low

Treat as adjacent context, not a core eval-method reference.

Human Feedback Signal

Not explicit in abstract metadata

Evaluation Signal

Detected

HFEPX Fit

Adjacent candidate

If you are doing eval pipeline work, start here:

Human Eval Hub LLM-as-Judge Hub Pairwise Preference Hub Tool-Use Eval Hub

Human Data Lens

Uses human feedback: No
Feedback types: None
Rater population: Unknown
Unit of annotation: Unknown
Expertise required: General
Extraction source: Persisted extraction

Evaluation Lens

Evaluation modes: Automatic Metrics
Agentic eval: None
Quality controls: Not reported
Confidence: 0.45
Flags: low_signal, possible_false_positive

Protocol And Measurement Signals

Benchmarks / Datasets

No benchmark or dataset names were extracted from the available abstract.

Reported Metrics

accuracyrecall

Research Brief

Deterministic synthesis

We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. HFEPX signals include Automatic Metrics with confidence 0.45. Updated from current HFEPX corpus.

Generated Mar 4, 2026, 3:33 PM · Grounded in abstract + metadata only

Key Takeaways

We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with…
These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA.

Researcher Actions

Treat this as method context, then pivot to protocol-specific HFEPX hubs.
Identify benchmark choices from full text before operationalizing conclusions.
Validate metric comparability (accuracy, recall).

Caveats

Generated from title, abstract, and extracted metadata only; full-paper implementation details are not parsed.
Low-signal flag detected: protocol relevance may be indirect.

Recommended Queries

human-eval protocol design pairwise preference data quality inter-rater agreement adjudication

Research Summary

Contribution Summary

We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic.
These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA.
RAG-style fine-tuning further boosts both factual recall and contextual reasoning, improving results on both in-domain and out-of-domain benchmarks.

Why It Matters For Eval

We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic.
These resources enable systematic evaluation of retrieval-augmented generation (RAG) for MQA.

Researcher Checklist

Gap: Human feedback protocol is explicit

No explicit human feedback protocol detected.
Pass: Evaluation mode is explicit

Detected: Automatic Metrics
Gap: Quality control reporting appears

No calibration/adjudication/IAA control explicitly detected.
Gap: Benchmark or dataset anchors are present

No benchmark/dataset anchor extracted from abstract.
Pass: Metric reporting is present

Detected: accuracy, recall

Category-Adjacent Papers (Broader Context)

These papers are nearby in arXiv category and useful for broader context, but not necessarily protocol-matched to this paper.

FlashEvaluator: Expanding Search Space with Parallel Evaluation Category Neighbor

Citations: 0 Relevance: 6.05
- Shared arXiv category (cs.CL, cs.IR)
- Shared metric mentions
- Shared terminology (accuracy)
Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration Category Neighbor

Citations: 0 Relevance: 4.45
- Shared arXiv category (cs.CL, cs.AI)
- Shared metric mentions
- Shared terminology (accuracy)
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs Category Neighbor

Citations: 0 Relevance: 4.45
- Shared arXiv category (cs.CL, cs.AI)
- Shared metric mentions
- Shared terminology (accuracy)
Contextualized Privacy Defense for LLM Agents Category Neighbor

Citations: 0 Relevance: 3.20
- Shared arXiv category (cs.CL, cs.AI)
Guideline-Grounded Evidence Accumulation for High-Stakes Agent Verification Category Neighbor

Citations: 0 Relevance: 3.20
- Shared arXiv category (cs.CL, cs.AI)
Cross-Family Speculative Prefill: Training-Free Long-Context Compression with Small Draft Models Category Neighbor

Citations: 0 Relevance: 2.85
- Shared arXiv category (cs.CL)
- Shared metric mentions
- Shared terminology (accuracy)
Think, But Don't Overthink: Reproducing Recursive Language Models Category Neighbor

Citations: 0 Relevance: 2.85
- Shared arXiv category (cs.CL)
- Shared metric mentions
- Shared terminology (accuracy)
Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory Category Neighbor

Citations: 0 Relevance: 2.50
- Shared arXiv category (cs.CL)
- Shared terminology (diverse, question)

Need human evaluators for your AI research? Scale annotation with expert AI Trainers.

Post a Job Get a Quote