Human Feedback Types
missingNone explicit
No explicit feedback protocol extracted.
"This paper introduces the task of analytical question answering over large, semi-structured document collections."
HFEPX · Eval paper review
Zhanli Li, Yixuan Cao, Lvzhou Luo, Ping Luo
Published
Apr 24, 2026
Citations
0
Trust level
Moderate
Usefulness score
25/100 (Low)
Extraction confidence
55% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Apr 24, 2026
This paper is adjacent to HFEPX scope and is best used for background context, not as a primary protocol reference.
Use this for comparison and orientation, not as your only source.
Best use
Background context only
Use if you need
A benchmark-and-metrics comparison anchor.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Treat as adjacent context, not a core eval-method reference.
If you are doing eval pipeline work, start here
This paper introduces the task of analytical question answering over large, semi-structured document collections. We present MuDABench, a benchmark for multi-document analytical QA, where questions require extracting and synthesizing information across numerous documents to perform quantitative analysis. Unlike existing multi-document QA benchmarks that typically require information from only a few documents with limited cross-document reasoning, MuDABench demands extensive inter-document analysis and aggregation. Constructed via distant supervision by leveraging document-level metadata and annotated financial databases, MuDABench comprises over 80,000 pages and 332 analytical QA instances. We also propose an evaluation protocol that measures final answer accuracy and uses intermediate-fact coverage as an auxiliary diagnostic signal for the reasoning process. Experiments reveal that standard RAG systems, which treat all documents as a flat retrieval pool, perform poorly. To address these limitations, we propose a multi-agent workflow that orchestrates planning, extraction, and code generation modules. While this approach substantially improves both process and outcome metrics, a significant gap remains compared to human expert performance. Our analysis identifies two primary bottlenecks: single-document information extraction accuracy and insufficient domain-specific knowledge in current systems. MuDABench is available at https://github.com/Zhanli-Li/MuDABench.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
None explicit
No explicit feedback protocol extracted.
"This paper introduces the task of analytical question answering over large, semi-structured document collections."
Automatic Metrics
Includes extracted eval setup.
"This paper introduces the task of analytical question answering over large, semi-structured document collections."
Not reported
No explicit QC controls found.
"This paper introduces the task of analytical question answering over large, semi-structured document collections."
Mudabench
Useful for quick benchmark comparison.
"We present MuDABench, a benchmark for multi-document analytical QA, where questions require extracting and synthesizing information across numerous documents to perform quantitative analysis."
Accuracy
Useful for evaluation criteria comparison.
"We also propose an evaluation protocol that measures final answer accuracy and uses intermediate-fact coverage as an auxiliary diagnostic signal for the reasoning process."
Domain Experts
Helpful for staffing comparability.
"While this approach substantially improves both process and outcome metrics, a significant gap remains compared to human expert performance."
This paper introduces the task of analytical question answering over large, semi-structured document collections.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
No explicit human feedback protocol detected.
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
Detected: Mudabench
Metric reporting is present
Detected: accuracy