Human Feedback Types
strongExpert Verification
Directly usable for protocol triage.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
HFEPX · Eval paper review
Mohammad Mansoori, Amira Soliman, Farzaneh Etminani
Published
Oct 15, 2025
Citations
0
Trust level
Moderate
Usefulness score
65/100 (Medium)
Extraction confidence
70% (Moderate)
Derived from extracted protocol signals and abstract evidence.
Rater population
Domain Experts
Signals refreshed
Aug 21, 2026
This paper has useful evaluation signal, but protocol completeness is partial; pair it with related papers before deciding implementation strategy.
Use this for comparison and orientation, not as your only source.
Best use
Secondary protocol comparison source
Use if you need
A secondary eval reference to pair with stronger protocol papers.
What to verify
Validate the evaluation procedure and quality controls in the full paper before operational use.
Main weakness
No major weakness surfaced.
Useful as a secondary reference; validate protocol details against neighboring papers.
If you are doing eval pipeline work, start here
Clinical notes contain unstructured text provided by clinicians during patient encounters. These notes are usually accompanied by a sequence of diagnostic codes following the International Classification of Diseases (ICD). Correctly assigning and ordering ICD codes is essential for medical diagnosis and reimbursement. However, automating this task remains challenging. State-of-the-art methods treated this problem as a classification task, leading to ignoring the order of ICD codes that is essential for different purposes. In this work, as a first attempt, we approach this task from a retrieval system perspective to consider the order of codes, thus formulating this problem as a classification and ranking task. Our results and analysis show that the proposed framework has a superior ability to identify high-priority codes compared to other methods. For instance, our model's accuracy in correctly ranking primary diagnosis codes is 47%, compared to 20% for the state-of-the-art classifier. Additionally, in terms of classification metrics, the proposed model achieves a micro- and macro-F1 scores of 0.6065 and 0.2904, respectively, surpassing the previous best model with scores of 0.6035 and 0.2741.
These are the protocol signals we could actually recover from the available paper metadata. Use them to decide whether this paper is worth deeper reading.
Expert Verification
Directly usable for protocol triage.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
Automatic Metrics
Includes extracted eval setup.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
Not reported
No explicit QC controls found.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
Not extracted
No benchmark anchors detected.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
Accuracy, F1, F1 macro
Useful for evaluation criteria comparison.
"For instance, our model's accuracy in correctly ranking primary diagnosis codes is 47%, compared to 20% for the state-of-the-art classifier."
Domain Experts
Helpful for staffing comparability.
"Clinical notes contain unstructured text provided by clinicians during patient encounters."
No benchmark or dataset names were extracted from the available abstract.
Clinical notes contain unstructured text provided by clinicians during patient encounters.
Based on abstract + metadata only. Check the source paper before making high-confidence protocol decisions.
Human feedback protocol is explicit
Detected: Expert Verification
Evaluation mode is explicit
Detected: Automatic Metrics
Quality control reporting appears
No calibration/adjudication/IAA control explicitly detected.
Benchmark or dataset anchors are present
No benchmark/dataset anchor extracted from abstract.
Metric reporting is present
Detected: accuracy, f1, f1 macro