Impact of Leakage on Data Harmonization in Machine Learning Pipelines in Class Imbalance Across Sites
Abstract
Domain fit: Niche / domain-specific · No strong AI-core implementation/artifact signals were detected from current providers.
Machine learning (ML) models benefit from large datasets. Collecting data in biomedical domains is costly and challenging, hence, combining datasets has become a common practice. However, datasets obtained under different conditions could present undesired site-specific variability. Data harmonization methods aim to remove site-specific variance while retaining biologically relevant information. This study evaluates the effectiveness of popularly used ComBat-based methods for harmonizing data in scenarios where the class balance is not equal across sites. We find that these methods struggle with data leakage issues. To overcome this problem, we propose a novel approach PrettYharmonize, designed to harmonize data by pretending the target labels. We validate our approach using controlled datasets designed to benchmark the utility of harmonization. Finally, using real-world MRI and clinical data, we compare leakage-prone methods with PrettYharmonize and show that it achieves comparable performance while avoiding data leakage, particularly in site-target-dependence scenarios.
Results and benchmarks
Machine learning (ML) models benefit from large datasets.
Benchmark evidence is limited
Evidence graph: 2 refs, 1 links.
Utility signals: depth 45/100, grounding 58/100, status medium.
Implementation
No direct implementation yet
Maintained implementation evidence is not confirmed for this paper yet.
Use the implementation status and reproduction sections for the current action plan.
No verified maintained repo yet
There is no verified maintained implementation yet. Use this baseline plan to decide whether to prototype now or defer.
- No direct maintained implementation was found. Use the paper PDF and citation graph to design a baseline reproduction.
- Start from related paper: The harmonization of commercial law in Africa: The project related to telecommunications in the OHADA harmonization process.
- Track assumptions and missing details in an experiment log before coding.
Time to first repro: a few days
Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.
- Estimate is based on paper-only reproduction flow
Reproduction readiness
No repo
No verified implementation available
- No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Datasets
Spaces
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
1
Citations
0
References
Tasks
Harmonization, Leakage (economics), Pipeline transport, Class (philosophy), Computer science, Data mining, Engineering, Electrical and Electronic Engineering
Methods
None detected
Domains
Artificial intelligence
Related papers
- The harmonization of commercial law in Africa: The project related to telecommunications in the OHADA harmonization processSearch on Paper2Code
2006 · Semantic similarity
- Introduction: Basic Definitions and Analytical TasksSearch on Paper2Code
2011 · Semantic similarity
- An Economic Analysis of Harmonization Regimes: Full Harmonization, Minimum Harmonization or Optional Instrument?Search on Paper2Code
2011 · Semantic similarity
- Current International Harmonization EffortsSearch on Paper2Code
2009 · Semantic similarity
- The Harmonization of Higher Education in Southeast AsiaSearch on Paper2Code
2014 · Semantic similarity
- The (limited) role of regulatory harmonization in international goods and services marketsSearch on Paper2Code
1999 · Semantic similarity
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.