Nonparametric IPSS: Fast, flexible feature selection with false discovery control
Abstract
Domain fit: Niche / domain-specific · No strong AI-core implementation/artifact signals were detected from current providers.
Feature selection is a critical task in machine learning and statistics. However, existing feature selection methods either (i) rely on parametric methods such as linear or generalized linear models, (ii) lack theoretical false discovery control, or (iii) identify few true positives. Here, we introduce a general feature selection method with finite-sample false discovery control based on applying integrated path stability selection (IPSS) to arbitrary feature importance scores. The method is nonparametric whenever the importance scores are nonparametric, and it estimates q-values, which are better suited to high-dimensional data than p-values. We focus on two special cases using importance scores from gradient boosting (IPSSGB) and random forests (IPSSRF). Extensive nonlinear simulations with RNA sequencing data show that both methods accurately control the false discovery rate and detect more true positives than existing methods. Both methods are also efficient, running in under 20 seconds when there are 500 samples and 5000 features. We apply IPSSGB and IPSSRF to detect microRNAs and genes related to cancer, finding that they yield better predictions with fewer features than existing approaches.
Results and benchmarks
Feature selection is a critical task in machine learning and statistics.
Benchmark evidence is limited
Evidence graph: 2 refs, 1 links.
Utility signals: depth 45/100, grounding 58/100, status medium.
Implementation
No direct implementation yet
Maintained implementation evidence is not confirmed for this paper yet.
Use the implementation status and reproduction sections for the current action plan.
No verified maintained repo yet
There is no verified maintained implementation yet. Use this baseline plan to decide whether to prototype now or defer.
- No direct maintained implementation was found. Use the paper PDF and citation graph to design a baseline reproduction.
- Start from related paper: 29 False positives and false negatives in genome scans.
- Track assumptions and missing details in an experiment log before coding.
Time to first repro: a few days
Recommendation evidence is currently too limited for a maintained-repo choice. Use Implementation Status and Reproduction Path for a practical baseline plan.
- Estimate is based on paper-only reproduction flow
Reproduction readiness
No repo
No verified implementation available
- No maintained repository has been identified for this paper. Check adjacent implementations or HF artifacts below.
Hardware requirements
- Expect multi-day setup/compute for meaningful reproduction based on current guidance.
Validation caveat
Hugging Face artifacts
No trustworthy direct or curated related Hugging Face artifacts were found yet. Use targeted searches to quickly locate candidate models, datasets, and demos.
Tip: start with models, then check datasets and spaces if you need evaluation data or demos.
Research context
0
Citations
30
References
Tasks
Computer science, Nonparametric statistics, False positives and false negatives, False positive paradox, Feature selection, False discovery rate, Python (programming language), Source code
Methods
None detected
Domains
Artificial intelligence, Machine learning, Biochemistry, Genetics and Molecular Biology
Related papers
- 29 False positives and false negatives in genome scansSearch on Paper2Code
2001 · Semantic similarity
- Simultaneous control of false positives and false negatives in multiple hypotheses testingSearch on Paper2Code
2007 · Semantic similarity
- On the Operational Characteristics of the Benjamini and Hochberg False Discovery Rate ProcedureSearch on Paper2Code
2007 · Semantic similarity
- The principled control of false positives in neuroimagingSearch on Paper2Code
2009 · Semantic similarity
- Controlling the Proportion of False Positives in Multiple Dependent TestsSearch on Paper2Code
2004 · Semantic similarity
- Scaled False Discovery Proportion and Related Error MetricsSearch on Paper2Code
2013 · Semantic similarity
Open this paper in HFEPX to review benchmark signals, evaluation modes, and human-feedback protocol context.
Open in HFEPXJump to Paper2Code search queries derived from this paper's research context.