ReviewNature methods2024
Guiding questions to avoid data leakage in biological machine learning applications.
Review in Nature methods, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 53 papers, 2 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
53 citing papers in PubMed, 2 syntheses or guidelines pooled it.
- Precision phage therapy in the AI/ML era: a systematic review of discovery-to-clinical translation evidence.Frontiers in microbiology · 2026Pooled it
- Prediction of frailty in community older adults based on machine learning: a systematic review and meta-analysis.Frontiers in public health · 2025Pooled it
- Review
- Beyond predictive accuracy: A case for mechanism-informed, uncertainty-aware machine learning in food microbiology.iScience · 2026Review
- Codon optimality predicts mRNA half-life but does not transfer to lncRNAs.Molecular genetics and genomics : MGG · 2026Article
- Article
- Evaluating a Glioma Transcriptomic Signature Against a Clinical Reference Model and a Random-Signature Null Distribution: A Leakage-Controlled Internal Audit and a Survey of the Field.Diagnostics (Basel, Switzerland) · 2026Article
- Leakage-Free Benchmarking of Electronic Noses for Beef Freshness: A Signal-Richness Criterion for Model Selection.Foods (Basel, Switzerland) · 2026Article
- Benchmarking the impact of data leakage on the performance of knowledge graph embedding models for biomedical link prediction.Bioinformatics (Oxford, England) · 2026Article
- Pre-Existing Heterogeneity Predicts Rare Proteostasis-Stress Programs Across Diverse Perturbations.Biology · 2026Article
- Artificial intelligence and automation in enzyme engineering: evolution, advances, and future perspectives.Bioresources and bioprocessing · 2026Review
- A generalizable computational framework for integrating heterogeneous biomarkers into interpretable scalar risk representations.BMC medical informatics and decision making · 2026Article
- GATSBI: improving context-aware protein embeddings through biologically motivated data splits.Bioinformatics (Oxford, England) · 2026Article
- MD-Transformer: Multimodal Integration of ProtBERT Embeddings and Physicochemical Descriptors for Protein-Protein Interface Residue Prediction.International journal of molecular sciences · 2026Article
- Towards clinically interpretable machine learning in emergency surgery: feature importance and insights across clinical time points in abdominal pain cases.Langenbeck's archives of surgery · 2026Article
- Prediction of cancer-associated thrombosis by machine learning: results from the Vienna Cancer and Thrombosis Study.ESMO open · 2026Article
- Remote sensing data and machine learning models estimate sorghum grain yield in a plant breeding program.Plant phenomics (Washington, D.C.) · 2026Article
- Pocket-Surface Discrete Differential Geometry as a Leakage-Robust Feature Class for Protein-Ligand Binding Affinity Prediction.Molecules (Basel, Switzerland) · 2026Article
- Critical evaluation of drug response prediction models with DrEval.Nature communications · 2026Article
- Machine learning for evolutionary genetics and molecular evolution.Trends in genetics : TIG · 2026Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
7 authors.
Funding
Abstract
Machine learning methods for extracting patterns from high-dimensional data are very important in the biological sciences. However, in certain cases, real-world applications cannot confirm the reported prediction performance. One of the main reasons for this is data leakage, which can be seen as the illicit sharing of information between the training data and the test data, resulting in performance estimates that are far better than the performance observed in the intended application scenario. Data leakage can be difficult to detect in biological datasets due to their complex dependencies. With this in mind, we present seven questions that should be asked to prevent data leakage when constructing machine learning models in biological domains. We illustrate the usefulness of our questions by applying them to nontrivial examples. Our goal is to raise awareness of potential data leakage problems and to promote robust and reproducible machine learning-based research in biology.
Indexed as
Identifiers
39122953What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.