Evidence map›Paper›PMID 39122953›Full record

ReviewNature methods2024

Guiding questions to avoid data leakage in biological machine learning applications.

Judith Bernett, David B Blumenthal, Dominik G Grimm, Florian Haselbeck, Roman Joeres, Olga V Kalinina, Markus List

Abstract readReview
PubMed Publisher
In one paragraph

Review in Nature methods, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 53 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
53citing papers in PubMed, 2 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

53 citing papers in PubMed, 2 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Review
  4. Review
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Review
  12. Article
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Judith Bernett *TUM School of Life Sciences, Technical University of Munich, Freising, Germany.ORCID http://orcid.org/0000-0001-5812-8013
David B Blumenthal *Department Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany. david.b.blumenthal@fau.de.ORCID http://orcid.org/0000-0001-8651-750X
Dominik G Grimm *TUM Campus Straubing for Biotechnology and Sustainability, Technical University of Munich, Straubing, Germany. dominik.grimm@tum.de.ORCID http://orcid.org/0000-0003-2085-4591
Florian Haselbeck *TUM Campus Straubing for Biotechnology and Sustainability, Technical University of Munich, Straubing, Germany.ORCID http://orcid.org/0000-0002-5702-376X
Roman Joeres *Department of Chemistry and Molecular Biology, University of Gothenburg, Gothenburg, Sweden.
Olga V Kalinina *Helmholtz Institute for Pharmaceutical Research Saarland (HIPS), Helmholtz Centre for Infection Research (HZI), Saarbrücken, Germany. olga.kalinina@helmholtz-hips.de.ORCID http://orcid.org/0000-0002-9445-477X
Markus List *TUM School of Life Sciences, Technical University of Munich, Freising, Germany. markus.list@tum.de.ORCID http://orcid.org/0000-0002-0941-4168

Funding

Bundesministerium für Bildung und Forschung (Federal Ministry of Education and Research) 031L0305ABundesministerium für Bildung und Forschung (Federal Ministry of Education and Research) 031L0309ADeutsche Forschungsgemeinschaft (German Research Foundation) 51618818
6 · The paper itself

Abstract

Machine learning methods for extracting patterns from high-dimensional data are very important in the biological sciences. However, in certain cases, real-world applications cannot confirm the reported prediction performance. One of the main reasons for this is data leakage, which can be seen as the illicit sharing of information between the training data and the test data, resulting in performance estimates that are far better than the performance observed in the intended application scenario. Data leakage can be difficult to detect in biological datasets due to their complex dependencies. With this in mind, we present seven questions that should be asked to prevent data leakage when constructing machine learning models in biological domains. We illustrate the usefulness of our questions by applying them to nontrivial examples. Our goal is to raise awareness of potential data leakage problems and to promote robust and reproducible machine learning-based research in biology.

Indexed as

Machine LearningAlgorithmsComputational BiologyHumans

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.