Evidence map›Paper›PMID 39695276›Full record

ArticleNPJ digital medicine2024

Estimation of minimal data sets sizes for machine learning predictions in digital mental health interventions.

Kirsten Zantvoort, Barbara Nacke, Dennis Görlich, Silvan Hornstein, Corinna Jacobi, Burkhardt Funk

Abstract read
In one paragraph

Article in NPJ digital medicine, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 49 papers, 4 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
49citing papers in PubMed, 4 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

49 citing papers in PubMed, 4 syntheses or guidelines pooled it.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Pooled it
  5. Trial
  6. Big data and psychiatry: advances, constraints and future directions.World psychiatry : official journal of the World Psychiatric Association (WPA) · 2026
    Article
  7. New approach methodologies (NAMs) for preclinical and translational evaluation of mRNA-lipid nanoparticle (LNP) therapeutics.Journal of controlled release : official journal of the Controlled Release Society · 2026
    Review
  8. Article
  9. Article
  10. Article
  11. Machine learning-based prediction of cross-immunity.Briefings in bioinformatics · 2026
    Article
  12. Article
  13. Review
  14. Article
  15. Article
  16. Multimodal assessments of therapist characteristics are largely unrelated to patient outcomes: A preregistered analysis.Clinical psychological science : a journal of the Association for Psychological Science · 2026
    Article
  17. Article
  18. Article
  19. Review
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Kirsten ZantvoortInstitute of Information Systems, Leuphana University, Lüneburg, Germany. kirsten.zantvoort@leuphana.de.ORCID http://orcid.org/0000-0001-9876-054X
Barbara NackeDepartment of Clinical Psychology and Psychotherapy, Faculty of Psychology, Technische Universität Dresden, Dresden, Germany.
Dennis GörlichInstitute of Biostatistics and Clinical Research, University Münster, Münster, Germany.
Silvan HornsteinDepartment of Psychology, Humboldt-Universität zu Berlin, Berlin, Germany.ORCID http://orcid.org/0000-0002-0398-7096
Corinna JacobiDepartment of Clinical Psychology and Psychotherapy, Faculty of Psychology, Technische Universität Dresden, Dresden, Germany.
Burkhardt FunkInstitute of Information Systems, Leuphana University, Lüneburg, Germany.ORCID http://orcid.org/0000-0001-5855-2666

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Artificial intelligence promises to revolutionize mental health care, but small dataset sizes and lack of robust methods raise concerns about result generalizability. To provide insights on minimal necessary data set sizes, we explore domain-specific learning curves for digital intervention dropout predictions based on 3654 users from a single study (ISRCTN13716228, 26/02/2016). Prediction performance is analyzed based on dataset size (N = 100-3654), feature groups (F = 2-129), and algorithm choice (from Naive Bayes to Neural Networks). The results substantiate the concern that small datasets (N ≤ 300) overestimate predictive power. For uninformative feature groups, in-sample prediction performance was negatively correlated with dataset size. Sophisticated models overfitted in small datasets but maximized holdout test results in larger datasets. While N = 500 mitigated overfitting, performance did not converge until N = 750-1500. Consequently, we propose minimum dataset sizes of N = 500-1000. As such, this study offers an empirical reference for researchers designing or interpreting AI studies on Digital Mental Health Intervention data.

Identifiers

PMID39695276
PMCPMC11655521

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.