ArticleNPJ digital medicine2024
Estimation of minimal data sets sizes for machine learning predictions in digital mental health interventions.
Article in NPJ digital medicine, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 49 papers, 4 of them syntheses that pooled it.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
49 citing papers in PubMed, 4 syntheses or guidelines pooled it.
- Prognostic value of machine learning for brain computed tomography as a predictor of neurologic outcomes after cardiac arrest: a systematic review and meta-analysis.Scandinavian journal of trauma, resuscitation and emergency medicine · 2026Pooled it
- Relapse prediction and individualized treatment-effect modeling in relapsing multiple sclerosis: a systematic review.Frontiers in public health · 2026Pooled it
- Machine learning-based risk predictive models for depression in patients with diabetes: a systematic review and meta-analysis.Frontiers in endocrinology · 2026Pooled it
- Prediction models for treatment response in migraine: a systematic review and meta-analysis.The journal of headache and pain · 2025Pooled it
- Personalized prediction of response to three meditation practices: A randomized controlled trial.Behaviour research and therapy · 2026Trial
- Big data and psychiatry: advances, constraints and future directions.World psychiatry : official journal of the World Psychiatric Association (WPA) · 2026Article
- New approach methodologies (NAMs) for preclinical and translational evaluation of mRNA-lipid nanoparticle (LNP) therapeutics.Journal of controlled release : official journal of the Controlled Release Society · 2026Review
- Predicting Psychological Flourishing Among Psychologists: Integrating Self-Compassion, Traditional Regression, and Machine Learning Approaches.Healthcare (Basel, Switzerland) · 2026Article
- Predicting Craniospinal Surgery in Pediatric Achondroplasia: Benchmarking Statistical Inference and Explainable AI Under Rare-Disease Constraints.Annals of biomedical engineering · 2026Article
- Functional connectivity during positive mood and anxiety treatment response in children.Journal of affective disorders · 2026Article
- Machine learning-based prediction of cross-immunity.Briefings in bioinformatics · 2026Article
- Prediction of Clinically Meaningful Improvement After Internet-Delivered Cognitive Behavioral Therapy for Depression and Anxiety Disorders: Machine Learning-Based Predictive Model Development and Temporal Validation Study.Journal of medical Internet research · 2026Article
- Review
- Article
- Sample size calculation for training ensemble machine learning models on health data.Patterns (New York, N.Y.) · 2026Article
- Multimodal assessments of therapist characteristics are largely unrelated to patient outcomes: A preregistered analysis.Clinical psychological science : a journal of the Association for Psychological Science · 2026Article
- Using Machine Learning to Analyze the Predictors of Life Satisfaction: Focus on Lifestyle Attitudes and Psychological Factors.International journal of methods in psychiatric research · 2026Article
- Combined multi-omics and multi-spectral profiling of plasma extracellular vesicles reveals liquid biopsy biomarkers for glioma diagnosis.Cell reports. Medicine · 2026Article
- Is the molecular microenvironment of the latent HIV reservoir predictable using deep learning approaches?Journal of virology · 2026Review
- Machine Learning Model to Predict Postmastectomy Breast Reconstruction Complications.JAMA network open · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Artificial intelligence promises to revolutionize mental health care, but small dataset sizes and lack of robust methods raise concerns about result generalizability. To provide insights on minimal necessary data set sizes, we explore domain-specific learning curves for digital intervention dropout predictions based on 3654 users from a single study (ISRCTN13716228, 26/02/2016). Prediction performance is analyzed based on dataset size (N = 100-3654), feature groups (F = 2-129), and algorithm choice (from Naive Bayes to Neural Networks). The results substantiate the concern that small datasets (N ≤ 300) overestimate predictive power. For uninformative feature groups, in-sample prediction performance was negatively correlated with dataset size. Sophisticated models overfitted in small datasets but maximized holdout test results in larger datasets. While N = 500 mitigated overfitting, performance did not converge until N = 750-1500. Consequently, we propose minimum dataset sizes of N = 500-1000. As such, this study offers an empirical reference for researchers designing or interpreting AI studies on Digital Mental Health Intervention data.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.