Evidence map›Paper›PMID 38883233›Full record

ArticleArXiv2025

Distributional bias compromises leave-one-out cross-validation.

George I Austin, Itsik Pe'er, Tal Korem

Abstract readPreprint
In one paragraph

Article in ArXiv, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

3 authors.

George I Austin
Itsik Pe'er
Tal Korem

Funding

Training in Biomedical Informatics at Columbia UniversityT15LM007079 · NLM · COLUMBIA UNIV NEW YORK MORNINGSIDE · PI NOEMIE ELHADAD, GEORGE M HRIPCSAK · 1992 to 2026
$28.9M
A large scale investigation of the vaginal metagenome and metabolome and their role in spontaneous preterm birthR01HD106017 · NICHD · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI KOREM, TAL · 2021 to 2025
$3.6M
Columbia University Graduate Training Program in Computational and Systems BiologyT32GM158494 · NIGMS · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI Peter Alan Sims, Chaolin Zhang · 2025 to 2026
$486k
NICHD NIH HHS R01 HD106017NIGMS NIH HHS T32 GM158494NLM NIH HHS T15 LM007079
6 · The paper itself

Abstract

Cross-validation is a common method for estimating the predictive performance of machine learning models. In a data-scarce regime, where one typically wishes to maximize the number of instances used for training the model, an approach called "leave-one-out cross-validation" is often used. In this design, a separate model is built for predicting each data instance after training on all other instances. Since this results in a single test instance available per model trained, predictions are aggregated across the entire dataset to calculate common performance metrics such as the area under the receiver operating characteristic or R2 scores. In this work, we demonstrate that this approach creates a negative correlation between the average label of each training fold and the label of its corresponding test instance, a phenomenon that we term distributional bias. As machine learning models tend to regress to the mean of their training data, this distributional bias tends to negatively impact performance evaluation and hyperparameter optimization. We show that this effect generalizes to leave-P-out cross-validation and persists across a wide range of modeling and evaluation approaches, and that it can lead to a bias against stronger regularization. To address this, we propose a generalizable rebalanced cross-validation approach that corrects for distributional bias for both classification and regression. We demonstrate that our approach improves cross-validation performance evaluation in synthetic simulations, across machine learning benchmarks, and in several published leave-one-out analyses.

Identifiers

PMID38883233
PMCPMC11177965

What OpenQuestion holds

Textmetadata
LicenceCC BY-SA
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.