ArticlebioRxiv : the preprint server for biology2026
Performance of IBD machine learning classifiers varies across microbiome training data independent of geographic diversity.
Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors.
Funding
Abstract
Microbiome-based machine learning classifiers show increasing promise for disease identification across gastrointestinal, metabolic, and immune-mediated conditions. Inflammatory bowel disease (IBD), a chronic immune-mediated disorder associated with disruption of the gut microbiome, has been a particularly successful application area. However, while many predictive models achieve high performance within individual datasets, their ability to generalize across independent populations and geographic contexts remains unclear. Here, we tested whether model class and training dataset composition influence model generalizability across geographically diverse evaluation studies. We compiled seven publicly available shotgun metagenomic studies spanning five geographic regions, comprising 697 individuals with IBD or healthy controls. We trained 246,986 model configurations across seven model classes and five distinct training dataset combinations and evaluated top-performing models on independent studies from the USA, Ireland, Germany, Israel and China Extreme gradient boosting and random forest models showed the highest and most consistent performance across training datasets, a ranking that was maintained on independent evaluation studies. However, models trained on geographically diverse datasets did not outperform those trained on USA-only datasets. Instead, model performance was strongly dependent on the evaluation study itself, with consistent differences in achievable accuracy across studies. Despite most models achieving similar AUC scores, there was limited overlap in the key microbial species identified. Furthermore, even for the small set of disease predictive microbes shared between models, the direction of enrichment between IBD or healthy subjects often varied in opposing directions across study populations. These findings suggest that study-specific factors constrain generalization and may help explain the lack of consistent microbiome-based biomarkers for IBD.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.