Evidence map›Paper›PMID 41648004›Full record

ArticleFrontiers in microbiology2025

Beyond microbial abundance: metadata integration enhances disease prediction in human microbiome studies.

Andre R Goncalves, Hiranmayi Ranganathan, Camilo Valdes, Haonan Zhu, Boya Zhang, Car Reen Kok, Jose Manuel Martí, Nisha J Mulakken, James B Thissen, Crystal Jaing and 1 more

Abstract read
In one paragraph

Article in Frontiers in microbiology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Andre R GoncalvesComputational Engineering Division, Engineering Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Hiranmayi RanganathanComputational Engineering Division, Engineering Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Camilo ValdesBiosciences and Biotechnology Division, Physical and Life Sciences Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Haonan ZhuComputational Engineering Division, Engineering Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Boya ZhangComputational Engineering Division, Engineering Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Car Reen KokBiosciences and Biotechnology Division, Physical and Life Sciences Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Jose Manuel MartíGlobal Security Computing Applications Division, Computing Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Nisha J MulakkenGlobal Security Computing Applications Division, Computing Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
James B ThissenBiosciences and Biotechnology Division, Physical and Life Sciences Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Crystal JaingBiosciences and Biotechnology Division, Physical and Life Sciences Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.
Nicholas A BeBiosciences and Biotechnology Division, Physical and Life Sciences Directorate, Lawrence Livermore National Laboratory, Livermore, CA, United States.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Multiple studies have highlighted the interaction of the human microbiome with physiological systems such as the gut, immune, liver, and skin, via key axes. Advances in sequencing technologies and high-performance computing have enabled the analysis of large-scale metagenomic data, facilitating the use of machine learning to predict disease likelihood from microbiome profiles. However, challenges such as compositionality, high dimensionality, sparsity, and limited sample sizes have hindered the development of actionable models. One strategy to improve these models is by incorporating key metadata from both the human host and sample collection/processing protocols. This remains challenging due to sparsity and inconsistency in metadata annotation and availability. In this paper, we introduce a machine learning-based pipeline for predicting human disease states by integrating host and protocol metadata with microbiome abundance profiles from 68 different studies, processed through a consistent pipeline. Our findings indicate that metadata can enhance machine learning predictions, particularly at higher taxonomic ranks like Kingdom and Phylum, though this effect diminishes at lower ranks. Our study leverages a large collection of microbiome datasets comprising 11,208 samples, therefore enhancing the robustness and statistical confidence of our findings. This work is a critical step toward utilizing microbiome and metadata for predicting diseases such as gastrointestinal infections, diabetes, cancer, and neurological disorders.

Indexed as

host disease predictionhost metadatahuman microbiomemachine learningmeta-analysis

Identifiers

PMID41648004
PMCPMC12869998

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.