Evidence map›Paper›PMID 42437941›Full record

ArticleEnvironmental microbiome2026

Predicting the seed microbiome using phylogeny-driven machine learning.

Julia Herbinger, Dinesh Kumar Ramakrishnan, Jannik Reißfelder, Majharulislam Babor, Marina Höhne, Ahmed Abdelfattah

Abstract read
In one paragraph

Article in Environmental microbiome, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Julia Herbinger *Department of Data Science, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany.ORCID http://orcid.org/0000-0003-0430-8523
Dinesh Kumar Ramakrishnan *Department of Microbiome Biotechnology, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany.ORCID http://orcid.org/0009-0009-0339-4029
Jannik ReißfelderDepartment of Data Science, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany.
Majharulislam BaborDepartment of Data Science, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany.ORCID http://orcid.org/0000-0002-5440-7573
Marina HöhneDepartment of Data Science, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany.ORCID http://orcid.org/0000-0003-3090-6279
Ahmed AbdelfattahDepartment of Microbiome Biotechnology, Leibniz Institute for Agricultural Engineering and Bioeconomy (ATB), Max-Eyth-Allee 100, 14469, Potsdam, Germany. aabdelfattah@atb-potsdam.de.ORCID https://orcid.org/0000-0001-6090-7200

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe composition of the seed-associated bacterial microbiome can reflect host evolutionary relationships, a pattern consistent with phylosymbiosis. While machine learning offers new opportunities to predict microbial community composition, existing models often require prior microbial profiles or environmental variables, limiting their application to unsampled hosts. Here, we tested whether plant nuclear internal transcribed spacer (ITS) sequences, used as a marker of host relatedness, can predict species-level seed-associated bacterial communities using 16S rRNA data from 61 plant species.

resultsWe introduced customized machine learning models that use sequence-based Hamming distances to capture plant host relatedness. Among the tested models, the Hamming Distance-based k-Nearest Neighbor model (HD-KNN) achieved the highest overall predictive accuracy, yielding an average Jensen-Shannon divergence (JSD) of 0.276 between observed and predicted microbiome profiles. HD-KNN performed particularly well within densely sampled host groups, including Brassicaceae and Poaceae, where closely related reference species were available. In contrast, Hamming Distance-based Gaussian Process Regression (HD-GPR) showed slightly better performance for phylogenetically isolated species, suggesting that model performance depends on host representation within the training dataset.

conclusionsOur framework demonstrates that plant nuclear ITS-derived host relatedness carries a partial predictive signal for seed-associated bacterial microbiome composition. These results provide a foundation for low-input predictive modelling of seed-associated bacteria and may help prioritise microbiome predictions for unsampled plant species when closely related reference species are available. However, our conclusions are strictly limited to seed-associated bacterial communities and should not be directly generalized to fungal communities or other plant compartments, such as the rhizosphere or phyllosphere, which may be shaped by different environmental filtering mechanisms.

Indexed as

Co-evolutionMachine learningMicrobial inheritanceMicrobiome predictionPhylosymbiosisSeed microbiomeVertical transmission

Identifiers

PMID42437941
PMCPMC13366936

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.