Evidence map›Paper›PMID 39650139›Full record

ArticleaBIOTECH2024

Impact of database choice and confidence score on the performance of taxonomic classification using Kraken2.

Yunlong Liu, Morteza H Ghaffari, Tao Ma, Yan Tu

Abstract read
In one paragraph

Article in aBIOTECH, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 23 papers.

0numbers the graph read from it
0cells of the map it votes in
23citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

23 citing papers in PubMed.

  1. Trial
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
  15. Review
  16. Absolute Quantification of Bacteria in the Microbiome and Its Application.Methods in molecular biology (Clifton, N.J.) · 2026
    Article
  17. Review
  18. Detection ofFood chemistry. Molecular sciences · 2025
    Article
  19. bioRxiv : the preprint server for biology · 2025
    Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Yunlong LiuKey Laboratory of Feed Biotechnology of the Ministry of Agricultural and Rural Affairs, Institute of Feed Research, Chinese Academy of Agricultural Sciences, Beijing, 100081 China.
Morteza H GhaffariInstitute of Animal Science, Physiology Unit, University of Bonn, Bonn, 53115 Germany.
Tao MaKey Laboratory of Feed Biotechnology of the Ministry of Agricultural and Rural Affairs, Institute of Feed Research, Chinese Academy of Agricultural Sciences, Beijing, 100081 China.ORCID 0000-0003-4821-836X
Yan TuKey Laboratory of Feed Biotechnology of the Ministry of Agricultural and Rural Affairs, Institute of Feed Research, Chinese Academy of Agricultural Sciences, Beijing, 100081 China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Accurate taxonomic classification is essential to understanding microbial diversity and function through metagenomic sequencing. However, this task is complicated by the vast variety of microbial genomes and the computational limitations of bioinformatics tools. The aim of this study was to evaluate the impact of reference database selection and confidence score (CS) settings on the performance of Kraken2, a widely used k-mer-based metagenomic classifier. In this study, we generated simulated metagenomic datasets to systematically evaluate how the choice of reference databases, from the compact Minikraken v1 to the expansive nt- and GTDB r202, and different CS (from 0 to 1.0) affect the key performance metrics of Kraken2. These metrics include classification rate, precision, recall, F1 score, and accuracy of true versus calculated bacterial abundance estimation. Our results show that higher CS, which increases the rigor of taxonomic classification by requiring greater k-mer agreement, generally decreases the classification rate. This effect is particularly pronounced for smaller databases such as Minikraken and Standard-16, where no reads could be classified when the CS was above 0.4. In contrast, for larger databases such as Standard, nt and GTDB r202, precision and F1 scores improved significantly with increasing CS, highlighting their robustness to stringent conditions. Recovery rates were mostly stable, indicating consistent detection of species under different CS settings. Crucially, the results show that a comprehensive reference database combined with a moderate CS (0.2 or 0.4) significantly improves classification accuracy and sensitivity. This finding underscores the need for careful selection of database and CS parameters tailored to specific scientific questions and available computational resources to optimize the results of metagenomic analyses. Supplementary Information: The online version contains supplementary material available at 10.1007/s42994-024-00178-0.

Indexed as

Confidence scoreKraken2MetagenomeReference databaseTaxonomic classification

Identifiers

PMID39650139
PMCPMC11624175

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.