Evidence map›Paper›PMID 42747298›Full record

ArticleBioinformatics (Oxford, England)2026

Perseus: lineage-aware refinement of Kraken2 taxonomic classification for long read metagenomes.

Matthew H Nguyen, Michael C Schatz

Abstract read
In one paragraph

Article in Bioinformatics (Oxford, England), 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

2 authors.

Matthew H NguyenDepartment of Computer Science, Johns Hopkins University, Baltimore, MD 21218, United States.ORCID 0000-0002-4312-5992
Michael C SchatzDepartment of Computer Science, Johns Hopkins University, Baltimore, MD 21218, United States.ORCID 0000-0002-4118-4446

Funding

Implementing the Genomic Data Science Analysis, Visualization, and Informatics Lab-space (AnVIL)U24HG010263 · NHGRI · JOHNS HOPKINS UNIVERSITY · PI Enis Afgan, VINCENT JAMES CAREY · 2018 to 2026
$23.8M
Democratization of Data Analysis in Life Sciences Through GalaxyU41HG006620 · NHGRI · PENNSYLVANIA STATE UNIVERSITY, THE · PI NEKRUTENKO, ANTON, SCHATZ, MICHAEL · 2012 to 2020
$14.4M
An in integrated platform for multiomic analyses of pathogen and host data using scalable public infrastructureU24AI183870 · NIAID · PENNSYLVANIA STATE UNIVERSITY, THE · PI Kelsey M Beavers, Maximilian Haeussler · 2024 to 2026
$10.2M
NHGRI NIH HHS U24 HG010263NHGRI NIH HHS U41 HG006620NIAID NIH HHS U24 AI183870NIH DBI-2419522NIH NSF
6 · The paper itself

Abstract

motivationLong-read metagenomic sequencing improves assembly contiguity and enables genome-resolved analysis of complex microbial communities, but accurate taxonomic classification of long reads and assembled contigs remains challenging. Highly scalable k-mer-based classifiers such as Kraken2 frequently over-assign fine-rank taxonomic labels when applied to long-read data, producing high false positive classification rates driven by sparse or localized k-mer matches, particularly in microbiomes with extensive taxonomic novelty.

resultsWe present Perseus, a lineage-aware confidence estimation framework for taxonomic classification that models the spatial distribution and hierarchical consistency of k-mer evidence along sequences. This formulation reframes taxonomic classification as a hierarchical confidence estimation problem rather than a single-rank prediction task. Perseus refines k-mer-level taxonomic signals from Kraken2 using a multi-headed convolutional neural network that estimates calibrated confidence scores for taxonomic correctness at each canonical rank. Using these estimates, Perseus confirms assignments, backs off to higher taxonomic ranks, or abstains when evidence is insufficient, prioritizing correctness and lineage consistency over overly specific assignments. Across simulations of taxonomic novelty and real-world metagenomic datasets, Perseus consistently and substantially reduces the false assignment rate while improving precision and lineage-consistent accuracy. These improvements are most pronounced for long reads and assembled contigs, where spatial context enables reliable discrimination between consistent taxonomic signal and spurious matches. AVAILABILITY AND IMPLEMENTATION: Perseus integrates with existing Kraken2 workflows and is available at https://github.com/matnguyen/perseus.

Indexed as

MetagenomeMetagenomicsSoftwareAlgorithmsClassification AlgorithmsSequence Analysis, DNA

Identifiers

PMID42747298
PMCPMC13633669

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.