Evidence map›Paper›PMID 41867861›Full record

ArticlebioRxiv : the preprint server for biology2026

A comprehensive assessment of tandem repeat genotyping methods for Nanopore long-read genomes.

Elbay Aliyev, Akshay Avvaru, Wouter De Coster, Garrison M Arner, Denis M Nyaga, Sophia B Gibson, Ben Weisburd, Bida Gu, Claudia Gonzaga-Jauregui, 1000 Genomes Long-Read Sequencing Consortium and 4 more

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

14 authors.

Elbay AliyevDepartment of Biomedical Informatics, University of Colorado Anschutz, Aurora, CO, USA.ORCID 0000-0002-6469-1854
Akshay AvvaruDepartment of Biomedical Informatics, University of Colorado Anschutz, Aurora, CO, USA.ORCID 0000-0003-4698-9363
Wouter De CosterDepartment of Biomedical Sciences, University of Antwerp, Antwerp, Belgium.ORCID 0000-0002-5248-8197
Garrison M ArnerDepartment of Biomedical Informatics, University of Colorado Anschutz, Aurora, CO, USA.
Denis M NyagaLiggins Institute, University of Auckland, New Zealand.ORCID 0000-0001-6240-4017
Sophia B GibsonDepartment of Genome Sciences, University of Washington, Seattle, WA, USA.ORCID 0000-0001-9839-9045
Ben WeisburdProgram in Medical and Population Genetics, Broad Institute of MIT and Harvard, Cambridge, MA, USA.ORCID 0000-0001-9898-9109
Bida GuDepartment of Quantitative and Computational Biology, University of Southern California, Los Angeles, CA, USA.ORCID 0000-0001-8575-997X
Claudia Gonzaga-JaureguiInternational Laboratory for Human Genome Research, Laboratorio Internacional de Investigación sobre el Genoma Humano, Universidad Nacional Autónoma de México, Querétaro, Mexico.
1000 Genomes Long-Read Sequencing Consortium
Mark J P ChaissonDepartment of Quantitative and Computational Biology, University of Southern California, Los Angeles, CA, USA.ORCID 0000-0001-5395-1457
Danny E MillerDepartment of Genome Sciences, University of Washington, Seattle, WA, USA.ORCID 0000-0001-6096-8601
Elizabeth OstrowskiLiggins Institute, University of Auckland, New Zealand.ORCID 0000-0002-9784-5979
Harriet DashnowDepartment of Biomedical Informatics, University of Colorado Anschutz, Aurora, CO, USA.ORCID 0000-0001-8433-6270

Funding

Long-read DNA and RNA sequencing to identify disease-causing genetic variation and streamline testingDP5OD033357 · OD · UNIVERSITY OF WASHINGTON · PI MILLER, DANNY ERWIN · 2022 to 2025
$1.9M
Pathways in Genomics Research Experiences for Undergraduates from Underrepresented GroupsR25HG012994 · NHGRI · METROPOLITAN STATE UNIVERSITY OF DENVER · PI Kristy L. Duran, Audrey E Hendricks · 2023 to 2026
$1.4M
Revealing new short tandem repeat variation in the human population across sequencing technologies: towards rare disease diagnosis and discoveryR00HG012796 · NHGRI · UNIVERSITY OF COLORADO DENVER · PI Harriet Dashnow · 2024 to 2026
$747k
NHGRI NIH HHS R00 HG012796NHGRI NIH HHS R25 HG012994NIH HHS DP5 OD033357
6 · The paper itself

Abstract

Background: Tandem repeats (TRs) play critical roles in human disease and phenotypic diversity but are among the most challenging classes of genomic variation to measure accurately. While it is possible to identify TR expansions using short-read sequencing, these methods are limited because they often cannot accurately determine repeat length or sequence composition. Long-read sequencing (LRS) has the potential to accurately characterize long TRs, including the identification of non-canonical motifs and complex structures. However, while there are an increasing number of genotyping methods available, no systematic effort has been undertaken to evaluate their length and sequence-level accuracy, performance across motifs from STRs to VNTRs and across allele lengths, and, critically, how usable these tools are in practice. Results: We reviewed 25 available bioinformatic tools, and selected seven that are actively maintained for benchmarking using publicly available Oxford Nanopore genome sequencing data from more than 100 individuals. Our benchmarking catalog included ~43k TR loci genome-wide, selected to represent a range of simple and challenging TR loci. As no "truth" exists for this purpose, we used four complementary strategies to assess accuracy: concordance with high-quality haplotype-resolved Human Pangenome Reference Consortium (HPRC) assemblies, Mendelian consistency in Genome in a Bottle trios, cross-tool consistency, and sensitivity in individuals with pathogenic TR expansions confirmed by molecular methods. For all comparisons, we assess both total allele length and full sequence similarity using the Levenshtein distance. We also evaluated installation, documentation, computational requirements, and output characteristics to reflect real-world use. We provide a complete analysis workflow for all tools to support community reuse.Tool performance varied substantially across both accuracy and usability. Most methods achieved high concordance with HPRC assemblies, with higher accuracy when using the R10 ONT pore chemistry. Accuracy generally declined with increasing allele length, and most tools performed worse on homopolymers, likely reflecting underlying sequencing accuracy. Tools generally performed worse at heterozygous loci and at alleles that differed from the reference genome. Interestingly, concordance with assembly in population samples did not predict sensitivity to pathogenic expansions, with different genotypers performing best in each category. Similarly, Mendelian consistency was highest in the tool that performed worst in assembly concordance. Conclusions: No single genotyper emerged as consistently best across all assessments, but strong contenders emerged in each. Our results demonstrate that length accuracy (a typical benchmarking approach) alone overestimates TR genotyping performance. Sequence-level benchmarking is essential for selecting tools best-suited for population studies and clinical diagnostics. This work provides practical guidance for tool selection and highlights key priorities for future long-read TR genotyping method development.

Identifiers

PMID41867861
PMCPMC13001349

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.