Evidence map›Paper›PMID 37232476›Full record

ArticleSystematic biology2023

Online Phylogenetics with matOptimize Produces Equivalent Trees and is Dramatically More Efficient for Large SARS-CoV-2 Phylogenies than de novo and Maximum-Likelihood Implementations.

Alexander M Kramer, Bryan Thornlow, Cheng Ye, Nicola De Maio, Jakob McBroome, Angie S Hinrichs, Robert Lanfear, Yatish Turakhia, Russell Corbett-Detig

Abstract read
In one paragraph

Article in Systematic biology, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 24 papers.

0numbers the graph read from it
0cells of the map it votes in
24citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

24 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
  4. Review
  5. Article
  6. A Pandemic-Scale Ancestral Recombination Graph for SARS-CoV-2.bioRxiv : the preprint server for biology · 2025
    Article
  7. UShER-TB: Scalable, Comprehensive, Accessible Phylogenomic Analysis ofmedRxiv : the preprint server for health sciences · 2025
    Article
  8. Article
  9. Article
  10. Article
  11. Challenges in Assembling the Dated Tree of Life.Genome biology and evolution · 2024
    Article
  12. Review
  13. Article
  14. Article
  15. Article
  16. Article
  17. Article
  18. Article
  19. Article
  20. On parsimony and clustering.PeerJ. Computer science · 2023
    Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

9 authors.

Alexander M KramerDepartment of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA 95064, USA.ORCID 0000-0003-3630-3209
Bryan ThornlowDepartment of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA 95064, USA.ORCID 0000-0001-6334-5186
Cheng YeDepartment of Electrical and Computer Engineering, University of California San Diego, San Diego, CA 92093, USA.
Nicola De MaioEuropean Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Cambridge CB10 1SD, UK.ORCID 0000-0002-1776-8564
Jakob McBroomeDepartment of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA 95064, USA.ORCID 0000-0002-5002-5156
Angie S HinrichsGenomics Institute, University of California Santa Cruz, Santa Cruz, CA 95064, USA.ORCID 0000-0002-1697-1130
Robert LanfearDepartment of Ecology and Evolution, Research School of Biology, Australian National University, Canberra, ACT 2601, Australia.
Yatish TurakhiaDepartment of Electrical and Computer Engineering, University of California San Diego, San Diego, CA 92093, USA.ORCID 0000-0001-5600-2900
Russell Corbett-DetigDepartment of Biomolecular Engineering, University of California Santa Cruz, Santa Cruz, CA 95064, USA.

Funding

Origins, Functional, and Evolutionary Consequences of Genomic VariationR35GM128932 · NIGMS · UNIVERSITY OF CALIFORNIA SANTA CRUZ · PI Russell Corbett-Detig · 2018 to 2026
$3.2M
UC Santa Cruz Training Program In Genomic SciencesT32HG008345 · NHGRI · UNIVERSITY OF CALIFORNIA SANTA CRUZ · PI BROOKS, ANGELA NORIE, GREEN, RICHARD EDWARD · 2015 to 2020
$1.4M
Evolutionary dynamics of tRNA genesF31HG010584 · NHGRI · UNIVERSITY OF CALIFORNIA SANTA CRUZ · PI THORNLOW, BRYAN PATRICK · 2019 to 2021
$86k
NHGRI NIH HHS F31 HG010584NHGRI NIH HHS T32 HG008345NIGMS NIH HHS R35 GM128932NIH HHS R35GM128932
6 · The paper itself

Abstract

Phylogenetics has been foundational to SARS-CoV-2 research and public health policy, assisting in genomic surveillance, contact tracing, and assessing emergence and spread of new variants. However, phylogenetic analyses of SARS-CoV-2 have often relied on tools designed for de novo phylogenetic inference, in which all data are collected before any analysis is performed and the phylogeny is inferred once from scratch. SARS-CoV-2 data sets do not fit this mold. There are currently over 14 million sequenced SARS-CoV-2 genomes in online databases, with tens of thousands of new genomes added every day. Continuous data collection, combined with the public health relevance of SARS-CoV-2, invites an "online" approach to phylogenetics, in which new samples are added to existing phylogenetic trees every day. The extremely dense sampling of SARS-CoV-2 genomes also invites a comparison between likelihood and parsimony approaches to phylogenetic inference. Maximum likelihood (ML) and pseudo-ML methods may be more accurate when there are multiple changes at a single site on a single branch, but this accuracy comes at a large computational cost, and the dense sampling of SARS-CoV-2 genomes means that these instances will be extremely rare because each internal branch is expected to be extremely short. Therefore, it may be that approaches based on maximum parsimony (MP) are sufficiently accurate for reconstructing phylogenies of SARS-CoV-2, and their simplicity means that they can be applied to much larger data sets. Here, we evaluate the performance of de novo and online phylogenetic approaches, as well as ML, pseudo-ML, and MP frameworks for inferring large and dense SARS-CoV-2 phylogenies. Overall, we find that online phylogenetics produces similar phylogenetic trees to de novo analyses for SARS-CoV-2, and that MP optimization with UShER and matOptimize produces equivalent SARS-CoV-2 phylogenies to some of the most popular ML and pseudo-ML inference tools. MP optimization with UShER and matOptimize is thousands of times faster than presently available implementations of ML and online phylogenetics is faster than de novo inference. Our results therefore suggest that parsimony-based methods like UShER and matOptimize represent an accurate and more practical alternative to established ML implementations for large SARS-CoV-2 phylogenies and could be successfully applied to other similar data sets with particularly dense sampling and short branch lengths.

Indexed as

COVID-19SARS-CoV-2GenomicsHumansPhylogenyProbabilitymaximum likelihoodoptimizationparsimonyphylogeneticsSARS-CoV-2

Identifiers

PMID37232476
PMCPMC10627557

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.