Evidence map›Paper›PMID 39882309›Full record

ArticleVirus evolution2025

SARS-CoV-2 CoCoPUTs: analyzing GISAID and NCBI data to obtain codon statistics, mutations, and free energy over a multiyear period.

Nigam H Padhiar, Tigran Ghazanchyan, Sarah E Fumagalli, Michael DiCuccio, Guy Cohen, Alexander Ginzburg, Brian Rikshpun, Almog Klein, Luis Santana-Quintero, Sean Smith and 2 more

Abstract read
In one paragraph

Article in Virus evolution, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Review
  5. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Nigam H PadhiarHemostasis Branch 1, Division of Hemostasis, Office of Plasma Protein Therapeutics CMC, Office of Therapeutic Products, Center for Biologics Evaluation and Research, Food and Drug Administration, 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.
Tigran GhazanchyanHigh-performance Integrated Virtual Environment (HIVE), Center for Biologics Evaluation and Research (CBER), Food and Drug Administration (FDA), 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.
Sarah E FumagalliHemostasis Branch 1, Division of Hemostasis, Office of Plasma Protein Therapeutics CMC, Office of Therapeutic Products, Center for Biologics Evaluation and Research, Food and Drug Administration, 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.
Michael DiCuccioRockville, MD 20853, USA.
Guy CohenAfeka Tel-Aviv Academic College of Engineering, Mivtsa Kadesh St 38, Tel Aviv-Yafo 6998812, Israel.
Alexander GinzburgAfeka Tel-Aviv Academic College of Engineering, Mivtsa Kadesh St 38, Tel Aviv-Yafo 6998812, Israel.
Brian RikshpunAfeka Tel-Aviv Academic College of Engineering, Mivtsa Kadesh St 38, Tel Aviv-Yafo 6998812, Israel.
Almog KleinAfeka Tel-Aviv Academic College of Engineering, Mivtsa Kadesh St 38, Tel Aviv-Yafo 6998812, Israel.
Luis Santana-QuinteroHigh-performance Integrated Virtual Environment (HIVE), Center for Biologics Evaluation and Research (CBER), Food and Drug Administration (FDA), 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.
Sean SmithHigh-performance Integrated Virtual Environment (HIVE), Center for Biologics Evaluation and Research (CBER), Food and Drug Administration (FDA), 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.
Anton A KomarDepartment of Biological, Geological and Environmental Sciences, Center for Gene Regulation in Health and Disease, Cleveland State University, 2121 Euclid Avenue, SR 259, Cleveland, OH 44115, USA.
Chava Kimchi-SarfatyHemostasis Branch 1, Division of Hemostasis, Office of Plasma Protein Therapeutics CMC, Office of Therapeutic Products, Center for Biologics Evaluation and Research, Food and Drug Administration, 10903 New Hampshire Ave, Silver Spring, MD 20993, USA.ORCID https://orcid.org/0000-0002-9355-8585

Funding

Safer and more effective FIX therapeutics: impact of codon optimizationR01HL151392 · NHLBI · CLEVELAND STATE UNIVERSITY · PI KOMAR, ANTON A. · 2020 to 2023
$1.5M
NHLBI NIH HHS R01 HL151392
6 · The paper itself

Abstract

A consistent area of interest since the beginning of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) pandemic has been the sequence composition of the virus and how it has changed over time. Many resources have been developed for the storage and analysis of SARS-CoV-2 data, such as GISAID (Global Initiative on Sharing All Influenza Data), NCBI, Nextstrain, and outbreak.info. However, relatively little has been done to compile codon usage data, codon-level mutation data, and secondary structure data into a single database. Here, we assemble the aforementioned data and many additional virus attributes in a new database entitled SARS-CoV-2 CoCoPUTs. We begin with an overview of the composition and overlap between two of the largest sources of SARS-CoV-2 sequence data: GISAID and NCBI Virus (GenBank). We then evaluate different types of sequence curation strategies to reduce the dataset of millions of sequences to only one sequence per Pango lineage variant. We then performed specific analyses on the coding sequences (CDSs), including calculating codon usage, codon pair usage, dinucleotides, junction dinucleotides, mutations, GC content, effective number of codons (ENCs), and effective number of codon pairs (ENCPs). We have also performed whole-genome secondary RNA structure prediction calculations for each variant, using the LinearPartition software and modified selective 2'-hydroxyl acylation analyzed by primer extension (SHAPE) data that are available online. Finally, we compiled all the data into our resource, SARS-CoV-2 CoCoPUTs, and paired many of the resulting statistics with variant proportion data over time in order to derive trends in viral evolution. Although the overall codon usage of SARS-CoV-2 did not change drastically, in line with the previous literature on this subject, we did observe that while overall GC% content decreased, GC% of the third position in the codon was more positive relative to overall GC% content between February 2021 and July 2023. Over the same interval, we noted that both synonymous and nonsynonymous mutations increased in number, with nonsynonymous mutations outpacing synonymous mutations at a rate of 3:1. We noted that the predicted whole-genome secondary structures nearly all contained the previously described virus-activated inhibitor of translation (VAIT) stem loops, validating for the first time their existence in a whole-genome secondary structure prediction for many SARS-CoV-2 variants (as opposed to previous local secondary structure predictions). We also separately produced a synonymous mutation-deprived set of SARS-CoV-2 variant sequences and repeated the secondary structure calculations on this set. This revealed an interesting trend of reduced ensemble free energy compared to the unaltered variant structures, indicating that synonymous mutations play a role in increasing the free energy of viral RNA molecules. These data both validate previous studies describing increases in viral free energy in human viruses over time and indicate a possible role for synonymous mutations in viral biology.

Indexed as

bioinformaticscodon usageSARS-CoV-2secondary structureVAIT

Identifiers

PMID39882309
PMCPMC11776705

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.