Evidence map›Paper›PMID 39589438›Full record

ArticleGigaScience2024

AltaiR: a C toolkit for alignment-free and temporal analysis of multi-FASTA data.

Jorge M Silva, Armando J Pinho, Diogo Pratas

Abstract read
In one paragraph

Article in GigaScience, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Jorge M SilvaIEETA/LASI, Institute of Electronics and Informatics Engineering of Aveiro, University of Aveiro, Aveiro, Portugal.ORCID 0000-0002-6331-6091
Armando J PinhoIEETA/LASI, Institute of Electronics and Informatics Engineering of Aveiro, University of Aveiro, Aveiro, Portugal.ORCID 0000-0002-9164-0016
Diogo PratasIEETA/LASI, Institute of Electronics and Informatics Engineering of Aveiro, University of Aveiro, Aveiro, Portugal.ORCID 0000-0003-1176-552X

Funding

EC 101081813Foundation for Science and Technology UIDB/00127/2020
6 · The paper itself

Abstract

backgroundMost viral genome sequences generated during the latest pandemic have presented new challenges for computational analysis. Analyzing millions of viral genomes in multi-FASTA format is computationally demanding, especially when using alignment-based methods. Most existing methods are not designed to handle such large datasets, often requiring the analysis to be divided into smaller parts to obtain results using available computational resources.

findingsWe introduce AltaiR, a toolkit for analyzing multiple sequences in multi-FASTA format using exclusively alignment-free methodologies. AltaiR enables the identification of singularity and similarity patterns within sequences and computes static and temporal dynamics without restrictions on the number or size of input sequences. It automatically filters low-quality, biased, or deviant data. We demonstrate AltaiR's capabilities by analyzing more than 1.5 million full severe acute respiratory virus coronavirus 2 sequences, revealing interesting observations regarding viral genome characteristics over time, such as shifts in nucleotide composition, decreases in average Kolmogorov sequence complexity, and the evolution of the smallest sequences not found in the human host.

conclusionsAltaiR can identify temporal characteristics and trends in large numbers of sequences, making it ideal for scenarios involving endemic or epidemic outbreaks with vast amounts of available sequence data. Implemented in C with multithreading and methodological optimizations, AltaiR is computationally efficient, flexible, and dependency-free. It accepts any sequence in FASTA format, including amino acid sequences. The complete toolkit is freely available at https://github.com/cobilab/altair.

Indexed as

Computational BiologyCOVID-19Genome, ViralSARS-CoV-2SoftwareAlgorithmsHumansPandemicsSequence Alignmentalignment-free toolkitdata compressionmulti-FASTArelative absent wordstemporal patternsviral genomes

Identifiers

PMID39589438
PMCPMC11590114

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.