ArticleBMC genomics2023
A new profiling approach for DNA sequences based on the nucleotides' physicochemical features for accurate analysis of SARS-CoV-2 genomes.
Article in BMC genomics, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
6 citing papers in PubMed, 11 citations in OpenAlex.
- Deciphering 6-mer Spectra Distribution Rules in Coronavirus Genomes: Application to Comparative Genomic Analysis.International journal of molecular sciences · 2026Article
- Predicting VNN resistance in European sea bass using machine learning on high dimensional low sample size data.Frontiers in bioinformatics · 2026Article
- PC-mer: An Ultra-fast memory-efficient tool for metagenomics profiling and classification.PloS one · 2024Article
- Bioinformatics tools for the sequence complexity estimates.Biophysical reviews · 2023Review
- BGRS: bioinformatics of genome regulation and data integration.Journal of integrative bioinformatics · 2023Article
- The published trend of studies on COVID-19 and diabetes: bibliometric analysis.Frontiers in endocrinology · 2023Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors at 1 institution in 1 country.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe prevalence of the COVID-19 disease in recent years and its widespread impact on mortality, as well as various aspects of life around the world, has made it important to study this disease and its viral cause. However, very long sequences of this virus increase the processing time, complexity of calculation, and memory consumption required by the available tools to compare and analyze the sequences.
resultsWe present a new encoding method, named PC-mer, based on the k-mer and physic-chemical properties of nucleotides. This method minimizes the size of encoded data by around 2
conclusionsPC-mer achieves 100% accuracy despite the use of very simple classification algorithms based on Machine Learning. Assuming dynamic programming-based pairwise alignment as the ground truth approach, we achieved a degree of convergence of more than 98% for coronavirus genus-level sequences and 93% for SARS-CoV-2 sequences using PC-mer in the alignment-free classification method. This outperformance of PC-mer suggests that it can serve as a replacement for alignment-based approaches in certain sequence analysis applications that rely on similarity/dissimilarity scores, such as searching sequences, comparing sequences, and certain types of phylogenetic analysis methods that are based on sequence comparison.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.