ReviewMolecular biology and evolution2025
K-mer-based Approaches to Bridging Pangenomics and Population Genetics.
Review in Molecular biology and evolution, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
15 citing papers in PubMed.
- Building and applying pangenome references to capture genetic diversity.Nature reviews. Genetics · 2026Review
- Advances in gene cloning and functional genomics approaches for wheat (Triticum aestivum L.) improvement.The plant genome · 2026Review
- Leveraging AI and integrated genomic-enviromic prediction for intelligent sugarcane breeding.Plant communications · 2026Review
- ac4C modification sites prediction in human mRNA: a complete review.Briefings in bioinformatics · 2026Review
- Introgression across ploidies contributes to genetic diversity in introduced urbanbioRxiv : the preprint server for biology · 2026Article
- A foundational quantum framework for multi-pattern string matching in k-mer detection.Frontiers in bioinformatics · 2026Article
- Identification of genome size and heterozygosity in 510 Jujube (Frontiers in plant science · 2026Article
- Article
- Quantum implementation of multi-pattern string matching for k-mer detection.bioRxiv : the preprint server for biology · 2025Article
- Augmenting small tabular health data for training prognostic ensemble machine learning models using generative models.BMC medical informatics and decision making · 2025Article
- Strainify: Strain-Level Microbiome Profiling for Low-Coverage Short-Read Metagenomic Datasets.bioRxiv : the preprint server for biology · 2025Article
- Natural variation in regulatory code revealed through Bayesian analysis of plant pan-genomes and pan-transcriptomes.bioRxiv : the preprint server for biology · 2025Article
- Computational function prediction of bacteria and phage proteins.Microbiology and molecular biology reviews : MMBR · 2025Review
- Article
- Independent domestication and cultivation histories of two West African indigenous fonio millet crops.Nature communications · 2025Article
Corrections and comments
- Update of
Authors and funding
4 authors.
Funding
Abstract
Many commonly studied species now have more than one chromosome-scale genome assembly, revealing a large amount of genetic diversity previously missed by approaches that map short reads to a single reference. However, many species still lack multiple reference genomes and correctly aligning references to build pangenomes can be challenging for many species, limiting our ability to study this missing genomic variation in population genetics. Here, we argue that k-mers are a very useful but underutilized tool for bridging the reference-focused paradigms of population genetics with the reference-free paradigms of pangenomics. We review current literature on the uses of k-mers for performing three core components of most population genetics analyses: identifying, measuring, and explaining patterns of genetic variation. We also demonstrate how different k-mer-based measures of genetic variation behave in population genetic simulations according to the choice of k, depth of sequencing coverage, and degree of data compression. Overall, we find that k-mer-based measures of genetic diversity scale consistently with pairwise nucleotide diversity (π) up to values of about π=0.025 (R2=0.97) for neutrally evolving populations. For populations with even more variation, using shorter k-mers will maintain the scalability up to at least π=0.1. Furthermore, in our simulated populations, k-mer dissimilarity values can be reliably approximated from counting bloom filters, highlighting a potential avenue to decreasing the memory burden of k-mer-based genomic dissimilarity analyses. For future studies, there is a great opportunity to further develop methods to identifying selected loci using k-mers.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.