Evidence map›Paper›PMID 40111256›Full record

ReviewMolecular biology and evolution2025

K-mer-based Approaches to Bridging Pangenomics and Population Genetics.

Miles D Roberts, Olivia Davis, Emily B Josephs, Robert J Williamson

Abstract readReview
In one paragraph

Review in Molecular biology and evolution, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 15 papers.

0numbers the graph read from it
0cells of the map it votes in
15citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

15 citing papers in PubMed.

  1. Review
  2. Review
  3. Review
  4. Review
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Computational function prediction of bacteria and phage proteins.Microbiology and molecular biology reviews : MMBR · 2025
    Review
  14. Evolution letters · 2025
    Article
  15. Article
4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

4 authors.

Miles D RobertsGenetics and Genome Sciences Program, Michigan State University, East Lansing, MI 48824, USA.ORCID 0000-0001-9854-701X
Olivia DavisDepartment of Computer Science and Software Engineering, Rose-Hulman Institute of Technology, Terre Haute, IN 47803, USA.ORCID 0009-0007-5571-9708
Emily B JosephsDepartment of Plant Biology, Michigan State University, East Lansing, MI 48824, USA.ORCID 0000-0002-0889-1130
Robert J WilliamsonDepartment of Computer Science and Software Engineering, Rose-Hulman Institute of Technology, Terre Haute, IN 47803, USA.ORCID 0000-0001-9732-0964

Funding

Determining the evolutionary forces shaping genotype-by-environment interactionsR35GM142829 · NIGMS · MICHIGAN STATE UNIVERSITY · PI JOSEPHS, EMILY · 2021 to 2025
$1.9M
National Science Foundation IOS-1934384NIGMS NIH HHS R35 GM142829NIGMS NIH HHS T32-GM110523NIH HHS GM142829
6 · The paper itself

Abstract

Many commonly studied species now have more than one chromosome-scale genome assembly, revealing a large amount of genetic diversity previously missed by approaches that map short reads to a single reference. However, many species still lack multiple reference genomes and correctly aligning references to build pangenomes can be challenging for many species, limiting our ability to study this missing genomic variation in population genetics. Here, we argue that k-mers are a very useful but underutilized tool for bridging the reference-focused paradigms of population genetics with the reference-free paradigms of pangenomics. We review current literature on the uses of k-mers for performing three core components of most population genetics analyses: identifying, measuring, and explaining patterns of genetic variation. We also demonstrate how different k-mer-based measures of genetic variation behave in population genetic simulations according to the choice of k, depth of sequencing coverage, and degree of data compression. Overall, we find that k-mer-based measures of genetic diversity scale consistently with pairwise nucleotide diversity (π) up to values of about π=0.025 (R2=0.97) for neutrally evolving populations. For populations with even more variation, using shorter k-mers will maintain the scalability up to at least π=0.1. Furthermore, in our simulated populations, k-mer dissimilarity values can be reliably approximated from counting bloom filters, highlighting a potential avenue to decreasing the memory burden of k-mer-based genomic dissimilarity analyses. For future studies, there is a great opportunity to further develop methods to identifying selected loci using k-mers.

Indexed as

Genetics, PopulationGenomicsAnimalsGenetic VariationHumansbloom filterk-merpangenomicspopulation genomics

Identifiers

PMID40111256
PMCPMC11925024

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.