Evidence map›Paper›PMID 40923439›Full record

ArticleJournal of animal science2025

A vision of how low-coverage sequence data should contribute to genetic evaluation in the future.

Richard Mark Thallman, Jacqueline E Borgert, Bailey N Engle, John W Keele, Warren M Snelling, Cedric Gondro, Larry A Kuehn

Abstract read
In one paragraph

Article in Journal of animal science, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Review
  2. Article
  3. Review
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Richard Mark ThallmanUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.ORCID 0000-0002-5114-9724
Jacqueline E BorgertUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.
Bailey N EngleUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.ORCID 0000-0003-2360-1012
John W KeeleUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.ORCID 0000-0002-8697-4564
Warren M SnellingUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.ORCID 0000-0001-6282-9728
Cedric GondroDepartment of Animal Science, Michigan State University, East Lansing, MI, USA.
Larry A KuehnUSDA, ARS, U.S. Meat Animal Research Center, Clay Center, NE, USA.ORCID 0000-0002-7573-5365

Funding

USDA ARS CRIS 3040-31000-104-00D
6 · The paper itself

Abstract

Low-coverage sequencing refers to sequencing DNA of individuals to a low depth of coverage (e.g., 0.5X) and imputing that sequence to a genomic sequence based on reference haplotypes from individuals sequenced to a high depth of coverage (e.g., ≥10X). It has been proposed as an alternative to genotyping by Single-nucleotide polymorphisms (SNP) arrays. At least one commercial product based on it is available for agricultural species. Concerns limiting adoption in its current form are: 1) the cost of storing the huge volume of data it generates and 2) whether that additional data will result in improved accuracy of genetic evaluation. This work envisions future implementation of low-coverage sequencing to reduce storage costs and enhance genetic evaluations by leveraging the additional information in the full sequence of the pangenome to account for more genetic variation. We propose addressing the storage issue by representing genomic sequence of an individual in a pair of haplotype arrays with each element pointing to an enumerated haplotype of the sequence within one of approximately 50,000 defined genome segments. Assuming 60 million genomic variants, the infrastructure required to translate the identifier of any enumerated haplotype into its genomic sequence would require less than 10 gigabytes of binary storage. Each haplotype array element would require 2 bytes, so the marginal binary storage required to represent the genomic sequence of an individual would be about 200 kilobytes (KB), similar to the genotypes from a SNP array with 200,000 markers. This assumes no pedigree and no ambiguity of the imputation, though the latter is unrealistic. Strategies to minimize, and when necessary, to manage and efficiently represent ambiguity are proposed. The genomic sequence of an individual could be stored in about 1 KB (binary) if both parents have unambiguous sequences stored as described above. The proposed system for representing the pangenome includes algorithms for read mapping and imputation intended to leverage all known genetic variation in the target population. It is also designed to use sequencing reads generated for imputing the genomic sequence of new individuals to identify unrecognized mutations, crossovers, and structural variants, thus continuously improving the genome representation, especially if widespread use of low-coverage sequencing in livestock industries is realized. This could make improved genetic merit and management of livestock feasible without computational burden.

Indexed as

GenomicsSequence Analysis, DNAAnimalsHaplotypesPolymorphism, Single Nucleotidegenomic evaluationgenotypinglow-coverage sequencinglow-pass sequencingpangenomeskim sequencing

Identifiers

PMID40923439
PMCPMC12559786

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.