Evidence map›Paper›PMID 40873134›Full record

ReviewJournal of animal science2026

Opportunities and computational challenges in large-scale whole-genome sequencing data analysis.

Hafedh Ben Zaabza, Mohammad H Ferdosi, Ismo Strandén, Beatriz C D Cuyabano, Mahesh Neupane, Ignacy Misztal, Daniela Lourenco, Cedric Gondro

Abstract readReview
In one paragraph

Review in Journal of animal science, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Hafedh Ben ZaabzaDepartment of Animal Science, Michigan State University, East Lansing, MI 48824.
Mohammad H FerdosiAnimal Genetics and Breeding Unit, a joint venture between the NSW Department of Primary Industries and Regional Development, University of New England, Armidale, New South Wales 2351, Australia.
Ismo StrandénNatural Resources Institute Finland (Luke), FI-31600 Jokioinen, Finland.ORCID 0000-0003-0161-2618
Beatriz C D CuyabanoINRAE, AgroParisTech, GABI, Université Paris Saclay, 78350 Jouy-en-Josas, France.
Mahesh NeupaneAnimal Genomics and Improvement, Agricultural Research Service, US Department of Agriculture, Beltsville, MD 20705.
Ignacy MisztalDepartment of Animal and Dairy Science, University of Georgia, Athens, GA 30602.ORCID 0000-0002-0382-1897
Daniela LourencoDepartment of Animal and Dairy Science, University of Georgia, Athens, GA 30602.ORCID 0000-0003-3140-1002
Cedric GondroDepartment of Animal Science, Michigan State University, East Lansing, MI 48824.

Funding

National Institute of Food and Agriculture 2019-67015-29323National Institute of Food and Agriculture 2021-67015-33411
6 · The paper itself

Abstract

Genomic selection has been used in animal breeding for c. 15 yr and continues to be an important tool in predicting genetic merit in livestock populations. The dairy cattle industry was the first to adopt genomic selection, initially based on some 50K single-nucleotide polymorphism (SNP) arrays for thousands of animals. Later advances in genome-scanning technologies have enabled inexpensive genotyping and sequencing, leading to wider adoption, and constantly increasing amounts of genomic data, both as to the number of genotyped animals and variants genotyped per animal. Full sequence data are expected to supersede SNP chips in the coming years. We review the methods and computational approaches used with sequence data and the impact of the methods and model assumptions on genomic prediction accuracy. The modeling, development, and applicability of these methods to sequence data are discussed, as well as the computational resources required. Sequence data should, in principle, provide full information on genetic variability, which should lead to higher prediction accuracy. In practice, there is limited evidence of additional benefit from using sequence data over medium- or high-density SNP panels. This is particularly true for small effective population sizes (Ne) such as cattle populations, where animals within a breed have many common ancestors and thus longer chromosome segments with high linkage disequilibrium accurately trackable with a relatively small number of markers. A population with a small N has long haplotype blocks, from 1 to 5 Mb, making it hard to identify causal variants within blocks. However, in major cattle breeds, a medium-density SNP panel is sufficient to tag the blocks themselves, and prediction with large datasets is highly accurate. Clearly, sequence data should not be used directly for genomic prediction, but for identifying putative causal variants to improve the accuracy and stability of subsequent predictions. We show that the best strategy to deal with any large data with high SNP densities is to use only a subset of (important) markers and determine the most appropriate model for exploiting the preselected variants in the genomic evaluation. Novel prediction methods that subset trait-specific informative markers could offer the advantage of using sequence data by potentially linking individuals through underlying functional variants rather than simply through shared haplotype blocks inherited from ancestors. Further research is required to clarify this aspect.

Indexed as

Computational BiologyGenomeGenomicsWhole Genome SequencingAnimalsBreedingCattlePolymorphism, Single Nucleotidecomputationsgenomic predictiongenomic selectionsequence data

Identifiers

PMID40873134
PMCPMC13220042

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.