Evidence map›Paper›PMID 41415428›Full record

ArticlebioRxiv : the preprint server for biology2026

Machine learning-based prediction of human structural variation and characterization of associated sequence determinants.

Daven Lim, Runyang Nicolas Lou, Nilah Ioannidis, Peter H Sudmant

Abstract readPreprint
In one paragraph

Article in bioRxiv : the preprint server for biology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Daven LimDepartment of Biosystems Science and Engineering, ETH Zürich, Zürich, Switzerland.
Runyang Nicolas LouDepartment of Integrative Biology, University of California Berkeley, Berkeley, CA, USA.
Nilah IoannidisDepartment of Applied Math, University of California Santa Cruz, Santa Cruz, CA, USA.
Peter H SudmantDepartment of Integrative Biology, University of California Berkeley, Berkeley, CA, USA.ORCID 0000-0002-9573-8248

Funding

The evolution and diversity of mutation, molecular fidelity, and genome structureR35GM142916 · NIGMS · UNIVERSITY OF CALIFORNIA BERKELEY · PI Peter Heshedahl Sudmant · 2021 to 2026
$2.5M
A compendium of complete primate reference genomes to facilitate conservation, genomics, and ecologyR01HG013017 · NHGRI · UNIVERSITY OF CALIFORNIA BERKELEY · PI Erik Garrison, Matthew William Mitchell · 2023 to 2026
$1.9M
NHGRI NIH HHS R01 HG013017NIGMS NIH HHS R35 GM142916
6 · The paper itself

Abstract

Structural variants (SVs) represent a major source of genetic diversity and play key roles in human disease and evolution. Yet, the extent to which local sequence context shapes the likelihood of structural variant formation remains poorly quantified. Here, we develop machine learning models to predict the occurrence of SVs across the human genome and characterize genomic determinants associated with their formation. We developed both a sequence only-based convolutional neural network (CNN) model as well as a random forest approach integrating diverse genomic annotations. Both models achieve high predictive performance individually (>90% AUROC) which can be further improved in an ensemble. The predictive ability of these models demonstrates that SV-prone regions can be accurately inferred from sequence context. Model interpretability techniques reveal key genomic contributors to SVs, including effects of sequence motifs such as microhomology and non-canonical DNA structures, as well as the presence of SV hotspots. We find that different classes of SVs exhibit distinct sequence determinants, with transposable elements and inversions displaying particularly unique signatures. Moreover, predicted SV probability correlates with allele frequency and gene functional constraint, indicating the potential utility of the model for variant effect prediction. These findings demonstrate that machine learning models trained on local sequence features can identify unstable genomic regions and provide a framework for quantifying SV susceptibility and SV variant effects in personalized genomics.

Identifiers

PMID41415428
PMCPMC12710826

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.