Evidence map›Paper›PMID 41867208›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Language models reveal evidence gaps in variants of uncertain significance.

Weijiang Li, Vineel Bhat, Tian Yu, Matthew Lebo, Marinka Zitnik, Christopher Cassa

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Weijiang LiBrigham and Women's Hospital / Harvard Medical School, Boston, MA, USA.ORCID 0009-0009-1304-3513
Vineel BhatBrigham and Women's Hospital / Harvard Medical School, Boston, MA, USA.
Tian YuBrigham and Women's Hospital / Harvard Medical School, Boston, MA, USA.
Matthew LeboMass General Brigham Personalized Medicine, Boston, MA, USA.
Marinka ZitnikHarvard Medical School, Boston, MA, USA.
Christopher CassaBrigham and Women's Hospital / Harvard Medical School, Boston, MA, USA.

Funding

From Text to Translation: Using Language Models to Resolve and Classify VariantsR21HG014015 · NHGRI · BRIGHAM AND WOMEN'S HOSPITAL · PI Christopher Cassa · 2025 to 2026
$506k
NHGRI NIH HHS R21 HG014015
6 · The paper itself

Abstract

Background: Most rare coding variants in monogenic disease genes remain classified as Variants of Uncertain Significance (VUS), limiting their use in clinical care. Many variant classifications have been submitted to ClinVar, often with rich free-text summaries of the evidence underlying each classification. These narratives are not standardized and are difficult to mine systematically, making it challenging to identify variants that might be reclassified as new evidence becomes available. Methods: We developed a two-stage language-model pipeline that (i) detects whether functional, population, or computational evidence is described in ClinVar and ClinGen variant summaries, and (ii) classifies whether it is evidence of pathogenicity or benignity. We first constructed Variant Evidence Text Annotations (VETA), a dataset of 44,522 ACMG/AMP keyword-description pairs derived from 18,678 ClinVar and ClinGen variant summaries using an LLM-based consensus annotation procedure. We then fine-tuned BioBERT-large models for each evidence type and stage, and validated performance using independent ClinGen expert-curated summaries as well as orthogonal variantlevel evidence, including functional screening, computational scores, and population estimates of disease impact. Results: Across evidence types, our models accurately identify whether functional, population, and computational evidence is present and whether it leans toward a pathogenic or benign impact. We find high agreement with ClinGen expert annotations and highly significant separation of validation scores between model-predicted benign and pathogenic groups (functional assays Conclusions: Transforming unstructured variant summaries into a structured, evidence-type matrix enables scalable detection of evidence gaps, allowing for the systematic integration of new data sources, and prioritization of VUS that are most likely to be reclassified. This language model-enabled pipeline provides a generalizable digital approach to identify clinical evidence gaps as functional screens, biobank resources, and computational predictors continue to evolve.

Indexed as

ClinVargenetic diagnosticslarge language modelsvariant classification

Identifiers

PMID41867208
PMCPMC13004083

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.