Evidence map›Paper›PMID 42669697›Full record

ArticleNature communications2026

An embedding-based framework enables statistical testing of gene-set function hypotheses inferred by large language models.

Yanhao Tan, Li-Ju Wang, Tianyuzhou Liang, Ying-Ju Lai, Chien-Hung Shih, Yibing Guo, Tyler M Yasaka, George C Tseng, Yu-Chiao Chiu

Abstract read
In one paragraph

Article in Nature communications, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Yanhao TanUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Li-Ju WangUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Tianyuzhou LiangUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Ying-Ju LaiUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Chien-Hung ShihUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Yibing GuoUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.
Tyler M YasakaUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA.ORCID 0000-0002-8482-0369
George C TsengDepartment of Biostatistics and Health Data Science, University of Pittsburgh, Pittsburgh, PA, USA.ORCID 0000-0002-5447-1014
Yu-Chiao ChiuUPMC Hillman Cancer Center, University of Pittsburgh School of Medicine, Pittsburgh, PA, USA. YUC250@pitt.edu.ORCID 0000-0003-1647-8634

Funding

VECTOR CORE FACILITYP30CA047904 · NCI · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI CHRISTOPHER J. BAKKENIST · 1988 to 2026
$158.0M
Cellular Approaches to Tissue Engineering/RegenerationT32EB001026 · NIBIB · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI DUNCAN, ANDREW W, MONGA, SATDARSHAN SINGH · 2003 to 2024
$5.5M
Single-cell congruence evaluation and selection of cancer models towards precision medicineR01CA285337 · NCI · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Adrian V Lee, George C. Tseng · 2025 to 2026
$1.2M
Novel computational approaches for pharmacogenomics of complex diseasesR35GM154967 · NIGMS · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Yu-Chiao Chiu · 2024 to 2026
$1.2M
Enhancing AI-readiness of multi-omics data for cancer pharmacogenomicsR00CA248944 · NCI · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI CHIU, YU-CHIAO · 2022 to 2024
$1.1M
High-Throughput Computing for Genomics and Bioinformatics ResearchS10OD028483 · OD · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI LEE, ADRIAN V · 2021 to 2021
$574k
An Integrated In Silico and In Vivo Genetic Screening Approach to Identify Subtype-specific Hepatocellular Carcinoma Genetic DependenciesF30CA298277 · NCI · UNIVERSITY OF PITTSBURGH AT PITTSBURGH · PI Tyler Yasaka · 2025 to 2026
$105k
NCI NIH HHS F30 CA298277NCI NIH HHS P30 CA047904NCI NIH HHS R00 CA248944NCI NIH HHS R01 CA285337NIBIB NIH HHS T32 EB001026NIGMS NIH HHS R35 GM154967NIH HHS S10 OD028483U.S. Department of Health & Human Services | National Institutes of Health (NIH) F30CA298277, T32EB001026U.S. Department of Health & Human Services | National Institutes of Health (NIH) R00CA248944, R35GM154967, P30CA047904U.S. Department of Health & Human Services | National Institutes of Health (NIH) R01CA285337
6 · The paper itself

Abstract

Emerging large language models (LLMs) can infer gene functions directly from gene lists, enabling hypothesis generation without predefined gene sets. However, these LLM-derived predictions are qualitative, and principled statistical validation is lacking. Here, we develop an embedding-based statistical framework that transforms gene and function descriptions into vector representations, enabling statistical testing of gene-gene and gene-function relationships and quantitative prioritization of de novo functional hypotheses inferred by LLMs. We benchmark seven state-of-the-art embedding models using curated and retrieval-augmented literature-derived gene descriptions across diverse biological contexts. OpenAI's text-embedding-3-large and Google's gemini-embedding-001 perform best, capturing gene-gene functional relationships in 88.7-92.5% of Gene Ontology biological processes and approximately 98.6% of canonical pathways. In gene-function association analyses, these models achieve high sensitivity (95.2-98.4%) and specificity (72.7-84.3%). Through contamination analysis and evaluation using experimentally informed protein assembly gene sets, our framework distinguishes biologically meaningful LLM-inferred hypotheses from noise, outperforming confidence-based inference and conventional enrichment analysis. We further develop the open-source R package DEGEmbedR and demonstrate its utility for interpreting a drug perturbation-derived differentially expressed gene (DEG) signature lacking significant conventional enrichment results. Together, these results establish LLM-derived embeddings as a quantitative foundation for functional genomics and the statistical validation of LLM-based gene function inference.

Indexed as

Computational BiologyLarge Language ModelsGene OntologyHumans

Identifiers

PMID42669697
PMCPMC13526818

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.