Evidence map›Paper›PMID 41553353›Full record

ArticleGigaScience2026

Federated knowledge retrieval elevates large language model performance on biomedical benchmarks.

Janet Joy, Andrew I Su

Abstract read
In one paragraph

Article in GigaScience, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

2 authors.

Janet JoyDepartment of Integrative Structural and Computational Biology, Scripps Research, 10550 N Torrey Pines Rd, La Jolla, CA, 92037, USA.ORCID 0000-0002-8871-0765
Andrew I SuDepartment of Integrative Structural and Computational Biology, Scripps Research, 10550 N Torrey Pines Rd, La Jolla, CA, 92037, USA.ORCID 0000-0002-9859-4104

Funding

Scripps Clinical and Translational Science HubUM1TR004407 · NCATS · SCRIPPS RESEARCH INSTITUTE, THE · PI Eric Jeffrey Topol · 2023 to 2026
$24.2M
BioThings Explorer: A platform for distributed knowledge integration across biomedical APIsOT2TR003427 · NCATS · SCRIPPS RESEARCH INSTITUTE, THE · PI SU, ANDREW I · 2020 to 2024
$5.2M
Compound repositioning for Alzheimer's Disease using knowledge graphs, insurance claims data, and gene expression complementarityR01AG066750 · NIA · SCRIPPS RESEARCH INSTITUTE, THE · PI SU, ANDREW I · 2020 to 2024
$4.4M
DOGSURF: Database Optimized for Graph Search Using Robust FederationOT2TR005710 · NCATS · SCRIPPS RESEARCH INSTITUTE, THE · PI ANDREW I SU, Chunlei Wu · 2025 to 2026
$3.5M
NCATS NIH HHS 1OT2TR003427NCATS NIH HHS 1OT2TR005710NCATS NIH HHS OT2 TR003427NCATS NIH HHS OT2 TR005710NCATS NIH HHS UM1 TR004407NIA NIH HHS R01 AG066750NIA NIH HHS R01AG066750Scripps Research Translational Institute UM1TR004407
6 · The paper itself

Abstract

backgroundLarge language models (LLMs) have significantly advanced natural language processing in biomedical research; however, their reliance on implicit, statistical representations often results in factual inaccuracies or hallucinations, posing significant concerns in high-stakes biomedical contexts.

resultsTo overcome these limitations, we developed BioThings Explorer-Retrieval-Augmented Generation (BTE-RAG), a Retrieval-Augmented Generation framework that integrates the reasoning capabilities of advanced language models with explicit mechanistic evidence sourced from BTE, an API federation of more than sixty authoritative biomedical knowledge sources. We systematically evaluated BTE-RAG in comparison to traditional LLM-only methods across three benchmark datasets that we created from DrugMechDB. These datasets specifically targeted gene-centric mechanisms (798 questions), metabolite effects (201 questions), and drug-biological process relationships (842 questions). On the gene-centric task, BTE-RAG increased accuracy from 51 to 75.8% for GPT-4o mini and from 69.8 to 78.6% for GPT-4o. In metabolite-focused questions, the proportion of responses with cosine similarity scores of at least 0.90 rose by 82% for GPT-4o mini and 77% for GPT-4o. While overall accuracy was consistent in the drug-biological process benchmark, the retrieval method enhanced response concordance, producing a greater than 10% increase in high-agreement answers (from 129 to 144) using GPT-4o. We additionally evaluated BTE-RAG alongside GeneGPT-based models on the GeneTuring gene-disease association benchmark and on our mechanistic gene benchmark, demonstrating that the BTE-RAG layer consistently improves accuracy relative to alternative approaches.

conclusionFederated knowledge retrieval provides transparent improvements in accuracy for LLMs, establishing BTE-RAG as a valuable and practical tool for mechanistic exploration and translational biomedical research.

Indexed as

Biomedical ResearchLarge Language ModelsNatural Language ProcessingBenchmarkingHumansbiomedical knowledge graphsBioThings ExplorerDrugMechDB benchmarkingfederated knowledge retrievalhallucination mitigationlarge language models (LLMs)mechanistic reasoningRetrieval-augmented generation (RAG)

Identifiers

PMID41553353
PMCPMC12888809

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.