Evidence map›Paper›PMID 41818295›Full record

ArticlePLOS digital health2026

Cardiology knowledge assessment of retrieval-augmented open versus proprietary large language models.

Constantine Tarabanis, Shaan Khurshid, Areti Karamanou, Rodo Piperaki, Lucas A Mavromatis, Aris Hatzimemos, Dimitrios Tachmatzidis, Constantinos Bakogiannis, Vassilios Vassilikos, Patrick T Ellinor and 2 more

Abstract read
In one paragraph

Article in PLOS digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Validating LLM judges for automated oversight of patient communication.medRxiv : the preprint server for health sciences · 2026
    Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Constantine TarabanisCardiology Division, Heart and Vascular Institute, Mass General Brigham, Boston, Massachusetts, United States of America.ORCID https://orcid.org/0000-0001-7563-2430
Shaan KhurshidCardiovascular Disease Initiative, Broad Institute of MIT and Harvard, Cambridge, Massachusetts, United States of America.
Areti KaramanouInformation Systems Laboratory, University of Macedonia, Thessaloniki, Greece.
Rodo PiperakiInformation Systems Laboratory, University of Macedonia, Thessaloniki, Greece.
Lucas A MavromatisNew York University School of Medicine, New York City, New York, United States of America.
Aris HatzimemosNew York University School of Medicine, New York City, New York, United States of America.
Dimitrios Tachmatzidis3rd Cardiology Department, Hippokrateion University Hospital, Aristotle University of Thessaloniki, Thessaloniki, Greece.
Constantinos Bakogiannis3rd Cardiology Department, Hippokrateion University Hospital, Aristotle University of Thessaloniki, Thessaloniki, Greece.
Vassilios Vassilikos3rd Cardiology Department, Hippokrateion University Hospital, Aristotle University of Thessaloniki, Thessaloniki, Greece.
Patrick T EllinorCardiovascular Disease Initiative, Broad Institute of MIT and Harvard, Cambridge, Massachusetts, United States of America.
Lior JankelsonLeon H. Charney Division of Cardiology, NYU Langone Health, New York University School of Medicine, New York City, New York, United States of America.
Evangelos KalampokisInformation Systems Laboratory, University of Macedonia, Thessaloniki, Greece.

Funding

IDENTIFICATION OF COMMON GENETIC VARIANTS FOR ATRIAL FIBRILLATION AND PR INTERVALR01HL092577 · NHLBI · MASSACHUSETTS GENERAL HOSPITAL · PI BENJAMIN, EMELIA J., ELLINOR, PATRICK THOMAS · 2009 to 2025
$20.6M
Using Electrocardiogram Genetics to Inform Arrhythmia RiskR01HL157635 · NHLBI · MASSACHUSETTS GENERAL HOSPITAL · PI ELLINOR, PATRICK THOMAS, MIRSHAHI, TOORAJ · 2022 to 2025
$2.9M
From genetic basis to mechanisms for Heart FailureR01HL177209 · NHLBI · BROAD INSTITUTE, INC. · PI Patrick Thomas Ellinor, Ling Xiao · 2025 to 2026
$1.6M
Electrocardiogram-based deep learning and decision analysis to improve atrial fibrillation risk estimationK23HL169839 · NHLBI · MASSACHUSETTS GENERAL HOSPITAL · PI Shaan Khurshid · 2023 to 2026
$860k
NHLBI NIH HHS K23 HL169839NHLBI NIH HHS R01 HL092577NHLBI NIH HHS R01 HL157635NHLBI NIH HHS R01 HL177209
6 · The paper itself

Abstract

To evaluate the performance of open-weight and proprietary LLMs, with and without Retrieval-Augmented Generation (RAG), on cardiology board-style questions and benchmark them against the human average. We tested 14 LLMs (6 open-weight, 8 proprietary) on 449 multiple-choice questions from the American College of Cardiology Self-Assessment Program (ACCSAP). Accuracy was measured as percent correct. RAG was implemented using a knowledge base of 123 guideline and textbook documents. The open-weight model DeepSeek R1 achieved the highest accuracy at 86.9% (95% CI: 83.4-89.7%), outperforming proprietary models and the human average of 78%. GPT 4o (80.9%, 95% CI: 77.0-84.2%) and the commercial platform OpenEvidence (81.3%, 95% CI: 77.4-84.7%) demonstrated similar performance. A positive correlation between model size and performance was observed within model families, but across families, substantial variability persisted among models with similar parameter counts. After RAG, all models improved, and open-weight models like Mistral Large 2 (78.0%, 95% CI: 73.9-81.5) performed comparably to proprietary alternatives like GPT 4o. Large language models (LLMs) are increasingly integrated into clinical workflows, yet their performance in cardiovascular medicine remains insufficiently evaluated. Open-weight models can match or exceed proprietary systems in cardiovascular knowledge, with RAG particularly beneficial for smaller models. Given their transparency, configurability, and potential for local deployment, open-weight models, strategically augmented, represent viable, lower-cost alternatives for clinical applications. Open-weight LLMs demonstrate competency in cardiovascular medicine comparable to or exceeding that of proprietary models, with and without RAG depending on the model.

Identifiers

PMID41818295
PMCPMC12981508

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.