Evidence map›Paper›PMID 41989882›Full record

SynthesisJMIR cardio2026

Large Language Models in Cardiology: Systematic Review.

Moran Gendler, Girish N Nadkarni, Karin Sudri, Michal Cohen-Shelly, Benjamin S Glicksberg, Orly Efros, Shelly Soffer, Eyal Klang

Abstract readSystematic Review
In one paragraph

Synthesis in JMIR cardio, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Moran Gendler *Azrieli Faculty of Medicine, Bar-Ilan University, Henrietta Szold St 8, Safed, Israel, Safed, 1311502, Israel, 972 542354444.ORCID http://orcid.org/0009-0004-7789-3801
Girish N NadkarniWindreich Department of AI and Human Health, Mount Sinai Medical Center, Mount Sinai, New York, NY, United States.ORCID http://orcid.org/0000-0001-6319-4314
Karin SudriSagol AI Hub, ARC Innovation Center, Sheba Medical Center, Ramat Gan, Israel.ORCID http://orcid.org/0000-0002-5118-8619
Michal Cohen-ShellySagol AI Hub, ARC Innovation Center, Sheba Medical Center, Ramat Gan, Israel.ORCID http://orcid.org/0009-0007-6532-9252
Benjamin S GlicksbergWindreich Department of AI and Human Health, Mount Sinai Medical Center, Mount Sinai, New York, NY, United States.ORCID http://orcid.org/0000-0003-4515-8090
Orly EfrosSchool of Medicine, Tel Aviv University, Tel Aviv, Israel.ORCID http://orcid.org/0000-0002-6024-9110
Shelly Soffer *School of Medicine, Tel Aviv University, Tel Aviv, Israel.ORCID http://orcid.org/0000-0002-7853-2029
Eyal Klang *Windreich Department of AI and Human Health, Mount Sinai Medical Center, Mount Sinai, New York, NY, United States.ORCID http://orcid.org/0000-0002-4567-3108

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) are increasingly used in health care, but their role in cardiology has not yet been systematically evaluated. Objective: This review aimed to assess the applications, performance, and limitations of LLMs across diverse cardiology tasks, including chronic and progressive conditions, acute events, education, and diagnostic testing. Methods: A systematic search was conducted in PubMed and Scopus for studies published up to April 14, 2024, using keywords related to LLMs and cardiology. Studies evaluating LLM outputs in cardiology-related tasks were included. Data were extracted across 5 predefined domains and the risk of bias was assessed using an adapted QUADAS-2 tool (developed by Whiting et al at the University of Bristol). The review protocol was registered in PROSPERO (CRD42024556397). Results: A total of 33 studies contributed quantitative outcome data to a descriptive synthesis. Across chronic conditions, ChatGPT-3.5 (OpenAI) answered 91% (43/47) heart failure questions accurately, although readability often required college-level comprehension. In acute scenarios, Bing Chat omitted key myocardial infarction first aid steps in 25% (5/20) to 45% (9/20) of cases, while cardiac arrest information was rated highly (mean 4.3/5, SD 0.7) but written above recommended reading levels. In physician education tasks, ChatGPT-4 (OpenAI) demonstrated higher accuracy than ChatGPT-3.5, improving from 38% (33/88) to 66% (58/88). In patient education studies, ChatGPT-4 provided scientifically adequate explanations (5.0-6.0/7) comparable to hospital materials but at higher reading levels (11th vs 7th grade). In diagnostic testing, ChatGPT-4 interpreted 91% (36/40) electrocardiogram vignettes correctly, significantly better than emergency physicians (31/40, 77%; P< .001), but showed lower performance in echocardiography. Conclusions: LLMs show meaningful potential in cardiology, especially for education and electrocardiogram interpretation, but performance varies across clinical tasks. Limitations in emergency guidance and readability, as well as small in silico study designs, highlight the need for multimodal models and prospective validation.

Indexed as

CardiologyLarge Language ModelsHumansartificial intelligencecardiologygenerative AIlarge language modelsLLMsnatural language processing

Identifiers

PMID41989882
PMCPMC13085985

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.