Evidence map›Paper›PMID 42211774›Full record

ArticleEuropean heart journal. Digital health2026

Large language models approach clinician performance in ESC cardiovascular risk stratification: a vignette-based benchmark study.

José Ferreira Santos, Regina de Brito Duarte, Inês Mota, Rita Carvalheira Santos, José Maria Moreira, Joana Campos, Nuno André Silva, Bernardo Neves, Ricardo Ladeiras-Lopes, Francisca Leite and 1 more

Abstract read
In one paragraph

Article in European heart journal. Digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

José Ferreira SantosCatólica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.ORCID https://orcid.org/0000-0001-8690-3523
Regina de Brito DuarteInstituto Superior Técnico, Universidade de Lisboa, Lisboa, Portugal.
Inês MotaHospital da Luz Learning Health, Luz Saúde, Lisboa, Portugal.
Rita Carvalheira SantosCardiology Department, Setúbal, Hospital da Luz Setúbal, Luz Saúde, Estrada Nacional 10, Km 37, 2900-722 Setúbal, Portugal.
José Maria MoreiraHospital da Luz Learning Health, Luz Saúde, Lisboa, Portugal.
Joana CamposInstituto Superior Técnico, Universidade de Lisboa, Lisboa, Portugal.
Nuno André SilvaHospital da Luz Learning Health, Luz Saúde, Lisboa, Portugal.
Bernardo NevesCatólica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.
Ricardo Ladeiras-LopesCardiovascular Research and Development Centre-UnIC@RISE, Department of Surgery and Physiology, Faculty of Medicine of the University of Porto, Porto, Portugal.ORCID https://orcid.org/0000-0002-5260-5613
Francisca LeiteCatólica Medical School, Sintra Campus, Estrada OctávioPato, 2635-631 Rio de Mouro, Lisboa, Portugal.
Hélder DoresHospital da Luz Lisboa, Luz Saúde, Lisboa, Portugal.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Aims: Guideline-based cardiovascular risk stratification requires three distinct competencies: extracting risk factor data from clinical text, computing a validated risk score, and applying guideline-defined thresholds to assign a final risk category. We evaluated contemporary large language models (LLMs) on each of these tasks within the European Society of Cardiology (ESC) SCORE2 framework and compared LLM performance against a pooled individual clinician benchmark to contextualize findings against real-world human reproducibility. Methods and results: Eleven LLMs were evaluated using 30 simulated outpatient clinical vignettes presented in both Portuguese and English. For each vignette, models extracted cardiovascular risk factors, determined SCORE2 applicability, generated 10-year risk estimates where appropriate, and assigned a final three-class ESC risk category. A committee of three cardiologists established the reference standard; eight independent clinicians provided an individual-level human benchmark. Traditional risk-factor extraction was near-perfect across all models (micro-F1 0.97-0.99). Agreement with expert-assigned final risk categories was moderate and variable (best: GPT-4o, quadratic-weighted κw 0.69, 95% CI 0.44-0.84), with 10 of 11 models more often underestimating than overestimating risk. To isolate the source of classification error, Conclusion: Contemporary LLMs reliably extract cardiovascular risk information from clinical text, and the best-performing systems achieved agreement within the range of average individual clinicians on this structured task. Their principal limitation lies in downstream computation and rule application.

Indexed as

Artificial intelligenceCardiovascular preventionClinical decision supportLarge language modelsRisk stratificationSCORE2

Identifiers

PMID42211774
PMCPMC13215473

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.