Evidence map›Paper›PMID 40105654›Full record

ArticleJournal of the American Medical Informatics Association : JAMIA2025

Utilizing large language models for detecting hospital-acquired conditions: an empirical study on pulmonary embolism.

Cheligeer Cheligeer, Danielle A Southern, Jun Yan, Guosong Wu, Jie Pan, Seungwon Lee, Elliot A Martin, Hamed Jafarpour, Cathy A Eastwood, Yong Zeng and 1 more

Abstract read
In one paragraph

Article in Journal of the American Medical Informatics Association : JAMIA, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Cheligeer CheligeerCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0009-0007-0810-3011
Danielle A SouthernCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0000-0002-0006-0033
Jun YanConcordia Institute for Information Systems Engineering, Concordia University, Montreal H3G 2W1, Canada.ORCID 0000-0002-5148-1399
Guosong WuCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.
Jie PanCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0000-0001-6398-1756
Seungwon LeeCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.
Elliot A MartinCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0000-0001-5127-4333
Hamed JafarpourConcordia Institute for Information Systems Engineering, Concordia University, Montreal H3G 2W1, Canada.ORCID 0009-0007-9410-7675
Cathy A EastwoodCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0000-0002-4569-8014
Yong ZengConcordia Institute for Information Systems Engineering, Concordia University, Montreal H3G 2W1, Canada.ORCID 0000-0001-6678-271X
Hude QuanCentre for Health Informatics, Cumming School of Medicine, University of Calgary, Calgary T2N 4N1, Canada.ORCID 0000-0002-7848-7256

Funding

OLFACTORY RECEPTORSF32DC000190 · NIDCD · CALIFORNIA INSTITUTE OF TECHNOLOGY · PI ZHANG, YINONG · 1995 to 1998
–
CIHR DC0190GPNIDCD NIH HHS F32 DC000190
6 · The paper itself

Abstract

objectivesAdverse event detection from Electronic Medical Records (EMRs) is challenging due to the low incidence of the event, variability in clinical documentation, and the complexity of data formats. Pulmonary embolism as an adverse event (PEAE) is particularly difficult to identify using existing approaches. This study aims to develop and evaluate a Large Language Model (LLM)-based framework for detecting PEAE from unstructured narrative data in EMRs. MATERIALS AND

methodsWe conducted a chart review of adult patients (aged 18-100) admitted to tertiary-care hospitals in Calgary, Alberta, Canada, between 2017-2022. We developed an LLM-based detection framework consisting of three modules: evidence extraction (implementing both keyword-based and semantic similarity-based filtering methods), discharge information extraction (focusing on six key clinical sections), and PEAE detection. Four open-source LLMs (Llama3, Mistral-7B, Gemma, and Phi-3) were evaluated using positive predictive value, sensitivity, specificity, and F1-score. Model performance for population-level surveillance was assessed at yearly, quarterly, and monthly granularities.

resultsThe chart review included 10 066 patients, with 40 cases of PEAE identified (0.4% prevalence). All four LLMs demonstrated high sensitivity (87.5-100%) and specificity (94.9-98.9%) across different experimental conditions. Gemma achieved the highest F1-score (28.11%) using keyword-based retrieval with discharge summary inclusion, along with 98.4% specificity, 87.5% sensitivity, and 99.95% negative predictive value. Keyword-based filtering reduced the median chunks per patient from 789 to 310, while semantic filtering further reduced this to 9 chunks. Including discharge summaries improved performance metrics across most models. For population-level surveillance, all models showed strong correlation with actual PEAE trends at yearly granularity (r=0.92-0.99), with Llama3 achieving the highest correlation (0.988). DISCUSSION: The results of our method for PEAE detection using EMR notes demonstrate high sensitivity and specificity across all four tested LLMs, indicating strong performance in distinguishing PEAE from non-PEAE cases. However, the low incidence rate of PEAE contributed to a lower PPV. The keyword-based chunking approach consistently outperformed semantic similarity-based methods, achieving higher F1 scores and PPV, underscoring the importance of domain knowledge in text segmentation. Including discharge summaries further enhanced performance metrics. Our population-based analysis revealed better performance for yearly trends compared to monthly granularity, suggesting the framework's utility for long-term surveillance despite dataset imbalance. Error analysis identified contextual misinterpretation, terminology confusion, and preprocessing limitations as key challenges for future improvement.

conclusionsOur proposed method demonstrates that LLMs can effectively detect PEAE from narrative EMRs with high sensitivity and specificity. While these models serve as effective screening tools to exclude non-PEAE cases, their lower PPV indicates they cannot be relied upon solely for definitive PEAE identification. Further chart review remains necessary for confirmation. Future work should focus on improving contextual understanding, medical terminology interpretation, and exploring advanced prompting techniques to enhance precision in adverse event detection from EMRs.

Indexed as

Electronic Health RecordsIatrogenic DiseaseNatural Language ProcessingPulmonary EmbolismAdolescentAdultAgedAged, 80 and overAlbertaFemaleHumansLarge Language ModelsMaleMiddle AgedSensitivity and SpecificityYoung Adultadverse event detectionclinical text mininglarge language modelspulmonary embolism

Identifiers

PMID40105654
PMCPMC12012340

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.