Evidence map›Paper›PMID 41232578›Full record

ArticleJournal of clinical epidemiology2026

Scalable medication extraction and discontinuation identification from electronic health records using large language models.

Chong Shao, Douglas Snyder, Chiran Li, Bowen Gu, Kerry Ngan, Chun-Ting Yang, Jiageng Wu, Richard Wyss, Kueiyu Joshua Lin, Jie Yang

Abstract read
In one paragraph

Article in Journal of clinical epidemiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Observational
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Chong ShaoDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA; Harvard T.H. Chan School of Public Health, Harvard University, Boston, MA, USA.
Douglas SnyderHarvard T.H. Chan School of Public Health, Harvard University, Boston, MA, USA.
Chiran LiHarvard T.H. Chan School of Public Health, Harvard University, Boston, MA, USA.
Bowen GuDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Kerry NganDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Chun-Ting YangDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Jiageng WuDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Richard WyssDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Kueiyu Joshua LinDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA.
Jie YangDivision of Pharmacoepidemiology and Pharmacoeconomics, Department of Medicine, Brigham and Women's Hospital, Harvard Medical School, Boston, MA, USA; Broad Institute of MIT and Harvard, Cambridge, MA, USA; Harvard Data Science Initiative, Harvard University, Cambridge, MA, USA. Electronic address: jyang66@bwh.harvard.edu.

Funding

Optimizing validity of comparative effectiveness research in Alzheimer's disease and related dementias using large language modelsRF1AG090405 · NIA · BRIGHAM AND WOMEN'S HOSPITAL · PI LIN, JOSHUA K, YANG, JIE · 2025 to 2025
$3.7M
Developing scalable algorithms to incorporate unstructured electronic health records for causal inference based on real-world dataR01LM013204 · NLM · BRIGHAM AND WOMEN'S HOSPITAL · PI LIN, JOSHUA K · 2020 to 2024
$2.7M
Developing large language models for drug safety and effectiveness causal analysisR01LM014667 · NLM · BRIGHAM AND WOMEN'S HOSPITAL · PI Jie Yang · 2025 to 2026
$806k
NIA NIH HHS RF1 AG090405NLM NIH HHS R01 LM013204NLM NIH HHS R01 LM014667
6 · The paper itself

Abstract

objectivesIdentifying medication discontinuations in electronic health records (EHRs) is vital for patient safety but is often hindered by information being buried in unstructured notes. This study aims to evaluate the capabilities of advanced open-sourced and proprietary large language models in extracting medications and classifying their medication status from EHR notes, focusing on their scalability for medication information extraction without human annotation. STUDY DESIGN AND

settingWe collected three EHR datasets from diverse sources to build the evaluation benchmark: 1 publicly available dataset (Reannotated Clinical Acronym Sense Inventory dataset [Re-CASI]), 1 we annotated based on public MIMIC notes (MIMIC-IV Medication Snippet dataset [MIV-Med]), and 1 internally annotated on clinical notes from Mass General Brigham (MGB-Med). We evaluated 12 advanced LLMs, including general-domain open-sourced models (eg, Llama-3.1-70B-Instruct, Qwen2.5-72B-Instruct), medical-specific models (eg, MeLLaMA-70B-chat), and a proprietary model (GPT-4o). We explored multiple LLM prompting strategies, including zero-shot, 5-shot, and Chain-of-Thought (CoT) approaches. Performance on medication extraction, medication status classification, and their joint task (extraction then classification) was systematically compared across all experiments.

resultsLLMs showed promising performance on medication extraction, while discontinuation classification and joint tasks were more challenging. GPT-4o consistently achieved the highest average F1 scores in all tasks under zero-shot setting - 94.0% for medication extraction, 78.1% for discontinuation classification, and 72.7% for the joint task. Open-sourced models followed closely, with Llama-3.1-70B-Instruct achieving the highest performance in medication status classification on the MIV-Med dataset (68.7%) and in the joint task on both the Re-CASI (76.2%) and MIV-Med (60.2%) datasets. Medical-specific LLMs demonstrated lower performance compared to advanced general-domain LLMs. Few-shot learning generally improved performance, while CoT reasoning showed inconsistent gains. Notably, open-sourced models occasionally surpassed GPT-4o performance, underscoring their potential in privacy-sensitive clinical research.

conclusionLLMs demonstrate strong potential for medication extraction and discontinuation identification on EHR notes, with open-sourced models offering scalable alternatives to proprietary systems and few-shot learning further improving LLMs' capability. PLAIN LANGUAGE SUMMARY: Stopping a medicine can affect safety and treatment decisions, yet this detail is often buried in long electronic health record notes. We evaluated whether large language models, which read and summarize text, can automatically find medication names and decide whether each medicine is still being taken, has been stopped, or neither. We tested 12 models, including open-source options suitable for secure hospital use, on three collections of clinical notes and compared three simple instruction styles: giving no examples, showing a few examples, and asking for step-by-step reasoning. All models produced usable results. The strongest systems scored about 94 for finding medication names and about 78 for deciding continued or stopped status, on a standard 0 to 100 measure that balances completeness and correctness. Showing a few examples usually helped more than step-by-step prompts, and several open-source models performed close to a leading proprietary system. These tools could help hospitals and researchers monitor medications at scale to support drug-safety studies, adherence tracking, and clinical decision support, with local validation and safeguards before clinical use.

Indexed as

Electronic Health RecordsNatural Language ProcessingHumansLarge Language ModelsElectronic health recordsInformation extractionLarge language modelMedication discontinuation

Identifiers

PMID41232578
PMCPMC12714491

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.