Evidence map›Paper›PMID 42574561›Full record

ArticleJournal of medical Internet research2026

Intelligent Framework for Adverse Drug Event Identification Using Large Language Models and Retrieval-Augmented Generation: Development and Evaluation Study.

Junlong Ma, Xuehong Wu, Zeying Feng, Yun Kuang, Zhendong Ding, Min Li, Guoping Yang

Abstract readEvaluation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Junlong Ma *Department of Pharmacy, Xiangya Hospital, Central South University, Changsha, Hunan, China, 1 0731 88618933.ORCID http://orcid.org/0000-0003-1949-1201
Xuehong Wu *School of Computer Science and Engineering, Central South University, Changsha, Hunan, China.ORCID http://orcid.org/0000-0003-4861-5648
Zeying FengClinical Trial Institution Office, Liuzhou Hospital of Guangzhou Women and Children's Medical Center, Liuzhou, Guangxi, China.ORCID http://orcid.org/0000-0002-7283-9564
Yun KuangCenter of Clinical Pharmacology, The Third Xiangya Hospital, Central South University, No 138 Tongzipo Road, Yuelu District,, Changsha, Hunan, 410013, China, 86 0731 88618933.ORCID http://orcid.org/0009-0009-5825-7232
Zhendong DingDepartment of Anesthesiology, The Third Xiangya Hospital, Central South University, Changsha, Hunan, China.ORCID http://orcid.org/0000-0001-5178-8983
Min LiSchool of Computer Science and Engineering, Central South University, Changsha, Hunan, China.ORCID http://orcid.org/0009-0003-7935-7300
Guoping YangCenter of Clinical Pharmacology, The Third Xiangya Hospital, Central South University, No 138 Tongzipo Road, Yuelu District,, Changsha, Hunan, 410013, China, 86 0731 88618933.ORCID http://orcid.org/0000-0001-5930-586X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Adverse drug events (ADEs) pose significant public health challenges and economic burdens. While substantial ADE information is documented in unstructured clinical notes, its extraction remains difficult due to semantic complexity. Large language models (LLMs) offer promising text comprehension capabilities but are often hindered by domain-specific hallucinations. Objective: This study aims to evaluate the effectiveness of retrieval-augmented generation (RAG) in improving the identification of ADEs using LLMs from Chinese clinical narratives and to establish a paradigm for this task. Methods: We collected and preprocessed 19,983 Chinese clinical notes, retaining 18,432 high-quality records. Following a rigorous annotation and deduplication process, we established a gold-standard reference dataset (n=2510) and an ADE knowledge base (n=5144) using a standardized JSON schema. We evaluated 3 state-of-the-art LLMs (DeepSeek-V3 [DeepSeek], ERNIE 3.5-8K [Baidu], and GPT-4o [OpenAI]) under 3 prompt strategies: nonaugmented generation (NAG), static-augmented generation (SAG), and RAG. Performance was comprehensively assessed using precision, recall, and F1-score across 3 recognition matching levels (L1 exact, L2 sentence, and L3 overlap) via 1000 bootstrap resamples. Model robustness was further validated from real-world clinical progress notes, reflecting real-world ADE prevalence. Results: We successfully constructed and publicly released the first Chinese ADE corpus derived from clinical notes. Across the tested LLMs, RAG yielded higher F1-scores than NAG and SAG at the L3 level. The optimal configuration, DeepSeek-V3 with RAG, achieved an overall L3-level F1-score of 0.9638 (95% CI 0.9541-0.9727). Notably, the RAG approach increased the recall of GPT-4o from 0.6419 under NAG to 0.9241 under RAG (FDR P=.003). Evaluation on real-world datasets demonstrated clinical utility, with the RAG prompt maintaining high discriminatory capability (specificity: 0.9821; F2-score: 0.8885). Error analysis revealed that RAG successfully resolved common identification errors, both omissions and commissions, that were intractable for nonaugmented models. Conclusions: Synergizing a curated domain-specific knowledge base with LLMs via a RAG architecture is an effective strategy for accurately identifying ADEs in unstructured Chinese clinical notes. This approach can mitigate hallucinations in LLMs, providing a foundational open-source benchmark and a robust technical framework to advance pharmacovigilance, drug safety research, and clinical decision support.

Indexed as

Drug-Related Side Effects and Adverse ReactionsLarge Language ModelsHumansadverse drug eventChinese clinical noteslarge language modelpharmacovigilanceretrieval-augmented generation

Identifiers

PMID42574561
PMCPMC13456140

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.