Evidence map›Paper›PMID 41093296›Full record

SynthesisJournal of the American Medical Informatics Association : JAMIA2026

Using natural language processing to extract information from clinical text in electronic medical records for populating clinical registries: a systematic review.

Leibo Liu, Victoria Blake, Matthew Barman, Blanca Gallego, Timothy Churches, Georgina Kennedy, Sze-Yuan Ooi, Geoffrey P Delaney, Louisa Jorm

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of the American Medical Informatics Association : JAMIA, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 14 papers.

0numbers the graph read from it
0cells of the map it votes in
14citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

14 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Article
  13. Article
  14. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Leibo LiuCentre for Big Data Research in Health, University of New South Wales, Sydney, NSW 2052, Australia.
Victoria BlakeCentre for Big Data Research in Health, University of New South Wales, Sydney, NSW 2052, Australia.
Matthew BarmanCentre for Big Data Research in Health, University of New South Wales, Sydney, NSW 2052, Australia.
Blanca GallegoCentre for Big Data Research in Health, University of New South Wales, Sydney, NSW 2052, Australia.ORCID 0000-0002-3704-7975
Timothy ChurchesIngham Institute for Applied Medical Research, Liverpool, NSW 2170, Australia.
Georgina KennedyIngham Institute for Applied Medical Research, Liverpool, NSW 2170, Australia.
Sze-Yuan OoiSchool of Clinical Medicine, University of New South Wales, Sydney, NSW 2052, Australia.
Geoffrey P DelaneyIngham Institute for Applied Medical Research, Liverpool, NSW 2170, Australia.
Louisa JormCentre for Big Data Research in Health, University of New South Wales, Sydney, NSW 2052, Australia.ORCID 0000-0003-0390-661X

Funding

Australian Medical Research Futures FundAustralian Medical Research Futures Fund MRFFRD000154Research Data Infrastructure MRFFRD000154
6 · The paper itself

Abstract

objectiveClinical registries advance healthcare by tracking patient outcomes and intervention safety. Manually extracting information from clinical text for registries is labor- and resource-intensive and often inaccurate. Therefore, this systematic review aims to evaluate the use and effectiveness of natural language processing (NLP) methods in extracting information from clinical text for populating clinical registries. MATERIALS AND

methodsPubMed, Embase, Scopus, Web of Science, and ACM Digital Library were systematically searched. Studies were included if they used NLP techniques to populate clinical registries. The extracted data included details of the registry, the clinical text, the registry data elements extracted, the NLP methods used, and how their performance was evaluated.

resultsFifteen articles were included in the review. Since 2020, the use of NLP methods for extracting information to populate clinical registries has been increasing steadily. Initially, rule-based NLP methods dominated the field, but machine learning-based approaches have gradually gained popularity. However, only one of the included studies employed generative large language models (LLMs). The diversity of clinical text and extracted data elements posed challenges to the generalizability of the NLP methods.

conclusionTo date, the application of NLP methods to clinical text for populating clinical registries has been limited in both the number of published studies and the scope of implementation. The NLP methods used thus far face significant challenges in effectively managing the complexity and diversity of clinical text and data elements. Moreover, the performance of the NLP methods varied significantly. This review underscores the need for a robust and adaptable NLP framework. Generative LLMs may provide direction for future research, but their use must account for challenges such as accuracy, cost, privacy, and limited supporting evidence.

Indexed as

Data MiningElectronic Health RecordsInformation Storage and RetrievalNatural Language ProcessingRegistriesHumansMachine Learningclinical registriesclinical textelectronic medical recordsinformation extractionnatural language processing

Identifiers

PMID41093296
PMCPMC12844598

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.