Evidence map›Paper›PMID 38042599›Full record

SynthesisArtificial intelligence in medicine2023

Natural language processing with machine learning methods to analyze unstructured patient-reported outcomes derived from electronic health records: A systematic review.

Jin-Ah Sim, Xiaolei Huang, Madeline R Horan, Christopher M Stewart, Leslie L Robison, Melissa M Hudson, Justin N Baker, I-Chan Huang

Open access · greenAbstract readSystematic Review
In one paragraph

Synthesis in Artificial intelligence in medicine, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 40 papers, 3 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
40citing papers in PubMed, 3 pooled it
11.0field-weighted citation impact, top 1% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

40 citing papers in PubMed, 3 syntheses or guidelines pooled it, 63 citations in OpenAlex.

  1. Pooled it
  2. Pooled it
  3. Pooled it
  4. Article
  5. Article
  6. Observational
  7. Review
  8. AI-based augmentation of oncology clinical trials.Nature reviews. Clinical oncology · 2026
    Review
  9. Review
  10. Observational
  11. Article
  12. Review
  13. Article
  14. Article
  15. Review
  16. Review
  17. Article
  18. Review
  19. Review
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors at 3 institutions in 2 countries.

Jin-Ah SimDepartment of Epidemiology and Cancer Control, St. Jude Children's Research Hospital, Memphis, TN, United States; School of AI Convergence, Hallym University, Chuncheon, Republic of Korea.
Xiaolei HuangDepartment of Computer Science, University of Memphis, Memphis, TN, United States.
Madeline R HoranDepartment of Epidemiology and Cancer Control, St. Jude Children's Research Hospital, Memphis, TN, United States.
Christopher M StewartInstitute for Intelligent Systems, University of Memphis, Memphis, TN, United States.
Leslie L RobisonDepartment of Epidemiology and Cancer Control, St. Jude Children's Research Hospital, Memphis, TN, United States.
Melissa M HudsonDepartment of Epidemiology and Cancer Control, St. Jude Children's Research Hospital, Memphis, TN, United States; Department of Oncology, St. Jude Children's Research Hospital, Memphis, TN, United States.
Justin N BakerDepartment of Pediatrics, Stanford University, Stanford, CA, United States.
I-Chan HuangDepartment of Epidemiology and Cancer Control, St. Jude Children's Research Hospital, Memphis, TN, United States. Electronic address: i-chan.huang@stjude.org.
St. Jude Children's Research Hospital · USUniversity of Memphis · USStanford University · US

Funding

The St. Jude Lifetime CohortU01CA195547 · NCI · ST. JUDE CHILDREN'S RESEARCH HOSPITAL · PI HUDSON, MELISSA M, NESS, KIRSTEN KIMBERLIE · 2015 to 2024
$14.9M
Patient-Generated Health Data to Predict Childhood Cancer Survivorship OutcomesR01CA258193 · NCI · ST. JUDE CHILDREN'S RESEARCH HOSPITAL · PI I-Chan Huang, YUTAKA YASUI · 2021 to 2026
$3.7M
Patient-Reported Outcomes Version of CTCAE involving Childhood Cancer SurvivorsR01CA238368 · NCI · ST. JUDE CHILDREN'S RESEARCH HOSPITAL · PI BAKER, JUSTIN N, HUANG, I-CHAN · 2019 to 2023
$3.5M
Training in Pediatric Cancer Survivorship Outcomes and InterventionsT32CA225590 · NCI · ST. JUDE CHILDREN'S RESEARCH HOSPITAL · PI Kevin R Krull · 2018 to 2026
$2.4M
NCI NIH HHS R01 CA238368NCI NIH HHS R01 CA258193NCI NIH HHS T32 CA225590NCI NIH HHS U01 CA195547
6 · The paper itself

Abstract

objectiveNatural language processing (NLP) combined with machine learning (ML) techniques are increasingly used to process unstructured/free-text patient-reported outcome (PRO) data available in electronic health records (EHRs). This systematic review summarizes the literature reporting NLP/ML systems/toolkits for analyzing PROs in clinical narratives of EHRs and discusses the future directions for the application of this modality in clinical care.

methodsWe searched PubMed, Scopus, and Web of Science for studies written in English between 1/1/2000 and 12/31/2020. Seventy-nine studies meeting the eligibility criteria were included. We abstracted and summarized information related to the study purpose, patient population, type/source/amount of unstructured PRO data, linguistic features, and NLP systems/toolkits for processing unstructured PROs in EHRs.

resultsMost of the studies used NLP/ML techniques to extract PROs from clinical narratives (n = 74) and mapped the extracted PROs into specific PRO domains for phenotyping or clustering purposes (n = 26). Some studies used NLP/ML to process PROs for predicting disease progression or onset of adverse events (n = 22) or developing/validating NLP/ML pipelines for analyzing unstructured PROs (n = 19). Studies used different linguistic features, including lexical, syntactic, semantic, and contextual features, to process unstructured PROs. Among the 25 NLP systems/toolkits we identified, 15 used rule-based NLP, 6 used hybrid NLP, and 4 used non-neural ML algorithms embedded in NLP.

conclusionsThis study supports the potential utility of different NLP/ML techniques in processing unstructured PROs available in EHRs for clinical care. Though using annotation rules for NLP/ML to analyze unstructured PROs is dominant, deploying novel neural ML-based methods is warranted.

Indexed as

Electronic Health RecordsMachine LearningNatural Language ProcessingPatient Outcome AssessmentHumansPubMedElectronic health recordsMachine learningNatural language processingPatient-reported outcomesUnstructured clinical narrative

Identifiers

PMID38042599
PMCPMC10693655
OpenAlexW4388116325

What OpenQuestion holds

Textmetadata
LicenceTDM
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.