Evidence map›Paper›PMID 37550244›Full record

ArticleJournal of the American Medical Informatics Association : JAMIA2023

LeafAI: query generator for clinical cohort discovery rivaling a human programmer.

Nicholas J Dobbins, Bin Han, Weipeng Zhou, Kristine F Lan, H Nina Kim, Robert Harrington, Özlem Uzuner, Meliha Yetisgen

Abstract read
In one paragraph

Article in Journal of the American Medical Informatics Association : JAMIA, 2023. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed.

  1. Review
  2. Review
  3. Article
  4. Article
  5. Article
  6. CACER: Clinical concept Annotations for Cancer Events and Relations.Journal of the American Medical Informatics Association : JAMIA · 2024
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Nicholas J DobbinsDepartment of Biomedical Informatics & Medical Education, University of Washington, Seattle, Washington, USA.ORCID 0000-0002-3598-8747
Bin HanInformation School, University of Washington, Seattle, Washington, USA.
Weipeng ZhouDepartment of Biomedical Informatics & Medical Education, University of Washington, Seattle, Washington, USA.
Kristine F LanDepartment of Medicine, University of Washington, Seattle, Washington, USA.
H Nina KimDepartment of Medicine, University of Washington, Seattle, Washington, USA.
Robert HarringtonDepartment of Medicine, University of Washington, Seattle, Washington, USA.
Özlem UzunerDepartment of Information Sciences and Technology, George Mason University, Fairfax, Virginia, USA.ORCID 0000-0001-8011-9850
Meliha YetisgenDepartment of Biomedical Informatics & Medical Education, University of Washington, Seattle, Washington, USA.

Funding

Transform Dissemination and Implementation Science in CTSA ProgramsUL1TR002319 · NCATS · UNIVERSITY OF WASHINGTON · PI John K. Amory · 2017 to 2026
$100.0M
Leveraging Unlabeled and Pseudo Data for Clinical Information ExtractionR15LM013209 · NLM · GEORGE MASON UNIVERSITY · PI UZUNER, OZLEM · 2019 to 2022
$840k
NCATS NIH HHS UL1 TR002319NIH HHS UL1TR002319NLM NIH HHS R15 LM013209NLM NIH HHS R15LM013209
6 · The paper itself

Abstract

objectiveIdentifying study-eligible patients within clinical databases is a critical step in clinical research. However, accurate query design typically requires extensive technical and biomedical expertise. We sought to create a system capable of generating data model-agnostic queries while also providing novel logical reasoning capabilities for complex clinical trial eligibility criteria. MATERIALS AND

methodsThe task of query creation from eligibility criteria requires solving several text-processing problems, including named entity recognition and relation extraction, sequence-to-sequence transformation, normalization, and reasoning. We incorporated hybrid deep learning and rule-based modules for these, as well as a knowledge base of the Unified Medical Language System (UMLS) and linked ontologies. To enable data-model agnostic query creation, we introduce a novel method for tagging database schema elements using UMLS concepts. To evaluate our system, called LeafAI, we compared the capability of LeafAI to a human database programmer to identify patients who had been enrolled in 8 clinical trials conducted at our institution. We measured performance by the number of actual enrolled patients matched by generated queries.

resultsLeafAI matched a mean 43% of enrolled patients with 27 225 eligible across 8 clinical trials, compared to 27% matched and 14 587 eligible in queries by a human database programmer. The human programmer spent 26 total hours crafting queries compared to several minutes by LeafAI.

conclusionsOur work contributes a state-of-the-art data model-agnostic query generation system capable of conditional reasoning using a knowledge base. We demonstrate that LeafAI can rival an experienced human programmer in finding patients eligible for clinical trials.

Indexed as

Natural Language ProcessingUnified Medical Language SystemClinical Trials as TopicHumansKnowledge Basesclinical trialscohort definitionelectronic health recordsmachine learningnatural language processing

Identifiers

PMID37550244
PMCPMC10654856

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.