Evidence map›Paper›PMID 39225779›Full record

ArticleJournal of the American Medical Informatics Association : JAMIA2024

CACER: Clinical concept Annotations for Cancer Events and Relations.

Yujuan Velvin Fu, Giridhar Kaushik Ramachandran, Ahmad Halwani, Bridget T McInnes, Fei Xia, Kevin Lybarger, Meliha Yetisgen, Özlem Uzuner

Abstract read
In one paragraph

Article in Journal of the American Medical Informatics Association : JAMIA, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Yujuan Velvin FuDepartment of Biomedical Informatics & Medical Education, University of Washington, Seattle, WA 98195, United States.
Giridhar Kaushik RamachandranDepartment of Information Sciences and Technology, George Mason University, Fairfax, VA 22030, United States.ORCID 0000-0002-9800-6149
Ahmad HalwaniHuntsman Cancer Institute, University of Utah, Salt Lake City, UT 84112, United States.
Bridget T McInnesDepartment of Computer Science, Virginia Commonwealth University, Richmond, VA 23284, United States.
Fei XiaDepartment of Linguistics, University of Washington, Seattle, WA 98195, United States.
Kevin LybargerDepartment of Information Sciences and Technology, George Mason University, Fairfax, VA 22030, United States.ORCID 0000-0001-5798-2664
Meliha YetisgenDepartment of Biomedical Informatics & Medical Education, University of Washington, Seattle, WA 98195, United States.
Özlem UzunerDepartment of Information Sciences and Technology, George Mason University, Fairfax, VA 22030, United States.ORCID 0000-0001-8011-9850

Funding

Large scale clinical and economic impact analysis of potentially malignant incidental findings in radiology reportsR01CA248422 · NCI · UNIVERSITY OF WASHINGTON · PI GUNN, MARTIN, YETISGEN, MELIHA · 2021 to 2024
$2.6M
Leveraging Unlabeled and Pseudo Data for Clinical Information ExtractionR15LM013209 · NLM · GEORGE MASON UNIVERSITY · PI UZUNER, OZLEM · 2019 to 2022
$840k
Extraction of Symptom Burden from Clinical Narratives of Cancer Patients using Natural Language ProcessingR21CA258242 · NCI · UNIVERSITY OF WASHINGTON · PI YETISGEN, MELIHA · 2021 to 2022
$719k
NCI NIH HHS 1R01CA248422-01A1NCI NIH HHS R01 CA248422NCI NIH HHS R21 CA258242NIHNIH HHSNIH HHS 1R01CA248422-01A1NLM NIH HHS 2R15LM013209-02A1NLM NIH HHS R15 LM013209
6 · The paper itself

Abstract

objectiveClinical notes contain unstructured representations of patient histories, including the relationships between medical problems and prescription drugs. To investigate the relationship between cancer drugs and their associated symptom burden, we extract structured, semantic representations of medical problem and drug information from the clinical narratives of oncology notes. MATERIALS AND

methodsWe present Clinical concept Annotations for Cancer Events and Relations (CACER), a novel corpus with fine-grained annotations for over 48 000 medical problems and drug events and 10 000 drug-problem and problem-problem relations. Leveraging CACER, we develop and evaluate transformer-based information extraction models such as Bidirectional Encoder Representations from Transformers (BERT), Fine-tuned Language Net Text-To-Text Transfer Transformer (Flan-T5), Large Language Model Meta AI (Llama3), and Generative Pre-trained Transformers-4 (GPT-4) using fine-tuning and in-context learning (ICL).

resultsIn event extraction, the fine-tuned BERT and Llama3 models achieved the highest performance at 88.2-88.0 F1, which is comparable to the inter-annotator agreement (IAA) of 88.4 F1. In relation extraction, the fine-tuned BERT, Flan-T5, and Llama3 achieved the highest performance at 61.8-65.3 F1. GPT-4 with ICL achieved the worst performance across both tasks. DISCUSSION: The fine-tuned models significantly outperformed GPT-4 in ICL, highlighting the importance of annotated training data and model optimization. Furthermore, the BERT models performed similarly to Llama3. For our task, large language models offer no performance advantage over the smaller BERT models.

conclusionsWe introduce CACER, a novel corpus with fine-grained annotations for medical problems, drugs, and their relationships in clinical narratives of oncology notes. State-of-the-art transformer models achieved performance comparable to IAA for several extraction tasks.

Indexed as

Electronic Health RecordsNatural Language ProcessingNeoplasmsAntineoplastic AgentsData MiningHumansSemanticsAntineoplastic Agentscancer patientsdata miningelectronic health recordsinformation extractionmachine learningnatural language processing

Identifiers

PMID39225779
PMCPMC11491616

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.