Evidence map›Paper›PMID 38875593›Full record

ArticleJMIR AI2024

Sample Size Considerations for Fine-Tuning Large Language Models for Named Entity Recognition Tasks: Methodological Study.

Zoltan P Majdik, S Scott Graham, Jade C Shiva Edward, Sabrina N Rodriguez, Martha S Karnes, Jared T Jensen, Joshua B Barbour, Justin F Rousseau

Abstract read
In one paragraph

Article in JMIR AI, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 9 papers.

0numbers the graph read from it
0cells of the map it votes in
9citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

9 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Zoltan P MajdikDepartment of Communication, North Dakota State University, Fargo, ND, United States.ORCID https://orcid.org/0000-0002-3851-4694
S Scott GrahamDepartment of Rhetoric & Writing, The University of Texas at Austin, Austin, TX, United States.ORCID https://orcid.org/0000-0003-1569-2428
Jade C Shiva EdwardDepartment of Rhetoric & Writing, The University of Texas at Austin, Austin, TX, United States.ORCID https://orcid.org/0009-0002-2834-4174
Sabrina N RodriguezDepartment of Neurology, The Dell Medical School, The University of Texas at Austin, Austin, TX, United States.ORCID https://orcid.org/0009-0005-8024-1625
Martha S KarnesDepartment of Rhetoric & Writing, University of Arkansas Little Rock, Little Rock, AR, United States.ORCID https://orcid.org/0000-0002-5432-7856
Jared T JensenDepartment of Rhetoric & Writing, The University of Texas at Austin, Austin, TX, United States.ORCID https://orcid.org/0000-0003-3602-6854
Joshua B BarbourDepartment of Communication, The University of Illinois at Urbana-Champaign, Urbana, IL, United States.ORCID https://orcid.org/0000-0001-8384-7175
Justin F RousseauStatistical Planning and Analysis Section, Department of Neurology, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0002-2817-9124

Funding

Closing the loop with an automatic referral population and summarization systemR01LM014306 · NLM · WEILL MEDICAL COLL OF CORNELL UNIV · PI Yifan Peng, Justin Frederick Rousseau · 2023 to 2026
$2.7M
A Network Science Approach to Conflicts of Interest: Metrics, Policies, and Communication DesignR01GM141476 · NIGMS · UNIVERSITY OF TEXAS AT AUSTIN · PI BARBOUR, JOSHUA BEN, GRAHAM, SAMUEL · 2021 to 2024
$832k
NIGMS NIH HHS R01 GM141476NLM NIH HHS R01 LM014306
6 · The paper itself

Abstract

backgroundLarge language models (LLMs) have the potential to support promising new applications in health informatics. However, practical data on sample size considerations for fine-tuning LLMs to perform specific tasks in biomedical and health policy contexts are lacking.

objectiveThis study aims to evaluate sample size and sample selection techniques for fine-tuning LLMs to support improved named entity recognition (NER) for a custom data set of conflicts of interest disclosure statements.

methodsA random sample of 200 disclosure statements was prepared for annotation. All "PERSON" and "ORG" entities were identified by each of the 2 raters, and once appropriate agreement was established, the annotators independently annotated an additional 290 disclosure statements. From the 490 annotated documents, 2500 stratified random samples in different size ranges were drawn. The 2500 training set subsamples were used to fine-tune a selection of language models across 2 model architectures (Bidirectional Encoder Representations from Transformers [BERT] and Generative Pre-trained Transformer [GPT]) for improved NER, and multiple regression was used to assess the relationship between sample size (sentences), entity density (entities per sentence [EPS]), and trained model performance (F

resultsFine-tuned models ranged in topline NER performance from F

conclusionsRelatively modest sample sizes can be used to fine-tune LLMs for NER tasks applied to biomedical text, and training data entity density should representatively approximate entity density in production data. Training data quality and a model architecture's intended use (text generation vs text processing or classification) may be as, or more, important as training data volume and model parameter size.

Indexed as

annotationconflict of interestdisclosuredisclosuresexpert annotationfine-tuninglanguage modellarge language modelsmachine learningnamed-entity recognitionnatural language processingsamplesample sizestatementstatementstransfer learning

Identifiers

PMID38875593
PMCPMC11140272

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.