Evidence map›Paper›PMID 42102716›Full record

ArticleFamily medicine2026

Validation of the Use of a Large Language Model for Detecting Sentiment in Student Course Evaluation.

Kate Rowland, Ling Wang, Kirstie Bash, Lori DeShetler, Sarah Vick, Emma Nguyen, Michelle Rogers-Johnson, Lauren Anderson, Michelle Sweet, Kimberly Fasula and 1 more

Abstract readValidation Study
In one paragraph

Article in Family medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Kate RowlandRush University Medical Center, Chicago, IL.
Ling WangMichigan State University, East Lansing, MI.
Kirstie BashUniversity of Nebraska Medical Center, Omaha, NE.
Lori DeShetlerDepartment of Medical Education, University of Toledo, Toledo, OH.
Sarah VickUniversity of Kentucky, Lexington, KY.
Emma NguyenUniversity of Kansas, Kansas City, KS.
Michelle Rogers-JohnsonEastern Virginia Medical School, Old Dominion University, Norfolk, VA.
Lauren AndersonRush University Medical Center, Chicago, IL.
Michelle SweetRush University Medical Center, Chicago, IL.
Kimberly FasulaChicago Medical School, Rosalind Franklin University of Medicine and Science, North Chicago, IL.
Stefanie CarterDr.Kiran C. Patel College of Allopathic Medicine, Nova Southeastern University, Fort Lauderdale, FL.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

BACKGROUND AND

objectivesThe use of large language models and natural language processing (NLP) in medical education has expanded rapidly in recent years. Because of the documented risks of bias and errors, these artificial intelligence (AI) tools must be validated before being used for research or education. Traditional and novel conceptual frameworks can be used. This study aimed to validate the application of an NLP method, bidirectional encoder representations from transformers (BERT) model, to identify the presence and patterns of sentiment in end-of-course evaluations from M3 (medical school year 3) core clerkships at multiple institutions.

methodsWe used the Patino framework, designed for the use of artificial intelligence in health professions education, as a guide for validating the NLP. Written comments from de-identified course evaluations at four schools were coded by teams of two human coders, and human-human interrater reliability statistics were calculated. Humans identified key terms to train the BERT model. The trained BERT model predicted the sentiments of a set of comments, and human-NLP interrater reliability statistics were calculated.

resultsA total of 364 discrete comments were evaluated in the human phase. The range of positive (30.6%-61.0%), negative (4.9%-39.5%), neutral (9.8%-19.0%), and mixed (1.7%-27.5%) sentiments varied by school. Human-human and human-AI interrater reliability also varied by school. Human-human and human-AI reliability were comparable.

conclusionsSeveral conceptual frameworks offer models for validation of AI tools in health professions education. A BERT model, with training, can detect sentiment in medical student course evaluations with an interrater reliability similar to human coders.

Indexed as

Natural Language ProcessingStudents, MedicalArtificial IntelligenceClinical ClerkshipHumansLarge Language ModelsReproducibility of Results

Identifiers

PMID42102716
PMCPMC12969553

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.