Evidence map›Paper›PMID 42222237›Full record

ArticleNeurology. Education2026

Education Research: Quality of Narrative Feedback Generated by a Large Language Model Compared With Expert Faculty for Case-Based Learning in Neurology Education.

Hannah Fruitman, Sasha Severin, Atikul Miah, Christina Gao, Haelynn Gim, Carolyn Qian, Sang-O Park, Kelly Hou, Edward L Kong, Benjamin Cook and 16 more

Abstract read
In one paragraph

Article in Neurology. Education, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

26 authors.

Hannah FruitmanNew York Medical College, Valhalla, NY.ORCID https://orcid.org/0009-0009-1245-0731
Sasha SeverinNew York Medical College, Valhalla, NY.ORCID https://orcid.org/0000-0002-0021-2295
Atikul MiahNew York Medical College, Valhalla, NY.
Christina GaoAdelaide Medical School, The University of Adelaide, South Australia, Australia.ORCID https://orcid.org/0009-0005-0033-3352
Haelynn GimHarvard Medical School, Harvard University, Boston, MA.
Carolyn QianHarvard Medical School, Harvard University, Boston, MA.
Sang-O ParkHarvard Medical School, Harvard University, Boston, MA.
Kelly HouAdelaide Medical School, The University of Adelaide, South Australia, Australia.ORCID https://orcid.org/0009-0008-2765-6954
Edward L KongHarvard Medical School, Harvard University, Boston, MA.
Benjamin CookAdelaide Medical School, The University of Adelaide, South Australia, Australia.
Jasmin LeAdelaide Medical School, The University of Adelaide, South Australia, Australia.
Brandon StrettonAdelaide Medical School, The University of Adelaide, South Australia, Australia.
John MaddisonLyell McEwin Hospital, Elizabeth Vale, South Australia, Australia.ORCID https://orcid.org/0000-0001-8692-8878
Liam G McCoyDivision of Neurology, Faculty of Medicine and Dentistry, University of Alberta, Edmonton, Canada.ORCID https://orcid.org/0000-0002-4468-2256
Luke CollinsAdelaide Medical School, The University of Adelaide, South Australia, Australia.
Andrew VanlintAdelaide Medical School, The University of Adelaide, South Australia, Australia.
Rudy GohAdelaide Medical School, The University of Adelaide, South Australia, Australia.ORCID https://orcid.org/0000-0003-1645-2258
Matthew ArnoldAdelaide Medical School, The University of Adelaide, South Australia, Australia.ORCID https://orcid.org/0009-0005-9510-0948
Aye ThantWashington University in St. Louis, School of Medicine, MO.ORCID https://orcid.org/0009-0000-7217-0082
Rani Priyanka VasireddyUniversity of Texas at Tyler School of Medicine.
Doris KungBaylor College of Medicine, Houston, TX.ORCID https://orcid.org/0000-0001-8458-7838
Ashley M PaulJohns Hopkins University School of Medicine, Johns Hopkins University, Baltimore, MD.ORCID https://orcid.org/0000-0001-9124-4024
Haatem RedaHarvard Medical School, Harvard University, Boston, MA.
Tamara B KaplanHarvard Medical School, Harvard University, Boston, MA.ORCID https://orcid.org/0000-0002-6262-6960
Adam KarpNew York Medical College, Valhalla, NY.ORCID https://orcid.org/0000-0001-9187-7980
Galina GheihmanHarvard Medical School, Harvard University, Boston, MA.ORCID https://orcid.org/0000-0003-1599-3271

Funding

Medical Scientist Training ProgramT32GM144273 · NIGMS · HARVARD MEDICAL SCHOOL · PI David Shumway Jones, Jacqueline A. Lees · 2022 to 2026
$14.7M
NIGMS NIH HHS T32 GM144273
6 · The paper itself

Abstract

Background and Objectives: Neurology learners often receive limited feedback in clinical settings because of workflow constraints, variability in supervision, and competing clinical demands. Artificial intelligence, including large language models (LLMs) may help address these gaps and provide clinical learners with effective formative feedback by generating real-time, case-specific feedback during neurology case-based learning (CBL). The aim of this study was to examine how the quality of LLM-generated feedback compares with human expert-generated feedback in neurology CBL. Methods: In this exploratory quantitative study, student participants undertook LLM-enabled interactive cases on the TEACHABLE platform, which included history gathering, physical examination elements, and ordering diagnostic testing. Participants were clinical-level students recruited from 2 medical institutions. Case transcripts were recorded and analyzed for feedback generation, which was provided by an LLM and human experts in 2 components: history taking/physical examination elements (H&P) and assessment and plan (A&P). Feedback characteristics including sentence count, word count, and reference to case key learning points were summarized and compared. Feedback quality was scored by blinded experts using the QuAL and EFeCT instruments. Results were compared for the H&P and A&P components of the case interactions. Results: Four student participants completed 5 interactive cases each, generating 20 total transcripts for feedback. Word and sentence number were similar among LLM-generated and expert-generated feedback, except for a greater word length in expert-generated A&P feedback. Regarding H&P, the LLM commented on the key learning points in 20/20 (100%) of the cases as compared with 39/60 (65%) for the human experts. For A&P, the LLM feedback discussed key points in 20/20 (100%) cases as compared with 39/40 (97.5%) for the human experts. The LLM feedback had no medical inaccuracies. QuAL and EFeCT scores were significantly greater for the LLM as compared with human experts for the H&P component, but not significantly different for the A&P component. Discussion: LLMs provided with key learning points can generate timely, quality feedback on case-based interactions in a manner comparable with human experts. A hybrid framework combining LLM-generated feedback with faculty input may offer high-quality and equitably accessible formative feedback at scale. These pilot findings are limited by a small sample size and experimental setting.

Identifiers

PMID42222237
PMCPMC13220965

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.