Evidence map›Paper›PMID 41411647›Full record

ArticleJMIR AI2025

Physical Examination Identification in Medical Education Videos: Zero-Shot Multimodal AI With Temporal Sequence Optimization Study.

Shinyoung Kang, Michael Holcomb, David Hein, Ameer Hamza Shakur, Thomas Dalton, Andrew Jamieson

Abstract read
In one paragraph

Article in JMIR AI, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Shinyoung KangLyda Hill Department of Bioinformatics, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0009-0003-8876-1196
Michael HolcombLyda Hill Department of Bioinformatics, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0002-0595-1476
David HeinLyda Hill Department of Bioinformatics, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0002-8625-9528
Ameer Hamza ShakurLyda Hill Department of Bioinformatics, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0002-2695-8378
Thomas DaltonDepartment of Internal Medicine, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0002-8351-8336
Andrew JamiesonLyda Hill Department of Bioinformatics, The University of Texas Southwestern Medical Center, Dallas, TX, United States.ORCID https://orcid.org/0000-0001-5416-9379

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundObjective structured clinical examinations (OSCEs) are widely used for assessing medical student competency, but their evaluation is resource-intensive, requiring trained evaluators to review 15-minute videos. The physical examination (PE) component typically constitutes only a small portion of these recordings; yet, current automated approaches struggle with processing long medical videos due to computational constraints and difficulties maintaining temporal context.

objectiveThis study aims to determine whether multimodal large language models (MM-LLMs) can effectively segment PE periods within OSCE videos without previous training, potentially reducing the evaluation burden on both human graders and automated assessment systems.

methodsWe analyzed 500 videos from 5 OSCE stations at University of Texas Southwestern Simulation Center, each 15 minutes long, by using hand-labeled PE periods as ground truth. Frames were sampled at 1, 2, or 3 seconds. A pose detection preprocessing step filtered frames without people. Six MM-LLMs performed frame-level classification into encounter states by using a standardized prompt. To enforce temporal consistency, we used a hidden Markov model with Viterbi decoding, merging states into 3 primary activities (consulting/notes, physical examination, and no doctor) and adding a brief edge buffer to avoid truncating true PE segments. Performance was computed per video and averaged across the dataset by using recall, precision, intersection over union (IOU), and predicted PE length with 95% CIs.

resultsAt 1-second sampling, GPT-4o achieved recall of 0.998 (95% CI 0.994-1.000), IOU of 0.784 (95% CI 0.765-0.803), and precision of 0.792 (95% CI 0.774-0.811), identifying a mean of 175 (SD 83) seconds of content per video as PE versus a mean labeled PE of 126 (SD 61) seconds, yielding an 81% reduction in video needing review (from 900 to 175 seconds). Across stations, recall remained high, with expected IOU variability linked to examination format and camera geometry. Increasing the sampling interval modestly decreased recall while slightly improving IOU and precision. Comparative baselines (eg, Gemini 2.0 Flash, Gemma 3, and Qwen2.5-VL variants) demonstrated trade-offs between recall and overselection; GPT-4o offered the best balance among high-recall models. Error analysis highlighted false negatives during occluded or verbally guided maneuvers and false positives during preparatory actions, suggesting opportunities for camera placement optimization and multimodal fusion (eg, audio cues).

conclusionsIntegrating zero-shot MM-LLMs with minimal-supervision temporal modeling effectively segments PE periods in OSCE videos without requiring extensive training data. This approach significantly reduces review time while maintaining clinical assessment integrity, demonstrating that artificial intelligence methods combining zero-shot capabilities and light supervision can be optimized for medical education's specific requirements. This technique establishes a foundation for more efficient and scalable clinical skill assessment across diverse medical education settings.

Indexed as

AIartificial intelligencemedical educationmultimodal large language modelsvideo segmentation

Identifiers

PMID41411647
PMCPMC12757708

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.