Evidence map›Paper›PMID 41081240›Full record

ArticleQuantitative imaging in medicine and surgery2025

Multimodal large language models in ultrasound diagnosis of breast masses: a multicenter comparative analysis based on GPT-4o, radiologists, and convolutional neural network (CNN).

Jia-Qian Yao, Rui Zhang, Ze-Bang Yang, Bo Zhang, Shuang-Quan Jiang, Lin Jiang, Xiao-Er Zhang, Xiao-Yan Xie, Tong-Yi Huang, Ming Xu

Abstract read
In one paragraph

Article in Quantitative imaging in medicine and surgery, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Jia-Qian Yao *Department of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.ORCID https://orcid.org/0000-0003-4225-5143
Rui Zhang *Department of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Ze-Bang YangDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Bo ZhangDepartment of Ultrasound, China-Japan Friendship Hospital, National Center for Respiratory Medicine, National Clinical Research Center for Respiratory Diseases, Institute of Respiratory Medicine of Chinese Academy of Medical Sciences, Beijing, China.
Shuang-Quan JiangDepartment of Ultrasound Medicine, The Second Affiliated Hospital of Harbin Medical University, Harbin, China.
Lin JiangDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Xiao-Er ZhangDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Xiao-Yan XieDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Tong-Yi HuangDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.
Ming XuDepartment of Medical Ultrasonics, Institute of Diagnostic and Interventional Ultrasound, The First Affiliated Hospital of Sun Yat-sen University, Guangzhou, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: An advanced version of large language model (LLM), ChatGPT 4o (GPT-4o), has shown capacity in image-text pair interpretation, yet the performance in medical image analysis remains unclear. This study aimed to evaluate the diagnostic capacity of GPT-4o in breast ultrasound (US) datasets. Methods: US exams including breast images and original reports were respectively included from January 2021 to December 2023 in three hospitals throughout China. The diagnostic performance in distinguishing benign or malignant breast masses of GPT-4o was assessed through two approaches: image-strategy and image-combined-text-strategy. Fleiss kappa was calculated to determine intra-LLM consistency. Thereafter, diagnostic accuracy was evaluated and compared with the convolutional neural network (CNN) model and 95 human experts with various levels of expertise from 60 institutions in China. Responses from GPT-4o were rated by diagnostic confidence and radiologist's evaluation. Results: The observations of 80 breast masses (37 malignant, 43 benign) from 80 patients [median age, 42.5 years; interquartile range (IQR), 37.0-53.0 years] were enrolled. GPT-4o with image-strategy exhibited a fair consistency [0.25, 95% confidence interval (CI): 0.07-0.43], whereas the agreement of image-combined-text-strategy was excellent (0.81, 95% CI: 0.67-0.91). Diagnostic accuracy improved when deploying the image-combined-text-strategy compared to only image [58% (46 of 80) Conclusions: The effectiveness of GPT-4o in interpreting real-world US images was limited, yet improved in image-combined-text-strategy. Deploying LLM warrants scrutiny from radiologists.

Indexed as

artificial intelligence (AI)breast massLarge language model (LLM)ultrasound (US)

Identifiers

PMID41081240
PMCPMC12514734

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.