ArticleQuantitative imaging in medicine and surgery2025
Multimodal large language models in ultrasound diagnosis of breast masses: a multicenter comparative analysis based on GPT-4o, radiologists, and convolutional neural network (CNN).
Article in Quantitative imaging in medicine and surgery, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
4 citing papers in PubMed.
- Multimodal Large Language Models for Breast Ultrasound Report Auditing: Workflow-Error Detection, False-Positive Burden, and Limits of Key-Image Interpretation.Journal of imaging informatics in medicine · 2026Article
- Leveraging large language models for information extraction from free-text liver MRI reports and assessment of clinical utility: a multicenter study.Quantitative imaging in medicine and surgery · 2026Article
- Fine-tuned multimodal GPT-4o for generating diagnostic impressions in breast magnetic resonance imaging: insights into non-mass enhancement lesions.Quantitative imaging in medicine and surgery · 2026Article
- Clinical Safety and Reliability of Large Language Models in Answering Hemorrhoid-Related Patient Questions: A Comparative Study of ChatGPT, Gemini, and DeepSeek.Healthcare (Basel, Switzerland) · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
10 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: An advanced version of large language model (LLM), ChatGPT 4o (GPT-4o), has shown capacity in image-text pair interpretation, yet the performance in medical image analysis remains unclear. This study aimed to evaluate the diagnostic capacity of GPT-4o in breast ultrasound (US) datasets. Methods: US exams including breast images and original reports were respectively included from January 2021 to December 2023 in three hospitals throughout China. The diagnostic performance in distinguishing benign or malignant breast masses of GPT-4o was assessed through two approaches: image-strategy and image-combined-text-strategy. Fleiss kappa was calculated to determine intra-LLM consistency. Thereafter, diagnostic accuracy was evaluated and compared with the convolutional neural network (CNN) model and 95 human experts with various levels of expertise from 60 institutions in China. Responses from GPT-4o were rated by diagnostic confidence and radiologist's evaluation. Results: The observations of 80 breast masses (37 malignant, 43 benign) from 80 patients [median age, 42.5 years; interquartile range (IQR), 37.0-53.0 years] were enrolled. GPT-4o with image-strategy exhibited a fair consistency [0.25, 95% confidence interval (CI): 0.07-0.43], whereas the agreement of image-combined-text-strategy was excellent (0.81, 95% CI: 0.67-0.91). Diagnostic accuracy improved when deploying the image-combined-text-strategy compared to only image [58% (46 of 80) Conclusions: The effectiveness of GPT-4o in interpreting real-world US images was limited, yet improved in image-combined-text-strategy. Deploying LLM warrants scrutiny from radiologists.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.