Evidence map›Paper›PMID 42783886›Full record

ArticleJournal of imaging2026

Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses.

Giulia Soccio, Stefania Di Napoli, Paolo Trerotoli, Vera Loizzi, Laura Grazia Zompì, Giuseppe Colonna, Daniele La Forgia, Gennaro Cormio, Francesca Arezzo

Abstract read
In one paragraph

Article in Journal of imaging, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Giulia SoccioDepartment of Interdisciplinary Medicine (DIM), University of Bari "Aldo Moro", 70124 Bari, Italy.
Stefania Di NapoliDepartment of Interdisciplinary Medicine (DIM), University of Bari "Aldo Moro", 70124 Bari, Italy.
Paolo TrerotoliDepartment of Interdisciplinary Medicine (DIM), University of Bari "Aldo Moro", 70124 Bari, Italy.ORCID 0000-0001-7102-9108
Vera LoizziGynecologic Oncology Unit, IRCCS Istituto Tumori "Giovanni Paolo II", 70124 Bari, Italy.ORCID 0000-0002-4006-6421
Laura Grazia ZompìDepartment of Interdisciplinary Medicine (DIM), University of Bari "Aldo Moro", 70124 Bari, Italy.
Giuseppe ColonnaGynecologic Oncology Unit, IRCCS Istituto Tumori "Giovanni Paolo II", 70124 Bari, Italy.ORCID 0009-0008-4672-5003
Daniele La ForgiaSSD Radiologia Senologia, IRCCS Istituto Tumori "Giovanni Paolo II", 70124 Bari, Italy.ORCID 0000-0002-3902-4523
Gennaro CormioGynecologic Oncology Unit, IRCCS Istituto Tumori "Giovanni Paolo II", 70124 Bari, Italy.ORCID 0000-0001-5745-372X
Francesca ArezzoGynecologic Oncology Unit, IRCCS Istituto Tumori "Giovanni Paolo II", 70124 Bari, Italy.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Ovarian cancer (OC) is the second most common gynecological malignancy and remains one of the leading causes of gynecological cancer-related mortality worldwide. A major clinical challenge is the lack of an accurate and widely applicable strategy for identifying patients at high risk of malignancy at an early stage. In this context, artificial intelligence (AI) has emerged as a promising tool to improve diagnostic performance. Among AI technologies, large language models (LLMs) have recently shown considerable potential in healthcare applications. In this study, we evaluated the diagnostic performance of ChatGPT (GPT-5) in classifying 300 adnexal masses as benign or malignant and compared its performance with that of the IOTA Simple Rules, the ADNEX model, and expert subjective assessment. We also assessed ChatGPT's ability to predict the most likely histological diagnosis for each lesion. All adnexal masses were described using the International Ovarian Tumor Analysis (IOTA) terminology, and histopathological examination served as the reference standard. Our findings showed that expert subjective assessment achieved the highest overall diagnostic performance for both benign/malignant classification (accuracy 87.3%; 95% CI, 83.0-90.9%) and prediction of the presumed histological diagnosis. ChatGPT A and ChatGPT B reached a sensitivity of 72.3% and 73.5%, a specificity of 74.5% and 75.9%, a positive predictive value of 75.2% and 76.5%, and a negative predictive value of 71.5% and 72.8%, respectively (inconclusive responses counted as misclassifications), with an overall accuracy of 73.3% and 74.7%. After adequate validation, large language models might complement existing decision-support tools for less experienced examiners, without replacing expert evaluation. Their ease of use and reliance on standardized ultrasound descriptors make them accessible to ultrasonographers with varying levels of expertise.

Indexed as

artificial intelligenceChatGPTlarge language modelsovarian cancerultrasound

Identifiers

PMID42783886
PMCPMC13608372

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.