Evidence map›Paper›PMID 42620269›Full record

ArticleResearch square2026

GPT Fusion: Reasoning-Based Integration of Specialized Convolutional Neural Networks for Melanoma Diagnosis.

Katie L Frederickson, Dong Li, Oren D Edrich, Ashish C Bhatia, Samuel E Adunyah, Qingguo Wang

Abstract readPreprint
In one paragraph

Article in Research square, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Katie L FredericksonDepartment of Biochemistry, Cancer Biology, Neurosciences and Pharmacology, School of Medicine, Meharry Medical College, Nashville, TN 37208, USA.ORCID 0009-0007-9318-6535
Dong LiHarris School of Public Policy, University of Chicago, Chicago, IL 60637, USA.
Oren D EdrichDepartment of Biochemistry, Cancer Biology, Neurosciences and Pharmacology, School of Medicine, Meharry Medical College, Nashville, TN 37208, USA.
Ashish C BhatiaNorthwestern University Feinberg School of Medicine, Chicago, IL 60637, USA.
Samuel E AdunyahDepartment of Biochemistry, Cancer Biology, Neurosciences and Pharmacology, School of Medicine, Meharry Medical College, Nashville, TN 37208, USA.
Qingguo WangDepartment of Biochemistry, Cancer Biology, Neurosciences and Pharmacology, School of Medicine, Meharry Medical College, Nashville, TN 37208, USA.ORCID 0000-0002-5125-3724

Funding

The RCMI Program in Health Disparities Research at Meharry Medical College - SupplementU54MD007586 · NIMHD · MEHARRY MEDICAL COLLEGE · PI SANIKA SAMUEL CHIRWA · 2017 to 2026
$48.2M
NIMHD NIH HHS U54 MD007586
6 · The paper itself

Abstract

Background: Malignant melanoma (MM) is the most aggressive form of metastatic skin cancer, for which early detection is critical to reducing morbidity and mortality while improving patient outcomes. Recent advances in multimodal large language models (LLMs) have demonstrated considerable potential for melanoma detection and clinical decision support. However, their diagnostic accuracy remains inferior to that of specialized convolutional neural networks (CNNs). Whether LLMs can effectively integrate predictions from specialized CNN models to improve melanoma diagnosis has not been systematically investigated. Methods: We developed a hybrid diagnostic approach, termed GPT Fusion, in which GPT-5.5 serves as a reasoning-based fusion framework rather than an image classifier. Instead of directly analyzing lesion images, GPT-5.5 received structured outputs from two independently developed CNN models: a multimodal ResNet-50 model trained on the MILK10K dataset for multiclass skin lesion classification and the first-place 90-model SIIM-ISIC ensemble optimized for melanoma detection. GPT-5.5 integrated these complementary predictions to generate the final diagnosis. The framework was evaluated using the publicly available Derm7pt dataset. Performance was assessed for melanoma prediction (melanoma vs. non-melanoma), malignancy prediction (malignant vs. benign lesions), primary diagnosis, and top-3 differential diagnosis. Results: GPT Fusion substantially outperformed image-based GPT-5.5 across all evaluation tasks. For melanoma prediction, accuracy improved from 63.4% to 86.9%, while ROC AUC increased from 0.750 to 0.923 and PR AUC from 0.547 to 0.853. For malignancy prediction, GPT Fusion not only exceeded the performance of image-based GPT-5.5 but also surpassed the specialized SIIM-ISIC ensemble in sensitivity (83.9%), F1 score (73.6%), ROC AUC (0.894), and PR AUC (0.823), while maintaining comparable overall accuracy. GPT Fusion further improved primary diagnosis accuracy and achieved the highest top-3 differential diagnosis accuracy (89.0%). Conclusions: These findings demonstrate that LLMs can provide greater clinical value as reasoning-based integrators of specialized AI systems than as standalone image classifiers. By combining the complementary strengths of specialized CNN models, GPT-5.5 substantially improved diagnostic performance while providing interpretable reasoning. This hybrid strategy offers a promising direction for developing more accurate, transparent, and clinically useful AI systems for melanoma diagnosis.

Indexed as

ChatGPTconvolutional neural networkdermoscopylarge language modelmelanoma diagnosis

Identifiers

PMID42620269
PMCPMC13484442

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.