ArticleResearch square2026
GPT Fusion: Reasoning-Based Integration of Specialized Convolutional Neural Networks for Melanoma Diagnosis.
Article in Research square, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
Abstract
Background: Malignant melanoma (MM) is the most aggressive form of metastatic skin cancer, for which early detection is critical to reducing morbidity and mortality while improving patient outcomes. Recent advances in multimodal large language models (LLMs) have demonstrated considerable potential for melanoma detection and clinical decision support. However, their diagnostic accuracy remains inferior to that of specialized convolutional neural networks (CNNs). Whether LLMs can effectively integrate predictions from specialized CNN models to improve melanoma diagnosis has not been systematically investigated. Methods: We developed a hybrid diagnostic approach, termed GPT Fusion, in which GPT-5.5 serves as a reasoning-based fusion framework rather than an image classifier. Instead of directly analyzing lesion images, GPT-5.5 received structured outputs from two independently developed CNN models: a multimodal ResNet-50 model trained on the MILK10K dataset for multiclass skin lesion classification and the first-place 90-model SIIM-ISIC ensemble optimized for melanoma detection. GPT-5.5 integrated these complementary predictions to generate the final diagnosis. The framework was evaluated using the publicly available Derm7pt dataset. Performance was assessed for melanoma prediction (melanoma vs. non-melanoma), malignancy prediction (malignant vs. benign lesions), primary diagnosis, and top-3 differential diagnosis. Results: GPT Fusion substantially outperformed image-based GPT-5.5 across all evaluation tasks. For melanoma prediction, accuracy improved from 63.4% to 86.9%, while ROC AUC increased from 0.750 to 0.923 and PR AUC from 0.547 to 0.853. For malignancy prediction, GPT Fusion not only exceeded the performance of image-based GPT-5.5 but also surpassed the specialized SIIM-ISIC ensemble in sensitivity (83.9%), F1 score (73.6%), ROC AUC (0.894), and PR AUC (0.823), while maintaining comparable overall accuracy. GPT Fusion further improved primary diagnosis accuracy and achieved the highest top-3 differential diagnosis accuracy (89.0%). Conclusions: These findings demonstrate that LLMs can provide greater clinical value as reasoning-based integrators of specialized AI systems than as standalone image classifiers. By combining the complementary strengths of specialized CNN models, GPT-5.5 substantially improved diagnostic performance while providing interpretable reasoning. This hybrid strategy offers a promising direction for developing more accurate, transparent, and clinically useful AI systems for melanoma diagnosis.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.