ArticleBMC medical imaging2026
Multimodal deep learning for papillary thyroid carcinoma diagnosis using ultrasound and cytology.
Article in BMC medical imaging, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
BACKGROUND AND
objectivePapillary thyroid carcinoma (PTC) is the most common thyroid malignancy, and pre-operative diagnosis depends on integrating ultrasound (US) imaging with fine-needle aspiration cytology (FNAC). Pang et al. (2025) recently released a paired US/cytology dataset and demonstrated that classical radiomics combined with classifiers such as support vector machines, random forests, and XGBoost can reach AUROC ≈ 0.99 on a single random split. We re-examine this result under a stricter evaluation protocol and contribute a calibrated multimodal deep learning model for PTC diagnosis.
methodsUsing 384 patients from the Pang cohort (220 PTC, 164 benign) we created a stratified, untouched 20% holdout (n = 77). On the development set (n = 307) we performed 5-fold cross-validation and trained a multimodal model (v2) combining a ConvNeXt-Tiny ultrasound encoder, a domain-pretrained CTransPath cytology encoder, gated-attention multiple-instance learning (MIL) over cytology patches, and bidirectional cross-attention fusion. We compared this against a self-attention multimodal baseline (v1), unimodal ablations, and a modified Pang-style classical comparator that, because lesion masks were not released with the public dataset, used whole-image radiomics-style features rather than ROI-based features. Evaluation included bootstrap and Wilson confidence intervals, paired DeLong tests, McNemar tests, and a four-method calibration analysis (none, temperature, Platt, isotonic) using ensemble out-of-fold predictions and equal-mass binning.
resultsMultimodal v2 achieved holdout AUROC 0.977 (95% CI 0.949-0.996), Brier 0.042 (95% CI 0.014-0.079), sensitivity 0.977 (Wilson 0.882-0.996), and specificity 0.939 (Wilson 0.804-0.983) at the cross-validation Youden threshold. Across three random seeds (42, 7, 123), AUROC was 0.977 ± 0.0004 (mean ± SD). Paired DeLong showed our model statistically outperformed Pang's reimplemented Random Forest (p = 0.017) and XGBoost (p = 0.022) on identical holdout patients. Paired DeLong vs. the v1 baseline showed identical patient ranking (p = 1.0), but v2 showed numerically better operating-point and probability quality despite the identical ranking: Brier was roughly halved (0.042 vs. 0.083), MCC at threshold 0.5 increased by 0.105 (0.894 vs. 0.789), and uncalibrated expected calibration error fell from 0.114 to 0.042. The difference in paired binary decisions between v1 and v2 was, however, not statistically significant (McNemar p = 0.221). Calibration analysis revealed that temperature scaling, the de facto standard, was inappropriate for cross-validated ensembles, while isotonic regression on out-of-fold predictions reduced ECE by 53% on the v1 baseline.
conclusionsA multimodal model with domain-pretrained encoders and proper MIL aggregation achieves discriminative performance comparable to a modified Pang-style classical comparator and shows improved probability quality (lower Brier and calibration error), although these gains over the simpler v1 baseline did not reach statistical significance on the present holdout. We argue that AUROC alone is insufficient for evaluating clinical AI: discrimination and calibration are distinct properties, and operating-point performance is more directly relevant to deployment. We acknowledge that the present ultrasound branch operates on whole-image crops without explicit lesion localization and therefore should not yet be regarded as a lesion-specific clinical decision tool. Code and analysis scripts are released on GitHub.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.