ArticleFrontiers in digital health2026
A patient-aware benchmarking of CNN and transformer architectures for breast cancer histopathology classification.
Article in Frontiers in digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Introduction: Breast cancer diagnosis using histopathological imaging remains a critical yet challenging task, requiring automated systems that generalize reliably across patients and varying imaging conditions. While deep learning models have shown strong performance, many prior studies employ image-wise data splits that introduce patient-level data leakage, resulting in overly optimistic and potentially misleading evaluations. This study aims to address this limitation by establishing a rigorous, leakage-free benchmarking framework for binary breast cancer histopathology classification. Methods: A comprehensive evaluation of nine deep learning architectures was conducted on the BreaKHis dataset, comprising 7,909 images from 82 patients. The models include six convolutional neural networks (ResNet50, MobileNetV2, VGG16, DenseNet121, Xception, and EfficientNetB0), one modern convolutional architecture (ConvNeXt), and two transformer-based models (Swin-Small and Swin-Base). A strict 5-fold patient-aware cross-validation protocol was implemented to ensure that images from the same patient were not shared between training and validation sets. All models were trained under identical experimental conditions. Performance was assessed using accuracy, precision, recall, and F1-score, reported as mean ± standard deviation. Statistical significance was evaluated using paired Results: All evaluated architectures demonstrated comparable performance, achieving mean accuracies in the range of 0.91-0.93. ResNet50 achieved the highest mean accuracy (0.9267 ± 0.0435) and F1-score (0.9472), although differences among models were marginal. Statistical analysis confirmed that no pairwise differences were statistically significant ( Discussion: The findings highlight that, under a rigorously controlled and leakage-free evaluation protocol, architectural differences among modern deep learning models do not lead to statistically significant performance variations. Instead, evaluation design plays a more critical role in determining reliable outcomes. The proposed patient-aware benchmarking framework enhances reproducibility and provides a robust foundation for future research, supporting the development of clinically translatable AI systems for breast cancer diagnosis.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.