Evidence map›Paper›PMID 41927948›Full record

ArticleDiscover oncology2026

Explainable vision transformer framework for multi-class classification and prognostic interpretation of oral cancer in histopathology images.

Chandrakanta Mahanty, Chin-Shiuh Shieh, Mong-Fong Horng, Anusha Nallamalla, S Gopal Krishna Patro, Shafat Khan, Trmesgen Engida

Abstract read
In one paragraph

Article in Discover oncology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Article
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Chandrakanta MahantyDepartment of Computer Science and Engineering, GITAM Deemed to Be University, Visakhapatnam, India. cmahanty@gitam.edu.ORCID http://orcid.org/0000-0002-3084-308X
Chin-Shiuh ShiehDepartment of Electronic Engineering, National Kaohsiung University of Science and Technology (NKUST), Kaohsiung, Taiwan.
Mong-Fong HorngDepartment of Electronic Engineering, National Kaohsiung University of Science and Technology (NKUST), Kaohsiung, Taiwan.
Anusha NallamallaDepartment of Computer Science, GITAM School of Science, GITAM Deemed to Be University, Visakhapatnam, India.
S Gopal Krishna PatroSchool of Engineering, Sreenidhi University, Hyderabad, 501301, India. gopal.s@suh.edu.in.
Shafat KhanDepartment of Computer Science, College of Computer Science, King Khalid University, Abha, 61421, Saudi Arabia.
Trmesgen EngidaComputational and Natural Sciences College, Dilla University, Dilla, Ethiopia. temueng@yahoo.com.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Oral cancer, particularly Oral Squamous Cell Carcinoma (OSCC), remains a major global health concern due to its high prevalence, late diagnosis, and limited prognostic precision in conventional histopathological evaluation. Although deep learning has been showing promising results in automated cancer classification, most current models, especially CNN-based architectures, generally lack interpretability, generalization capability, and prognostic insight, hence limiting their clinical applicability. To address these shortcomings, this work introduces an Explainable Vision Transformer framework (MMX-ViT) for multi-class classification and prognostic interpretation of oral cancer in histopathology images. The proposed model fuses convolutional feature extraction with transformer-based global attention using an Adaptive Cross-Fusion Module (ACFM), allowing efficient multi-scale learning of cellular and tissue-level features. The MMX-ViT model was trained and evaluated on a publicly available oral cancer histopathology dataset, extended in this study into four diagnostic categories, and compared with eight state-of-the-art architectures. It reached a high classification performance of 98.45%, with an AUC of 0.99, thus surpassing all the baseline methods. Explainability analysis based on Grad-CAM + + , SHAP, and Transformer Attention Rollout techniques demonstrated that biologically relevant areas of attention were identified by the model, such as dysplastic nuclei, keratin pearls, and invasion zones in stroma, with an XCI (Explainability Consistency Index) value of 94%. The model proposed here represents a major progress towards the establishment of reliable and interpretable AI-based diagnosis of oral cancer.

Indexed as

Explainable artificial intelligenceHistopathology image analysisMulti-class cancer classificationOral squamous cell carcinomaVision transformer

Identifiers

PMID41927948
PMCPMC13172119

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.