Evidence map›Paper›PMID 40630532›Full record

ArticleResearch square2025

Evaluating Vision and Pathology Foundation Models for Computational Pathology: A Comprehensive Benchmark Study.

Olivier Gevaert, Rohan Bareja, Francisco Carrillo-Perez, Yuanning Zheng, Marija Pizurica, Tarak Nandi, Jeanne Shen, Ravi Madduri

Abstract readPreprint
In one paragraph

Article in Research square, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

5 · Who and what money

Authors and funding

8 authors.

Olivier GevaertStanford University School of Medicine.ORCID 0000-0002-9965-5466
Rohan BarejaStanford University School of Medicine.
Francisco Carrillo-PerezStanford University School of Medicine.
Yuanning ZhengStanford University.ORCID 0000-0002-0018-3252
Marija PizuricaGhent University.
Tarak NandiData Science and Learning Division, Argonne National Laboratory.
Jeanne ShenStanford University.ORCID 0000-0002-1519-0308
Ravi MadduriArgonne National Laboratory.ORCID 0000-0003-2130-2887

Funding

Multi-scale modeling of glioma for the prediction of treatment response, treatment monitoring and treatment allocationR01CA260271 · NCI · STANFORD UNIVERSITY · PI GEVAERT, OLIVIER · 2021 to 2025
$3.1M
NCI NIH HHS R01 CA260271
6 · The paper itself

Abstract

To advance precision medicine in pathology, robust AI-driven foundation models are increasingly needed to uncover complex patterns in large-scale pathology datasets, enabling more accurate disease detection, classification, and prognostic insights. However, despite substantial progress in deep learning and computer vision, the comparative performance and generalizability of these pathology foundation models across diverse histopathological datasets and tasks remain largely unexamined. In this study, we conduct a comprehensive benchmarking of 31 AI foundation models for computational pathology, including general vision models (VM), general vision-language models (VLM), pathology-specific vision models (Path-VM), and pathology-specific vision-language models (Path-VLM), evaluated over 41 tasks sourced from TCGA, CPTAC, external benchmarking datasets, and out-of-domain datasets. Our study demonstrates that Virchow2, a pathology foundation model, delivered the highest performance across TCGA, CPTAC, and external tasks, highlighting its effectiveness in diverse histopathological evaluations. We also show that Path-VM outperformed both Path-VLM and VM, securing top rankings across tasks despite lacking a statistically significant edge over vision models. Our findings reveal that model size and data size did not consistently correlate with improved performance in pathology foundation models, challenging assumptions about scaling in histopathological applications. Lastly, our study demonstrates that a fusion model, integrating top-performing foundation models, achieved superior generalization across external tasks and diverse tissues in histopathological analysis. These findings emphasize the need for further research to understand the underlying factors influencing model performance and to develop strategies that enhance the generalizability and robustness of pathology-specific vision foundation models across different tissue types and datasets. PathBench : https://pathbench.stanford.edu/.

Identifiers

PMID40630532
PMCPMC12236927

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.