Evidence map›Paper›PMID 40044818›Full record

ArticleNature biomedical engineering2025

A data-efficient strategy for building high-performing medical foundation models.

Yuqi Sun, Weimin Tan, Zhuoyao Gu, Ruian He, Siyuan Chen, Miao Pang, Bo Yan

Abstract read
PubMed Publisher
In one paragraph

Article in Nature biomedical engineering, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 13 papers.

0numbers the graph read from it
0cells of the map it votes in
13citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

13 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
  4. Article
  5. Review
  6. Article
  7. Article
  8. A Generative Foundation Model for Scalable Cytology Image Synthesis in AI-Powered Diagnostics.Clinical cancer research : an official journal of the American Association for Cancer Research · 2026
    Article
  9. Article
  10. 2025 in review.Nature biomedical engineering · 2025
    Article
  11. Synthetic data boosts medical foundation models.Nature biomedical engineering · 2025
    Article
  12. Review
  13. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Yuqi Sun *Shanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.
Weimin Tan *Shanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.ORCID http://orcid.org/0000-0001-7677-4772
Zhuoyao GuShanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.
Ruian HeShanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.ORCID http://orcid.org/0000-0001-9598-3043
Siyuan ChenShanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.
Miao PangShanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China.
Bo YanShanghai Key Laboratory of Intelligent Information Processing, School of Computer Science, Fudan University, Shanghai, China. byan@fudan.edu.cn.ORCID http://orcid.org/0000-0001-5692-3486

Funding

National Natural Science Foundation of China (National Science Foundation of China) 62372117National Natural Science Foundation of China (National Science Foundation of China) 62472102National Natural Science Foundation of China (National Science Foundation of China) U2001209Natural Science Foundation of Shanghai (Natural Science Foundation of Shanghai Municipality) 21ZR1406600
6 · The paper itself

Abstract

Foundation models are pretrained on massive datasets. However, collecting medical datasets is expensive and time-consuming, and raises privacy concerns. Here we show that synthetic data generated via conditioning with disease labels can be leveraged for building high-performing medical foundation models. We pretrained a retinal foundation model, first with approximately one million synthetic retinal images with physiological structures and feature distribution consistent with real counterparts, and then with only 16.7% of the 904,170 real-world colour fundus photography images required in a recently reported retinal foundation model (RETFound). The data-efficient model performed as well or better than RETFound across nine public datasets and four diagnostic tasks; and for diabetic-retinopathy grading, it used only 40% of the expert-annotated training data used by RETFound. We also support the generalizability of the data-efficient strategy by building a classifier for the detection of tuberculosis on chest X-ray images. The text-conditioned generation of synthetic data may enhance the performance and generalization of medical foundation models.

Indexed as

RetinaAlgorithmsDatabases, FactualDiabetic RetinopathyHumansImage Processing, Computer-AssistedRadiography, ThoracicTuberculosis

Identifiers

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.