Evidence map›Paper›PMID 41951858›Full record

ArticleNPJ digital medicine2026

ClinicRealm: Re-evaluating large language models with conventional machine learning for non-generative clinical prediction tasks.

Yinghao Zhu, Junyi Gao, Zixiang Wang, Weibin Liao, Xiaochen Zheng, Lifang Liang, Miguel O Bernabeu, Yasha Wang, Lequan Yu, Chengwei Pan and 2 more

Abstract read
In one paragraph

Article in NPJ digital medicine, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Yinghao Zhu *School of Artificial Intelligence, Beihang University, Beijing, China.
Junyi Gao *Centre for Medical Informatics, The University of Edinburgh, Edinburgh, UK.
Zixiang Wang *National Engineering Research Center for Software Engineering, Peking University, Beijing, China.
Weibin Liao *National Engineering Research Center for Software Engineering, Peking University, Beijing, China.
Xiaochen ZhengETH Zurich, Zurich, Switzerland.
Lifang LiangNational Engineering Research Center for Software Engineering, Peking University, Beijing, China.
Miguel O BernabeuCentre for Medical Informatics, The University of Edinburgh, Edinburgh, UK.
Yasha WangNational Engineering Research Center for Software Engineering, Peking University, Beijing, China.
Lequan YuSchool of Computing and Data Science, The University of Hong Kong, Hong Kong SAR, China.
Chengwei PanSchool of Artificial Intelligence, Beihang University, Beijing, China. pancw@buaa.edu.cn.
Ewen M HarrisonCentre for Medical Informatics, The University of Edinburgh, Edinburgh, UK. ewen.harrison@ed.ac.uk.
Liantao MaNational Engineering Research Center for Software Engineering, Peking University, Beijing, China. malt@pku.edu.cn.

Funding

Beijing Natural Science Foundation L244063Health Data Research UK-The Alan Turing Institute Wellcome PhD Programme in Health Data Science 218529/Z/19/ZNational Natural Science Foundation of China 62402017Peking University Medicine plus X Pilot Program-Key Technologies R&D Project 2024YXXLHGG007Xuzhou Scientific Technological Projects KC23143
6 · The paper itself

Abstract

Large Language Models (LLMs) are increasingly deployed in medicine. However, their utility for non-generative clinical prediction is under-evaluated, and they are often assumed to be inferior to specialized models, creating potential for misuse and misunderstanding. To address this, our ClinicRealm benchmark systematically evaluates 15 GPT-style LLMs, 5 BERT-style models, and 11 traditional methods on unstructured clinical notes and structured Electronic Health Records (EHR) across predictive performance, reasoning, fairness, etc. Our findings reveal a significant shift: on clinical notes, leading zero-shot LLMs (e.g., DeepSeek-V3.1-Think, GPT-5) now decisively outperform finetuned BERT models. On structured EHRs, while specialized models excel with ample data, advanced LLMs demonstrate potent zero-shot capabilities, often surpassing conventional models in data-scarce settings. Notably, leading open-source LLMs match or exceed their proprietary counterparts. This provides compelling evidence that modern LLMs are competitive tools for clinical prediction, necessitating a re-evaluation of model selection strategies by health data scientists and developers.

Identifiers

PMID41951858
PMCPMC13079879

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.