Evidence map›Paper›PMID 42391223›Full record

ArticlePloS one2026

From knowledge to judgment: A three-year longitudinal analysis of artificial intelligence large language model performance on the Chinese national nurse licensing examination.

Xinju Zhan, Weihua Yu, Jianshu Cai, Jionghuang Chen

Abstract read
In one paragraph

Article in PloS one, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Xinju ZhanNursing Department, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, Hangzhou, China.
Weihua YuDepartment of General Surgery, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, Hangzhou, China.
Jianshu CaiNursing Department, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, Hangzhou, China.
Jionghuang ChenDepartment of General Surgery, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, Hangzhou, China.ORCID https://orcid.org/0009-0007-2618-6500

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe rapid advancement of Large Language Models (LLMs) presents unprecedented opportunities for healthcare education and professional credentialing. However, comprehensive longitudinal analyses of their evolving capabilities in nursing contexts remain limited.

objectiveTo conduct a three-year longitudinal performance analysis of major international and Chinese-native LLMs on the Chinese National Nurse Licensing Examination (NNLE) from July 2022 to June 2025, examining performance trajectories, comparative effectiveness, and domain-specific competencies.

methodsWe curated a comprehensive corpus of 9,800 multiple-choice questions from NNLE examinations (2022-2025) through validated educational resources. Fifteen leading LLMs were evaluated using standardized zero-shot prompting protocols, with temporal fidelity ensuring models were tested only on examinations administered after their release dates. Performance was measured as raw accuracy and benchmarked against the approximate 300-point passing threshold. Statistical analyses included trend analysis, comparative performance testing, and qualitative error categorization.

resultsLLM performance demonstrated a steep upward trajectory, with top-tier models achieving accuracy rates from 47.0% in 2022 to 78.8% in 2025. Chinese-native models consistently outperformed international counterparts. The mean Chinese-native advantage decreased from 6.1 percentage points in 2023 to 3.0 percentage points in 2025, while the top-model advantage remained present but non-monotonic, measuring 4.5, 3.0, and 3.8 percentage points in 2023, 2024, and 2025, respectively. Models exhibited superior performance in the knowledge-oriented Professional Practice section (81.6% average accuracy) versus the application-oriented Practical Skills section (70.9% average accuracy). Clinical reasoning failures, particularly in nursing intervention prioritization, constituted 43% of errors among top-performing models.

conclusionWhile state-of-the-art LLMs demonstrate substantial codified nursing knowledge sufficient to achieve approximate passing thresholds on professional licensing examinations, significant deficiencies in complex clinical judgment persist, defining the current boundary between artificial intelligence capabilities and human professional competence. Critically, examination performance should not be interpreted as evidence of clinical readiness or autonomous practice capability.

Indexed as

Artificial IntelligenceEducational MeasurementLicensure, NursingChinaHumansJudgmentLarge Language ModelsLongitudinal Studies

Identifiers

PMID42391223
PMCPMC13327122

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.