Evidence map›Paper›PMID 41184207›Full record

ArticleJMIR medical informatics2025

Performance of the Large Language Models on the Chinese National Nurse Licensure Examination: Cross-Sectional Evaluation Study.

Longhui Xu, Xiao Cong, Renxiu Wang, Na Li, Xinru Liu, Ronghui Wang, Cuiping Xu

Abstract read
In one paragraph

Article in JMIR medical informatics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Longhui XuSchool of Nursing, Shandong University of Traditional Chinese Medicine, Jinan, China.ORCID 0009-0008-2790-8701
Xiao CongSchool of Nursing, Shandong University of Traditional Chinese Medicine, Jinan, China.ORCID 0009-0009-5066-0676
Renxiu WangDepartment of Nursing, Shandong Provincial QianFoShan Hospital, No.16766 Jingshi Road, Jinan, 250014, China, 86 13791126826.ORCID 0009-0006-0932-5786
Na LiSchool of Nursing, Shandong University of Traditional Chinese Medicine, Jinan, China.ORCID 0009-0004-7582-3113
Xinru LiuSchool of Nursing, Shandong University of Traditional Chinese Medicine, Jinan, China.ORCID 0009-0008-7676-7925
Ronghui WangSchool of Nursing, Shandong University of Traditional Chinese Medicine, Jinan, China.ORCID 0009-0007-4447-1081
Cuiping XuDepartment of Nursing, Shandong Provincial QianFoShan Hospital, No.16766 Jingshi Road, Jinan, 250014, China, 86 13791126826.ORCID 0000-0002-8868-4878

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs) are increasingly explored in nursing education, but their capabilities in specialized, high-stakes, culturally specific examinations, such as the Chinese National Nurse Licensure Examination (CNNLE), remain underevaluated, making rigorous evaluation crucial before their adoption in nursing training and practice. Objective: This study aimed to evaluate the performance, accuracy, repeatability, confidence, and robustness of 4 LLMs on the CNNLE. Methods: Four LLMs (Sider Fusion [Vidline Inc], GPT-4o [OpenAI], Gemini 2.0 Pro [Google DeepMind], and DeepSeek V3) were tested on 237 multiple-choice questions from the 2024 CNNLE. Accuracy and repeatability were assessed using 2 prompting strategies. Confidence was evaluated via self-ratings (1-10 scale) and robustness via repeated adversarial prompting. Results: DeepSeek V3 and Gemini 2.0 Pro demonstrated significantly higher overall accuracy (ranging from 199/237 to 209/237; >83%) compared to GPT-4o and Sider Fusion (ranging from 151/237 to 166/237; <71%). However, all LLMs showed suboptimal repeatability (highest at 206/237; <87% consistency). Critically, poor confidence calibration was evident; most models showed high confidence often mismatching actual accuracy (Sider Fusion: P=.01; GPT-4o: P=.03; and Gemini 2.0 Pro: P=.049). A stability-flexibility trade-off paradox was also observed. Conclusions: While some LLMs show promising accuracy on the CNNLE, fundamental reliability limitations (poor confidence calibration and inconsistent repeatability) hinder safe application in nursing education and practice. Future LLM development must prioritize trustworthiness and calibrated reliability over surface accuracy.

Indexed as

Educational MeasurementLanguageLicensure, NursingChinaCross-Sectional StudiesHumansLarge Language ModelsReproducibility of Resultsaccuracyartificial intelligenceconfidencelarge language modelsnursing educationreliabilityrobustness

Identifiers

PMID41184207
PMCPMC12582878

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.