Evidence map›Paper›PMID 42676381›Full record

ArticleFrontiers in artificial intelligence2026

Evaluating the performance of DeepSeek-V3, Doubao, ERNIE 4.5 Turbo and iFLYTEK Spark in the Chinese national nursing licensing examination: a cross-sectional comparative study.

Min Yang, Li Zhu, Zeju Zhang, Qigui Xu

Abstract read
In one paragraph

Article in Frontiers in artificial intelligence, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Min YangSchool of Nursing, Chongqing Medical and Pharmaceutical College, Chongqing, China.
Li ZhuSchool of Nursing, Chongqing Medical and Pharmaceutical College, Chongqing, China.
Zeju ZhangSchool of Nursing, Chongqing Medical and Pharmaceutical College, Chongqing, China.
Qigui XuSchool of Pharmacy, Chongqing Medical University, Chongqing, China.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: While large language models (LLMs) have been widely adopted in nursing education and several studies have evaluated their performance on the Chinese National Nursing Licensing Examination (NNLE), systematic comparisons focusing on the latest generation of Chinese LLMs such as DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark remain limited. Furthermore, no study has specifically examined the multimodal capabilities of these models using image-based nursing questions. Objective: This study aimed to compare the performance of DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark LLMs on the NNLE and evaluate their potential for nursing education. Methods: This cross-sectional study assessed DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark using questions from the 2025 NNLE, including both authentic text-based items ( Results: On text-based NNLE items, all four models achieved accuracy rates above 90% (DeepSeek-V3: 92.9, 95% CI [89.0, 95.6%]; Doubao: 94.6, 95% CI [91.0, 96.9%]; ERNIE 4.5 Turbo: 93.3, 95% CI [89.5, 95.9%]; iFLYTEK Spark: 92.1, 95% CI [88.0, 95%]), meeting the NNLE passing requirement. No significant differences in overall accuracy were observed among models (all pairwise comparisons Conclusion: This study demonstrates that DeepSeek-V3, Doubao, ERNIE 4.5 Turbo, and iFLYTEK Spark perform well on Chinese nursing examinations questions in text format, indicating their potential as auxiliary resources for nursing education and exam preparation. However, their performance on image-based items is substantially lower, and the findings do not support claims of clinical applicability. Further research is needed to assess their utility in real-world clinical reasoning or patient care.

Indexed as

AIartificial intelligenceChinese nursing licensing examinationcross-sectional designlarge language models

Identifiers

PMID42676381
PMCPMC13526601

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.