Evidence map›Paper›PMID 40854301›Full record

ArticleJMIR formative research2025

Performance Assessment of ChatGPT-4.0 and ChatGLM Series in Traditional Chinese Medicine for Metabolic Associated Fatty Liver Disease: Comparative Study.

Xionghui Wang, Tianxiao Zheng, Bo Liu, Zhi Pei, Kaihan Meng, Changquan Ling

Abstract readComparative Study
In one paragraph

Article in JMIR formative research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Review
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Xionghui Wang *Department of Gastroenterology, No. 967 Hospital of PLA Joint Logistics Support Force, Dalian, China.ORCID 0009-0004-1522-2607
Tianxiao Zheng *School of Traditional Chinese Medicine, Naval Medical University, No. 800, Xiangyin Road, Yangpu District, Shanghai, 200433, China, 86 02181871561.ORCID 0009-0002-3792-8833
Bo Liu *Medical Administration Office, No. 967 Hospital of PLA Joint Logistics Support Force, Dalian, China.ORCID 0009-0000-5296-6964
Zhi PeiSchool of Traditional Chinese Medicine, Naval Medical University, No. 800, Xiangyin Road, Yangpu District, Shanghai, 200433, China, 86 02181871561.ORCID 0009-0007-0358-6034
Kaihan MengSchool of Traditional Chinese Medicine, Naval Medical University, No. 800, Xiangyin Road, Yangpu District, Shanghai, 200433, China, 86 02181871561.ORCID 0009-0007-2436-2371
Changquan LingSchool of Traditional Chinese Medicine, Naval Medical University, No. 800, Xiangyin Road, Yangpu District, Shanghai, 200433, China, 86 02181871561.ORCID 0000-0002-0037-4059

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: ChatGPT-4.0 and the ChatGLM series are novel conversational large language models (LLMs). ChatGLM includes 3 versions: ChatGLM4 (with internet connectivity but no knowledge base pretraining), ChatGLM4+Knowledge base (combining internet search capabilities with knowledge base pretraining), ChatGLM3-6B (offline knowledge base pretraining but no internet connectivity). The ability of ChatGPT-4.0 and ChatGLM to apply medical knowledge in the Chinese environment has been preliminarily verified, but the potential of the 2 models for clinical assistance in traditional Chinese medicine (TCM) is still unknown. Objective: This study aims to explore the performance of ChatGPT-4.0, ChatGLM4, ChatGLM4+Knowledge base, and ChatGLM3-6B in providing AI-assisted diagnosis and treatment for metabolic dysfunction-associated fatty liver disease within a TCM clinical framework, thereby assessing their potential as TCM clinical decision support tools. Methods: This study evaluated 4 LLMs by providing them with medical records of 87 metabolic dysfunction-associated fatty liver disease cases treated with TCM and querying them about TCM treatment plans. The answering texts from 4 LLMs were evaluated using predefined scoring criteria, focusing on 3 critical dimensions: ability in syndrome differentiation and treatment principles, confusion of concepts between TCM and Western medicine, and comprehensive evaluation of question-answering texts (comprising 6 components: ability to integrate Chinese and Western medicine, ability to formulate treatment plans, health management capacity, disease monitoring ability, self-positioning awareness, and medication safety). Results: In the evaluation module of "Ability in syndrome differentiation and treatment principles," the performance ranking of the 4 models was: (1) ChatGLM4+ Knowledge Base, (2) ChatGLM4, (3) ChatGLM3-6B, and (4) ChatGPT-4.0. Regarding the assessment of confusion between TCM and Western medicine concepts, ChatGPT-4.0 exhibited conceptual confusion in 32 out of 87 cases, while the ChatGLM series of LLMs showed no such confusion (except for ChatGLM3-6B, which had 1 instance). In the "Comprehensive evaluation of question-answering texts" module (comprising 6 components: ability to integrate Chinese and Western medicine, ability to formulate treatment plans, health management capacity, disease monitoring ability, self-positioning awareness, and medication safety), the ranking was: (1) ChatGLM4+ Knowledge Base, (2) ChatGPT-4.0, (3) ChatGLM4, and (4) ChatGLM3-6B. Conclusions: Our study results demonstrated that real-time internet connectivity played a critical role in LLM-assisted TCM diagnosis and treatment, while offline models showed significantly reduced performance in clinical decision support. Furthermore, pretraining LLMs with TCM-specific knowledge bases while maintaining internet search capabilities substantially enhanced their diagnostic and therapeutic performance in TCM applications. Importantly, general-purpose LLMs required both domain-specific medical fine-tuning and culturally sensitive adaptation to meet the rigorous standards of TCM clinical practice.

Indexed as

Decision Support Systems, ClinicalFatty LiverMedicine, Chinese TraditionalAdultFemaleGenerative Artificial IntelligenceHumansKnowledge BasesMaleMiddle AgedAI medical assistantsartificial intelligenceChinese language modelsinternet-enabled LLMsknowledge base fine-tuninglarge language modelstraditional Chinese medicine

Identifiers

PMID40854301
PMCPMC12377871

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.