Evidence map›Paper›PMID 41436487›Full record

ArticleScientific data2025

TCMEval-PA: a question-answering benchmark dataset for the prescription audit of Traditional Chinese Medicine.

Baifeng Wang, Yiwei Lu, Zhe Wang, Peixin Ge, Guanjie Wang, Keyu Yao, Suyuan Peng, Yan Zhu

Abstract readDataset
In one paragraph

Article in Scientific data, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Review
  2. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Baifeng WangDongzhimen Hospital, Beijing University of Chinese Medicine, Beijing, 100700, China.ORCID http://orcid.org/0009-0009-6812-0256
Yiwei LuInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, 100700, China.
Zhe WangInstitute of Basic Medical Sciences, Chinese Academy of Medical Sciences, Beijing, 100005, China.ORCID http://orcid.org/0009-0000-2387-784X
Peixin GeSchool of Medical Informatics, Changchun University of Traditional Chinese Medicine, Changchun, 130117, China.
Guanjie WangDepartment of Clinical Pharmacist, Weifang Hospital of Traditional Chinese Medicine, Weifang, 261041, China.
Keyu YaoInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, 100700, China.
Suyuan PengInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, 100700, China. peng.suyuan@bjmu.edu.cn.ORCID http://orcid.org/0000-0002-8221-7574
Yan ZhuInstitute of Information on Traditional Chinese Medicine, China Academy of Chinese Medical Sciences, Beijing, 100700, China. zhuyan166@126.com.ORCID http://orcid.org/0000-0002-5592-8258

Funding

China Academy of Chinese Medical Sciences (CACMS) ZZ170320China Academy of Chinese Medical Sciences (CACMS) ZZ18XRZ069Natural Science Foundation of Beijing Municipality (Beijing Natural Science Foundation) 7252253Natural Science Foundation of Beijing Municipality (Beijing Natural Science Foundation) 7254504
6 · The paper itself

Abstract

Although large language models (LLMs) have witnessed rapid development in medical applications, their capacities to support rational medication use and guarantee prescription safety remain insufficiently investigated-especially in tasks such as prescription audit, which plays a critical role in safeguarding both. This paper presents TCMEval-PA, a benchmark dataset for assessing the capabilities of LLMs in prescription audit of Chinese herbal medicines. The dataset comprises 328 choice questions, including 297 single-choice and 31 multiple-choice. All questions were designed and compiled through rule extraction from official documents and reviewed by licensed TCM physicians. TCMEval-PA comprehensively encompasses the key dimensions of prescription safety, including normativity (e.g., dispensing, decoction requirements, and regulations for special medicines) and appropriateness (e.g., contraindicated combinations and excessive dosages). The present study employed TCMEval-PA to assess several prevalent Chinese and English LLMs. This dataset can be utilized for the evaluation of LLMs and other artificial intelligence (AI) systems in TCM prescription safety scenarios and promotes research in intelligent auditing and decision support.

Indexed as

Drug PrescriptionsDrugs, Chinese HerbalMedicine, Chinese TraditionalArtificial IntelligenceBenchmarkingHumansLanguageDrugs, Chinese Herbal

Identifiers

PMID41436487
PMCPMC12827263

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.