ArticleScientific data2025
TCMEval-PA: a question-answering benchmark dataset for the prescription audit of Traditional Chinese Medicine.
Article in Scientific data, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
2 citing papers in PubMed.
- Beyond transparency: why Traditional Chinese Medicine (TCM) need explainable artificial intelligence (XAI).Chinese medicine · 2026Review
- A roadmap for medical large language models: a review of foundations, applications, and challenges.Military Medical Research · 2026Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
8 authors.
Funding
Abstract
Although large language models (LLMs) have witnessed rapid development in medical applications, their capacities to support rational medication use and guarantee prescription safety remain insufficiently investigated-especially in tasks such as prescription audit, which plays a critical role in safeguarding both. This paper presents TCMEval-PA, a benchmark dataset for assessing the capabilities of LLMs in prescription audit of Chinese herbal medicines. The dataset comprises 328 choice questions, including 297 single-choice and 31 multiple-choice. All questions were designed and compiled through rule extraction from official documents and reviewed by licensed TCM physicians. TCMEval-PA comprehensively encompasses the key dimensions of prescription safety, including normativity (e.g., dispensing, decoction requirements, and regulations for special medicines) and appropriateness (e.g., contraindicated combinations and excessive dosages). The present study employed TCMEval-PA to assess several prevalent Chinese and English LLMs. This dataset can be utilized for the evaluation of LLMs and other artificial intelligence (AI) systems in TCM prescription safety scenarios and promotes research in intelligent auditing and decision support.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.