Evidence map›Paper›PMID 41397693›Full record

ArticleJMIR AI2025

Effectiveness of ChatGPT, Google Gemini, and Microsoft Copilot in Answering Thai Drug Information Queries: Cross-Sectional Study.

Suphannika Pornwattanakavee, Nattawut Leelakanok, Teerarat Todsarot, Gabrielle Angele Tatta Guinto, Ratchanon Takun, Assadawut Sumativit, Marisa Senngam

Abstract read
In one paragraph

Article in JMIR AI, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Suphannika PornwattanakaveeDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0000-0002-1386-0902
Nattawut LeelakanokDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0000-0003-3533-4775
Teerarat TodsarotDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0009-0002-1396-2172
Gabrielle Angele Tatta GuintoDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0009-0002-4398-6415
Ratchanon TakunDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0009-0004-5996-4119
Assadawut SumativitDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0009-0008-1409-2547
Marisa SenngamDivision of Clinical Pharmacy, Faculty of Pharmaceutical Sciences, Burapha University, Chonburi, Thailand.ORCID https://orcid.org/0000-0002-2936-8066

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundChatGPT-4o, Google Gemini, and Microsoft Copilot have shown potential in generating health care-related information. However, their accuracy, completeness, and safety for providing drug-related information in Thai contexts remain underexplored.

objectiveThis study aims to evaluate the performance of artificial intelligence (AI) systems in responding to drug-related questions in Thai.

methodsAn analytical cross-sectional study was conducted using 76 public drug-related questions compiled from medical databases and social media between November 1, 2019, and December 31, 2024. All questions were categorized into 19 distinct categories, each comprising 4 questions. ChatGPT-4o, Google Gemini, and Microsoft Copilot were queried in a single session on March 1, 2025, by using input in Thai. All responses were evaluated for correctness, completeness, risk, and reproducibility independently by clinical pharmacists using standardized evaluation criteria.

resultsAll 3 AI models provided generally complete responses (P=.08). ChatGPT-4o yielded the highest proportion of fully correct responses (P=.08). The overall risk levels of high-risk answers were not significantly different (P=.12). Response correctness was influenced by the category of the drug-related questions (P=.002) but not completeness (P=.23). The correctness of Google Gemini and Microsoft Copilot was higher than that of ChatGPT for pharmacology queries. The type of questions also statistically significantly affected the risk level of the answers (P=.04). In particular, the pregnancy and lactation category had the highest high-risk response rate (1/76, 1% per system). All 3 AI models demonstrated consistent response patterns when the same questions were re-queried after 1, 7, and 14 days.

conclusionsThe evaluated AI chatbots were able to answer the queries with generally complete content; however, we found limited accuracy and occasional high-risk errors in responding to drug-related questions in Thai. All models exhibited good reproducibility.

Indexed as

accuracyartificial intelligencechatbotChatGPT-4ocompletenesscorrectnessdrug informationGoogle Geminimedication informationMicrosoft Copilotrisk assessment

Identifiers

PMID41397693
PMCPMC12750067

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.