Evidence map›Paper›PMID 41366422›Full record

ArticleBMC medical ethics2025

Performance of large language models in non-English medical ethics-related multiple choice questions: comparison of ChatGPT performance across versions and languages.

Yoongu Kim, Soan Shin, Sang-Ho Yoo

Abstract readComparative Study
In one paragraph

Article in BMC medical ethics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 6 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
6citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

6 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. Article
  4. Article
  5. Evaluating the efficacy and readability of advanced large language models in responding to patients' frequently asked questions about chronic rhinosinusitis: a comparative analysis.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Article
  6. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Yoongu KimDepartment of Medical Humanities and Ethics, College of Medicine, Hanyang University, 222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Korea.ORCID 0000-0003-2256-7473
Soan ShinDepartment of Medical Humanities and Ethics, College of Medicine, Hanyang University, 222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Korea.ORCID 0009-0000-1048-1475
Sang-Ho YooDepartment of Medical Humanities and Ethics, College of Medicine, Hanyang University, 222, Wangsimni-ro, Seongdong-gu, Seoul, 04763, Korea. karmaboy@hanyang.ac.kr.ORCID 0000-0001-9030-1365

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAs large language models (LLMs) evolve, assessing their competence in ethically sensitive domains such as medical ethics has become increasingly important. Since medical ethics is a universal component of medical education, disparities in AI performance across languages may result in unequal benefits for learners. Therefore, it is essential to examine performances in non-English contexts. While previous studies have evaluated performance of Chat Generative Pre-trained Transformer(ChatGPT) on English-language multiple-choice questions (MCQs) in medical ethics, none have examined version-based improvements across non-English contexts. This study therefore evaluated ChatGPT versions 3.5, 4.0, and 4.5 for MCQs on Korean medical ethics and their English translations, with a focus on performance trends across versions and languages.

methodsWe selected 36 MCQs from the Korean National Medical Licensing Examination and the Comprehensive Clinical Medicine Evaluation databases. Each question was entered ten times per ChatGPT version (3.5, 4.0, 4.5) and language (Korean, English) for a total of 60 trials. Additionally, to assess the model's capacity to identify the ethical core without relying on the options provided, 31 of the 36 questions were modified by masking the correct choice. Accuracy was analyzed using independent sample t-tests and Mann Whitney U test, and consistency was assessed using Krippendorff's alpha.

resultsOverall, the accuracy and consistency of ChatGPT improved with each version. Version 4.5 achieved near-perfect scores and high reliability in both languages, while version 3.5 showed limited performance, particularly in the Korean test. Performance gaps between languages decreased with model upgrades but remained statistically significant in version 4.5 for some questions. In the masked-answer condition, all versions showed notable drops in accuracy and consistency, with version 4.5 still outperforming earlier versions. However, the performance remained below 50%, indicating limitations in the model's autonomous ethical reasoning.

conclusionsChatGPT demonstrated substantial improvements in medical ethics MCQ performance across versions, particularly in terms of consistency and accuracy. However, performance disparities between languages and reduced accuracy under masked answer conditions highlight the ongoing limitations of non-English ethical reasoning and context recognition. These findings emphasize the need for further research on language-sensitive fine-tuning and the evaluation of LLMs in specialized ethical domains. The findings suggest that advanced LLMs may serve as valuable supplementary tools in medical education and clinical ethics training. At the same time, the observed language disparities call for context-sensitive adaptations to prevent inequities in practice.

Indexed as

Educational MeasurementEthics, MedicalLanguageGenerative Artificial IntelligenceHumansLarge Language ModelsRepublic of KoreaSurveys and QuestionnairesArtificial intelligenceChatGPTLarge language modelsMedical educationMedical ethicsMultiple-choice questions

Identifiers

PMID41366422
PMCPMC12687472

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.