ArticleJournal of patient-reported outcomes2025
Can machine translation match human expertise? Quantifying the performance of large language models in the translation of patient-reported outcome measures (PROMs).
Article in Journal of patient-reported outcomes, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
3 citing papers in PubMed.
- Article
- Article
- Use of AI within COA linguistic validation and eCOA migration processes: analysis and good practice recommendations.Journal of patient-reported outcomes · 2026Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe rise in artificial intelligence tools, especially those competent at language interpretation and translation, enables opportunities to enhance patient-centered care. One might be the ability to rapidly and inexpensively create accurate translations of English language patient-reported outcome measures (PROMs) to facilitate global uptake. Currently, it is unclear if machine translation (MT) tools can produce sufficient translation quality for this purpose. METHODOLOGY: We used Generative Pretrained Transformer (GPT)-4, GPT-3.5, and Google Translate to translate the English versions of selected scales from the Breast-Q and Face-Q, two widely used PROMs assessing outcomes following breast and face reconstructive surgery, respectively. We used MT to forward and back translate the scales from English into Arabic, Vietnamese, Italian, Hungarian, Malay, and Dutch. We compared translation quality using the Metrics for Evaluation of Translation with Explicit Ordering (METEOR). We compared the scores between different translation versions using the Kruskal-Wallis test or analysis of variance as appropriate.
resultsIn forward translations, the METEOR scores significantly varied depending on target languages for all MT tools (p < 0.001), with GPT-4 having the highest scores in most languages. We detected significantly different scores among translators for all languages (p < .05), except for Italian (p = 0.59). In backward translations, MTs (GPT-4: 0.81 ± 0.10; GPT-3.5: 0.78 ± 0.12; Google Translate: 0.80 ± 0.06) received higher or compatible scores to human translations (0.76 ± 0.11) for all languages. The differences in backward translation scores by different forward translators were significant for all languages (p < 0.01; except for Italian, p = 0.2). The scores between different languages were also significantly different for all translators (p < 0.001).
conclusionsOur findings suggest that large language models provide high-quality PROM translations to support human translations to reduce costs. However, substituting human translation with MT is not advisable at the current stage.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.