Evidence map›Paper›PMID 40711496›Full record

ArticleJournal of patient-reported outcomes2025

Can machine translation match human expertise? Quantifying the performance of large language models in the translation of patient-reported outcome measures (PROMs).

Sheng-Chieh Lu, Cai Xu, Manraj Kaur, Maria Orlando Edelen, Andrea Pusic, Chris Gibbons

Abstract read
In one paragraph

Article in Journal of patient-reported outcomes, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Sheng-Chieh LuDepartment of Symptom Research, The University of Texas MD Anderson Cancer Center, 6565 MD Anderson Blvd, Houston, Texas, USA. slu4@mdanderson.org.ORCID http://orcid.org/0000-0002-6685-1524
Cai XuDepartment of Bioinformatics & Border Biomedical Research Center, The University of Texas at El Paso, El Paso, Texas, USA.
Manraj KaurDepartment of Surgery, Brigham and Women's Hospital, Harvard Medical School, Boston, Massachusetts, USA.
Maria Orlando EdelenDepartment of Surgery, Brigham and Women's Hospital, Harvard Medical School, Boston, Massachusetts, USA.
Andrea PusicDepartment of Surgery, Brigham and Women's Hospital, Harvard Medical School, Boston, Massachusetts, USA.
Chris GibbonsDepartment of Symptom Research, The University of Texas MD Anderson Cancer Center, 6565 MD Anderson Blvd, Houston, Texas, USA.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundThe rise in artificial intelligence tools, especially those competent at language interpretation and translation, enables opportunities to enhance patient-centered care. One might be the ability to rapidly and inexpensively create accurate translations of English language patient-reported outcome measures (PROMs) to facilitate global uptake. Currently, it is unclear if machine translation (MT) tools can produce sufficient translation quality for this purpose. METHODOLOGY: We used Generative Pretrained Transformer (GPT)-4, GPT-3.5, and Google Translate to translate the English versions of selected scales from the Breast-Q and Face-Q, two widely used PROMs assessing outcomes following breast and face reconstructive surgery, respectively. We used MT to forward and back translate the scales from English into Arabic, Vietnamese, Italian, Hungarian, Malay, and Dutch. We compared translation quality using the Metrics for Evaluation of Translation with Explicit Ordering (METEOR). We compared the scores between different translation versions using the Kruskal-Wallis test or analysis of variance as appropriate.

resultsIn forward translations, the METEOR scores significantly varied depending on target languages for all MT tools (p < 0.001), with GPT-4 having the highest scores in most languages. We detected significantly different scores among translators for all languages (p < .05), except for Italian (p = 0.59). In backward translations, MTs (GPT-4: 0.81 ± 0.10; GPT-3.5: 0.78 ± 0.12; Google Translate: 0.80 ± 0.06) received higher or compatible scores to human translations (0.76 ± 0.11) for all languages. The differences in backward translation scores by different forward translators were significant for all languages (p < 0.01; except for Italian, p = 0.2). The scores between different languages were also significantly different for all translators (p < 0.001).

conclusionsOur findings suggest that large language models provide high-quality PROM translations to support human translations to reduce costs. However, substituting human translation with MT is not advisable at the current stage.

Indexed as

Artificial IntelligenceLanguagePatient Reported Outcome MeasuresTranslatingTranslationsFemaleHumansLarge Language ModelsBreast-QFace-QLarge language modelsMachine translationPatient-reported outcome measure

Identifiers

PMID40711496
PMCPMC12297096

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.