Evidence map›Paper›PMID 42536971›Full record

SynthesisJournal of medical Internet research2026

Impact of Large Language Model-Based AI Tools on Physician-Patient Communication: Systematic Review and Meta-Analysis.

Sven Richter, Clara Helene Buszello, Markus Prem, Sophia Willkommen, Elida Hasani, Ortrud Uckermann, Tareq A Juratli, Ilker Y Eyüpoglu, Witold H Polanski

Abstract readSystematic ReviewMeta-Analysis
In one paragraph

Synthesis in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Sven RichterDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0000-0003-1648-5754
Clara Helene BuszelloDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0009-0008-8230-1064
Markus PremDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0009-0005-9888-1380
Sophia WillkommenDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0009-0003-8993-7720
Elida HasaniDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0009-0002-0560-6741
Ortrud UckermannDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0000-0001-8196-604X
Tareq A JuratliDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0000-0003-2236-6719
Ilker Y EyüpogluDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0000-0002-8185-7764
Witold H PolanskiDepartment of Neurosurgery, Medical Faculty and University Hospital Carl Gustav Carus, Technische Universität Dresden, Fetscherstrasse 74, Dresden, 01307, Germany, 49 3514582883.ORCID http://orcid.org/0000-0002-9214-150X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Recent advances in large language models (LLMs) such as GPT-3/4 have spurred the development of artificial intelligence (AI) chatbots and advisory tools in medicine. These systems are posited to assist or augment physician-patient communication, potentially improving empathy, clarity, and responsiveness. However, their actual impact on communication outcomes remains uncertain. Objective: This study aimed to systematically review and meta-analyze peer-reviewed studies (2020-2025) evaluating how LLM-based interventions affect physician-patient communication, including empathy, clarity, trust, and patient understanding. Methods: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, we searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for studies published from 2020 to 2025 examining LLM or chatbot applications in clinical communication contexts. Eligible designs included randomized, observational, cross-sectional, and qualitative studies. Two reviewers (WHP and SR) independently screened titles or abstracts, assessed full texts, and extracted data on study design, population, LLM type, communication measures, and outcomes. We conducted a qualitative synthesis and random-effects meta-analysis, reporting pooled standardized mean differences or odds ratios with 95% CIs. Results: From 312 records, 10 studies were included, all quantitative and predominantly cross-sectional. Populations ranged from patients with chronic conditions to health care professionals and laypersons. Outcomes assessed included empathy (8 studies), clarity or information quality (6 studies), satisfaction or usefulness (4 studies), and trust perceptions (2 studies). In 6 direct comparisons of AI- versus physician-generated responses, LLMs were rated significantly higher in empathy in 5 studies. One large study found that chatbot replies were judged empathetic in 45.1% of cases versus 4.6% for physician replies (odds ratio approximately 9.8, P<.001). Similarly, ChatGPT-4 answers scored higher in empathy on a 5-point scale than human-written responses (mean 4.18 vs 2.70, P<.001). One neurology study showed higher empathy scores (Consultation and Relational Empathy Scale +1.38, P<.01) for ChatGPT answers. Only 1 study found no significant empathy difference. LLM content was also longer and more information-rich, improving patient-perceived clarity and understanding. On the other hand, GPT-4 simplified pathology reports, increasing patient comprehension scores (7.98 vs 5.23/10, P<.001) and reducing consultation time by 70%. However, AI replies were sometimes less concise or less readable for low-literacy patients. In pooled analyses (k=4 studies; total evaluations N=2604), LLM assistance showed a large positive effect on empathy (standardized mean difference 1.02, 95% CI 0.44-1.60; random-effects model). Patient satisfaction results were mixed. No study directly assessed long-term trust. Conclusions: Current evidence suggests that LLM-based chatbots can enhance physician-patient communication by producing more empathetic, detailed, and understandable responses. These improvements may positively influence patient experience and engagement. However, LLMs may also generate overly lengthy or occasionally inaccurate advice, emphasizing the need for physician oversight. While meta-analytic findings are promising, robust randomized controlled trials, real-world and longitudinal studies are needed to confirm benefits, assess trust outcomes, and define optimal clinical integration strategies.

Indexed as

Artificial IntelligenceCommunicationPhysician-Patient RelationsEmpathyGenerative Artificial IntelligenceHumansLarge Language Modelsartificial intelligence in health careChatGPTempathylarge language modelsLLMsphysician-patient communication

Identifiers

PMID42536971
PMCPMC13427064

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.