Evidence map›Paper›PMID 42144967›Full record

SynthesisJournal of medical Internet research2026

Concerns of Using Large Language Models in Health Care Research and Practice: Umbrella Review.

Feyza Yarar, Pauline Addis, Megan Fairweather, Dawn Craig, Hannah O'Keefe

Abstract readSystematic Review
In one paragraph

Synthesis in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Feyza YararPopulation Health Sciences Institute, Faculty of Medical Sciences, Newcastle University, Framlington Place, Newcastle-Upon-Tyne, England, NE2 4HH, United Kingdom, 44 7826034122.ORCID http://orcid.org/0000-0002-4355-3786
Pauline Addis *Population Health Sciences Institute, Faculty of Medical Sciences, Newcastle University, Framlington Place, Newcastle-Upon-Tyne, England, NE2 4HH, United Kingdom, 44 7826034122.ORCID http://orcid.org/0000-0003-2392-070X
Megan Fairweather *Population Health Sciences Institute, Faculty of Medical Sciences, Newcastle University, Framlington Place, Newcastle-Upon-Tyne, England, NE2 4HH, United Kingdom, 44 7826034122.ORCID http://orcid.org/0009-0003-7147-5054
Dawn CraigPopulation Health Sciences Institute, Faculty of Medical Sciences, Newcastle University, Framlington Place, Newcastle-Upon-Tyne, England, NE2 4HH, United Kingdom, 44 7826034122.ORCID http://orcid.org/0000-0002-5808-0096
Hannah O'KeefePopulation Health Sciences Institute, Faculty of Medical Sciences, Newcastle University, Framlington Place, Newcastle-Upon-Tyne, England, NE2 4HH, United Kingdom, 44 7826034122.ORCID http://orcid.org/0000-0002-0107-711X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Large language models (LLMs), such as ChatGPT (OpenAI), are rapidly evolving, and their applications in health care are increasing. There is a growing demand for automation of routine tasks and a drive to use LLMs or similar to support research. Objective: This umbrella review examines concerns of health care professionals and researchers related to the use of LLMs in health care research and practice. We aimed to identify common issues raised and the implications for patient care, policy, and practice. Methods: A protocol was registered on PROSPERO (CRD420250640997). Searches were conducted in 7 databases (Ovid MEDLINE, Ovid Embase, Scopus, Web of Science, JBI Database of Systematic Reviews and Implementation Reports, Cochrane Database of Systematic Reviews, and Epistemonikos) in February 2025 and updated in February 2026. Screening was conducted in 2 stages, with independent screening by 2 reviewers. Studies published in the English language after January 2017 with at least one outcome expressing concerns of LLM or generative artificial intelligence use in health care research were included. The included studies were quality appraised for risk of bias and certainty of the evidence using AMSTAR-2 (A Measurement Tool to Assess Systematic Reviews) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively. Data was extracted using a piloted form and narratively synthesized following SWiM guidelines and the PRIOR (Preferred Reporting Items for Overviews of Reviews) checklist. Results: The search retrieved 448 systematic reviews, of which 42 met the inclusion criteria. Further, 12 distinct populations were identified, including researchers and clinicians in various medical specialties. The included reviews were assessed to be of very poor quality, and the level of overlap between primary studies could not be determined. Additionally, 15 reviews focused on ChatGPT, a further 15 on two or more LLMs, and 12 on generic artificial intelligence. Thus, 3 main themes emerged from the narrative synthesis. In order of most to least frequently discussed: (1) technical capability; (2) ethical, legal, and societal; and (3) costs. Conclusions: To our knowledge, this is the first umbrella review to address the concerns of LLMs in health care research and practice. Thematic analyses provided insight into the complexity of different perspectives, and by using a whole population approach, it demonstrates common narratives. However, the poor quality of the included studies and potential overlap of results are substantial limitations. Data quality is at the heart of these concerns, and combative action must ensure health care professionals and researchers have the resources required to overcome these apprehensions. Ethical, legal, and societal implications of artificial intelligence use were also commonly raised. As technology accelerates and demands on health care increase, we must adapt and embrace change with equity, diversity, inclusion, and safety at the core.

Indexed as

Health Services ResearchLarge Language ModelsGenerative Artificial IntelligenceHumansartificial intelligenceconcernshealth and social carelife sciencesumbrella review

Identifiers

PMID42144967
PMCPMC13181733

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.