SynthesisFrontiers in digital health2026
Large language models in healthcare: applications, evaluation frameworks, and governance pathways - a scoping review and multidimensional framework.
Synthesis in Frontiers in digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
2 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Large language models (LLMs) are increasingly evaluated for healthcare applications spanning clinical documentation, decision support, patient communication, research assistance, and operational workflows. Despite rapid adoption interest, evidence remains heterogeneous and standardised approaches for evaluation and governance are not yet consistently applied. Objective: To map healthcare applications of LLMs, synthesise reported outcomes and risks, and propose a multidimensional evaluation and governance framework oriented toward digital-health implementation. Methods: A scoping review was conducted following the methodological guidance of Arksey and O'Malley, Levac et al., and the Joanna Briggs Institute, with reporting aligned to the PRISMA-ScR checklist. The protocol was prospectively registered on the Open Science Framework https://doi.org/10.17605/OSF.IO/SWP78). Searches were performed in PubMed/MEDLINE, Embase, Scopus, Web of Science, IEEE Xplore, ACM Digital Library, and the ACL Anthology, supplemented by medRxiv/bioRxiv preprint searches and grey-literature scanning of WHO, FDA, and EU AI-Office guidance, covering 1 January 2019-30 April 2026. Title/abstract and full-text screening were carried out in duplicate; inter-rater agreement was Cohen's κ = 0.78. Data were extracted in duplicate using a piloted form. Synthesis followed Braun and Clarke's six-phase reflexive thematic analysis. The complete list of 78 included studies is provided in Supplementary File S4. Results: Seventy-eight studies met inclusion criteria; 68% were published between 2023 and 2025. Five application domains were identified: (1) clinical documentation and summarisation; (2) clinical decision support and reasoning assistance; (3) patient communication and health-literacy support; (4) biomedical research and knowledge synthesis; and (5) administrative and operational use cases. The evidence base was dominated by benchmark and simulated-workflow studies (89%), with limited prospective workflow-embedded evaluations (11%). Reported benefits concentrated on documentation efficiency, text quality, and knowledge synthesis; safety-relevant risks included hallucinated content, omission of clinically critical information, demographic bias, privacy vulnerabilities, limited explainability, and automation bias. Studies were geographically concentrated in North America and East Asia, with limited representation from Sub-Saharan Africa, South Asia, and Latin America. Conclusions: Current evidence supports cautious deployment of LLMs in selected healthcare tasks under structured oversight. Translational progress depends on prospective evaluation, standardised reporting, equity-focused audits, and lifecycle governance with continuous monitoring. The proposed five-dimensional framework (technical performance, clinical validity, equity, workflow integration, governance) coupled with a three-tier risk model is intended to support researchers and healthcare organisations in assessing readiness and implementing LLM-enabled tools responsibly.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.