Evidence map›Paper›PMID 40624264›Full record

ArticleNPJ digital medicine2025

Framework for bias evaluation in large language models in healthcare settings.

Tara Templin, Sophia Fort, Prasanna Padmanabham, Pratyush Seshadri, Ram Rimal, Junier Oliva, Kristin Hassmiller Lich, Sean Sylvia, Nasa Sinnott-Armstrong

Abstract read
In one paragraph

Article in NPJ digital medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 25 papers.

0numbers the graph read from it
0cells of the map it votes in
25citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

25 citing papers in PubMed.

  1. Trial
  2. Article
  3. Article
  4. Article
  5. From scoring to stress testing: strengthening safety validation of large language model answers in otolaryngology.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Article
  6. Article
  7. Article
  8. "MELMA" in otolaryngology: Medical evaluation of large language model answers. Clinician-rated scoring (MELMA-Q) and web-based auditing (MELMA-W) novel tools for AI assessment.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Article
  9. Ethical Considerations in Personal Health Large Language Models.Journal of medical Internet research · 2026
    Article
  10. Article
  11. Review
  12. Article
  13. Article
  14. Article
  15. Ethical considerations for clinical adoption of ambient digital scribe technology.Journal of the American Medical Informatics Association : JAMIA · 2026
    Article
  16. Article
  17. Review
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Tara TemplinDepartment of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA. ttemplin@unc.edu.
Sophia Fort *Department of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Prasanna Padmanabham *Herbold Computational Biology Program, Fred Hutchinson Cancer Center, Seattle, WA, USA.
Pratyush Seshadri *Department of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Ram RimalDepartment of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Junier OlivaDepartment of Computer Science, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Kristin Hassmiller LichDepartment of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Sean SylviaDepartment of Health Policy and Management, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA.
Nasa Sinnott-ArmstrongHerbold Computational Biology Program, Fred Hutchinson Cancer Center, Seattle, WA, USA. nasa@fredhutch.org.

Funding

Understanding the Drivers of Antibiotic use in the Treatment of Childhood Diarrhea and Relationship to Antibiotic Resistance in ChinaK01AI159233 · NIAID · UNIV OF NORTH CAROLINA CHAPEL HILL · PI SYLVIA, SEAN Y · 2022 to 2025
$517k
National Science Foundation 2026498NIAID NIH HHS K01 AI159233NIAID NIH HHS K01AI159233Public Health Sciences Klorfine Pilot Award N/A
6 · The paper itself

Abstract

A critical gap in the adoption of large language models for AI-assisted clinical decisions is the lack of a standardized audit framework to evaluate models for accuracy and bias. Our framework introduces a five-step framework that guides practitioners through stakeholder engagement, model calibration to specific patient populations, and rigorous testing through clinically relevant scenarios. We provide open-access tools for stakeholder engagement and an example of an audit. As the regulation of models becomes more critical, we believe adoption of an audit framework that tests model outputs, rather than regulating specific hyperparameters or inputs, will encourage the responsible use of AI in clinical settings.

Identifiers

PMID40624264
PMCPMC12234702

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.