ArticleNPJ digital medicine2025
Framework for bias evaluation in large language models in healthcare settings.
Article in NPJ digital medicine, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 25 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
25 citing papers in PubMed.
- Human medical documentation significantly outperforms ChatGPT-4o in critical clinical dimensions: A blinded comparative assessment in paediatric orthopaedics.Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA · 2026Trial
- The Rise of Small Language Models in Healthcare: A Comprehensive Survey.Computer science review · 2026Article
- Can students identify AI? - A cross-sectional quantitative study about AI recognition in tablet-based MCQ assessment among fifth-year undergraduate medical students at Saarland University, Germany.BMC medical education · 2026Article
- Sex and gender bias in large language models: an old problem at a new scale.BMJ health & care informatics · 2026Article
- From scoring to stress testing: strengthening safety validation of large language model answers in otolaryngology.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026Article
- Multiagent Large Language Model Framework for Psychotherapy Fidelity Assessment in Motivational Interviewing and Cognitive Behavioral Therapy Training: Cross-Sectional, Simulation-Based Evaluation Study.JMIR medical education · 2026Article
- Article
- "MELMA" in otolaryngology: Medical evaluation of large language model answers. Clinician-rated scoring (MELMA-Q) and web-based auditing (MELMA-W) novel tools for AI assessment.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026Article
- Ethical Considerations in Personal Health Large Language Models.Journal of medical Internet research · 2026Article
- Participatory Digital Twins for Chronic Care: From Predictive Models to Shared Sensemaking.Journal of participatory medicine · 2026Article
- Foundation models in healthcare: a comprehensive review from technical advances to clinical translation.Journal of translational medicine · 2026Review
- From promise to proof: making multimodal LLMs in reproductive psychiatry audit-ready, privacy-verifiable, and globally deployable.Archives of women's mental health · 2026Article
- Generative artificial intelligence-driven chatbots and medical misinformation: an accuracy, referencing and readability audit.BMJ open · 2026Article
- The Alberta Quality Assessment Tool: Risk of Bias (AQAT:RoB) for the Evaluation of Medical Large Language Model Question-Answer Studies: Development and Pilot Validation.Journal of medical Internet research · 2026Article
- Ethical considerations for clinical adoption of ambient digital scribe technology.Journal of the American Medical Informatics Association : JAMIA · 2026Article
- Is Artificial Intelligence Ready for Emergency Department Triage? A Retrospective Evaluation of Multiple Large Language Models in 39,375 Patients at a University Emergency Department.Journal of clinical medicine · 2026Article
- Implementing generative artificial intelligence in precision oncology: safety, governance, and significance.Journal of hematology & oncology · 2026Review
- A multidimensional hierarchical framework for sources of bias in real-world healthcare evidence: a scoping review.Journal of biomedical informatics · 2026Article
- Bridging the mentorship divide: how large language models could reshape medical workforce equity.NPJ digital medicine · 2026Article
- Same child, different risk: demographic bias in childhood obesity attribution by large language models.Frontiers in public health · 2026Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
9 authors.
Funding
Abstract
A critical gap in the adoption of large language models for AI-assisted clinical decisions is the lack of a standardized audit framework to evaluate models for accuracy and bias. Our framework introduces a five-step framework that guides practitioners through stakeholder engagement, model calibration to specific patient populations, and rigorous testing through clinically relevant scenarios. We provide open-access tools for stakeholder engagement and an example of an audit. As the regulation of models becomes more critical, we believe adoption of an audit framework that tests model outputs, rather than regulating specific hyperparameters or inputs, will encourage the responsible use of AI in clinical settings.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.