Evidence map›Paper›PMID 41334074›Full record

ArticleMayo Clinic proceedings. Digital health2025

Identifying Bias at Scale in Clinical Notes Using Large Language Models.

Donald U Apakama, Kim-Anh-Nhi Nguyen, Daphnee Hyppolite, Shelly Soffer, Aya Mudrik, Emilia Ling, Akini Moses, Ivanka Temnycky, Allison Glasser, Rebecca Anderson and 17 more

Abstract read
In one paragraph

Article in Mayo Clinic proceedings. Digital health, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

27 authors.

Donald U ApakamaThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Kim-Anh-Nhi NguyenInstitute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, NY.
Daphnee HyppoliteOffice of the Chief Medical Officer, Mount Sinai Health System, New York, NY.
Shelly SofferInstitute of Hematology, Davidoff Cancer Center, Rabin Medical Center Petah-Tikva, Israel.
Aya MudrikBen-Gurion University of the Negev, Be'er Sheva, Israel.
Emilia LingDepartment of Emergency Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Akini MosesDepartment of Emergency Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Ivanka TemnyckyDepartment of Emergency Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Allison GlasserOffice of the Chief Medical Officer, Mount Sinai Health System, New York, NY.
Rebecca AndersonOffice of the Chief Medical Officer, Mount Sinai Health System, New York, NY.
Prathamesh ParchureInstitute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, NY.
Evajoyce WoullardInstitute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, NY.
Masoud EdalatiInstitute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, NY.
Lili ChanThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Clair KronkInstitute for Health Equity Research, Icahn School of Medicine at Mount Sinai, New York, NY.
Robert FreemanThe Charles Bronfman Department of Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Arash KiaInstitute for Healthcare Delivery Science, Icahn School of Medicine at Mount Sinai, New York, NY.
Prem TimsinaThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Matthew A LevinThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Rohan KheraDepartment of Medicine, Yale University School of Medicine New Haven, CT.
Patricia KovatchDepartment of Genetics and Genomic Sciences, Icahn School of Medicine at Mount Sinai, New York, NY.
Alexander W CharneyThe Charles Bronfman Department of Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Brendan G CarrDepartment of Emergency Medicine, Icahn School of Medicine at Mount Sinai, New York, NY.
Lynne D RichardsonInstitute for Health Equity Research, Icahn School of Medicine at Mount Sinai, New York, NY.
Carol R HorowitzInstitute for Health Equity Research, Icahn School of Medicine at Mount Sinai, New York, NY.
Eyal KlangThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.
Girish N NadkarniThe Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai, New York, NY.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: To evaluate whether generative pretrained transformer (GPT)-4 can detect and revise biased language in emergency department (ED) notes, against human-adjudicated gold-standard labels, and to identify modifiable factors associated with biased documentation. Patients and Methods: We randomly sampled 50,000 ED medical and nursing notes from the Mount Sinai Health System (January 1, 2023, to December 31, 2023). We also randomly sampled 500 discharge notes from the Medical Information Mart for Intensive Care IV database. The GPT-4 flagged 4 types of bias: discrediting, stigmatizing/labeling, judgmental, and stereotyping. Two human reviewers verified model detections. We used multivariable logistic regression to examine associations between bias and health care utilization, presenting problems (eg, substance use), shift timing, and provider type. We then asked physicians to rate GPT-4's proposed language revisions on a 10-point scale. Results: The GPT-4 showed 97.6% sensitivity and 85.7% specificity compared with the human review. Biased language appeared in 6.5% (3229 of 50,000) of Mount Sinai notes and 7.4% (37 of 500) of Medical Information Mart for Intensive Care IV notes. In adjusted models, frequent health care utilization (adjusted odds ratio [aOR], 2.85; 95% CI, 1.95-4.17), substance use presentations (aOR, 3.09; 95% CI, 2.51-3.80), and overnight shifts (aOR, 1.37; 95% CI, 1.23-1.52) showed elevated odds of biased documentation. Physicians were more likely to include bias than nurses (aOR, 2.26; 95% CI, 2.07-2.46); GPT-4's recommended revisions received mean physician ratings above 9 of 10. Conclusion: The study showed that GPT-4 accurately detects biased language in clinical notes, identifies modifiable contributors to that bias, and delivers physician-endorsed revisions. This approach may help mitigate documentation bias and reduce disparities in care.

Identifiers

PMID41334074
PMCPMC12666851

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.