Evidence map›Paper›PMID 39659993›Full record

ArticleJAMIA open2024

Decoding disparities: evaluating automatic speech recognition system performance in transcribing Black and White patient verbal communication with nurses in home healthcare.

Maryam Zolnoori, Sasha Vergez, Zidu Xu, Elyas Esmaeili, Ali Zolnour, Krystal Anne Briggs, Jihye Kim Scroggins, Seyed Farid Hosseini Ebrahimabad, James M Noble, Maxim Topaz and 6 more

Abstract read
In one paragraph

Article in JAMIA open, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 26 papers.

0numbers the graph read from it
0cells of the map it votes in
26citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

26 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Review
  5. Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. Review
  11. Article
  12. Article
  13. Review
  14. Article
  15. Ethical considerations for clinical adoption of ambient digital scribe technology.Journal of the American Medical Informatics Association : JAMIA · 2026
    Article
  16. Article
  17. Article
  18. Review
  19. Article
  20. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

16 authors.

Maryam ZolnooriColumbia University Irving Medical Center, New York, NY 10032, United States.
Sasha VergezCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.
Zidu XuSchool of Nursing, Columbia University, New York, NY 10032, United States.ORCID https://orcid.org/0000-0002-6122-8426
Elyas EsmaeiliColumbia University Irving Medical Center, New York, NY 10032, United States.
Ali ZolnourColumbia University Irving Medical Center, New York, NY 10032, United States.
Krystal Anne BriggsDepartment of Computer Science, Columbia University, New York, NY 10027, United States.
Jihye Kim ScrogginsSchool of Nursing, Columbia University, New York, NY 10032, United States.
Seyed Farid Hosseini EbrahimabadDepartment of Automatic Control and Computer Science, Politehnica University of Bucharest, Bucharest RO-060042, Romania.
James M NobleColumbia University Irving Medical Center, New York, NY 10032, United States.
Maxim TopazColumbia University Irving Medical Center, New York, NY 10032, United States.ORCID https://orcid.org/0000-0002-2358-9837
Suzanne BakkenSchool of Nursing, Columbia University, New York, NY 10032, United States.ORCID https://orcid.org/0000-0001-6202-6001
Kathryn H BowlesCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.
Ian SpensCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.
Nicole OnoratoCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.
Sridevi SridharanCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.
Margaret V McDonaldCenter for Home Care Policy & Research, VNS Health, New York, NY 10017, United States.

Funding

Research Education CoreP30AG066462 · NIA · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI PHILIP L DE JAGER · 2020 to 2026
$30.1M
Technology Identification and Training CoreP30AG073105 · NIA · UNIVERSITY OF PENNSYLVANIA · PI DEMIRIS, GEORGE, KARLAWISH, JASON H · 2021 to 2025
$21.2M
Development of a Screening Algorithm for Timely Identification of Patients with Mild Cognitive Impairment and Early Dementia in Home HealthcareK99AG076808 · NIA · COLUMBIA UNIVERSITY HEALTH SCIENCES · PI ZOLNOORI, MARYAM · 2023 to 2024
$230k
NIA NIH HHS K99 AG076808NIA NIH HHS P30 AG066462NIA NIH HHS P30 AG073105
6 · The paper itself

Abstract

Objectives: As artificial intelligence evolves, integrating speech processing into home healthcare (HHC) workflows is increasingly feasible. Audio-recorded communications enhance risk identification models, with automatic speech recognition (ASR) systems as a key component. This study evaluates the transcription accuracy and equity of 4 ASR systems-Amazon Web Services (AWS) General, AWS Medical, Whisper, and Wave2Vec-in transcribing patient-nurse communication in US HHC, focusing on their ability in accurate transcription of speech from Black and White English-speaking patients. Materials and Methods: We analyzed audio recordings of patient-nurse encounters from 35 patients (16 Black and 19 White) in a New York City-based HHC service. Overall, 860 utterances were available for study, including 475 drawn from Black patients and 385 from White patients. Automatic speech recognition performance was measured using word error rate (WER), benchmarked against a manual gold standard. Disparities were assessed by comparing ASR performance across racial groups using the linguistic inquiry and word count (LIWC) tool, focusing on 10 linguistic dimensions, as well as specific speech elements including repetition, filler words, and proper nouns (medical and nonmedical terms). Results: The average age of participants was 67.8 years (SD = 14.4). Communication lasted an average of 15 minutes (range: 11-21 minutes) with a median of 1186 words per patient. Of 860 total utterances, 475 were from Black patients and 385 from White patients. Amazon Web Services General had the highest accuracy, with a median WER of 39%. However, all systems showed reduced accuracy for Black patients, with significant discrepancies in LIWC dimensions such as "Affect," "Social," and "Drives." Amazon Web Services Medical performed best for medical terms, though all systems have difficulties with filler words, repetition, and nonmedical terms, with AWS General showing the lowest error rates at 65%, 64%, and 53%, respectively. Discussion: While AWS systems demonstrated superior accuracy, significant disparities by race highlight the need for more diverse training datasets and improved dialect sensitivity. Addressing these disparities is critical for ensuring equitable ASR performance in HHC settings and enhancing risk prediction models through audio-recorded communication.

Indexed as

automatic speech recognition (ASR)health disparitieshome healthcarelinguistic inquiry and word count (LIWC)speech to textword error rate (WER)

Identifiers

PMID39659993
PMCPMC11631515

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.