Evidence map›Paper›PMID 42712338›Full record

ArticleBMJ digital health & AI2026

Benchmarking large language models and clinicians using locally generated primary healthcare vignettes in Kenya.

Paul Mwaniki, Wilkister Musau, Lynda Isaaka, Conrad Wanyama, Vaishnavi Menon, Alastair K Denniston, Xiaoxuan Liu, Mira Emmanuel-Fabula, Gwydion Williams, Bilal Akhter Mateen and 1 more

Abstract read
In one paragraph

Article in BMJ digital health & AI, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 2 papers.

0numbers the graph read from it
0cells of the map it votes in
2citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

2 citing papers in PubMed.

  1. Trial
  2. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Paul MwanikiKEMRI-Wellcome Trust Research Programme, Nairobi, Kenya.
Wilkister MusauPATH, Nairobi, Kenya.
Lynda IsaakaKEMRI-Wellcome Trust Research Programme, Nairobi, Kenya.
Conrad WanyamaKEMRI-Wellcome Trust Research Programme, Nairobi, Kenya.
Vaishnavi MenonUniversity of Birmingham, Birmingham, UK.
Alastair K DennistonUniversity of Birmingham, Birmingham, UK.
Xiaoxuan LiuUniversity of Birmingham, Birmingham, UK.
Mira Emmanuel-FabulaPATH, Geneva, Switzerland.
Gwydion WilliamsPATH, London, UK.
Bilal Akhter MateenUniversity of Birmingham, Birmingham, UK.ORCID 0000-0003-4423-6472
Ambrose AgweyuKEMRI-Wellcome Trust Research Programme, Nairobi, Kenya.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Objective: To determine the ability of large language models (LLMs) to answer open-ended healthcare questions (rather than multiple-choice questions) accurately in a low-resource setting. Methods and analysis: We benchmarked five LLMs (GPT-4.1, Gemini-2.5-Flash, DeepSeek-R1, MedGemma and o3) against Kenyan clinicians, using a randomly subsampled dataset of 507 vignettes (from a larger pool of 5107 clinical scenarios) spanning 12 nursing competency categories. Blinded physician panels rated responses using a 5-point Likert scale on an 11-domain rubric covering accuracy, safety, contextual appropriateness, and communication. We summarised mean scores and used Bayesian ordinal logistic regression to estimate probabilities of high-quality ratings (≥4) and to perform pairwise comparisons between LLMs and clinicians. Results: Clinician mean ratings were lower than those for LLMs in 9/11 domains: 2.86 vs 4.25-4.72 (guideline alignment), 2.76 vs 4.25-4.73 (expert knowledge), 2.96 vs 4.30-4.73 (logical coherence) and 2.58 vs 4.16-4.68 (low omission of critical information). On safety-related domains, LLMs received higher ratings: minimal extent of possible harm 3.16 vs 4.29-4.68; low likelihood of harm 3.68 vs 4.54-4.81. Performance was similar for low inclusion of irrelevant content (4.28 vs 4.25-4.35) and for avoidance of demographic bias (4.86 vs 4.91-4.94). In Bayesian models, LLMs had >90% probability of ratings ≥4 in most domains, whereas clinicians exceeded 90% only for contextual relevance and demographic/socioeconomic bias. Pairwise contrasts showed broadly overlapping credible intervals among LLMs, with o3 leading numerically most domains except contextual relevance, demographic/socio-economic bias and relevance to the question. Generating all LLM responses cost US$3.86-US$8.68 per model (US$0.008-US$0.017 per vignette), compared with US$3.35 per clinician-generated vignette. Conclusions: In controlled vignette-based tasks, LLMs produced responses that were more accurate, safer and more structured than clinicians, suggesting LLMs may have potential as supplementary knowledge and safety support tools. Findings support further evaluation in real patient encounters to determine effectiveness, safety and integration into clinical workflows, particularly in resource-constrained health systems.

Indexed as

Artificial intelligencePrimary Health CareTelemedicine

Identifiers

PMID42712338
PMCPMC13528382

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.