Evidence map›Paper›PMID 41927104›Full record

ArticleBMJ health & care informatics2026

Acceptable accuracy for medical AI: a survey of physicians and the general population in Sweden.

Rasmus Arvidsson, Jonathan Widén, Lina Al-Naasan, Ronny Kent K Gunnarsson, Peter Nymberg, Charlotte R Blease, Anna Moberg, Pär-Daniel Sundvall, Carl Wikberg, David Sundemo

Abstract read
In one paragraph

Article in BMJ health & care informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Rasmus ArvidssonGeneral Practice/Family Medicine, School of Public Health and Community Medicine, Institute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden rasmus.arvidsson@gu.se.ORCID http://orcid.org/0009-0006-0387-3108
Jonathan WidénResearch, Education, Development and Innovation, Primary Health Care, Västra Götalandsregionen, Gothenburg, Västra Götaland County, Sweden.ORCID http://orcid.org/0009-0005-8584-0862
Lina Al-NaasanInstitute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden.
Ronny Kent K GunnarssonGeneral Practice/Family Medicine, School of Public Health and Community Medicine, Institute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden.ORCID http://orcid.org/0000-0001-9183-3072
Peter NymbergCenter for Primary Health Care Research, Department of Clinical Sciences Malmö, Lund University, Malmö, Sweden.ORCID http://orcid.org/0000-0001-9901-0580
Charlotte R BleaseWomen's and Children's Health, Uppsala Universitet, Uppsala, Sweden.ORCID http://orcid.org/0000-0002-0205-1165
Anna MobergDepartment of Health, Medicine and Caring Sciences, Linköping University, Linköping, Sweden.ORCID http://orcid.org/0000-0001-5431-8469
Pär-Daniel SundvallGeneral Practice/Family Medicine, School of Public Health and Community Medicine, Institute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden.ORCID http://orcid.org/0000-0001-9889-509X
Carl WikbergGeneral Practice/Family Medicine, School of Public Health and Community Medicine, Institute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden.ORCID http://orcid.org/0000-0002-6494-5922
David SundemoGeneral Practice/Family Medicine, School of Public Health and Community Medicine, Institute of Medicine, University of Gothenburg Sahlgrenska Academy, Gothenburg, Västra Götaland County, Sweden.ORCID http://orcid.org/0000-0002-5871-1636

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

objectivesTo identify the lowest sensitivity and specificity that physicians and the general population consider acceptable for medical artificial intelligence (AI), relative to current human performance.

methodsIn a nationwide, cross-sectional survey in Sweden, 2025, random samples of 500 physicians and 500 adults from the general population were mailed a questionnaire presenting three vignettes (chest pain triage, sore throat triage, ECG myocardial infarction detection) with the corresponding human performance. Participants reported the maximum number of cases an AI should be allowed to miss or over-refer.

resultsResponse rates were 45% among physicians and 31% in the general population. Both groups demanded higher AI accuracy than the human benchmark for all cases. In the chest pain triage vignette, the nurse correctly referred 84 of 100 true emergencies; physicians required the AI to correctly refer 11 additional patients (95% sensitivity) and the general population demanded referral of 16 additional patients (100% sensitivity) (p<0.001 for both groups). Among 100 patients not requiring referral, the nurse would mistakenly refer 66. Both groups required the AI to reduce unnecessary referrals by 16 (50% specificity) (p<0.001). A similar pattern was observed in the other vignettes. DISCUSSION: The accuracy thresholds required by the respondents exceed the performance of many existing systems, although emerging AI research shows promise in narrowing the gap.

conclusionPhysicians and the general population require medical AI systems to outperform human clinicians. When implementing AI in healthcare settings, early engagement with both groups may be necessary to align expectations with real-world system performance.

Indexed as

Artificial IntelligencePhysiciansAdultChest PainCross-Sectional StudiesFemaleHumansMaleMiddle AgedReferral and ConsultationSensitivity and SpecificitySurveys and QuestionnairesSwedenTriageArtificial intelligenceDecision Support Systems, Clinical

Identifiers

PMID41927104
PMCPMC13052720

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.