Evidence map›Paper›PMID 42416102›Full record

ReviewJournal of medical education and curricular development

Comparing Artificial intelligence to physicians' competences in the domain of clinical reasoning: A systematic review and meta-analysis.

Mona Mlika, Mohamed Majdi Zorgati, Imen Ben Ismail, Sarra Cheikhrouhou, Paul Hofman, Iheb Labbene

Abstract readReview
In one paragraph

Review in Journal of medical education and curricular development. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Mona MlikaDepartment of Pathology, Center of Traumatology and Major Burns, Ben Arous, Tunis, Tunisia.ORCID https://orcid.org/0000-0003-2470-0012
Mohamed Majdi ZorgatiDepartment of Biology, Military Hospital, Doha, Qatar.
Imen Ben IsmailFaculty of Medicine of Tunis, University Tunis El Manar, Tunis, Tunisia.
Sarra CheikhrouhouFaculty of Medicine of Tunis, University Tunis El Manar, Tunis, Tunisia.
Paul HofmanDepartment of Pathology, IHU Côte d'Azur, RespirERA, Nice, France.
Iheb LabbeneFaculty of Medicine of Tunis, University Tunis El Manar, Tunis, Tunisia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Clinical reasoning is a complex process that plays a crucial role in order to solve the patients' problems. Our aim was to compare AI competences to the physicians' competences in clinical reasoning. Methods: We performed a meta-analysis under the guidelines of the AMSTAR (version 2). All eligible articles were retrieved from Pubmed, Embase and Cochrane databases. The articles included were rated according to the MMERSQI scores. Binary diagnostic accuracy was used as the primary outcome. Continuous performance scores were considered as secondary outcome in order to avoid excluding studies using continuous scores instead of concordance rates. The Review Manager software 5.4 (free version) was used to conduct this meta-analysis. The OR with the 95% CI were calculated for studies reporting binary accuracy. The SMD with the 95% CI were calculated for studies using continuous scores. Q test and I Results: Considering the studies using binary scores and comparing odds ratios, 1609 clinical vignettes were used to compare the different LLM to the human clinical reasoning. The combined OR reached 0.65 with 95% CI [0.38, 1.12]. No significant difference between both groups was observed (p=0.12). When considering studies using continuous scoring systems, the combined standard mean difference reached 0.08 with 95% CI [-0.19, 0.35]. No significant difference between both groups was observed. In order to explain the heterogeneity that was noticed, we performed a sub-group analysis taking into account the nature of the clinical cases (real-world or published), the LLM system used (Chat-GPT v4) and the expertise of the respondants (novices or experts). The heterogeneity was observed in all subgroups excluding the subgroup of the studies using binary scoring systems and comparing LLM's scores to novices' scores. The combined OR reached 0.62 with 95% CI [0.35, 1.1]. No significant difference was observed between both groups (p=0.1) and the heterogeneity I-square was evaluated to 0% and Tau Conclusion: Even if this meta-analysis showed the absence of difference between AI and human clinical reasoning, these reaults have to be taken with caution because of the important heterogeneity that wasn't resolved by subgroup analyses.

Indexed as

artificial intelligenceclinical reasoningclinical settings

Identifiers

PMID42416102
PMCPMC13338529

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.