Evidence map›Paper›PMID 41287740›Full record

ReviewCureus2025

Accuracy and Reliability of Artificial Intelligence in Surgical Decision-Making: A Literature Review.

Nicolás Idárraga Ruiz, Israel Cardona Salazar, Lincoln Xavier Naranjo Palacio, Carolina Agudelo Agudelo, Alfonso Miguel Ledesma Parra, Julio Cesar Flores Rodriguez

Abstract readReview
In one paragraph

Review in Cureus, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 3 papers.

0numbers the graph read from it
0cells of the map it votes in
3citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

3 citing papers in PubMed.

  1. Article
  2. Article
  3. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

6 authors.

Nicolás Idárraga RuizSurgery, University of Manizales, Manizales, COL.
Israel Cardona SalazarSurgery, Instituto Mexicano del Seguro Social, Sonora, MEX.
Lincoln Xavier Naranjo PalacioGeneral Practice/Emergency, Clínica Internacional de Traumatología, Quito, ECU.
Carolina Agudelo AgudeloEmergency, Clínica Las Américas Auna, Medellín, COL.
Alfonso Miguel Ledesma ParraSurgery, Hospital de Especialidades del Centro Médico Nacional La Raza, Instituto Mexicano del Seguro Social (IMSS), Ciudad de México, MEX.
Julio Cesar Flores RodriguezAesthetic and Regenerative Medicine, Clínica Aura, San Pedro Garza García, MEX.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

This narrative literature review synthesized evidence to address gaps in knowledge regarding AI performance and its integration into surgical operations. The purpose of the review was to assess AI accuracy and reliability, benchmark real-time guidance technologies, identify data and ethical issues, compare model performance across different specialties, and review the role of AI in improving surgical accuracy and safety. It reviewed 28 studies conducted across various geographic and disciplinary contexts and discussed machine learning (ML) and deep learning (DL) as applied to major surgeries. Results show that AI models' overall performance is substantial in intraoperative (IOP) decision-making, with five of six studies reporting AUC values of 0.85-0.95, indicating significant discriminatory power. Moreover, the accuracy performance metric across 22 studies showed high predictive performance of AI models in surgical settings, with accuracies ranging from 80% to 99%, except for one study, which reported an accuracy below 70%. These findings emphasized the practical feasibility of AI in IOP decision-making. Hence, AI's role in IOP is promising, assisting surgeons' decision-making in the operating room. Therefore, ML and DL are highly precise in anatomic detection, surgical-phase detection, complication prediction, and real-time event detection. Developments in DL algorithms, such as convolutional neural networks and generative adversarial networks, have enabled more accurate surgical guidance and the prediction of IOP events, thereby increasing surgical accuracy and potentially reducing errors. However, the model's performance needs to be validated through long-term computational and real-time clinical study designs, ensuring appropriate strategies for data validation and model performance assessment. The narrative review study design focused solely on the narrative synthesis, rather than on data validation (internal or external) or quality assessment of the included studies. Therefore, future researchers should conduct a systematic review to validate the findings. The readers must be cautious when interpreting the findings. Hence, AI use in surgery training and workflow optimization has the potential to improve surgical performance and patient outcomes, but scalability and long-term outcomes have yet to be demonstrated. Although AI technologies can improve the accuracy and reliability of decisions made in IOP settings, it is critical to address methodological, infrastructural, and ethical constraints to enable safe and effective clinical application in major surgeries.

Indexed as

artificial intelligence in surgerydecision-makingintraoperativereliabilitysurgery

Identifiers

PMID41287740
PMCPMC12640677

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.