ArticleObesity surgery2026
Large Language Models and Metabolic Bariatric Surgery: A Pilot Concordance Study Between ChatGPT and Multidisciplinary Team Recommendations.
Article in Obesity surgery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introductionOperation selection in metabolic surgery is a complex decision making process led by a multidisciplinary team that integrates multiple anatomical, clinical, metabolic and psychosocial aspects. The ability of large language models (LLMs) has been proposed to provide capability to act as decision support tools, but their performance in replicating MDT level decision making in metabolic surgery remains uncertain.
methodsA retrospective pilot study was performed using anonymised data from 100 patients at a single high-volume UK NHS bariatric centre who underwent surgery up to August 2025. Preoperative demographic, anthropometric, clinical and psychosocial variables were extracted from electronic records. These were provided to ChatGPT-4 Auto using a single standardised prompt instructing the model to act collectively on behalf of the whole bariatric MDT and recommend the most appropriate metabolic operation. These recommendations were compared with formal MDT decisions and with the operation ultimately performed. Concordance was assessed using raw percentage agreement, Cohen's kappa and Stuart-Maxwell tests.
resultsChatGPT demonstrated 70% concordance with MDT recommendations. Concordance between MDT and operation performed was 83%, while concordance between ChatGPT and operation performed was 63%. Agreement beyond chance between ChatGPT and MDT recommendations was low (Cohen's kappa 0.036) reflecting class imbalance. Stuart-Maxwell test showed no significant difference in marginal distribution between ChatGPT and MDT recommendations. Both ChatGPT and MDT recommended bypass procedures more frequently than ultimately performed.
conclusionChatGPT overall demonstrated moderate crude agreement with bariatric MDT decision making in a real world UK NHS cohort of patients. However, limited agreement beyond chance and influence of unmeasured human factors may preclude its use in an autonomous fashion. This study establishes the requirement for larger scale evaluation of LLM clinical reasoning and supports the future exploration of them as adjunctive rather than autonomous decision support tools in metabolic surgery.
Indexed as
Identifiers
42240797What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.