Evidence map›Paper›PMID 38329525›Full record

ArticleEuropean archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery2024

Diagnosis of malignancy in oropharyngeal confocal laser endomicroscopy using GPT 4.0 with vision.

Matti Sievert, Marc Aubreville, Sarina Katrin Mueller, Markus Eckstein, Katharina Breininger, Heinrich Iro, Miguel Goncalves

Abstract read
PubMed Publisher
In one paragraph

Article in European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 17 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
17citing papers in PubMed, 2 pooled it
13.2field-weighted citation impact, top 1% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

17 citing papers in PubMed, 2 syntheses or guidelines pooled it, 23 citations in OpenAlex.

  1. Clinical decision support using large language models in otolaryngology: a systematic review.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2025
    Pooled it
  2. Pooled it
  3. Artificial intelligence in otolaryngology: current applications, limitations, and future perspectives.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026
    Review
  4. Clinical Applications of Multimodal Artificial Intelligence in Otolaryngology: A State-of-the-Art Review.Otolaryngology--head and neck surgery : official journal of American Academy of Otolaryngology-Head and Neck Surgery · 2026
    Review
  5. Article
  6. Article
  7. Review
  8. Article
  9. Article
  10. Article
  11. Characterization of irradiated mucosa using confocal laser endomicroscopy in the upper aerodigestive tract.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2025
    Article
  12. Article
  13. Article
  14. Review
  15. Review
  16. Applications of Large Language Models in Pathology.Bioengineering (Basel, Switzerland) · 2024
    Review
  17. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors at 3 institutions in 1 country.

Matti SievertDepartment of Otorhinolaryngology, Head and Neck Surgery, Friedrich Alexander University of Erlangen-Nuremberg, Erlangen University Hospital, Erlangen, Germany.
Marc AubrevilleTechnische Hochschule Ingolstadt, Ingolstadt, Germany.
Sarina Katrin MuellerDepartment of Otorhinolaryngology, Head and Neck Surgery, Friedrich Alexander University of Erlangen-Nuremberg, Erlangen University Hospital, Erlangen, Germany.
Markus EcksteinInstitute of Pathology, Friedrich-Alexander-Universität Erlangen-Nürnberg, University Hospital, Erlangen, Germany.
Katharina BreiningerDepartment of Artificial Intelligence in Biomedical Engineering, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
Heinrich IroDepartment of Otorhinolaryngology, Head and Neck Surgery, Friedrich Alexander University of Erlangen-Nuremberg, Erlangen University Hospital, Erlangen, Germany.
Miguel GoncalvesDepartment of Otorhinolaryngology, Plastic and Aesthetic Operations, University Hospital Würzburg, Joseph-Schneider-Straße 11, 97080, Würzburg, Germany. Goncalves_M@ukw.de.ORCID http://orcid.org/0000-0002-0036-4598
Friedrich-Alexander-Universität Erlangen-Nürnberg · DETechnische Hochschule Ingolstadt · DEUniversitätsklinikum Würzburg · DE

Funding

Deutsche Forschungsgemeinschaft 3182/2-1
6 · The paper itself

Abstract

purposeConfocal Laser Endomicroscopy (CLE) is an imaging tool, that has demonstrated potential for intraoperative, real-time, non-invasive, microscopical assessment of surgical margins of oropharyngeal squamous cell carcinoma (OPSCC). However, interpreting CLE images remains challenging. This study investigates the application of OpenAI's Generative Pretrained Transformer (GPT) 4.0 with Vision capabilities for automated classification of CLE images in OPSCC.

methodsCLE Images of histological confirmed SCC or healthy mucosa from a database of 12 809 CLE images from 5 patients with OPSCC were retrieved and anonymized. Using a training data set of 16 images, a validation set of 139 images, comprising SCC (83 images, 59.7%) and healthy normal mucosa (56 images, 40.3%) was classified using the application programming interface (API) of GPT4.0. The same set of images was also classified by CLE experts (two surgeons and one pathologist), who were blinded to the histology. Diagnostic metrics, the reliability of GPT and inter-rater reliability were assessed.

resultsOverall accuracy of the GPT model was 71.2%, the intra-rater agreement was κ = 0.837, indicating an almost perfect agreement across the three runs of GPT-generated results. Human experts achieved an accuracy of 88.5% with a substantial level of agreement (κ = 0.773).

conclusionsThough limited to a specific clinical framework, patient and image set, this study sheds light on some previously unexplored diagnostic capabilities of large language models using few-shot prompting. It suggests the model`s ability to extrapolate information and classify CLE images with minimal example data. Whether future versions of the model can achieve clinically relevant diagnostic accuracy, especially in uncurated data sets, remains to be investigated.

Indexed as

Head and Neck NeoplasmsHumansLasersMicroscopy, ConfocalReproducibility of ResultsSquamous Cell Carcinoma of Head and NeckConfocal laser endomicroscopyGPTHead and neck malignanciesOropharyngeal squamous cell carcinoma

Identifiers

PMID38329525
OpenAlexW4391653030

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.