Evidence map›Paper›PMID 39441598›Full record

ArticleJAMA network open2024

Performance of Multimodal Artificial Intelligence Chatbots Evaluated on Clinical Oncology Cases.

David Chen, Ryan S Huang, Jane Jomy, Philip Wong, Michael Yan, Jennifer Croke, Daniel Tong, Andrew Hope, Lawson Eng, Srinivas Raman

Abstract read
In one paragraph

Article in JAMA network open, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 26 papers.

0numbers the graph read from it
0cells of the map it votes in
26citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

26 citing papers in PubMed.

  1. Trial
  2. Review
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Article
  9. Article
  10. Article
  11. Article
  12. Review
  13. Article
  14. Article
  15. Moving toward precision and personalized treatment strategies in psychiatry.The international journal of neuropsychopharmacology · 2025
    Review
  16. Article
  17. Article
  18. Article
  19. Review
  20. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

David ChenRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Ryan S HuangRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Jane JomyRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Philip WongRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Michael YanRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Jennifer CrokeRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Daniel TongRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Andrew HopeRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.
Lawson EngDivision of Medical Oncology and Hematology, Department of Medicine, Princess Margaret Cancer Centre/University Health Network Toronto, Toronto, Ontario, Canada.
Srinivas RamanRadiation Medicine Program, Princess Margaret Hospital Cancer Centre, Toronto, Ontario, Canada.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Importance: Multimodal artificial intelligence (AI) chatbots can process complex medical image and text-based information that may improve their accuracy as a clinical diagnostic and management tool compared with unimodal, text-only AI chatbots. However, the difference in medical accuracy of multimodal and text-only chatbots in addressing questions about clinical oncology cases remains to be tested. Objective: To evaluate the utility of prompt engineering (zero-shot chain-of-thought) and compare the competency of multimodal and unimodal AI chatbots to generate medically accurate responses to questions about clinical oncology cases. Design, Setting, and Participants: This cross-sectional study benchmarked the medical accuracy of multiple-choice and free-text responses generated by AI chatbots in response to 79 questions about clinical oncology cases with images. Exposures: A unique set of 79 clinical oncology cases from JAMA Network Learning accessed on April 2, 2024, was posed to 10 AI chatbots. Main Outcomes and Measures: The primary outcome was medical accuracy evaluated by the number of correct responses by each AI chatbot. Multiple-choice responses were marked as correct based on the ground-truth, correct answer. Free-text responses were rated by a team of oncology specialists in duplicate and marked as correct based on consensus or resolved by a review of a third oncology specialist. Results: This study evaluated 10 chatbots, including 3 multimodal and 7 unimodal chatbots. On the multiple-choice evaluation, the top-performing chatbot was chatbot 10 (57 of 79 [72.15%]), followed by the multimodal chatbot 2 (56 of 79 [70.89%]) and chatbot 5 (54 of 79 [68.35%]). On the free-text evaluation, the top-performing chatbots were chatbot 5, chatbot 7, and the multimodal chatbot 2 (30 of 79 [37.97%]), followed by chatbot 10 (29 of 79 [36.71%]) and chatbot 8 and the multimodal chatbot 3 (25 of 79 [31.65%]). The accuracy of multimodal chatbots decreased when tested on cases with multiple images compared with questions with single images. Nine out of 10 chatbots, including all 3 multimodal chatbots, demonstrated decreased accuracy of their free-text responses compared with multiple-choice responses to questions about cancer cases. Conclusions and Relevance: In this cross-sectional study of chatbot accuracy tested on clinical oncology cases, multimodal chatbots were not consistently more accurate than unimodal chatbots. These results suggest that further research is required to optimize multimodal chatbots to make more use of information from images to improve oncology-specific medical accuracy and reliability.

Indexed as

Artificial IntelligenceMedical OncologyCross-Sectional StudiesHumansNeoplasms

Identifiers

PMID39441598
PMCPMC11581577

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.