Evidence map›Paper›PMID 41281672›Full record

ArticleActa informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH2025

The Accuracy And Clinical Relevance of Chat GPT-4 in Triple Negative Breast Cancer Research.

Ramakrishna Gummadi, Sai Kiran S S Pindiprolu, Chirravuri S Phani Kumar

Abstract read
In one paragraph

Article in Acta informatica medica : AIM : journal of the Society for Medical Informatics of Bosnia & Herzegovina : casopis Drustva za medicinsku informatiku BiH, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Ramakrishna GummadiAditya Pharmacy College, Surampalem, Andhrapradesh, 533437, India.
Sai Kiran S S PindiproluAditya Pharmacy College, Surampalem, Andhrapradesh, 533437, India.
Chirravuri S Phani KumarAditya Pharmacy College, Surampalem, Andhrapradesh, 533437, India.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Triple negative breast cancer (TNBC) is an aggressive subtype of breast cancer characterized by the lack of estrogen receptor(ER), progesterone receptor (PR), and human epidermal growth factor receptor 2 (HER2). The absence of these receptors reduces the effectiveness of targeted treatment approaches. With the increasing use of artificial intelligence (AI) in medical research and clinical decision- making, there is growing interest in evaluating the accuracy and reliability of large language models (LLMS), such as chatGPT-4, in oncology related applications. Objective: The research aims to systematically assess the reliability of ChatGPT-4 in addressing frequently asked questions related to TNBC in four critical areas: diagnosis, treatment, prognosis and survival, and quality of life. Expert evaluations and statistical analyses are employed to measure the accuracy of the models responses. Methods: A set of 100 questions related to TNBC was gathered from credible medical sources, including peer-reviewed journals and clinical oncology specialists evaluated the response generated by ChatGPT-4 using a structured assessment framework, classifying each answer into one of four accuracy levels, completely inaccurate, partially accurate, accurate but lacking depth and highly accurate. Results: To evaluate the consistency among reviewers, Cohen's kappa coefficient was calculated, and descriptive statistical analysis was conducted to identify overall accuracy patterns. The findings indicated that 73% of the responses were classified as either "Accurate" or "Highly Accurate", suggesting the potential of ChatGPT-4 as a supplementary resource for obtaining information on TNBC. However, 27% of the responses were categorized as "partially accurate" or 'Completely Inaccurate, "highlighting gaps in contextual understanding and instances of misinformation.. Cohen's kappa coefficient was recorded at 0.007, reflecting a week level of agreement among evaluators and highlighting the impact of subjective interpretation. The model demonstrated strong performance in well-established areas such as chemotherapy protocols and diagnostic procedures but faced challenges with emerging research topics, personalized treatment recommendations, and fertility related concerns. Conclusion: ChatGPT-4 exhibits significant potential in summarizing information on TNBC; however, the accuracy of its responses varies depending on the complexity and specificity of the queries. Due to inconsistencies and low inter-rater reliability, AI-generated medical content requires verification by medical professionals before being applied to patient care or clinical decision-making. Future developments in large language models should focus on reducing inaccuracies, incorporating the latest medical data, and improving adaptability to better support personalized medicine.

Indexed as

ChatGPT-4Cohen’s Kappa CoefficientLarge Language Models (LLLMs)Medical AI AccuracyOncology Decision SupportTriple Negative Brest Cancer (TNBC)

Identifiers

PMID41281672
PMCPMC12634090

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.