Evidence map›Paper›PMID 41663094›Full record

ArticleJMIR infodemiology2026

Leveraging AI for Analysis of Digital Health Information on Cancer Prevention Among Arab Youth and Adults: Content Analysis.

Alia Komsany, Obada Al Zoubi, Laetitia Sebaaly, Gabrielle Harrison, Orysya Soroka, Safa ElKefi, David Scales, Erica Phillips, Laura C Pinheiro, Israa Ismail and 1 more

Abstract read
In one paragraph

Article in JMIR infodemiology, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

11 authors.

Alia KomsanyDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0009-0008-1450-0213
Obada Al ZoubiIndependent Researcher, Boston, MA, United States.ORCID 0000-0003-0203-5779
Laetitia SebaalyIndependent Consultant, New York, NY, United States.ORCID 0009-0005-8933-7818
Gabrielle HarrisonIndependent Consultant, New York, NY, United States.ORCID 0009-0005-5083-2432
Orysya SorokaDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0000-0003-3282-2378
Safa ElKefiSchool of System Sciences and Industrial Engineering, Watson College of Engineering, Binghamton University, New York, NY, United States.ORCID 0000-0002-4293-0404
David ScalesDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0000-0001-5727-7148
Erica PhillipsDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0000-0002-6803-2442
Laura C PinheiroDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0000-0002-6920-8526
Israa IsmailDivision of General Internal Medicine, Weill Cornell Medicine, New York City, NY, United States.ORCID 0009-0006-6129-7217
Perla ChebliNYU Langone Health, New York City, NY, United States.ORCID 0000-0002-0344-442X

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAs TikTok (ByteDance) grows as a major platform for health information, the quality and accuracy of Arabic-language cancer prevention content remain unknown. Limited access to culturally relevant and evidence-based information may exacerbate disparities in cancer knowledge and prevention behaviors. Although large language models offer scalable approaches for analyzing online health content, their utility for short-form video data, especially in underrepresented languages, has not been well established.

objectiveWe aimed to characterize and evaluate the quality of Arabic-language TikTok videos on cancer prevention and explore the use of large language models for scalable content analysis.

methodsWe used the TikTok research application programming interface and a GPT-assisted keyword strategy to collect Arabic-language TikTok videos (2021-2024). From an initial collection of 1800 TikTok videos, 320 were eligible after preprocessing. Of these, the top 25% (N=30) most-viewed were analyzed and manually coded for content type, cancer type, uploader identity, tone and register, scientific citation, and disclaimers. Video quality was assessed using the Patient Education Materials Assessment Tool for Audiovisual Materials for understandability and actionability, and the Global Quality Scale (GQS). GPT-4 was used to generate artificial intelligence annotations, which were compared to human coding for select variables.

resultsThe top 25% (N=30) most-viewed videos amassed a total of 21.6 million views. Diet and alternative therapies were most common (15/30, 50%), which included recommendations to reduce hydrogenated oils, increase fruit and vegetable intake, and the use of traditional remedies such as garlic and black seed. Only 6.6% (2/30) of videos cited scientific literature. General cancer (15/30, 53%), breast (5/30, 17%), and cervical (4/30, 13%) cancers were most frequently mentioned. Doctors led 30% (9/30) of videos and were more likely to produce higher quality content, including significantly higher global quality scores (GQS=4, median 4, IQR 4-4 vs 3, median 3, IQR 2-3, P=.06). Over half of the videos had low understandability (16/30, 53%) and actionability (18/30, 60%). Emotionally framed content had the highest engagement across likes and shares, although this did not reach statistical significance (P=.08 and P=.05, respectively). However, emotional tone was significantly associated with lower GQS scores (P=.01). GPT-4 showed high agreement with human coders for cancer type (Cohen κ=1.0), strong agreement for GQS (κ=0.94), but low agreement for tone classification (κ=0.15), due to misclassification of emotional delivery from text-only input.

conclusionsArabic-language TikTok cancer prevention content is highly engaging but variable in quality, with emotionally framed videos attracting substantial attention despite lower informational value. Artificial intelligence-assisted tools show strong potential for scalable, multilingual health content analysis, but multimodal approaches are needed to accurately interpret tonal and audiovisual features.

Indexed as

ArabsArtificial IntelligenceConsumer Health InformationNeoplasmsAdolescentAdultDigital HealthDigital MediaFemaleHumansLarge Language ModelsMaleAI-driven content analysiscancer preventiondigital health communicationengagementTikTok

Identifiers

PMID41663094
PMCPMC12930147

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.