ArticleThe Yale journal of biology and medicine2024
Assessing the Efficacy of Large Language Models in Health Literacy: A Comprehensive Cross-Sectional Study.
Article in The Yale journal of biology and medicine, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 26 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
26 citing papers in PubMed, 36 citations in OpenAlex.
- Rethinking Pediatric Asthma Education Through Large Language Model Generation and Simplification: Randomized Double-Blind Study.Journal of medical Internet research · 2026Trial
- Evaluating large language model`s performance in answering principles of health course questions.Scientific reports · 2026Article
- Gonarthrosis Advisor vs ChatGPT-5: quality and readability of artificial intelligence-generated patient education for knee osteoarthritis.Acta orthopaedica et traumatologica turcica · 2026Article
- Applications of Large Language Models in Glaucoma: A Scoping Review.Vision (Basel, Switzerland) · 2026Review
- Utility of ChatGPT in generating accurate client handouts for common veterinary internal medicine diseases.Journal of veterinary internal medicine · 2026Article
- Effectiveness of large language models in preoperative and discharge education: a systematic review based on an evaluation framework.NPJ digital medicine · 2026Article
- Artificial intelligence as a source of psychotherapy-related psychoeducation: a cross-sectional content analysis of information on eye movement desensitization and reprocessing.Frontiers in psychiatry · 2026Article
- Can LLMs simplify operative notes? A comparative analysis in otorhinolaryngology.European archives of oto-rhino-laryngology : official journal of the European Federation of Oto-Rhino-Laryngological Societies (EUFOS) : affiliated with the German Society for Oto-Rhino-Laryngology - Head and Neck Surgery · 2026Article
- AI-generated explanations in kidney transplantation: accuracy vs. readability and implications for patient education.Frontiers in artificial intelligence · 2026Article
- Large language models as information providers for appropriate antimicrobial use: computational text analysis and expert-rated comparison of ChatGPT, Claude and Gemini.BMJ health & care informatics · 2025Article
- From Data to Decisions: Leveraging Retrieval-Augmented Generation to Balance Citation Bias in Burn Management Literature.European burn journal · 2025Article
- Artificial Intelligence in Peripheral Artery Disease Education: A Battle Between ChatGPT and Google Gemini.Cureus · 2025Article
- Assessing ChatGPT responses to patient questions on epidural steroid injections: A comparative study of general vs specific queries.Interventional pain medicine · 2025Article
- Assessment of patient information guides generated by LLMs for common cardiological procedures.Global cardiology science & practice · 2025Article
- ChatGPT 4.0's efficacy in the self-diagnosis of non-traumatic hand conditions.Journal of hand and microsurgery · 2025Article
- Article
- Comparative Study to Evaluate the Accuracy of Differential Diagnosis Lists Generated by Gemini Advanced, Gemini, and Bard for a Case Report Series Analysis: Cross-Sectional Study.JMIR medical informatics · 2024Article
- Comparative performance analysis of large language models: ChatGPT-3.5, ChatGPT-4 and Google Gemini in glucocorticoid-induced osteoporosis.Journal of orthopaedic surgery and research · 2024Article
- The Potential Impact of Large Language Models on Doctor-Patient Communication: A Case Study in Prostate Cancer.Healthcare (Basel, Switzerland) · 2024Article
- The Emerging Role of Large Language Models in Improving Prostate Cancer Literacy.Bioengineering (Basel, Switzerland) · 2024Article
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
4 authors at 1 institution in 1 country.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Enhanced health literacy in children has been empirically linked to better health outcomes over the long term; however, few interventions have been shown to improve health literacy. In this context, we investigate whether large language models (LLMs) can serve as a medium to improve health literacy in children. We tested pediatric conditions using 26 different prompts in ChatGPT-3.5, ChatGPT-4, Microsoft Bing, and Google Bard (now known as Google Gemini). The primary outcome measurement was the reading grade level (RGL) of output as assessed by Gunning Fog, Flesch-Kincaid Grade Level, Automated Readability Index, and Coleman-Liau indices. Word counts were also assessed. Across all models, output for basic prompts such as "Explain" and "What is (are)," were at, or exceeded, the tenth-grade RGL. When prompts were specified to explain conditions from the first- to twelfth-grade level, we found that LLMs had varying abilities to tailor responses based on grade level. ChatGPT-3.5 provided responses that ranged from the seventh-grade to college freshmen RGL while ChatGPT-4 outputted responses from the tenth-grade to the college senior RGL. Microsoft Bing provided responses from the ninth- to eleventh-grade RGL while Google Bard provided responses from the seventh- to tenth-grade RGL. LLMs face challenges in crafting outputs below a sixth-grade RGL. However, their capability to modify outputs above this threshold, provides a potential mechanism for adolescents to explore, understand, and engage with information regarding their health conditions, spanning from simple to complex terms. Future studies are needed to verify the accuracy and efficacy of these tools.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.