ArticleFrontiers in psychiatry2026
Assessing large language model responses to pediatric depression FAQs: a cross-sectional study on readability, accuracy, and sentiment.
Article in Frontiers in psychiatry, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Pediatric depression shows age-specific symptoms that hinder recognition and delay care, while parents and adolescents increasingly turn to online sources, including large language models, for mental health information and guidance. The quality of such information depends on readability, factual accuracy, completeness, and emotional tone. This study compared responses from 3 contemporary large language models (LLMs) to frequently asked questions about pediatric depression to assess their suitability as informational tools. Methods: A cross-sectional analytical study design was used. 15 standardized frequently asked questions covering definition, causes, clinical features, diagnosis, prevention, treatment, and prognosis of pediatric depression were submitted to ChatGPT-5, Microsoft Copilot GPT-5 in Smart Research mode, and DeepSeek 3.1V. Responses were collected verbatim. Readability was assessed using seven established indices. Accuracy and completeness were independently scored on a 0 to 6 scale using a predefined rubric. Sentiment was measured with sentiment scores. One-way analysis of variance (ANOVA) with Tukey Results: Readability was different among the various models. DeepSeek 3.1V achieved the highest Flesch Reading Ease Score of 54 to 55 and the lowest Flesch-Kincaid Grade Level of about 9.5 thus indicating easier comprehension. ChatGPT-5 showed intermediate readability with scores of 49 to 50 and grade level about 10.5. Copilot-5 had the lowest Reading Ease score of 43 to 44 and the highest grade level near 10.8. Accuracy on a 0 to 6 scale was highest for Copilot-5. ChatGPT-5 showed the greatest completeness, whereas other models had variable coverage in detailed clinical items. Conclusion: Large language models (LLMs) provide information on pediatric depression but show varying levels of readability, accuracy, and completeness. DeepSeek 3.1V provides greater linguistic accessibility, Microsoft Copilot GPT-5 shows stronger factual consistency, and ChatGPT-5 provides more comprehensive coverage. These artificial intelligence (AI) chatbot systems require human understanding before use in pediatric mental health education or guidance.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.