ArticleInternational urogynecology journal2026
ChatGPT Health and ChatGPT Plus in Urogynecology: A Blinded Comparison of Response Quality, Guideline Concordance, and Safety.
Article in International urogynecology journal, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
1 author.
Funding
No grant is acknowledged in the PubMed record.
Abstract
introduction and hypothesisChatGPT Health is a health-focused conversational artificial intelligence (AI) environment, but its performance in patient-oriented urogynecology has not been directly compared with standard ChatGPT using both validated quality assessment and objective guideline-based criteria. We hypothesized that ChatGPT Health would provide higher patient-facing response quality while maintaining comparable guideline concordance and safety.
methodsTwenty-four guideline-informed patient queries covering common urogynecological and lower urinary tract concerns were submitted independently to ChatGPT Plus and ChatGPT Health on 14 August 2026. Forty-eight responses were anonymized and independently evaluated by five blinded clinicians using the validated Quality Analysis of Medical Artificial Intelligence (QAMAI) instrument, a prespecified five-element guideline-concordance checklist, and a safety scale. Blinded head-to-head preference was assessed secondarily. Prompt-level paired comparisons used Wilcoxon signed-rank tests.
resultsMedian guideline concordance was 100% for ChatGPT Health and 98% for ChatGPT Plus (p = 0.298). ChatGPT Health had a higher blinded five-domain QAMAI subtotal (25.0 vs 24.2; p = 0.026) and greater clarity (p = 0.020), completeness (p = 0.016), and usefulness (p = 0.033), with all three remaining significant after correction for false discovery rate. Accuracy did not differ (p = 0.453). No clinically meaningful safety concerns were identified in 240 ratings. ChatGPT Health was preferred in 47.5% of 120 assessments, ChatGPT Plus in 18.3%, and no meaningful difference was reported in 34.2%.
conclusionsBoth platforms generated highly accurate, guideline-concordant, and safe responses. ChatGPT Health showed advantages primarily in clarity, completeness, usefulness, and overall quality of patient-facing communication rather than objective concordance of clinical content.
Indexed as
Identifiers
42827178What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.