Evidence map›Paper›PMID 42647856›Full record

SynthesisJournal of medical Internet research2026

Improvement of Clinical Practice Guideline Appraisal by Human Experts and AI Agents by Using Structured Guidance: Systematic Review, Meta-Analysis, and Validation Study.

Xingrun Mao, Junhao Wang, Zezhang Wang, Yuwei Zhang, Shiyu Qiu, Jingyu Ye, Shiyan He, Chongyang Wang, Yong Xia, Chengqi He and 2 more

Abstract readSystematic ReviewMeta-AnalysisValidation Study
In one paragraph

Synthesis in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

12 authors.

Xingrun Mao *Rehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0007-4074-0264
Junhao Wang *Center of Statistical Research, School of Statistics and Data Science, Southwestern University of Finance and Economics, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0004-4115-6981
Zezhang WangRehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0002-5932-132X
Yuwei ZhangCenter of Statistical Research, School of Statistics and Data Science, Southwestern University of Finance and Economics, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0005-3519-7602
Shiyu QiuWest China School of Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0004-7136-0481
Jingyu YeWest China School of Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0005-2138-196X
Shiyan HeRehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0009-0005-8000-5200
Chongyang WangRehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0000-0002-9819-088X
Yong XiaRehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0000-0002-7939-9995
Chengqi HeRehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0000-0002-5349-0571
Ke Li *Center of Statistical Research, School of Statistics and Data Science, Southwestern University of Finance and Economics, Chengdu, Sichuan, China.ORCID https://orcid.org/0000-0002-0252-898X
Siyi Zhu *Rehabilitation Medicine Center and Institute of Rehabilitation Medicine, West China Hospital, Sichuan University, Chengdu, Sichuan, China.ORCID https://orcid.org/0000-0001-8213-7622

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundRehabilitation clinical practice guidelines (CPGs) have increased rapidly, but inconsistent methodological quality limits their implementation. Although Appraisal of Guidelines for Research and Evaluation II (AGREE II) and Reporting Items for Practice Guidelines in Health Care (RIGHT) provide standardized appraisal frameworks, their application is time-consuming. Large language model (LLM)-based AI agents may offer a scalable alternative with uncertain reliability.

objectiveWe evaluated rehabilitation CPGs' methodological and reporting quality and determined whether structured guidance improves human expert-AI agent agreement.

methodsWe systematically reviewed English- and Chinese-language rehabilitation CPGs from Embase, Scopus, PubMed, China National Knowledge Infrastructure, Wanfang Data, National Institute for Health and Care Excellence, Scottish Intercollegiate Guidelines Network, and Guidelines International Network up to June 2026. Methodological and reporting quality were assessed using AGREE II and the RIGHT checklist. Factors associated with guideline quality were examined using regression and subgroup analyses. Two AI agents were compared with human consensus with and without a structured guideline appraisal workbook, followed by external validation using 6 anterior cruciate ligament reconstruction CPGs.

resultsWe included 227 CPGs (163 English-language, 64 Chinese-language). After introducing a structured guideline appraisal workbook, agreement among human experts improved markedly-mean intraclass correlation coefficients (ICCs) increased from -0.09 to 0.66 to 0.84-0.92 across AGREE II domains. Overall guideline quality remained low, with 35.9% (SD 18,8%) applicability and 52% (SD 17.2%) stakeholder involvement. English-language guidelines outperformed Chinese-language guidelines in scope and purpose (mean 74.64, SD 15.4 vs mean 68.88, SD 13.9; P=.004) and applicability (mean 39.14, SD 18.6 vs mean 27.54, SD 16.7; P<.001). Backward-elimination logistic regression revealed external review as an associated process characteristic (odds ratio 20.39, 95% CI 4.66-89.27; P<.001). RIGHT assessments showed consistent reliability (ICC=0.80-0.88). Reporting was highest for basic information (70.5%) and lowest for funding, declaration, and management of interests (44.1%). Meta-analysis of RIGHT reporting rates showed lower reporting among Chinese-language than English-language guidelines (risk difference [RD] -0.07, 95% CI -0.13 to -0.02, 95% prediction interval [PI] -0.39 to 0.24) and among guidelines published before vs after RIGHT release (RD -0.19, 95% CI -0.26 to -0.13, 95% PI -0.56 to 0.17). Without additional guidance, agent-human agreement was moderate (ICC=0.608-0.629). The workbook improved agreement for both models, with DeepSeek-R1's increasing from 0.613 to 0.709 and o1-mini's from 0.629 to 0.687. In validation beyond rehabilitation, DeepSeek-R1 maintained stable agreement (ICC=0.711) and completed appraisals in 5.44 minutes compared to 11.18 minutes for humans.

conclusionsRehabilitation CPGs, particularly Chinese-language CPGs, continue showing deficiencies in applicability and stakeholder involvement. LLM-based appraisal without structured guidance provides insufficient agreement. Structured guidance improved agent-human agreement, supporting AI-assisted guideline appraisal under human oversight. Although further validation across additional clinical specialties is needed, AI agents can serve as efficient assistants in guideline appraisal instead of replacing humans. Future synthesis requires human-AI integration guided by structured, expert-defined principles.

trial registrationPROSPERO CRD420251270676; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251270676.

Indexed as

Artificial IntelligencePractice Guidelines as TopicHumansLarge Language ModelsReproducibility of ResultsAGREE IIlarge language modelspractice guidelinerehabilitationRIGHT

Identifiers

PMID42647856
PMCPMC13559151

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.