ArticleJournal of medical Internet research2026
Authoritative Textbook-Augmented Large Language Models for High-Altitude Public Health Medical Education in the Xizang Autonomous Region: Cross-Sectional Comparative Evaluation Study.
Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
31 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
Background: Public health medical education is increasingly important in the low-resource, high-altitude Xizang Autonomous Region (Tibet). Traditional authoritative textbooks do not meet modern needs for accessibility and interactivity, whereas general large language models (LLMs) may hallucinate in specialized medical domains. Developing specialized LLMs for low-resource regions is also expensive and difficult. Objective: This study aimed to explore a novel approach to high-altitude public health medical education in the low-resource Xizang Autonomous Region that integrates modern LLMs and authoritative textbooks, using a comprehensive benchmark evaluation across multiple dimensions and retrieval-augmented generation (RAG) technology. Methods: We conducted a 2-stage cross-sectional comparative evaluation study to benchmark publicly available LLMs and evaluate the added value of textbook-augmented retrieval under standardized generation settings and blinded expert assessment. First, 4 publicly available LLMs (GPT-5.2 [OpenAI], Gemini 3.0 Pro [Google], DeepSeek R1 [DeepSeek], and Tencent HY 2.0 [Tencent]) were benchmarked using an 80-question benchmark on high-altitude public health medicine developed by authoritative medical specialists. Each question was asked 3 times, yielding 960 outputs; first responses (n=320) were scored under blinded conditions by 2 independent 8-member physician panels. A clinically weighted evaluation of multidimensional first-response scores (including comprehensiveness, accuracy, clarity, and relevance) and a composite consistency metric (including semantic similarity and algorithmic similarity) was administered. Second, 4 specific and prevalent authoritative textbooks on high-altitude public health medicine-Ward, Milledge and West's High Altitude Medicine and Physiology, High Altitude Medicine: A Case-Based Approach, High Altitude Medicine, and High Altitude Medical Protection-were deployed as the external knowledge base for the evaluation-optimized model. Statistical analyses included Spearman ρ, Cronbach α, intraclass correlation coefficients, Friedman tests with Dunn multiple comparisons, and paired Wilcoxon signed-rank tests. The significance threshold was set at α=.05. Results: DeepSeek R1 was selected as the optimal base model for achieving the highest weighted score (5.61/10.00), followed by GPT-5.2 (5.51/10.00), Gemini 3.0 Pro (5.39/10.00), and Tencent HY 2.0 (4.71/10.00). The deployed retrieval-augmented model integrating the authoritative textbooks and the optimal LLM DeepSeek R1, HPHME-Xplus-RAG, achieved remarkable improvement in multidimensional scores compared to baseline DeepSeek R1 (median 8.00 [IQR 7.88-8.00] vs median 7.63 [IQR 7.38-7.88]; P<.001, r_rb=0.68, indicating a large effect). Conclusions: Integrating authoritative textbooks with an evaluation-optimized general LLM through an RAG framework showed strong performance for medical education in the low-resource Xizang Autonomous Region. Unlike prior studies that mainly evaluated general LLMs or used clinical guidelines to build RAG systems for diagnosis and treatment, this study used authoritative textbooks for the broader, guideline-scarce field of public health medical education. This work provides a replicable workflow-domain-authoritative knowledge+RAG+model optimization and evaluation-for low-resource settings, with practical implications for medical instructors and students, hospitals, and public health services seeking cost-effective, convenient, and trustworthy educational support.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.