Evidence map›Paper›PMID 41554116›Full record

ArticleJournal of medical Internet research2026

Developing a Quality Evaluation Index System for Health Conversational Artificial Intelligence: Mixed Methods Study.

Weizhen Liao, Meng Li, Chengyu Ma, Youli Han, Dan Wang, Haopeng Liu, Yi Wang, Zijie Feng, Huichao Wang, Yiru Guan

Abstract read
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

10 authors.

Weizhen LiaoSchool of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0009-0001-1775-9967
Meng LiBytedance Xiaohe Health, Hainan, China.ORCID https://orcid.org/0009-0000-3545-8490
Chengyu Ma *School of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0000-0002-2458-8422
Youli Han *School of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0000-0001-8882-3500
Dan WangBytedance Xiaohe Health, Hainan, China.ORCID https://orcid.org/0009-0009-4303-3738
Haopeng LiuSchool of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0009-0003-4017-577X
Yi WangSchool of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0009-0004-0238-9598
Zijie FengSchool of Public Health, Capital Medical University, Beijing, China.ORCID https://orcid.org/0009-0006-2190-4350
Huichao WangBytedance Xiaohe Health, Hainan, China.ORCID https://orcid.org/0009-0006-6561-6414
Yiru GuanBytedance Xiaohe Health, Hainan, China.ORCID https://orcid.org/0009-0001-8575-7648

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundEffective communication is fundamental to health care; however, demographic transitions and a widening global health workforce gap are intensifying the imbalance between service demand and resource supply. Health conversational artificial intelligence (HCAI) based on large language models offers a potential pathway to improve the accessibility and personalization of care. Nevertheless, the lack of a rigorous, user-centered evaluation framework limits the systematic assessment of HCAI quality, raising concerns regarding safety, reliability, and clinical applicability.

objectiveThis study aims to establish a scientific and systematic quality evaluation index system for HCAI, providing both a theoretical foundation and a practical tool for the assessment and optimization of HCAI.

methodsBased on a literature review, industry standards, and expert group discussions, a preliminary framework for the index system was established. Two rounds of Delphi expert consultations were then conducted to collect expert opinions. The analytic hierarchy process (AHP) was applied to assign weights to indicators at each level, and the final content and structure of the index system were determined.

resultsBoth rounds of expert consultation achieved a 100% response rate. The authority coefficient of the experts was 0.84 in both rounds. Kendall W coefficient ranged from 0.14 to 0.20 in the first round and from 0.13 to 0.17 in the second round, with all values showing statistical significance (round one: importance P<.001, feasibility P<.001, sensitivity P<.001; round two: importance P=.001, feasibility P<.001, sensitivity P=.001). The final HCAI quality evaluation index system consisted of 3 primary indicators, 7 secondary indicators, and 28 tertiary indicators. According to AHP weight calculations, the primary indicators were ranked in descending order as follows: ethics and compliance (0.4781), health consultation capability (0.4112), and user experience (0.1107).

conclusionsThe evaluation index system constructed in this study demonstrates scientific validity and practical relevance. It provides a valuable reference for the quality assessment, model optimization, and regulatory oversight of HCAI systems.

Indexed as

Artificial IntelligenceCommunicationDelphi TechniqueHumansanalytic hierarchy processconversational AIDelphi methodevaluation index systemhealth consultation

Identifiers

PMID41554116
PMCPMC12865354

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.