Evidence map›Paper›PMID 41950508›Full record

ArticleJournal of medical Internet research2026

The Alberta Quality Assessment Tool: Risk of Bias (AQAT:RoB) for the Evaluation of Medical Large Language Model Question-Answer Studies: Development and Pilot Validation.

Carrie Ye, Joseph Ross Mitchell, Daniel C Baumgart, Zechen Ma, Angela Lim Fung, Daniela Garcia Orellana, Juel Chowdhury, Abdullah Abass, Steven Katz, Jacob L Jaremko and 15 more

Abstract readValidation Study
In one paragraph

Article in Journal of medical Internet research, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

25 authors.

Carrie YeUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-9858-5398
Joseph Ross MitchellUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-2340-4708
Daniel C BaumgartUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0003-2146-507X
Zechen MaUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0003-3064-8118
Angela Lim FungUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0009-0004-2904-0521
Daniela Garcia OrellanaUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0009-0006-4813-3541
Juel ChowdhuryUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0003-4329-2392
Abdullah AbassUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0009-0009-0643-0945
Steven KatzUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0001-6718-1393
Jacob L JaremkoUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0001-5314-2297
Pierre BoulangerUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-4219-1699
Claire E H BarberUniversity of Calgary, Calgary, AB, Canada.ORCID 0000-0002-3062-5488
Gillian LemermeyerUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0001-5658-3573
Hosna JabbariUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-7155-2297
Lili MouUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0001-7753-4295
Maryam MirzaeiNAIT Applied Research, Edmonton, AB, Canada.ORCID 0000-0001-5309-9786
Mary Waithera Beckett GithumbiUniversity of Toronto, Toronto, ON, Canada.ORCID 0009-0007-6324-7237
Puneeta TandonUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0003-0486-0174
Randy GoebelUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-0739-2946
Rhys ClarkAlberta Health Services, Edmonton, AB, Canada.ORCID 0009-0002-1162-0259
Whitney HungAlberta Health Services, Edmonton, AB, Canada.ORCID 0000-0001-9273-8798
Marjan AbbasiUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0001-7252-6798
Farhad MalekiUniversity of Calgary, Calgary, AB, Canada.ORCID 0000-0002-5673-8210
Scott KlarenbachUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-8611-1056
Mohamed AbdallaUniversity of Alberta, 8-130 Clinical Sciences Building, 11350 83 Ave NW, Edmonton, AB, T6G2G3, Canada, 1 7804927002, 1 7804926088.ORCID 0000-0002-2776-6036

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Despite the transformative potential of large language models (LLMs) in health care, the rapid development of these tools has outpaced their rigorous evaluation. While artificial intelligence-specific reporting guidelines have been developed to address standardized reporting of artificial intelligence studies, there is currently no specific tool available for risk of bias assessment of LLM question-answer (QA) studies. Existing risk-of-bias tools for medical research are not well suited to the unique challenges of evaluating LLM-QA studies, which creates a critical gap in assessing their safety and effectiveness. Objective: This study aims to develop the Alberta Quality Assessment Tool: Risk of Bias (AQAT:RoB) for LLM-QA studies to systematically evaluate the validity and risk of bias in LLM-QA studies. Methods: We conducted 2 literature reviews. The first was on quality assessment tools for LLM-QA studies, and the second was on LLM-QA studies, which informed the first draft of the AQAT:RoB. The draft AQAT:ROB was further refined through a prespecified iterative process of modified Delphi, consensus meeting, and validation. The first Delphi process occurred between May 1 and May 20, 2025, and the first consensus meeting was held on May 22. The first round of validation was completed by 4 evaluators, who were not part of the consensus meeting, on 16 randomly selected studies. As this first round of validation surpassed our a priori threshold of ≥80% agreement and a Cohen κ of ≥0.61 between evaluators, no further rounds of development and validation were undertaken. A second Delphi process occurred between February 20 and February 23, 2026, to vote on postpilot changes in response to peer review. Results: The AQAT:RoB consists of 5 high-level domains (Questions, Reference Answers, LLM Answers, Evaluators, Outcomes). These domains are subdivided into 9 subdomains. Each subdomain includes at least one "Support for Judgment" and at least one "Type of Bias" and is to be rated "low," "high," or "unclear" for risk of bias. A pilot evaluation was completed by internal validators who were not part of the consensus discussion and were asked to complete the AQAT:RoB form for each assigned study. Each of the 16 studies was evaluated by 2 evaluators independently. Pilot validation showed a percent agreement of 86.1% and a Cohen κ of 0.70 between assessors. Conclusions: The AQAT:RoB demonstrates promising initial reliability for assessing the validity or risk of bias in LLM-QA studies. The tool will benefit from future refinements, external validation, and periodic updates to keep pace with evolving technology.

Indexed as

BiasEducational MeasurementLarge Language ModelsPilot ProjectsAlberta Risk of Bias Assessment Tool for LLM-QA studiesAQAT: RoBartificial intelligencechatbotlarge language modelquality assessmentquestion-answer studiesrisk of bias

Identifiers

PMID41950508
PMCPMC13061365

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.