Evidence map›Paper›PMID 42428122›Full record

ArticlemedRxiv : the preprint server for health sciences2026

Proposed Context-of-Use Evaluation Framework for Medication Management Tasks Completed by Generative Artificial Intelligence.

Kelli Henry, Kaitlin Blotske, Brooke Smith, Tianle Li, Yanjun Gao, Xingmeng Zhao, Tianming Liu, Andrea Sikora

Abstract readPreprint
In one paragraph

Article in medRxiv : the preprint server for health sciences, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

8 authors.

Kelli HenryUniversity of Colorado School of Medicine, Department of Biomedical Informatics, Aurora, CO.ORCID 0000-0002-6686-1079
Kaitlin BlotskeUniversity of Colorado School of Medicine, Department of Biomedical Informatics, Aurora, CO.ORCID 0000-0002-2977-0266
Brooke SmithWellstar MCG Health, Department of Pharmacy, Augusta, GA.ORCID 0000-0003-3356-0648
Tianle LiATLAS Institute, University of Colorado Boulder, Boulder, CO, USA.
Yanjun GaoUniversity of Colorado School of Medicine, Department of Biomedical Informatics, Aurora, CO.ORCID 0000-0002-9341-7360
Xingmeng ZhaoUniversity of Colorado School of Medicine, Department of Biomedical Informatics, Aurora, CO.ORCID 0000-0002-7346-758X
Tianming LiuDepartment of Computer Science, University of Georgia, Athens, GA.ORCID 0000-0002-8132-9048
Andrea SikoraUniversity of Colorado School of Medicine, Department of Biomedical Informatics, Aurora, CO.ORCID 0000-0003-2020-0571

Funding

Developing and Evaluating Multi-Modal Clinical Diagnostic Reasoning Models for Automated Diagnosis GenerationR00LM014308 · NLM · UNIVERSITY OF COLORADO DENVER · PI Yanjun Gao · 2024 to 2026
$726k
AHRQ HHS R01 HS029009AHRQ HHS R21 HS028485NLM NIH HHS R00 LM014308
6 · The paper itself

Abstract

Background: Standardized evaluation of agentic artificial intelligence (AI) for medication management is lacking. Given the potential lethality of medication errors endorsed or missed by AI, performance evaluation constructs are essential. The purpose of this evaluation was to develop a standardized grading framework for performance evaluation of medication management tasks. Methods: A mixed-methods approach was undertaken that included literature evaluation for standards and best practices of comprehensive medication management (CMM), panel discussions, and iterative application to set of cases. The goal was to develop a grading framework that effectively evaluated domains like safety, factuality, and clinical relevance that can be employed for a broad range of medication domains (i.e., electrolyte replacement, antibiotic selection). Inter-rater reliability with intraclass Krippendorff's Alpha was the primary outcome. Results: A total of 5 panelists developed the CMM Evaluation Framework, which includes 4 dimensions: safety, factuality, completeness, and preference. These dimensions are applied to three CMM skills: collecting patient data, analyzing information, and designing regimens. Each dimension is rated from 1-5. An additional dimension evaluated the presence of hallucinations and errors with high harm scores (i.e., "absolute failure" criteria regardless of an overall score). The Krippendorff's Alpha was highest in the medication therapy problem and medication therapy format categories, for 50 pneumonia cases, run in triplicate (150 total). Conclusions: This framework is informed by national standards for CMM and the healthcare professionals dedicated to the provision of this service. These domains allow for the possibilities of practice variation via the preference domain while also having strong guardrails against the commission of medication errors. Further analyses beyond pilot testing are necessary.

Indexed as

Artificial intelligencebenchmarklarge language modelsmedication safety

Identifiers

PMID42428122
PMCPMC13345436

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.