Evidence map›Paper›PMID 42218453›Full record

ArticleBMC medical education2026

Artificial intelligence-assisted feedback in pharmacology education: a pilot evaluation of a custom generative model.

Edward Stephenson, Katie Morrigan, Ayesha Irfan, Kate Bascombe, Michael Okorie

Abstract read
In one paragraph

Article in BMC medical education, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

5 authors.

Edward StephensonDepartment of Medical Education, Brighton & Sussex Medical School, Brighton, UK. e.stephenson@bsms.ac.uk.
Katie MorriganDepartment of Medical Education, Brighton & Sussex Medical School, Brighton, UK.
Ayesha IrfanDepartment of Medical Education, Brighton & Sussex Medical School, Brighton, UK.
Kate BascombeDepartment of Medical Education, Brighton & Sussex Medical School, Brighton, UK.
Michael OkorieDepartment of Medical Education, Brighton & Sussex Medical School, Brighton, UK.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundTimely, individualised feedback is central to effective learning in medical education but remains resource intensive, particularly when based on short answer questions (SAQs). Large language models (LLMs) offer potential to support feedback provision, yet prospective evaluations of their accuracy and educational value remain limited. This pilot study evaluated the feasibility, accuracy and student perceptions of AI-generated pharmacology SAQ feedback.

methodsA prospective pilot study was conducted within a UK Physician Associate MSc programme between March and October 2025. Students were invited to complete voluntary four-question formative pharmacology SAQs. A bespoke custom Generative Pre-trained Transformer (GPT), built using GPT-4 family LLM, guided by a standardised rubric and references, generated structured feedback. All outputs underwent faculty moderation prior to release. Student perceptions were assessed using a 5-point Likert-scale survey. 25% of student responses were double marked and independently reviewed in a blinded comparison of AI and faculty feedback.

resultsTwenty students submitted complete responses, generating feedback for a total of 80 questions. Mean AI feedback generation time was 34 s per quiz, compared with 508 s for faculty marking, representing a 15-fold reduction. Twelve of 20 (60%) feedback files required no modification before release, while eight required amendments, including four major corrections, even with faculty modification there was substantial efficiency gains from AI generated feedback. Eleven students (55%) completed the survey, reporting favourable perceptions of clarity, actionability, confidence and overall usefulness (median 4 out of 5) for the moderated feedback. In an exploratory blinded review of 5 double-marked submissions, no statistically significant differences were detected between AI-generated and faculty-generated feedback across rated domains (p > 0.23 for all domains).

conclusionsA custom GPT-based LLM delivered rapid, structured pharmacology feedback with substantial efficiency gains and positive student perceptions. However, clinically important errors occurred, necessitating consistent faculty oversight. Generative models may augment formative assessment in medical education, but rigorous calibration, moderation and ethical safeguards remain essential for safe implementation.

Indexed as

Educational MeasurementFormative FeedbackPharmacologyGenerative Artificial IntelligenceHumansLarge Language ModelsPilot ProjectsProspective StudiesUnited KingdomAssessment feedbackGenerative AILarge language modelsPharmacology education

Identifiers

PMID42218453
PMCPMC13435439

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.