ArticleJMIR medical informatics2026
A Controlled Comparison of Human and AI-Assisted Automated Revision of Delphi Statements on RNA-Based Medicines: Parallel, 2-Arm Study.
Article in JMIR medical informatics, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
6 authors.
Funding
No grant is acknowledged in the PubMed record.
Abstract
backgroundThe Delphi method is widely used to derive expert consensus on complex clinical problems, yet it is slow and resource intensive. Recent advances in large language models and retrieval‑augmented generation (RAG) offer the possibility of accelerating consensus while maintaining methodological rigor. Large language models can retrieve and summarize evidence, but they frequently hallucinate and cannot reliably cite sources. At the same time, RNA‑based drugs and messenger RNA vaccines are rapidly moving from concept to clinic, generating a pressing need for timely, evidence‑based consensus on regulatory, manufacturing, and clinical issues.
objectiveWe evaluated whether a modular, RAG‑enabled, multi‑agent artificial intelligence (AI) pipeline could replicate the post-round 1 behavior of a human reviewer in a Delphi study. The primary objective was to determine whether AI‑assisted statement revision could rescue a greater proportion of subthreshold statements and achieve a consensus comparable to that obtained through human revision by round 2.
methodsA parallel, 2‑arm Delphi study was conducted on 28 statements about RNA medicines. In total, 50 international panelists (clinicians, researchers, and patient representatives) were randomized into human (arm A) and AI‑assisted (arm B) groups. After round 1, statements below the 75% agreement threshold were revised either manually by human reviewers or by an AI pipeline comprising the following software agents: (1) ReferenceDetector, to identify external citations; (2) Summarizer, to produce structured summaries of supporting PDFs; (3) a hybrid RAG module that combined dense and sparse retrieval with cross‑encoder reranking; and (4) Refiner, which generated revised statements, reasoning logs, and explicit citations. Two human reviewers with expertise in Delphi methodology and literature review who had contributed to statement development verified retrieved citations and approved or amended revisions. Agreement rates and vote distributions were compared across arms.
resultsArm A reached consensus on 71.4% (20/28) of the statements in round 1, whereas arm B reached consensus on 46.4% (13/28). After revision, consensus increased to 92.9% (26/28) of the statements in arm A and 85.7% (24/28) in arm B. The AI arm exhibited a larger mean improvement (absolute difference between rounds 1 and 2=39.3 percentage points) because more statements were initially below the threshold. Nonetheless, the absolute difference between arms after round 2 was modest (7.2 percentage points). AI‑assisted revisions were particularly effective for statements far below the threshold, but both arms failed to rescue 2 to 3 statements owing to substantive disagreements.
conclusionsA modular, citation‑anchored AI pipeline can closely approximate human performance in Delphi consensus procedures while substantially reducing manual workload. When paired with human oversight, AI assistance accelerated revision and closed most of the performance gap by the second round. Adoption of AI‑assisted workflows could accelerate consensus development on emerging technologies such as RNA therapeutics provided that transparency, rigorous retrieval, and human review are maintained.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.