Evidence map›Paper›PMID 40448034›Full record

ArticleBMC medical research methodology2025

Evaluating the performance of artificial intelligence in summarizing pre-coded text to support evidence synthesis: a comparison between chatbots and humans.

Kim Nordmann, Stefanie Sauter, Mirjam Stein, Johanna Aigner, Marie-Christin Redlich, Michael Schaller, Florian Fischer

Abstract readComparative Study
In one paragraph

Article in BMC medical research methodology, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers, 1 of them a synthesis that pooled it.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed, 1 pooled it
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed, 1 synthesis or guideline pooled it.

  1. Pooled it
  2. Article
  3. AI for scientific discovery is a social problem.Patterns (New York, N.Y.) · 2026
    Review
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Kim NordmannKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Stefanie SauterKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Mirjam SteinKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Johanna AignerKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Marie-Christin RedlichKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Michael SchallerKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany.
Florian FischerKempten University of Applied Sciences, Bavarian Research Center for Digital Health and Social Care, Kempten, Germany. florian.fischer@hs-kempten.de.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundWith the rise of large language models, the application of artificial intelligence in research is expanding, possibly accelerating specific stages of the research processes. This study aims to compare the accuracy, completeness and relevance of chatbot-generated responses against human responses in evidence synthesis as part of a scoping review.

methodsWe employed a structured survey-based research methodology to analyse and compare responses between two human researchers and four chatbots (ZenoChat, ChatGPT 3.5, ChatGPT 4.0, and ChatFlash) to questions based on a pre-coded sample of 407 articles. These questions were part of an evidence synthesis of a scoping review dealing with digitally supported interaction between healthcare workers.

resultsThe analysis revealed no significant differences in judgments of correctness between answers by chatbots and those given by humans. However, chatbots' answers were found to recognise the context of the original text better, and they provided more complete, albeit longer, responses. Human responses were less likely to add new content to the original text or include interpretation. Amongst the chatbots, ZenoChat provided the best-rated answers, followed by ChatFlash, with ChatGPT 3.5 and ChatGPT 4.0 tying for third. Correct contextualisation of the answer was positively correlated with completeness and correctness of the answer.

conclusionsChatbots powered by large language models may be a useful tool to accelerate qualitative evidence synthesis. Given the current speed of chatbot development and fine-tuning, the successful applications of chatbots to facilitate research will very likely continue to expand over the coming years.

Indexed as

Artificial IntelligenceGenerative Artificial IntelligenceHumansSurveys and QuestionnairesArtificial intelligenceChatbotChatFlashChatGPTLarge language modelZenoChat

Identifiers

PMID40448034
PMCPMC12123790

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.