Trial reportNature human behaviour2025
On the conversational persuasiveness of GPT-4.
Trial report in Nature human behaviour, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. An erratum has been issued. Cited by 23 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
23 citing papers in PubMed.
- Disentangling interaction and bias effects in opinion dynamics of large language models.Nature communications · 2026Article
- Review
- Article
- Persuading large language models to comply with objectionable requests.Proceedings of the National Academy of Sciences of the United States of America · 2026Article
- When collaboration fails: persuasion driven adversarial influence in multi agent large language model debate.Scientific reports · 2026Article
- It Is the Journey, Not the Destination: Moving From End Points to Trajectories When Assessing Chatbot Mental Health Safety.JMIR mental health · 2026Article
- Large language models have the potential to level the playing field in consumer financial complaints.Nature human behaviour · 2026Article
- The adoption and efficacy of large language models in US consumer financial complaints.Nature human behaviour · 2026Article
- Comparison of Emotional Content in Text Responses From Physicians and AI Chatbots to Patient Health Queries: Cross-Sectional Study.Journal of medical Internet research · 2026Article
- How latent and prompting biases in AI-generated historical narratives influence opinions.PNAS nexus · 2026Article
- Large reasoning models are autonomous jailbreak agents.Nature communications · 2026Article
- Words matter: measuring and titrating the communicative character of AI.Frontiers in artificial intelligence · 2026Article
- Metacognition of ChatGPT in confidence judgements.Frontiers in artificial intelligence · 2026Article
- The missing discipline in AI: a call for behavioural science.Wellcome open research · 2026Article
- Generative AI for climate governance and acceptability-constrained policy design.npj climate action · 2026Review
- Never say never: exploring the effects of knowledge availability on agent persuasiveness in controlled physiotherapy motivation dialogues.Frontiers in artificial intelligence · 2026Article
- A meta-analysis of the persuasive power of large language models.Scientific reports · 2025Article
- The impact of advanced AI systems on democracy.Nature human behaviour · 2025Review
- Article
- Addressing Autonomy Risks in Generative Chatbots with the Socratic Method.Science and engineering ethics · 2025Article
Corrections and comments
- Erratum issued
Authors and funding
4 authors.
Funding
Abstract
Early work has found that large language models (LLMs) can generate persuasive content. However, evidence on whether they can also personalize arguments to individual attributes remains limited, despite being crucial for assessing misuse. This preregistered study examines AI-driven persuasion in a controlled setting, where participants engaged in short multiround debates. Participants were randomly assigned to 1 of 12 conditions in a 2 × 2 × 3 design: (1) human or GPT-4 debate opponent; (2) opponent with or without access to sociodemographic participant data; (3) debate topic of low, medium or high opinion strength. In debate pairs where AI and humans were not equally persuasive, GPT-4 with personalization was more persuasive 64.4% of the time (81.2% relative increase in odds of higher post-debate agreement; 95% confidence interval [+26.0%, +160.7%], P < 0.01; N = 900). Our findings highlight the power of LLM-based persuasion and have implications for the governance and design of online platforms.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.