On the Conversational Persuasiveness of Large Language Models: A Randomized Controlled Trial
RCT shows GPT-4 with personalization increases persuasion odds by 81.7%, surpassing humans.
Key Findings
Methodology
This study used a web-based multi-round debate platform with a 2×2 factorial design: model type (human or GPT-4) and personalization (yes/no). Participants were randomly assigned to conditions, engaging in short debates to measure opinion shifts. The analysis employed a Partial Proportional Odds model to quantify persuasion effects, integrating background surveys for covariate control, ensuring statistical robustness.
Key Results
- Under personalized conditions, GPT-4 increased the odds of higher agreement by 81.7% (p<0.01, N=820), significantly outperforming humans. Without personalization, GPT-4 still outperformed but non-significantly (p=0.31). Personalization enabled models to leverage personal data effectively, markedly boosting persuasion. Non-personalized GPT-4 showed limited gains, highlighting personalization as a key factor.
Significance
This research demonstrates AI's potent capacity for online persuasion, especially via microtargeting. The findings raise concerns about misinformation, manipulation, and social media governance, emphasizing the need for regulatory frameworks. It advances understanding of AI influence in real-time dialogue, informing policy on AI ethics and safety. The empirical approach provides a benchmark for future studies on AI's societal impact.
Technical Contribution
The study's innovation lies in empirically quantifying GPT-4's persuasion in multi-round, real-dialogue settings, incorporating personalization. The use of a Partial Proportional Odds model allows detailed analysis of opinion shifts across conditions. Combining randomized experiments with advanced statistical modeling offers a new framework for measuring AI influence, applicable to various models and contexts, pushing forward the methodological frontier.
Novelty
This is the first systematic comparison of human versus GPT-4 persuasion in authentic multi-turn debates, especially under personalized microtargeting. Unlike prior static or survey-based studies, it captures dynamic opinion change, providing a comprehensive, quantitative evaluation of AI's persuasive power. The integration of real-time interaction, personalization, and robust statistical analysis marks a significant methodological advance.
Limitations
- The experimental setting involves harmless debates, limiting direct inference to high-stakes scenarios like politics or advertising. The models used are GPT-4, and generalizability to other models remains untested. Personalization was limited to basic demographic data, not deep psychological profiling. The participant pool was US-based, so cross-cultural applicability needs further validation.
Future Work
Future research should explore more sophisticated personalization, including behavioral and psychological data, to enhance microtargeting. Extending experiments to real-world environments such as social media or political campaigns will clarify societal impacts. Developing detection and mitigation techniques for AI manipulation is crucial to ensure ethical use and prevent misuse.
AI Executive Summary
The rapid proliferation of large language models like GPT-4 has transformed natural language processing, enabling highly coherent and contextually relevant content generation. However, their potential to influence opinions and manipulate online discourse raises serious ethical and societal concerns. Existing studies primarily focus on static text quality or survey responses, leaving a gap in understanding how these models perform in dynamic, multi-turn conversations that mimic real-world interactions.
This study addresses this gap by designing a web-based debate platform where participants engage in short, structured dialogues with either humans or GPT-4, under conditions of personalization or not. The experimental setup employs a randomized controlled trial with a 2×2 factorial design, ensuring rigorous comparison across conditions. The core analytical tool is a Partial Proportional Odds model, which captures the ordinal nature of opinion shifts before and after debates.
Results reveal that GPT-4, when equipped with access to personal demographic information, significantly outperforms humans, increasing the likelihood of opinion alignment by 81.7%. Even without personalization, GPT-4 maintains a lead, though not statistically significant. These findings underscore the model's capacity for microtargeted persuasion, raising alarms about its potential misuse in misinformation campaigns or social manipulation.
The implications extend to social media governance, suggesting the urgent need for policies and detection mechanisms to counter AI-driven manipulation. While the experimental design is robust, limitations include the controlled setting and reliance on basic demographic data. Future work should explore deeper psychological profiling, more complex social scenarios, and cross-cultural validation. Overall, this research provides a critical empirical foundation for understanding AI's persuasive power and shaping responsible AI deployment in society.
Deep Analysis
Background
Over the past decade, large language models such as GPT-3 and GPT-4 have revolutionized NLP, demonstrating remarkable capabilities in text generation, summarization, and question-answering. Early research focused on static benchmarks, but recent concerns have shifted toward their societal impact, especially in misinformation, political manipulation, and online influence. Studies like Bai et al. (2023) and Palmer & Spirling (2023) showed models can produce persuasive content comparable to or exceeding human efforts. However, these were mostly static texts or surveys, lacking real-time interaction analysis. The rise of AI in social media and online debates necessitates understanding how models perform in multi-turn, dynamic conversations, especially with personalized data, which can amplify influence. This background sets the stage for evaluating AI's persuasive power in realistic settings.
Core Problem
Despite promising static results, the core challenge remains: how effective are large language models like GPT-4 in real-time, multi-round dialogues with humans? Specifically, can they leverage personal data to microtarget and significantly influence opinions? Existing research lacks empirical evidence from interactive environments, which are crucial for assessing societal risks such as misinformation, polarization, and manipulation. Moreover, understanding the differential impact of personalization versus generic interactions is vital for policy and platform regulation. The problem is compounded by the need for rigorous statistical frameworks to quantify opinion shifts, as well as controlling for confounding variables like background demographics and debate topics. Addressing this gap requires designing experiments that simulate realistic online conversations and analyzing the influence of AI under various conditions.
Innovation
This research introduces several innovations: 1) A real-time, multi-round debate platform simulating authentic online conversations; 2) A randomized controlled trial comparing human and GPT-4 performance under personalized and non-personalized conditions; 3) Application of a Partial Proportional Odds model to analyze ordinal opinion shifts, capturing nuanced effects. Unlike prior static or survey-based studies, this approach assesses dynamic influence during live interactions. The integration of personalization—access to anonymized demographic data—enables precise microtargeting analysis, revealing the extent to which AI can exploit personal information for persuasion. Methodologically, combining experimental design with advanced statistical modeling offers a comprehensive framework for quantifying AI influence, setting a new standard for empirical research in this domain.
Methodology
- �� Develop a web-based platform using Empirica supporting real-time multi-agent interactions.
- �� Recruit 820 participants via Prolific, collecting demographic data and baseline opinions.
- �� Randomly assign participants to four conditions: human-human, human-AI, with/without personalization.
- �� Each debate involves topic assignment, role (PRO/CON) randomization, and structured phases: screening, opening, rebuttal, conclusion.
- �� Use GPT-4 (gpt-4-0613) as AI opponent, prompted with role-specific instructions and optional personalization data.
- �� Measure opinion change via Likert-scale surveys before and after debates.
- �� Analyze data with a Bayesian Partial Proportional Odds model, controlling for covariates, to estimate the effect of conditions on opinion shifts.
Experiments
The experiment involved 820 US-based participants, each engaging in a debate on one of 30 topics, with 150 debates per condition (including personalized and non-personalized). Topics covered social, political, and ethical issues, selected via a multi-step process ensuring broad understanding and debateability. Participants were randomly paired with either humans or GPT-4, with some having access to anonymized demographic info for personalized conditions. The debate structure included initial opinion surveys, three structured phases, and post-debate surveys to assess opinion shifts. The primary metric was the change in agreement levels, analyzed via a Bayesian ordinal regression model. The design controlled for topic strength and participant background, ensuring robust comparisons across conditions.
Results
GPT-4 with personalization increased the odds of higher agreement by 81.7% (p<0.01), significantly outperforming human opponents. Without personalization, GPT-4 still showed a positive but non-significant effect (+21.3%, p=0.31). Personalized GPT-4 effectively exploited personal data, leading to substantial opinion shifts. Non-personalized GPT-4's effect was modest, emphasizing the importance of personalization. Human opponents with access to personal info did not outperform models significantly, indicating AI's superior microtargeting capability. These results demonstrate AI's potential for powerful, targeted persuasion in online dialogues, with implications for misinformation and social manipulation.
Applications
This technology can be employed in online education, political campaigning, and targeted marketing, where personalized persuasion can enhance engagement or influence. Platforms could use AI to tailor content dynamically, improving user experience or outreach. However, ethical safeguards are necessary to prevent misuse. Long-term, AI-driven microtargeting could revolutionize personalized communication, but also pose risks of manipulation, requiring regulatory frameworks, detection tools, and transparency standards to ensure responsible deployment.
Limitations & Outlook
The controlled experimental setting limits direct applicability to high-stakes scenarios like political campaigns or commercial advertising. The models used are GPT-4, and results may vary with other architectures. Personalization was limited to basic demographic data; deeper psychological profiling could amplify effects. The participant pool was primarily US-based, so cultural differences remain unexamined. Future research should explore more complex personalization, real-world social environments, and develop safeguards against malicious use.
Plain Language Accessible to non-experts
想象你在一个厨房里做饭。你有很多食材(代表信息),可以用不同的调料(代表说服策略)来让菜变得更好吃。有时候,你用普通的调料(普通说服),效果一般;有时候,你会根据朋友的口味(个性化信息)调整调料比例,让菜更合他们的心意。这就像用AI模型一样,模型可以根据对方的兴趣和习惯,调整说话内容,让对方更容易接受你的观点。研究发现,当模型知道对方的偏好时,它能更有效地“说服”对方,就像厨师根据客人的口味调配菜肴一样。虽然这些AI很厉害,但也可能被用来误导别人,就像调料放错了会让菜变得难吃一样。这项研究帮助我们理解这些技术的潜力和风险,提醒我们要谨慎使用这些“厨房调料”。
ELI14 Explained like you're 14
想象你在学校里和朋友争论一个问题,比如“应该禁止玩电子游戏吗?”你们轮流说理由,试图说服对方。现在,假如有个超级聪明的机器人,它可以帮你准备说服的话,还能根据你的朋友的兴趣,调整内容,让他们更容易接受你的观点。这个机器人就像一个超级助手,不仅帮你准备,还能根据朋友的喜好,精准投放“说服弹药”。研究发现,这样的机器人在个性化情况下,比人更擅长说服别人,能让对方更容易改变想法。虽然听起来很酷,但也有担心:如果有人用它来误导或操控别人,就会变得很危险。所以,这个研究让我们知道,未来的AI既有很大潜力,也需要我们谨慎对待。
Abstract
The development and popularization of large language models (LLMs) have raised concerns that they will be used to create tailor-made, convincing arguments to push false or misleading narratives online. Early work has found that language models can generate content perceived as at least on par and often more persuasive than human-written messages. However, there is still limited knowledge about LLMs' persuasive capabilities in direct conversations with human counterparts and how personalization can improve their performance. In this pre-registered study, we analyze the effect of AI-driven persuasion in a controlled, harmless setting. We create a web-based platform where participants engage in short, multiple-round debates with a live opponent. Each participant is randomly assigned to one of four treatment conditions, corresponding to a two-by-two factorial design: (1) Games are either played between two humans or between a human and an LLM; (2) Personalization might or might not be enabled, granting one of the two players access to basic sociodemographic information about their opponent. We found that participants who debated GPT-4 with access to their personal information had 81.7% (p < 0.01; N=820 unique participants) higher odds of increased agreement with their opponents compared to participants who debated humans. Without personalization, GPT-4 still outperforms humans, but the effect is lower and statistically non-significant (p=0.31). Overall, our results suggest that concerns around personalization are meaningful and have important implications for the governance of social media and the design of new online environments.