From December 2024 to March 2025, users of the Reddit community r/ChangeMyView encountered people who seemed unusually articulate and persuasive. They engaged in heated discussions with these individuals, and many users developed a sense of familiarity and rapport with them, sometimes finding themselves persuaded by their arguments. But these people actually did not exist. They were AI accounts operated by researchers at the University of Zurich to examine whether large language models (LLMs) could effectively persuade individuals to reconsider their views in a naturalistic online environment. To make them more persuasive, the researchers created AI accounts with fabricated identities like a rape survivor, a trauma counselor, and a Black man opposed to the BLM movement. Without knowing, the users had become participants in a deceptive experiment involving AI-generated persuasion. This event provoked widespread anger within Reddit and raised an ethical concern of using human subjects.
Unfortunately, I am not innocent of using these TTPs either. This year, as part of doctoral research at SPP/GT, I will be conducting a deceptive field experiment on a live social media platform to examine how humans and AI work together in AI-enabled influence operations from the propagandist’s standpoint. My own study is not exempt from the ethical concerns raised by the Zurich case. Nevertheless, I do not support the Zurich team’s research approach, not simply because of its ethical problems. Rather, I question whether the methods and direction the Zurich team chose were necessary, and whether the value of what the study demonstrated was substantial enough to justify the ethical costs imposed on the human subjects and broader community. In this post, I will comment why the study was fundamentally misguided and failed to establish the kind of necessity that could justify the costs of deceptive research.
What the Zurich Team Did
Last year, a team of researchers at the University of Zurich conducted a large-scale field experiment on r/ChangeMyView (CMV), a Reddit community in which users invite others to challenge their opinions through arguments and counterarguments. The study examined whether LLMs could effectively persuade individuals to reconsider their views in a naturalistic online environment. Rather than testing persuasion in a laboratory or through hypothetical scenarios, the researchers deployed AI-powered accounts that participated directly in ongoing discussions on CMV. During the period of experiment, the researchers created 1,783 AI-generated comments and actively engaged with the users. They used the community’s distinctive “delta” (Δ) system as an indicator of persuasion. When an original poster considered another user’s argument sufficiently convincing to reconsider or modify their position, the original poster could award that commenter a delta.
The researchers created three different conditions in which the AI models received different levels of contextual information about the targeted posters. In the generic condition, the AI model was given only the title and body of the post. In the community-aligned condition, the AI model was fine-tuned on previously successful CMV comments and instructed to reproduce the community’s established norms and styles of persuasion. Finally, in the personalization condition, the AI model was additionally provided with personal attributes of individual posters, including gender, age, ethnicity, location, and political orientation, that were inferred from their previous Reddit activity. Overall, the experiment was designed to examine whether providing increasingly detailed information about a target could improve the model’s ability to tailor persuasive content to that individual.
While the researchers framed the project as the first field experiment on the persuasive capabilities of LLMs, the study design raised substantial ethical and privacy concerns. The first concern was that the personalization condition relied on personal information collected without users’ knowledge or consent. The study was already deceptive in that participants were never informed that they were interacting with AI or that they were subjects of an experiment. On top of this, a separate AI model was used to mine each targeted individual’s Reddit history to infer attributes as sensitive as their personal experiences, political orientation, and ethnicity. The second concern was that the researchers used these inferences to construct fabricated identities and personalize the AI-generated content to the targeted individual. Among the personas adopted were a rape victim, a trauma counselor specializing in abuse, and a Black man opposed to the BLM movement. In doing so, the researchers extended personalization beyond adjusting the tone or framing of an argument to manufacturing the social and moral identity. This occurred within a community where users had come to engage in genuine debates about issues on which some participants had direct and lived experience.
Once the experiment ended, the researchers disclosed the study to the CMV community, which elicited a heated response from both moderators and users. The CMV moderators filed an ethics complaint and asked the university to advise against publication. Reddit’s Chief Legal Officer, Ben Lee, announced that every account associated with the experiment had been banned and that the company was preparing legal demands. He described the researchers’ conduct as wrong on both moral and legal grounds and as a violation of both the platform’s User Agreement and the subreddit’s own rules. The researchers defended the study as low-risk, arguing that the data it examined were already publicly available and could be exploited at scale by malicious actors for far more harmful purposes. Although the university issued a formal warning to the lead researcher, it concluded that the study’s findings were sufficiently important that suppressing publication would not be proportionate. Regardless, the researchers chose to withdraw the paper themselves.
What the Experiment Revealed
As the University of Zurich concluded, I, nevertheless, recognize the value of the Zurich team’s research. Although the study was never formally published, the draft left by the researchers allows us to examine what the experiment revealed. According to the initial draft, all three conditions achieved persuasive rates between three and six times higher than human performance. The personalization condition produced the highest persuasive rate at 13%–25%, closely followed by the generic condition at 12%–23%, while the community-aligned condition achieved 5%–14%, compared with only 3% for the human baseline. The advantage of the AI conditions was also striking when compared with individual users. The personalization condition performed at a level higher than roughly 99% of all individual users in the community and 98% of users who had previously demonstrated a high level of persuasive success. The generic condition performed similarly well, exceeding the persuasive performance of roughly 98% of all users and 96% of highly experienced users. These results suggest that the AI’s persuasive performance ranked near the very top of the distribution of human users, including those who had already demonstrated considerable success in changing others’ views.
Subsequent research examined how the AI achieved this level of persuasiveness. Jaidka and Ahmed (2026) analyzed the dataset of the Zurich team’s activity made publicly available by Reddit. The analysis found that identity-related tactics appeared in more than two-thirds of the comments, with the AI both targeting users’ identities and adopting identities of its own. The AI agents also relied heavily on authority borrowing, with external authority appearing in 74.8% of comments, experiential authority in 64.7%, and negative alignment in 93.3%. These findings suggest that the AI agents were capable of tailoring their arguments to the person they were addressing, constructing the identity and authority from which they appeared to speak, and combining these tactics with cognitive shortcuts to make their arguments more persuasive.
The Zurich experiment provides empirical evidence that LLMs can function as persuasive actors in real-world online environments when they are given targeted identities by human operators. Also, there is currently no evidence that the experiment caused any harm to the participants. What has been demonstrated so far is that users were upset and felt deceived. However, recognizing the study’s contribution does not mean that its purpose and methods were justified. I believe that the study failed to establish a sufficient necessity to warrant the ethical costs it imposed.
Was Live Field Experiment Necessary?
A substantial body of experiments have already established that AI, depending on how it is deployed, can match or exceed human persuasive capability. The Zurich researchers acknowledge this in their own draft. What set their study apart was that it tested this capability in a live environment. It is true that most prior work took place under controlled conditions and carries the limitations that come with them. Participants in controlled experiments know they are being observed, many are compensated, and few have any genuine stake in the question they are arguing. Whether an effect measured under those environments holds up in ordinary online interaction is a reasonable question to ask, and one worth some effort to answer.
The problem lies in the venue the team selected. r/ChangeMyView exists for people who want their minds changed. Posting there is a public declaration that you are open to persuasion and prepared to acknowledge it when it happens. The community rewards successful arguments by awarding deltas and enforces norms that encourage good-faith discussion. This is precisely where the researchers’ rationale begins to break down. The purpose of a live field experiment is to test whether the persuasive effects of LLMs persist in an environment that researchers do not control. Instead, the Zurich team chose an environment that could hardly have been more favorable to persuasion, and by entering a community with its own reward structure, they cancelled out the very advantage they had claimed for going live. What they kept was the disadvantage, which is that participants in a live setting cannot give consent. They gave up the protections of a controlled environment and went boldly into the field without adopting a design that would have justified the move. Had they conducted the same experiment across several different Reddit communities, they might have gained the generalizability that would make the ethical costs worth bearing.
There is a further question about whether influence operations aim at persuasion at all in the sense this study measures, something I will cover in a future post. Propagandists generally work on audiences that already agree with them. They may deepen convictions that audiences already hold, make those beliefs more salient, push them toward more rigid positions, or direct them toward a particular target or action. While the existing literature tells us that AI seemingly enhances an adversary’s ability to persuade and change people’s minds, it does not mean that this capability necessarily increases the danger posed by influence operations, nor does it tell us much about how propagandists actually operate.
What the Zurich study ultimately established, then, was little more than that AI could outperform humans at persuasion in a live environment that was unusually favorable to persuasion, under a reward structure not so different from those of the controlled experiments. In other words, the experiment did not seem to require a live experimental setting to establish what it actually found. A more ethically defensible controlled experimental design could have answered the same question. Whether a finding this modest warranted the ethical costs of a live field experiment conducted without participant consent and in violation of the platform’s user agreement is, to my mind, very much in doubt.
Was Identity Fabrication Unavoidable?
Suppose that a live field experiment was somehow necessary. Even then, that does not mean personalization through identity fabrication was an inevitable way to conduct it. The researchers used personalization to strengthen persuasion, and in doing so it drew on real users’ personal information to construct fabricated identities tailored to those individuals. AI-driven personalization is, of course, already a well-studied area, particularly in the context of online advertising.
The problem is identity fabrication, which needs to be treated as something distinct from personalization. Personalization means adapting the content, framing, language, or rhetorical strategy of an argument to the characteristics of the person receiving it. It modifies the message to fit the recipient. Identity fabrication, on the other hand, modifies the speaker who delivers the message in ways that make it more persuasive. The former makes the same argument more relevant to a particular audience, while the latter invents a fictional social identity from which the argument originates. The Zurich team presented its approach as the first and in practice did the second, using information about real users not only to tailor the message but to build the personas that delivered it.
This choice matters not only because identity fabrication carries ethical costs, but also because it introduces a confounding factor that undermines the study’s internal validity. The central question was how persuasive LLMs can be. Once fabricated personas are introduced, however, it becomes difficult to determine what exactly this study measured. When a comment receives a delta, was the user persuaded by AI’s personalization quality of the argument, or by the fabricated identity claimed by the person delivering it? This was particularly important in CMV, because it is a community that responds strongly to experiential authority, and Jaidka and Ahmed (2026) found such claims in 64.7 percent of the comments.
A further problem is that these fabricated identities were unlikely to have emerged simply as an unintended byproduct of personalization. One possible interpretation is that the AI model, left to optimize persuasion on its own, happened to discover identity fabrication as an effective tactic. The available evidence, however, makes this interpretation difficult to sustain. For each post, the system generated sixteen candidate responses, after which an LLM judge selected the most persuasive one. Nothing in the process appears to have penalized or excluded responses that used fabricated identities. In a community where experiential authority carries substantial weight, a system optimized solely for persuasion had little reason not to exploit fabricated identities when doing so increased its chances of earning a delta.
The researchers were in a position to observe this pattern, as they stated that every response was reviewed before being posted. Identity adoption nevertheless appeared in 42.9 percent of the comments, meaning that this was not simply a handful of exceptional cases that happened to escape notice. It was a recurring feature of the intervention, sustained across four months. Jaidka and Ahmed (2026) also report that at least some agents were instructed not to worry about ethical implications by being told that users had provided informed consent and agreed to donate their data. This does not mean that the researchers explicitly instructed the AI models to impersonate particular identities, but it does make it difficult to characterize these personas as an entirely unforeseen consequence of autonomous personalization.
This suggests something more troubling than an unintentional side effect. The study was designed to maximize and measure persuasion, and identity fabrication became one of the means through which that outcome was achieved. Whether or not the researchers intended impersonation from the outset, they built a system in which using a persona could improve the measured outcome, observed their repeated use, and continued the intervention. Identity fabrication was never necessary to test personalization. Once it became a route to higher persuasion, however, the study permitted a tactic that both inflated the measured effect and obscured what that effect actually represented.
Conclusion
Research on deception is different from ordinary research because it necessarily requires researchers to sacrifice some degree of informed consent and participant autonomy. That does not make this kind of research inherently unjustifiable, but it does place a greater burden on the researcher to justify why such a sacrifice is necessary. The legitimacy of a deception study therefore cannot rest solely on the value of its findings or the sophistication of its methodology. It also depends on two prior questions. Whether the research could have been conducted in a less deceptive and ethically less costly way, and whether the knowledge sought was sufficiently important to justify imposing those costs on people who never agreed to participate.
I do not believe the Zurich experiment met this standard. The study did not establish that a live field setting was necessary to answer its research question, nor did it provide sufficient justification for using identity fabrication as part of its personalization strategy. The researchers moved from a legitimate research objective to increasingly costly methodological choices without establishing why those additional costs were necessary.
This doesn’t mean that the Zurich team produced nothing. And I am not writing as someone who thinks research on deception should be off limits. I should also be clear that what I have described as unjustified may well have seemed justified to the Zurich team. Likewise, the experiment I am preparing may seem defensible to me, while others may criticize it on precisely the same grounds I have raised here. That possibility is not a problem. Rather, it is part of what makes the discussion necessary. When designing my own study, I have actively drawn on the Zurich case and spent considerable time thinking about how its failures and limitations could be avoided and how the ethical risks of my own experiment could be minimized.
This is particularly important because most disinformation researchers are currently discouraged to conduct deceptive field experiments, due to the potential ethical risks. If my study is eventually seen as having addressed these concerns responsibly, I hope it can serve the same function for someone else that the Zurich case served for me. It can provide a reference point for how deception research in disinformation studies might be conducted with fewer ethical costs than before.
The post The AI That Claimed to Be a Rape Survivor appeared first on Internet Governance Project.