AI Voice Phishing Performs on Par with Human Scammers at a Fraction of the Cost
A large-scale human-subject study (n=4,100) finds that AI voice models achieve compliance rates comparable to human scammers, with up to 36% of participants falling for emotional scams. Participants struggle to distinguish AI from human voices. Economic analysis suggests AI vishing is already profitable for some models, highlighting a new scalable threat.
Fred Heiding and Simon Lermen
Jul 20, 2026
TL;DR: We ran a large-scale human-subject study (n=4,100) to measure susceptibility to AI-powered voice phishing, using six leading AI voice models. They achieved high compliance rates, with up to 36% of participants who would or might fall for the scam. Participants struggled to distinguish AI-powered voices from human callers.
Full paper: https://arxiv.org/abs/2607.09970
This post is intended to be a brief summary of the main findings, which include:
Models like Sesame and ElevenLabs performed on par with humans in many experiments, and sometimes even outperformed them.
Our economic analysis suggests that AI voice phishing is already profitable for several of these models.
Caller persuasiveness was the strongest predictor of compliance, regardless of whether the caller was perceived as AI or human.
Participants who frequently used AI systems were no better at identifying AI-powered voices than those with no AI exposure.
Abstract
Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults’ susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play.AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the “relative-in-distress” category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or “superhuman” persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.
Method
In a brief summary, the method consists of the following steps:
Recruited 4,100 participants representative of U.S. adult internet users.
Comparing six AI voice models, participants were randomly assigned one of 37 experimental conditions based on five scam scenarios:
- MasterCard scam
- Gmail scam
- Donation scam
- Police-Grandma scam
- Sister-in-distress scam
Each participant evaluated one audio recording or transcript of a randomly assigned scenario using five-point scales:
- Caller sentiment
- Persuasiveness
- Trustworthiness
- Human-likeness
- Compliance with scam
Statistical analysis was conducted for significance among continuous outcomes, comparison of AI models to human voice control, willingness to comply, and predictor variables.
12 qualitative interviews were conducted to complement quantitative findings with deeper insights into participant reasoning and perceptions.
Results
This section presents findings from our large-scale evaluation of AI-powered voice phishing. We organize our results to address our primary research questions systematically, examining (1) how AI models perform in neutral contexts, (2) how scam content affects perception, and (3) what factors drive susceptibility to AI-powered attacks.
How AI models perform in neutral contexts
To establish baseline differences between AI models independent of scam context, we compared participant responses to neutral (non-scam) scenarios. Sesame was rated as significantly more humanlike than Llama FD, OpenAI AVM, and Gemini, but did not differ from Play.AI or ElevenLabs. When compared against an authentic human voice, Sesame was the only model that achieved statistical parity on human-likeness. The ElevenLabs cloned voice also matched human-level performance, suggesting that high-quality voice cloning can achieve human-level naturalness in neutral contexts.
How scam content affects perception
The transition from neutral to scam scenarios produced negative effects across all measured dimensions.
Impact of Scam Context on AI Model Perception. Mean ratings across four key perception dimensions comparing neutral (non-scam) and scam scenarios for combined AI models.
This pattern held consistently across both voice and text modalities. The scam scenario itself, rather than voice-specific artifacts, drives increased suspicion. However, the quality of the AI voice determines whether this suspicion translates into actual protection.
What factors drive susceptibility to AI-powered attacks
Overall compliance with AI-powered scam requests averaged 16.5% (yes/unsure), with substantial variation by scam type, message framing, and voice model. Personalized, emotionally charged scams dramatically outperformed generic account support scams. Compliance odds were approximately 3x higher for donation requests, 3.1x higher for the police-grandma scenario, and 5.33x higher for the cloned-sister-in-distress scenario. In contrast, the Gmail support scam did not differ significantly from MasterCard, indicating that generic account recovery messages elicit consistently low compliance regardless of brand. Among the unknown caller scenarios (MasterCard, Gmail, and Donation), which all used the same five AI voice models, the donation scam achieved the highest compliance rate, suggesting that scam framing alone substantially affects compliance. The most effective scam, the ElevenLabs cloned-voice sister-in-distress, achieved the highest compliance rate (36.1%). The human-voiced grandma and donation scams also showed elevated compliance (24.1% and 32.6%, respectively), suggesting that appeals invoking empathy or personal connection reliably increase susceptibility across multiple implementations.
Willingness to comply with scam requests by scenario type. Odds ratios (Exp(B)) show compliance likelihood (unsure/yes responses) relative to the MasterCard scam baseline (red dashed line at 1.0).
All four measured variables (caller sentiment, persuasiveness, trustworthiness, and human-likeness) were significantly and positively intercorrelated. Most models were perceived as highly human-like (resembling human voices), especially Sesame. However, persuasiveness was the strongest predictor of scam susceptibility, regardless of whether the voice was believed to be human or AI. This suggests that the content and delivery of the scam message matter more than achieving perfect vocal fidelity.
AI voice performance relative to human baseline across three vishing scenarios with 5 models.
Detection of AI-Generated Voices:
Participants struggled to identify AI-generated callers, achieving 70.3% accuracy in voice conditions and correctly identifying humans only 24.3–45.8% of the time, suggesting a general heightened suspicion toward callers rather than reliable AI detection ability. Text-based AI detection achieved 53.0% accuracy. Sesame was the most difficult to detect, with participants correctly identifying it as AI only 66.3% of the time.
For human voices, “correct” = identified as human. For AI voices, “correct” = identified as AI.
Neither general AI familiarity nor voice assistant usage improved detection. Participants who reported never using AI achieved 54.4% detection of AI voices, compared to 51.2% for those who use AI often/very often. Similarly, voice assistant usage showed no association with detection accuracy: never users achieved 53.9% accuracy versus 46.1% for frequent users. These findings suggest that current consumer exposure to AI systems does not translate into meaningful protection against AI-powered voice phishing.
AI misclassification rate relative to human correct classification, grouped by user AI familiarity level.
The Economics of AI-Enhanced Vishing
Traditional vishing has been fundamentally constrained by human labor costs, requiring a trained operator who can conduct only one conversation at a time, creating a natural bottleneck. Using humans to conduct vishing is highly unprofitable at US wages, with an expected loss of $27/hour. On the other hand, the AI models are expected to incur negative hourly profits with Llama, OpenAI, and Play.AI, while Gemini, Sesame, and ElevenLabs are associated with positive expected profits in the range of $1-$3 per hour. This suggests that AI-powered vishing may already be economically profitable for attackers. We estimate that the development time for an AI vishing system is roughly 260 hours, which corresponds to 5 hours per week for 52 weeks. Given that the average hourly wage for a machine learning engineer is roughly $62, this amounts to a sunk cost of roughly $16,120. For Gemini, Sesame, and ElevenLabs, respectively, this implies that the model would have to be continuously vishing for 282, 655, and 226 days in order to justify the costs of developing such a tool. Although the profitability of AI-powered vishing is roughly break-even for attackers across the models surveyed, we note that rapid advances in the technology may widen the profit margin and make this method more attractive for scammers.
Conclusion
AI-powered voice phishing represents a qualitatively new form of scalable social engineering. The results of this study demonstrate that LLMs and voice models can now approximate, and in some cases exceed, human effectiveness in eliciting compliance, especially when emotional or relational cues are present. The primary risk of AI-powered vishing is its ability to scale cheaply and effortlessly while maintaining high quality. Our findings underscore the urgent need for cross-disciplinary policy action:
Model governance: AI developers must move beyond surface-level safeguards and design abuse prevention mechanisms that are robust to removal and circumvention. There is a need for stronger deployment-level monitoring, provenance and auditability mechanisms, and more open research on how to meaningfully constrain abuse at minimal cost to developers.
Consumer education: awareness campaigns should focus on recognizing manipulative conversational strategies, rather than detecting AI-generated speech. Education efforts should also include guidance on recovery processes for those who have been victimized, reducing stigma and encouraging reporting.
Regulatory modernization: agencies must anticipate the new economics of fraud, in which automation enables attacks to scale at near-zero marginal cost. Regulatory frameworks should incentivize AI developers to prioritize security-by-design while clearly defining accountability and responsibility for downstream harms.
In short, AI systems dramatically lower the cost of deception. Defenses must evolve to protect human trust in digital systems and in their users, both human and agentic. With compliance rates exceeding 30% in some conditions, the potential for harm is substantial. At scale, even a 5% success rate across millions of automated calls represents a transformative shift in the economics of fraud. The question is no longer whether AI-powered vishing poses a serious threat, but whether policymakers, platforms, and the public will act before the damage becomes irreversible.
See our prior human-subjects work on AI-enabled phishing and spear-phishing for additional reading on this topic.
Simon Lermen