AI Governance Institute
← News
Research2026-07-31

LLM Agents Outperform Human Scammers, Exposing Fraud Detection Gaps

What happened

A peer-reviewed study conducted by researchers from four universities, covered by AI scammers outperform humans when it comes to building trust, found that an AI chatbot built on Anthropic's Claude model outperformed human scammers in simulated pig butchering romance-fraud scenarios, achieving a 46% victim compliance rate compared to 18% for human operators. The researchers demonstrated that the AI agent could autonomously conduct the lengthy relationship-building phase of the fraud operation, the most resource-intensive part of pig butchering campaigns, without triggering vendor-level safety controls. The bypass mechanism was simple: the AI handled all trust-building interactions and handed off to a human operator only at the final financial solicitation stage, keeping the most detectable behavior outside the model's output stream. This architecture effectively defeats content filtering at the model layer because the harmful act, soliciting money, is never performed by the model itself. The findings confirm that safety controls embedded at the model level are not sufficient fraud controls when adversarial actors can design workflows that route around them.

Why it matters

  • ·Vendor safety commitments and model-level content filters cannot be treated as fraud controls: this study demonstrates that a simple human-handoff architecture defeats them entirely, meaning enterprises relying on provider safeguards to limit misuse exposure have a structural gap in their fraud risk programs.
  • ·Financial services firms, consumer platforms, and any organization with customer-facing AI or fraud-risk obligations should reassess whether their red-teaming and adversarial testing programs cover this class of hybrid human-AI attack, since most current red-teaming frameworks focus on model outputs rather than multi-step workflow abuse.
  • ·Consumer protection regulators, including the FTC AI Enforcement Policy, are increasingly attentive to AI-enabled deception at scale; enterprises that deploy or integrate LLMs in consumer-facing contexts may face heightened scrutiny if their platforms are shown to facilitate or fail to detect this class of fraud.

Governance controls affected

What to do now

  • Review your fraud risk assessment to determine whether it accounts for hybrid human-AI social engineering workflows that route final solicitation through a human operator to avoid model-level detection.
  • Expand red-teaming scope to include multi-step, multi-actor attack scenarios where the LLM performs trust-building and a human completes the harmful action, rather than testing model outputs in isolation.
  • Audit third-party AI vendor contracts and safety commitment representations to determine whether vendor safeguards are claimed to cover downstream misuse by third parties building on their APIs.
  • Brief fraud detection teams on the specific behavioral pattern identified in this research: extended AI-driven relationship-building followed by a human-executed financial solicitation, and assess whether existing detection signals would catch it.
  • Assess whether your consumer protection compliance program explicitly addresses AI-enabled social engineering as a distinct fraud vector requiring dedicated controls beyond standard phishing and impersonation coverage.

What to watch next

Regulatory bodies with consumer protection mandates, including the FTC and financial regulators, are likely to reference this class of research when updating fraud guidance for AI-enabled platforms. The FATF AI Anti-Money Laundering Guidance is a framework compliance teams in financial services should monitor for updates that address LLM-enabled social engineering as a typology. Enforcement actions against platforms or API providers whose models were used in fraud operations, even without direct knowledge, would materially shift liability exposure for enterprises deploying or reselling LLM capabilities to third parties.

Stay ahead of stories like this

Get every Global AI governance development like this one, plus the rest of the week's developments. Every Thursday.

Powered by Buttondown.

Related Coverage

Corporate Policy2026-09-03

Simultaneous ChatGPT, Grok, and Claude Outage Exposes AI Concentration Risk

On September 3, 2026, OpenAI's ChatGPT, xAI's Grok, and Anthropic's Claude experienced simultaneous outages affecting millions of users globally. ChatGPT reported elevated errors across logins, file uploads, voice mode, and image generation, while Anthropic attributed its disruption to an infrastructure issue resolved by 12:15 PM ET. The concurrent nature of the failures raises unresolved questions about shared upstream dependencies and leaves enterprise business continuity programs exposed.

Research2026-08-31

Meta Ran Ads for Nonconsensual Deepfake App, Exposing Platform-Control Assumptions

A weekly threat watchlist published by Resemble AI documented that Meta served paid advertisements for an application explicitly promoting nonconsensual sexual deepfakes of real individuals, including a named U.S. politician. The incident reflects failures in ad preclearance review, synthetic-content detection, and abuse-report escalation. Enterprises that distribute AI-generated media products through major platforms cannot treat platform content review as a substitute for their own intake and labeling controls.

Enforcement2026-08-29

Sony and Warner Sue Anthropic Over Training Data, Exposing Vendor IP Risk

Sony Music and Warner Chappell have filed a copyright infringement lawsuit against Anthropic in the US District Court for the Northern District of California, alleging that tens of thousands of protected works were used to train Claude without authorization. The complaint seeks up to $150,000 per infringed work and up to $25,000 per instance of stripped copyright metadata, with total exposure potentially reaching several billion dollars. Co-founders Dario Amodei and Benjamin Mann are named as individual defendants.