Back to feed
News Story
APriority76
InfoQ AI/ML/Data Eng
2 sources

WhatsApp Tests On-Device ML for Scam Detection with Privacy-Preserving Analytics

WhatsApp is testing a Scam Alert feature in a limited beta that uses on-device machine learning to detect potential scam messages from non-contacts. The system leverages confidential computing, Oblivious HTTP, differential privacy, and model transparency to protect user privacy while measuring performance. This approach keeps message content on the device, enhancing privacy and security.

SynthePulse Insight · AI deep reading

WhatsApp On-Device AI Anti-Scam: Balancing Privacy and Security

Version 1 · 1 source

WhatsApp is testing an on-device AI scam alert feature in a limited rollout, aiming to provide safety tips for messages from unknown senders while maintaining end-to-end encryption.

  • WhatsApp has begun limited testing of its Scam Alert feature, using on-device machine learning models to identify potential scam messages from non-contacts.
  • The feature is optional, with detection performed locally on the device; message content is not automatically sent to Meta, preserving end-to-end encryption.
  • If users mark a conversation as trusted, the warning disappears and does not reappear; users can voluntarily share the last five messages to improve the model.
  • Meta states it cannot selectively push AI models to specific users through this system; model versions are recorded on a third-party transparency ledger with verified signatures.
  • Performance data, such as warning frequency and subsequent user actions, is protected through aggregation, confidential computing, and differential privacy.
  • The feature remains experimental; Meta will continue testing and refining with security researchers and has not decided on a broader rollout.
Open section navigationFeature Overview and Testing Scope

Feature Overview and Testing Scope

WhatsApp is introducing the Scam Alert feature, an optional capability that uses on-device machine learning models to identify potential scams in messages from non-contacts. The feature is currently available only in a limited test, with Meta working alongside security researchers to test and gather feedback before a wider rollout.

The model looks for language and conversation patterns associated with known scams. If a potential scam is detected, WhatsApp displays a warning visible only to the recipient. Users can then block or report the sender, continue the conversation, or mark the chat as trusted.

Privacy-Preserving Design

Meta states that Scam Alert is designed to work without granting WhatsApp access to users' encrypted conversations. The detection model runs locally on the device, and message content is not automatically sent back to the company.

The company also says it cannot use this system to selectively send specific AI models to particular users. Each model version is recorded on a third-party transparency ledger before distribution, and devices verify the published signature and file hash before loading the model.

WhatsApp also collects limited performance data, such as how often warnings appear and what users do afterward. According to Meta, this data is aggregated and protected through confidential computing and differential privacy, rather than exposing individual conversations.

User Control and Transparency

If a warning is incorrect, marking the conversation as trusted removes the warning and prevents Scam Alert from flagging that chat again. Users can also voluntarily share the last five messages from a trusted conversation with WhatsApp to help improve the model.

Users will eventually be able to view Scam Alert activity in WhatsApp under 'Account > Request Info > Scam Alert Activity', including which messages were analyzed, the results, and the model version involved.

Addressing the Growing Scam Problem

The feature launches as scammers increasingly use messaging platforms to build trust, pressure victims, and solicit money or personal information. The U.S. Federal Trade Commission reports that consumer-reported losses from social media scams reached $2.1 billion in 2025, with $425 million linked to WhatsApp (as reported by CNET).

This makes an on-device approach particularly important: WhatsApp aims to add an automated layer of protection without requiring its servers to inspect private conversations.

Trade-offs and Limitations

Scam Alert can provide a useful second opinion before users reply to unknown senders, but it does not guarantee that a message is safe or fraudulent. Machine learning may miss sophisticated scams or incorrectly flag legitimate conversations.

Scam Alert remains an experiment rather than a finalized security feature. Meta says it will continue testing the system with security researchers, expand bug bounty coverage, and refine the model before deciding on a broader rollout.

If this approach works, WhatsApp could add an extra layer of scam protection without abandoning the privacy model that defines its service—but users should still treat any automated warning as guidance rather than evidence.

Credibility boundary

This report is based on a single source from TechRepublic, which cites Meta's statements and FTC data (as relayed by CNET). Feature details and privacy claims come from Meta's official statements and have not been independently verified. The FTC data is a secondary source and should be considered a source claim rather than confirmed fact.

Insight takeaway

WhatsApp's Scam Alert represents a careful balance between privacy and security: using on-device AI to provide scam tips while protecting encryption, but its effectiveness, false positive rate, and privacy protections still require further testing and validation.

Primary report

InfoQ AI/ML/Data Eng

Primary source

Same-event coverage

Also covered by 1 sources