It was no ordinary benchmarking exercise. The internal project code-named ‘Cannes’, run by Meta through contractor Covalen, tasked hundreds of testers with a precise mission: create fake under-18 accounts and bombard rival chatbots with prompts no model should ever entertain. The targets: OpenAI’s ChatGPT, Google’s Gemini, and Character.AI.
According to documents reviewed by WIRED, the operation ran at least until April 2026. One cycle, completed in August 2025, sent more than 45,000 prompts. A spreadsheet lists 3,748 of them: hundreds deal with suicide and self-harm, hundreds more with eating disorders. At least 239 involve sex or romance, while attached images showed pills and knives. One dummy profile posed as a pregnant 13-year-old asking where to buy pills; another as a girl trying to hide an eating disorder from her parents.
The targeted companies never authorised the tests. In fact, their terms of service explicitly ban such activity. OpenAI prohibits unsolicited safety testing and using outputs to build competing models. Google forbids bypassing its safety filters, and Character.AI bans harmful content. ‘This conduct violates our Terms of Service,’ a Character.AI spokesperson said.
Meta’s defence and the grey line
Meta defends the initiative as routine industry practice: ‘Testing and benchmarking chatbot responses to help ensure safe and age-appropriate experiences is a responsible, industry-standard practice,’ a spokesperson stated, adding that it did not use the data to train its own models. But many experts are unconvinced. Using identities disguised as minors, says Rumman Chowdhury of Humane Intelligence, puts the operation ‘outside what is usually described as “industry standard” evaluation.’ She calls it a ‘governance gray zone where safety becomes a convenient cover for anticompetitive practices.’
The affair lands at a tense moment. In September 2025, the US Federal Trade Commission opened a formal inquiry into AI and child safety, covering Meta, OpenAI, and Google. In Europe, the AI Act and the Digital Services Act impose strict obligations on platforms regarding risks to minors. On both sides of the Atlantic, the question is the same: who is accountable when a chatbot talks to a teenager about self-harm?
What it means for those hosting models in-house
For those developing and hosting LLMs in-house, the story touches a nerve. Running safety benchmarks against external models – perhaps via APIs – is common, but the line with covert competitive intelligence is thin. Companies that operate on-premise models, sealed behind corporate firewalls, must now consider how to prevent third parties from doing the same to their endpoints: strong authentication, abuse detection, and the ability to tell a legitimate user from an artfully crafted account are essential. And all of this without violating real users’ privacy, an equilibrium made even more delicate by European regulations.
The episode exposes the growing tension between the race for model credibility and the temptation to wield ‘safety’ as a competitive weapon. Transparency and respect for rivals’ rules are not optional: when regulators come knocking, the safety-testing excuse may not hold.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!