Anthropic: bio filter off for 11 months — 133 million chats
An internal flag disabled bioweapons classifiers and logging at once. Anthropic disclosed the incident itself in its latest risk report.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Period: May 2025 to April 2026 (around 11 months)
- Affected: ~133 million exchanges by ~50,000 contractors
- Cause: internal flag disabled classifiers and logging simultaneously
- Review: 1,197 flagged transcripts, 757 from internal teams; 62 examined, no clear misuse
- Source: Anthropic's own risk report (August 2026 update)
Anthropic's classifiers blocking bioweapons-related requests were disabled on its human-feedback platforms for around 11 months — from May 2025 to April 2026. Roughly 133 million exchanges by about 50,000 external contractors were affected. The company disclosed the incident itself in its latest risk report; The Next Web and The Decoder reported on it over the weekend.
The cause was a flag intended for internal use that simultaneously switched off both the blocking classifiers and logging. The flagged traffic "was not recorded or propagated to any review mechanisms," the report says — which meant the incident could only be found by going through raw transcripts.
The follow-up review ended mildly: Anthropic ran Claude Sonnet 5 over the affected data; 1,197 transcripts were flagged as potentially risky, 757 of them from Anthropic's own teams, and all but 62 of the rest were deliberate red-teaming exercises. In the remaining 62 reviewed cases the company found no clearly concerning misuse.
More unsettling is the report's own conclusion: the discovery "leads us to believe that there is an increased likelihood of other, similar issues unknown to us." The report also notes that many vendor platforms lacked screening processes capable of stopping even low-level threat actors.
For the industry the case is doubly instructive: it shows how fragile safety chains are in practice — and that transparent self-disclosure is possible. That disclosure culture is exactly what regulators increasingly demand from all frontier labs.
FAQ
What exactly failed?
Classifiers blocking bioweapons-related requests, plus the associated logging — on Anthropic's human-feedback platforms, not in the public Claude products.
Was there proven misuse?
Not according to Anthropic's review: of 1,197 flagged transcripts, 62 remained for close review and none showed clearly concerning misuse.
Why is the case still serious?
Because blocking and logging failed together, and the report itself warns of an increased likelihood of other unknown issues — a warning sign for the whole industry.


