Anthropic: AI classifier replaces human approval in Claude Code
From 14 August, auto mode becomes the default: instead of humans signing off on risky commands, an AI classifier reviews them. The reason is uncomfortable – humans caught dangerous commands only 13.6% of the time.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Source: Anthropic blog, 7 August 2026
- Auto mode becomes default from 14 August 2026 (Pro, Max, Team)
- Study: 1,053 testers
- Human detection of dangerous commands: 13.6% (~5% after 50 approvals)
- AI classifier: 89% detection
Anthropic is changing the safety logic of its developer tool Claude Code: from 14 August 2026, automatic mode becomes the default for the Pro, Max and Team plans. Instead of a human manually approving every potentially risky command, an AI classifier takes over the review.
The trigger was a sobering study with 1,053 testers. The result: humans detected dangerous commands only 13.6 percent of the time. The AI classifier reached 89 percent. Even more striking: after roughly 50 approvals in a row, the human hit rate dropped to about 5 percent – a classic case of 'click fatigue', where users wave warnings through reflexively.
The switch is a remarkable admission: for repetitive safety decisions, the human is not the better guardian but the weakest link. The AI does not tire, does not get sloppy, and judges the thousandth command as carefully as the first.
For the growing number of companies deploying coding agents and AI-driven automation, this is highly relevant. It shifts the human's role: away from nodding through individual actions, toward setting rules, sample auditing, and stepping in on exceptions.
At the same time, the fair counter-question remains: who watches the watcher? Anthropic bets that a specialized, transparently evaluated classifier is more reliable than overtired humans. The study supports the thesis – but the real test comes in the daily flow of millions of automated commands.
FAQ
What changes for Claude Code users?
From 14 August an AI classifier reviews risky commands by default, instead of every command being manually approved by a human.
Why the change?
A study showed humans detect dangerous commands only 13.6% of the time – the AI classifier 89%.
Can you still approve manually?
Auto mode becomes the default; users should check their plan's settings if they want more manual control.


