Anthropic Raises Misalignment Risk — Keeps 'Model 2' Internal
Anthropic's new risk report lifts its misalignment estimate from "very low" to "low" — and deliberately holds back a stronger internal model.

Illustration · AI-generated (AI IN LIFE)
At a glance
- Misalignment risk raised from "very low" to "low"
- Reasons: cyber incidents and reduced confidence in its own evaluations
- Task-based safety tests no longer capture capability gains, Anthropic says
- Internal "Model 2": better at coding/data generation; no external release planned
- OpenAI is meanwhile delaying its "Astra" model over cyber risks
Anthropic has published a new risk report, raising its estimate of misalignment risk in high-stakes situations from "very low" to "low." Per Axios, the company cites recent cybersecurity incidents as the driver — along with growing uncertainty about its own measurement methods.
The self-criticism stands out: its most concrete task-based evaluations "no longer capture" increases in model capabilities, Anthropic writes. Key safety benchmarks are considered saturated — the yardstick is growing more slowly than the models.
For the first time, the report also describes a stronger internal system, known internally as "Model 2." It shows noticeable improvement on internal tasks like coding and data generation, though the gains are smaller than in earlier generational jumps.
A release is not planned: "We do not currently have plans to release this model externally," the company states, stressing this is standard R&D practice — many exploratory model versions are trained and never shipped.
The contrast with competitors is striking: OpenAI is delaying its "Astra" model over cyber risks, while Anthropic continues development without announced pauses — even as it revises its own risk estimate upward.
FAQ
What does misalignment mean?
An AI system pursuing goals or behaviors that diverge from what its developers and users intend — in the extreme, with severe consequences.
Why isn't Model 2 being released?
Anthropic calls it normal R&D practice: many internal model versions never ship. The report cites no safety blockade as the official reason.
Is the upgrade an alarm signal?
It signals increased uncertainty rather than acute danger — partly because the company's own measurement tools are losing resolution.


