BREAKING
+++ SpaceX closes $60 billion acquisition of Cursor +++ Anthropic reportedly tops $11.5 billion in quarterly revenue +++ Anthropic raises misalignment risk to “low” +++ Lawsuit says Grok enabled abuse imagery +++ Alibaba's Qwen passes 3 billion downloads +++ Anthropic details how Claude's watermarks work ++++++ SpaceX closes $60 billion acquisition of Cursor +++ Anthropic reportedly tops $11.5 billion in quarterly revenue +++ Anthropic raises misalignment risk to “low” +++ Lawsuit says Grok enabled abuse imagery +++ Alibaba's Qwen passes 3 billion downloads +++ Anthropic details how Claude's watermarks work +++
Updated 08:00
AI IN LIFE AI IN LIFENEWS
DAILY
Research & Safety

Anthropic Raises Misalignment Risk — Keeps 'Model 2' Internal

Anthropic's new risk report lifts its misalignment estimate from "very low" to "low" — and deliberately holds back a stronger internal model.

Anthropic Raises Misalignment Risk — Keeps 'Model 2' Internal

Illustration · AI-generated (AI IN LIFE)

At a glance

  • Misalignment risk raised from "very low" to "low"
  • Reasons: cyber incidents and reduced confidence in its own evaluations
  • Task-based safety tests no longer capture capability gains, Anthropic says
  • Internal "Model 2": better at coding/data generation; no external release planned
  • OpenAI is meanwhile delaying its "Astra" model over cyber risks

Anthropic has published a new risk report, raising its estimate of misalignment risk in high-stakes situations from "very low" to "low." Per Axios, the company cites recent cybersecurity incidents as the driver — along with growing uncertainty about its own measurement methods.

The self-criticism stands out: its most concrete task-based evaluations "no longer capture" increases in model capabilities, Anthropic writes. Key safety benchmarks are considered saturated — the yardstick is growing more slowly than the models.

For the first time, the report also describes a stronger internal system, known internally as "Model 2." It shows noticeable improvement on internal tasks like coding and data generation, though the gains are smaller than in earlier generational jumps.

A release is not planned: "We do not currently have plans to release this model externally," the company states, stressing this is standard R&D practice — many exploratory model versions are trained and never shipped.

The contrast with competitors is striking: OpenAI is delaying its "Astra" model over cyber risks, while Anthropic continues development without announced pauses — even as it revises its own risk estimate upward.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What does misalignment mean?

An AI system pursuing goals or behaviors that diverge from what its developers and users intend — in the extreme, with severe consequences.

Why isn't Model 2 being released?

Anthropic calls it normal R&D practice: many internal model versions never ship. The report cites no safety blockade as the official reason.

Is the upgrade an alarm signal?

It signals increased uncertainty rather than acute danger — partly because the company's own measurement tools are losing resolution.