LIVE
+++ NVIDIA Pools Home PCs to Run AI Agents Locally +++ OpenAI agents hijacked a German developer wiki +++ Playco says GPT-6 Astra halved its manual game fixes +++ Three AI services went down at once, cause unexplained +++ NBA 2K27 Brings DLSS 5 to GeForce NOW's Cloud RTX 5080s +++ GPT-6 Astra Cuts a 41-Document Review to Minutes ++++++ NVIDIA Pools Home PCs to Run AI Agents Locally +++ OpenAI agents hijacked a German developer wiki +++ Playco says GPT-6 Astra halved its manual game fixes +++ Three AI services went down at once, cause unexplained +++ NBA 2K27 Brings DLSS 5 to GeForce NOW's Cloud RTX 5080s +++ GPT-6 Astra Cuts a 41-Document Review to Minutes +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
DE EN
MODELS

GPT-6 Astra Cuts a 41-Document Review to Minutes

The model caught all four planted errors and lifted performance by nearly 40 percent — so far the numbers rest on OpenAI's own account.

GPT-6 Astra Cuts a 41-Document Review to Minutes

Symbolic image: a document scanner pulls report pages through while abstract highlight blocks glow on two monitors behind it.

In a financial-statement review run by Legora on GPT-6 Astra, the model worked through 41 documents in minutes, found all four deliberately planted errors and improved performance by nearly 40 percent.

At a glance

  • 41 financial documents reviewed in minutes, all four planted errors found — figures from OpenAI's own Legora case study.
  • Performance gain in the review workflow: nearly 40 percent; OpenAI does not name the baseline.
  • API pricing per heise: $10 per million input tokens, $50 per million output tokens.
  • Generally available in GitHub Copilot since September 4, 2026 for Pro+, Max, Business and Enterprise.
  • FrontierMath Tier 4: 97.6 percent; Humanity's Last Exam: 57.2 percent — below Fable 5.1 at 65 percent.

Legora put GPT-6 Astra through a financial-statement review and reported the result together with OpenAI: 41 documents handled in minutes, all four deliberately planted errors caught, performance up by nearly 40 percent.

What the test actually measured

Seeding a known number of errors into a document stack and counting how many a review catches is a standard audit exercise. In OpenAI's short write-up of the run, Astra found every one of the four seeded errors.

The second figure, nearly 40 percent, describes workflow performance rather than a benchmark score. The write-up does not say what the comparison baseline was, how the gain was calculated, or how long the 41 documents ran.

How much weight the numbers carry

All of them come from the vendor. OpenAI's case-study page returned HTTP 403 when we tried to open it, so we could not check the claims against the original text, and we have no independent confirmation from Legora or a third party.

Treat the figures as a signal of what OpenAI considers a showcase result, not as evidence about error rates in your own review stack.

Astra has since gone broad

GitHub made the model generally available in Copilot on September 4, 2026 for the Pro+, Max, Business and Enterprise plans. It runs in Visual Studio Code, Visual Studio, the Copilot CLI, the coding agent, on github.com, in GitHub Mobile and in JetBrains IDEs, Xcode and Eclipse.

Usage is billed at provider list pricing under usage-based billing, and the rollout is gradual rather than instant. Business and Enterprise administrators get the model switched on by default under model policy and have to turn it off if they do not want it.

Strong in places, behind in others

heise online puts API pricing at $10 per million input tokens and $50 per million output tokens, with an optional fast mode running at 2.5x speed for double the money.

The benchmark picture is uneven: 97.6 percent on FrontierMath Tier 4 against 87.8 percent, 88.0 percent on SRE-Bench at the first attempt and 99.2 percent across four, and a first-ever 100 percent on ExploitBench. On Humanity's Last Exam, Astra's 57.2 percent trails Fable 5.1 at 65 percent.

heise also reports that Astra is the first OpenAI model classified as critical for cybersecurity under the Preparedness Framework, after surfacing two zero-day vulnerabilities during testing. OpenAI president Greg Brockman is quoted there placing the arrival of AGI, seen in hindsight, at roughly this moment and this model.

What review teams should take from it

The Legora story is not about automating sign-off. It is about triage: the model works through a stack and flags what looks wrong, while accountability for the outcome stays with a person.

Whether the gain survives contact with real files is an empirical question each team has to answer for itself. Rebuilding the test needs known errors in the stack, a human checking path, and a measured baseline of how the work is done today.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What did Legora test with GPT-6 Astra?

A financial-statement review: 41 documents with four errors planted in them. OpenAI says the model found all four within minutes.

How much does GPT-6 Astra cost via the API?

heise lists $10 per million input tokens and $50 per million output tokens; the optional fast mode costs twice as much and runs at 2.5x speed.

Where can I use GPT-6 Astra today?

Per heise, first at selected organizations and then in ChatGPT Plus, Pro, Business and Enterprise; in GitHub Copilot it has been generally available since September 4, 2026.