LIVE
+++ Z.ai delays GLM-5.3 weights over exploit capability +++ First Nvidia H200s reach ByteDance and Tencent +++ Cerebras CS-4: three wafers, up to 30x faster inference +++ Warp launches Factories for fleets of coding agents +++ Microsoft patches one-click Copilot flaw after 8 months +++ OpenAI grows slower than Anthropic as loss hits $12.3B ++++++ Z.ai delays GLM-5.3 weights over exploit capability +++ First Nvidia H200s reach ByteDance and Tencent +++ Cerebras CS-4: three wafers, up to 30x faster inference +++ Warp launches Factories for fleets of coding agents +++ Microsoft patches one-click Copilot flaw after 8 months +++ OpenAI grows slower than Anthropic as loss hits $12.3B +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
HARDWARE

Cerebras CS-4: three wafers, one rack — up to 30x faster

The freshly listed chipmaker's first multi-wafer system targets the inference market: over 1,000 tokens per second on trillion-parameter models.

Cerebras CS-4: three wafers, one rack — up to 30x faster

Illustration · AI-generated (AI IN LIFE)

At a glance

  • Three WSE-3 Turbo processors per rack, 750 petaflops (sparse FP16)
  • Over 1,000 tokens/s on models above 10 trillion parameters
  • Clock doubled from 1.4 to 2.8 GHz, 2-microsecond wafer-to-wafer latency
  • Partners: OpenAI, G42, AWS, AMD; shipping from Q3 2026
  • Cerebras Q2 revenue: $180.1M (+74% YoY), Nasdaq IPO in May 2026

What's new? Cerebras unveiled the CS-4, its first system bundling three wafer-scale processors into a single rack. The company promises inference up to 30 times faster than GPU systems and more than 1,000 tokens per second even on models beyond ten trillion parameters. First shipments are due this quarter.

How does it get there? Not with new silicon: the WSE-3 Turbo inside is the familiar chip with four trillion transistors and 900,000 cores — just clocked twice as fast, up from 1.4 to 2.8 gigahertz. Add 750 petaflops of sparse FP16 compute per system, two-microsecond wafer-to-wafer latency, and a claimed tenfold gain in throughput per watt over the CS-3.

Who is on board? Cerebras names OpenAI, G42, AWS, and AMD among its partners. The division of labor is notable: AMD Helios racks or AWS Trainium can handle prompt processing (prefill), while the CS-4 takes over latency-critical response generation. "In AI, speed is productivity," is how CEO Andrew Feldman sums up the positioning.

Why now? The market is shifting from training to inference — where agents and reasoning models burn compute on every request. Cerebras has been listed on Nasdaq since a $5.55 billion May debut and last reported $180.1 million in quarterly revenue, up 74 percent year over year.

What's still open? Cerebras discloses no pricing and no confirmed large CS-4 orders. And the 30x figure is a vendor claim under favorable conditions — independent benchmarks are pending. Only the direction is clear: the race for the fastest token output has a new benchmark.

This article was produced with AI assistance and editorially reviewed.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What sets the CS-4 apart from the CS-3?

The CS-4 is the first to couple three wafer processors in one rack and doubles the chip clock to 2.8 GHz. The silicon itself remains the familiar WSE-3 — the leap comes from architecture and power delivery.

Is the CS-4 really 30x faster than GPUs?

That is a vendor figure for selected inference scenarios. Independent benchmarks are not yet available.

Why is Cerebras betting on inference over training?

Agents and reasoning models consume compute on every request — the inference market grows structurally, and low latency becomes the selling point.