LIVE
+++ LazyLLM: a low-code kit for multi-agent LLM apps +++ Transformers 5.16.1 adds GLM-5.3-Flash support +++ BoxLite runs AI agents inside hardware-isolated VMs +++ ArcReel: Self-Hosted AI Video Workspace on Docker +++ QwenPaw: a self-hosted AI assistant for seven chat apps +++ Nvidia's NVHBM Moves the Memory Controller Into HBM ++++++ LazyLLM: a low-code kit for multi-agent LLM apps +++ Transformers 5.16.1 adds GLM-5.3-Flash support +++ BoxLite runs AI agents inside hardware-isolated VMs +++ ArcReel: Self-Hosted AI Video Workspace on Docker +++ QwenPaw: a self-hosted AI assistant for seven chat apps +++ Nvidia's NVHBM Moves the Memory Controller Into HBM +++
All news ›
AI IN LIFE AI IN LIFENEWS
DAILY
DE EN
CHIPS

Nvidia's NVHBM Moves the Memory Controller Into HBM

Nvidia moves its memory controller into the HBM base die, claiming up to 30 percent more bandwidth and 15 percent less memory power. Annapurna Labs signs on first.

Nvidia's NVHBM Moves the Memory Controller Into HBM

Illustrative image: a stack of memory dies on a dark substrate under the cold blue light of an open server rack — computer-generated, not a real component.

Nvidia is relocating the memory controller off the compute die and into the base die of the HBM stack, calling the design NVHBM and offering it to NVLink Fusion partners, with Amazon's Annapurna Labs named as the first collaborator.

At a glance

  • NVHBM places Nvidia's own memory controller in the HBM stack's base die instead of on the XPU compute die.
  • Vendor claim: up to 30 percent more memory bandwidth versus standard HBM4E.
  • Vendor claim: 15 percent lower HBM power draw and up to 25 percent more free area on the compute die.
  • First collaborator is Amazon's Annapurna Labs; NVLink Fusion support arrives with the Trainium4 generation.
  • Nvidia wants NVHBM as a standard offered by multiple memory suppliers — the August 26, 2026 post names none of them.

Nvidia is taking the memory controller off the compute die and putting it into the base die that HBM stacks sit on. It calls the design NVHBM and is opening it to partners on its NVLink Fusion rack platform. Amazon's chip unit Annapurna Labs is named as the first collaborator. The announcement came in a company blog post on August 26, 2026.

What actually changes

Today the logic that drives memory lives on the processor die itself. Under NVHBM, Nvidia's controller moves down into the base die beneath the memory stack. That hands back silicon area on the compute die, which can then hold more compute. In effect, the memory interface becomes part of the memory package rather than part of the processor.

Nvidia puts the payoff at up to 30 percent more memory bandwidth against standard HBM4E, 15 percent lower power in the HBM itself, and up to 25 percent more usable area on the compute die. All three figures come from the company. The post does not say how they were measured, and no third party has published its own numbers as of this writing.

What NVLink Fusion is for

NVLink Fusion is Nvidia's pitch to operators building their own accelerators: attach a custom XPU or CPU to Nvidia's rack architecture instead of designing one from scratch. Partners get NVLink chiplets, NVLink-C2C, NVLink Switches and the MGX systems and racks. Nvidia also points to a surrounding group of CPU partners, ASIC designers, system builders and technology providers. The platform ships alongside each generation of the rack-scale architecture.

Annapurna Labs says it will support NVLink Fusion with its next-generation Trainium silicon, starting at Trainium4, so Amazon parts and Nvidia GPUs can share one rack-scale design. Nafea Bshara, vice president at Annapurna Labs, describes NVHBM in the announcement as a new architectural approach to high-bandwidth memory. No ship date for Trainium4 appears in the post.

The standard is the real move

The second half of the announcement matters more than the percentages. Rather than treating NVHBM as a bespoke part per customer, Nvidia says it is defining a standard implementation that several memory vendors can build and validate. The stated goal is less engineering work to integrate and qualify parts across suppliers. Which suppliers, and on what timeline, is left unsaid.

Nvidia frames the whole thing around current workloads: with AI agents and trillion-parameter models becoming routine, it argues performance no longer rests on compute alone but on compute, memory, storage, networking and software being designed as one system. That is also the argument the company uses to sell the rack rather than the chip.

What this report cannot confirm

Exactly one source was available for this story: the vendor's own blog post. No independent newsroom has verified the claims here. That leaves the three percentage figures, the identity of the participating memory makers, every shipping or availability date, and pricing all unconfirmed. We are labeling this as a report of a company announcement, not as an established fact.

◈ AI-GENERATED REPORT · SOURCES LINKED

FAQ

What is NVHBM?

A memory design in which Nvidia's memory controller sits in the HBM stack's base die rather than on the compute die. Nvidia is making it available to partners on its NVLink Fusion platform.

How much faster is NVHBM than HBM4E?

Nvidia claims up to 30 percent more memory bandwidth versus standard HBM4E, plus 15 percent lower power. Those are vendor figures with no published test conditions and no independent verification.

Which chips will use NVLink Fusion with NVHBM?

Amazon's Annapurna Labs is named as the first collaborator and plans to support NVLink Fusion starting with Trainium4. The announcement gives no date for that part.