HBM memory-first visual atlas

Follow one byte from the workload to the physical machine.

Choose why you are here, choose a workload, then peel the system from the node down to HBM and back out to fabric, power, and cost. The facts stay fixed. Only the explanation changes.

Public interactive companion
Evidence reviewed July 25, 2026
Back to the complete HBM article
Source-backed architecture Public systems and course material support the named layers and relationships.
Illustrative workload mapping The highlighted route teaches causality. It is not a profiler trace or dispatch receipt.
Live run not observed No run ID, HBM counter, power sample, facility meter, or accepted-output receipt is claimed here.

Same machine. Three ways into it.

Audience selection changes the business-meaning panel. It never changes the workload, the physical path, or the evidence state.

Three stories. One physical map.

The HBM story explains the hardware. The coding-agent and Wan2.2 stories show why two workloads can use the same memory system differently.

Node to HBM, one causal layer at a time.

View 1 of 11Navigation state, not course completion

Keyboard: ← and → move one layer, ↑ and ↓ change story, 1 2 3 change audience, Home and End jump to the first or last layer.

Physical peel-back Click any labeled block
to inspect that layer
Compute node physical hierarchy An interactive map from a compute node through a GPU package, compute unit, scheduler, matrix engine, local memory, L2 cache, memory controller, HBM, fabric, and the power and facility boundary. COMPUTE NODE host · accelerators · memory · network · power · cooling boundary GPU / ACCELERATOR PACKAGE compute die · package routing · HBM-side interface SM / CU thread and wave execution resources SCHEDULER warps / waves MATRIX ENGINE Tensor Core / MFMA REGISTERS + SHARED / LDS small, fast, explicitly reused state L2 CACHE shared hot blocks MEMORY CONTROLLER queues · commands HBM STACK core dies · TSVs · base die FABRIC / OFFLOAD link · topology · destination tier POWER → ENERGY → HEAT → COOLING / WATER BOUNDARY → COST PER ACCEPTED OUTPUT
01 / 11 Compute node
The system boundary that joins the workload interval, hardware topology, verifier, and resource ledgers.

Code

What software can name here

request = {"question": "Where do hot AI bytes live?"}
receipt = {"run_id": null, "observed_hbm_bytes": null}

Original teaching pseudocode · not executed

Operation

What this layer does

Frames the workload, the acceptance rule, and the physical system boundary before choosing a memory tier.

Data movement

What moves, and what does not

No physical byte path is claimed yet. The node view only declares possible producers, consumers, tiers, and links.

Business meaning

The decision at this layer

A node is the first capital unit that can connect accelerator capacity to accepted work. Peak HBM bandwidth alone does not prove useful capacity or margin.

Evidence

What is known, illustrated, and missing

Architecture
Source-backed
Workload mapping
Illustrative
Live observation
Not observed

Public architecture supports the layer. This atlas does not contain a selected runtime, kernel artifact, counter trace, power sample, or accepted-output run.

Exact next proofJoin node inventory, accelerator topology, memory capacity, workload interval, and verifier under one real run ID.

Touch the math, then go deeper at the sources.

Start with the original deterministic micro-lab, then use the official public courses for the compute foundations. The atlas connects those foundations to HBM, full workload state, evidence, and accepted-output economics.

Rights boundary: this file contains no copied slides, videos, assignments, starter repositories, or course artwork. Public availability does not by itself grant redistribution rights. Follow the official links and each source's own terms.

The complete physical path remains visible.

The interaction enlarges one layer at a time. It does not remove the rest of the stack.

  1. Compute node
  2. GPU or accelerator package
  3. Streaming Multiprocessor or Compute Unit
  4. Warp or wave scheduler
  5. Tensor Core or MFMA matrix engine
  6. Registers and shared memory or LDS
  7. L2 cache
  8. Memory controller
  9. HBM stack, channels, banks, and cells
  10. Fabric, offload path, and destination tier
  11. Power, energy, heat, cooling, water boundary, and cost