HBM memory-first visual atlas
Follow one byte from the workload to the physical machine.
Choose why you are here, choose a workload, then peel the system from the node down to HBM and back out to fabric, power, and cost. The facts stay fixed. Only the explanation changes.
1. Choose the decision you own
Same machine. Three ways into it.
Audience selection changes the business-meaning panel. It never changes the workload, the physical path, or the evidence state.
2. Choose the state you want to follow
Three stories. One physical map.
The HBM story explains the hardware. The coding-agent and Wan2.2 stories show why two workloads can use the same memory system differently.
3. Peel back the machine
Node to HBM, one causal layer at a time.
View 1 of 11Navigation state, not course completion
Keyboard: ← and → move one layer, ↑ and ↓ change story, 1 2 3 change audience, Home and End jump to the first or last layer.
to inspect that layer
The system boundary that joins the workload interval, hardware topology, verifier, and resource ledgers.
Code
What software can name here
request = {"question": "Where do hot AI bytes live?"}
receipt = {"run_id": null, "observed_hbm_bytes": null}
Original teaching pseudocode · not executed
Operation
What this layer does
Frames the workload, the acceptance rule, and the physical system boundary before choosing a memory tier.
Data movement
What moves, and what does not
No physical byte path is claimed yet. The node view only declares possible producers, consumers, tiers, and links.
Business meaning
The decision at this layer
A node is the first capital unit that can connect accelerator capacity to accepted work. Peak HBM bandwidth alone does not prove useful capacity or margin.
Evidence
What is known, illustrated, and missing
- Architecture
- Source-backed
- Workload mapping
- Illustrative
- Live observation
- Not observed
Public architecture supports the layer. This atlas does not contain a selected runtime, kernel artifact, counter trace, power sample, or accepted-output run.
Public learning sources
Touch the math, then go deeper at the sources.
Start with the original deterministic micro-lab, then use the official public courses for the compute foundations. The atlas connects those foundations to HBM, full workload state, evidence, and accepted-output economics.
- Touchdown C-001 code-to-HBM micro-lab Step one pinned GLM-5.2 router-style matrix multiply through 46 deterministic states. Physical traffic, timing, power, energy, cost, and accepted outcome remain unobserved. Run the 46-state micro-lab
-
Accelerated Computing Academy, Fall 2025
GPU execution, memory behavior, tiling, scheduling, tensor cores, TMA, WGMMA, PTX/SASS, and measured kernel work.
Open the official course site
Open the official labs index - Stanford CS336: Language Modeling from Scratch Tokenization, model architecture, systems, scaling, data, evaluation, and the end-to-end language-model training path. Open the official CS336 site
Rights boundary: this file contains no copied slides, videos, assignments, starter repositories, or course artwork. Public availability does not by itself grant redistribution rights. Follow the official links and each source's own terms.
No-JavaScript and print reference
The complete physical path remains visible.
The interaction enlarges one layer at a time. It does not remove the rest of the stack.
- Compute node
- GPU or accelerator package
- Streaming Multiprocessor or Compute Unit
- Warp or wave scheduler
- Tensor Core or MFMA matrix engine
- Registers and shared memory or LDS
- L2 cache
- Memory controller
- HBM stack, channels, banks, and cells
- Fabric, offload path, and destination tier
- Power, energy, heat, cooling, water boundary, and cost