Short answer. High Bandwidth Memory, or HBM, is the fast working memory placed beside an AI accelerator. The accelerator repeatedly needs model weights and temporary working data. HBM stacks DRAM dies vertically and connects them through dense vertical wires, base-die logic, and a very wide package interface, so many bytes can move in parallel over a short physical path. HBM is excellent for data that must stay close and arrive quickly. It does not make every data access fast, remove capacity limits, or replace software decisions. The system still has to decide what belongs in HBM, what can live elsewhere, when data should move, and whether that movement helped a real task finish correctly.

July 23, 2026 update: AMD Helios is now part of the production memory story. At Advancing AI 2026, Lisa Su said the Helios reference design is in full production. AMD says shipments are scheduled to start at the end of Q3 and ramp through Q4. One Helios rack combines 72 launched MI455X accelerators across 18 four-GPU trays, with 432 GB of HBM4 and 23.3 TB/s of published peak HBM bandwidth per accelerator, or 31 TB and 1.7 PB/s of published aggregate peak HBM bandwidth at rack scope. That gives an operator more room for hot weights, KV state, activations, and communication buffers, but it does not end the memory problem. Useful performance still depends on placement, kernels, ROCm qualification, fabric behavior, power, cooling, packaging, supply, and a workload receipt. Read the complete Helios and MI455X system path, including what AMD proved, what remains vendor-rated, and what has not been observed on a Touchdown run. Primary sources: AMD's Helios launch update, AMD's MI400 launch update, and the official Advancing AI keynote.

July 25, 2026 update: MI350P and the RTX 6000 family now have a workload comparison. The name RTX 6000 can mean four different cards. The new comparison and input-driven calculator keeps RTX 6000 Ada, RTX PRO 6000 Blackwell Workstation, Max-Q, and Server Edition separate. It follows one agent workload through model weights, KV state, HBM3E or GDDR7, host DDR or LPDDR5X, PCIe movement, measured throughput, power, energy, and cost. Vendor specifications are prefilled. Tokens per second, tokens per watt, price, and cost per accepted task stay unknown until the reader supplies a matching receipt.

If you remember one idea, remember this: HBM is the GPU's fast working area, not the entire storage system. The whole point is not the largest bandwidth number on a specification sheet. It is more accepted AI work per dollar, watt, server, and minute without breaking quality or reliability.

Three ways to read this article

Path Who it is for Start with Decision to leave with
New to HBM Anyone who wants the complete mental model without assuming a hardware background Read this opening, then follow the nine-question map in order Explain what HBM is, why AI needs it, and why the workload still decides whether it helps
Business and investment CEO, CFO, investor, product leader, or operator Start with the investor concept map, follow one coding-agent task, then read energy and economics and the proof gate Identify which memory problem a product solves, where its limit moves next, and what receipt separates a real advantage from a specification-sheet claim
Engineering CTO, software, inference, kernel, GPU, memory, package, or facility engineer Start with useful bandwidth, choose a coding-agent or video trace, then use the technical appendix Name the exact software object, code path, hardware boundary, measurement, and missing proof

Every path ends with the same question: what receipt proves that this memory decision improved the accepted result?

The non-technical investor map: why each layer exists

Read this table from left to right. Each new layer solves one part of the memory problem and creates another boundary that still has to be measured. The point is not to memorize product names. The point is to understand why the products exist, what they actually own, and where an attractive claim can outrun its evidence.

Concept The problem that creates it What it actually does What it does not solve Follow the exact path
Memory wall Compute can request weights and working state faster than the complete memory system can deliver useful bytes Names the growing gap among arithmetic, capacity, latency, bandwidth, movement energy, and software locality It is not one bandwidth number and it is not fixed by buying one faster chip Start with the accepted task
HBM3E and HBM4 The hottest GPU state needs a wide, short, highly parallel path Stack DRAM beside the accelerator and connect it through vertical conductors, base-die interfaces, and dense package wiring HBM does not make capacity unlimited, make every access local, or turn peak bandwidth into application throughput Build one HBM stack
CPU and host memory Agent control flow, tokenization, retrieval, tools, files, networking, and operating-system work do not disappear when the GPU starts computing Hold control-heavy code and a larger system-memory tier, then coordinate work with the accelerator CPU memory is not local GPU HBM. Moving state between them has a topology, latency, bandwidth, energy, and software contract Follow the CPU tool loop
GPU, caches, memory controllers, and HBM A tensor must travel from a software allocation to the execution units and return as a useful result Turn operations into loads, stores, cache transactions, HBM commands, arithmetic, and synchronization A GPU label does not identify the selected kernel, physical traffic, sustained throughput, or accepted-task result Trace source code to useful bandwidth
Scale-up links and switches One accelerator may not hold all hot weights and state, so nearby accelerators must exchange data NVLink and NVSwitch, Infinity Fabric, or UALink-class paths connect defined accelerator domains Link bandwidth is not HBM bandwidth, collective latency, or one flat shared memory pool Separate the fabric levels
NCCL and RCCL Software needs collective operations such as all-reduce, all-gather, reduce-scatter, and all-to-all across accelerators Choose and execute communication algorithms over available transports and topologies They are software libraries, not physical wires, NICs, or proof that a named link carried the bytes Trace the collective layer
NIXL and movement policy Disaggregated inference needs to identify and move tensors or KV state across different storage and transport backends Provides a transfer abstraction that can orchestrate a selected backend NIXL is not NVLink, not a cache policy, and not the physical network Follow an offload path hop by hop
NICs and DPUs such as BlueField or Salina Network, storage, security, virtualization, and data-movement work can consume CPU time and complicate the accelerator path Terminate and process named infrastructure functions at the network or storage boundary A DPU does not add local HBM capacity, become an AI NIC by default, or prove lower task cost Place the DPU in the hierarchy
CMX, VAST, CXL, host DRAM, and NVMe tiers Reusable context can be too large or too cold to justify permanent HBM residency Hold state farther from the GPU and restore it before its reuse deadline when that is cheaper than recomputation A lower tier never becomes HBM. Capacity helps only when identity, transfer, restore, quality, p99 latency, energy, and failure behavior close Compare placement choices
eDRAM, custom HBM, XBM, memory-on-logic, near-memory compute, NAND, and HBF Standard HBM may still lose on capacity, energy, logic flexibility, package area, cost, or workload-specific movement Change the distance, interface, logic boundary, media, or granularity for a particular state role None is a universal HBM replacement. Each changes thermal, yield, repair, software, qualification, supply, and economic obligations Compare memory roles
The joined receipt Specifications, source code, demos, and vendor benchmarks can describe different systems and different proof levels Joins one workload, input, software path, hardware topology, counters, time interval, power boundary, output verifier, and cost denominator It does not generalize beyond the named run without further evidence Read the proof gate

Place the current product generations on that map

System boundary Memory role Why it belongs in this article Evidence boundary
Grace plus Blackwell B200 or GB200 Grace LPDDR5X holds CPU-side state; Blackwell HBM3E holds the hot GPU working set; NVLink-C2C and the rack fabric connect declared scopes This is the active NVIDIA teaching baseline for the current coding-agent path Product and architecture facts are source-backed. The public C-001 GPU run remains uncaptured
Vera plus Rubin Vera expands the control and CPU-memory side; Rubin moves the hot tier to HBM4; BlueField, CMX, NIXL, NVLink, and networking occupy separate movement and persistence roles This shows why an accelerator roadmap becomes a memory, fabric, CPU, storage, power, and cooling roadmap Preliminary or roadmap claims remain vendor-scoped until delivered systems and matched workload receipts exist
AMD MI455X plus Helios MI455X uses HBM4; Helios combines 72 accelerators with Venice CPUs and named scale-up and scale-out boundaries This is the current AMD production-reference comparison and keeps AMD first class instead of treating CUDA as the whole market AMD's product and rack ceilings are vendor-published. A joined C-001 or V-001 MI455X run remains uncaptured
Alternative and custom memory paths eDRAM, custom HBM, XBM, CXL, NAND, HBF, near-memory compute, and memory-on-logic target different state lifetimes and movement costs These paths test where the memory wall might move after standard HBM and software optimization Patent, paper, roadmap, fixture, and proposed states remain separate from shipping and measured evidence

Start here if this stack is new

Imagine asking an AI coding agent to fix a bug. Before it can return a useful patch, the system has to collect your instructions, read files, turn text and code into model inputs, load learned numbers called weights, create temporary working data, run GPU operations, call tools on the CPU, and check the final change. A video generator follows a different sequence, but it also carries large amounts of working data through repeated GPU operations before anyone can accept the result.

Modern GPUs can perform arithmetic faster than a slow memory path can supply the required numbers. When the next weight, token state, image tile, or temporary result arrives late, compute units wait. HBM exists to make the hottest part of that path wider, shorter, and more parallel.

This article uses two real open-source workload shapes so the hardware never floats away from the user:

  • C-001 is the primary example. A developer asks Hermes Agent, backed by GLM-5.2, to inspect a repository, change code, run tests, repair failures, and return a verified patch. Hermes is the agent system that connects the language model to files, tools, terminal commands, memory, skills, and subagents. GLM-5.2 is a sparse Mixture-of-Experts model. That means it contains many expert parameter blocks, while a router selects only some of them for each token. Its official FP8 configuration records that architecture and numeric format.
  • V-001 is the contrast example. Wan2.2 starts with a noisy video representation and repeatedly changes it until a decoder produces a clip. Its A14B variants use different experts during different parts of that denoising process. A quality gate is the named pass/fail check applied to the output. It is not a promise that every generated clip is useful.

By the end, you should be able to answer five questions without guessing:

  1. What data does the workload create, and how long must each object remain available?
  2. Which software layer requests the operation, and which kernel actually runs?
  3. Which bytes stay on the GPU, reach HBM, cross a fabric, or fall to a lower tier?
  4. Where do power, heat, cooling, water, supply, and money enter the same task?
  5. Which facts came from a vendor, which came from source code, which were measured, and which remain unknown?

You only need three technical nouns to start. A tensor is a typed multidimensional array, which is how software represents model weights and working data. A kernel is a device function launched on a GPU so many threads can perform one bounded operation in parallel. A verifier is the rule or program that checks the real output and returns accept or reject. The tensor definition follows PyTorch's tensor contract; the kernel definition follows the CUDA programming model. Verifier is this article's accepted-task contract, not a GPU component.

TEXT
user asks for useful work
-> software turns the request into model-readable state
-> a serving system schedules operations
-> kernels move and transform tensors
-> GPU compute and memory hardware do electrical work
-> the system may call tools, move state, retry, or wait
-> a verifier accepts or rejects the result

The article keeps that one chain intact. When the left side highlights a code line, an HBM stack, a fabric link, a cooling loop, or a cost field, it is showing another boundary of the same accepted task. It is not changing to an unrelated example.

The whole article in nine questions

Question Plain answer Go deeper
1. What does the user need? A patch or clip that passes a named acceptance check Start with the user
2. What must the machine remember? Weights, instructions, active working data, reusable state, tool results, and outputs Classify the state
3. What is memory physically? Electrical state stored, sensed, restored, moved, and refreshed by circuits Follow one DRAM bit
4. What makes HBM different? Stacked DRAM, dense vertical connections, many parallel channels, and a short package path beside the accelerator Build the HBM stack
5. What does software control? Allocation, layout, batching, reuse, kernels, caching, and movement policy determine useful bandwidth Trace useful bandwidth
6. How do real workloads differ? A coding agent grows reusable language-model state; video diffusion repeatedly transforms large temporary state Coding agent · Wan2.2 video
7. When should data move? Only when the capacity saved is worth the transfer time, energy, risk, and restore deadline Compare memory tiers
8. Where does the business cost appear? In hardware capacity, waiting, retries, energy, cooling, failed output, and human review Trace energy and economics
9. What proves the decision? One joined record of the workload, software, hardware, measurements, outcome, and caveats Read the proof gate

The smallest complete inference loop

Inference means running a trained model to produce an output. It is different from training, where the system changes model parameters. A real inference product contains more than a model call:

What the person sees What the machine does
The user submits a request The application gathers instructions, permissions, tools, files, retrieved material, and prior state
The system appears to think and act The tokenizer, serving engine, model, GPU kernels, memory system, CPU tools, storage, and network take turns doing work
The user receives a patch or clip A verifier checks the real output, records accept or reject, and decides whether another paid attempt is needed

The detailed engineering loop expands those three visible moments:

  1. The application collects the user request, system policy, tool definitions, retrieved documents, repository state, and prior conversation.
  2. A tokenizer maps text or code into integer token IDs. A token is a model unit, not necessarily a whole word.
  3. A serving engine admits the request, chooses a batch and placement, and schedules model work on one or more accelerators. The vLLM architecture overview is one concrete example of how an engine separates the API server, engine core, scheduler, KV-cache manager, and model executor.
  4. The model reads learned weights and transforms tensors. A tensor is a multidimensional array with a shape, data type, layout, and physical allocation.
  5. Prefill processes the supplied input tokens and creates reusable attention state. Decode produces later tokens incrementally, usually one generation position per sequence at a time even when many sequences are batched.
  6. The runtime may call a tool on CPUs, storage, a network, or a sandbox. The tool result returns as new context and may create another prefill and decode cycle.
  7. A verifier checks the actual outcome: tests pass, a patch is reviewable, or a clip passes the named quality gate.

Use this physical-location key before following the hardware path:

Term Physical meaning in this article
Host memory CPU-side system memory, normally DRAM. CUDA defines host memory as memory directly connected to the CPU. Unified-memory and system-on-chip designs can blur residency, so a receipt must still name where the bytes physically lived.
HBM Stacked DRAM connected close to the accelerator through dense package wiring. It is the hot device-memory tier in the accelerator examples, not an on-chip cache.
SM or CU A repeated GPU execution block. NVIDIA calls it a streaming multiprocessor, or SM. AMD calls its fundamental execution block a compute unit, or CU. These blocks contain schedulers, registers, execution pipelines, and local memory structures that execute kernel work.
Workspace An operation-specific temporary read/write buffer allocated by a runtime or library. In the GPU examples it commonly consumes device memory even though it is not a model weight, persistent KV block, or accepted output.

Primary definitions come from the CUDA programming model, AMD's GPU hardware model, Samsung's HBM overview, and the current NVIDIA cuDNN Backend API, whose operation-graph and execution-plan interfaces expose operation-specific workspace sizes.

The hardware path runs underneath those steps at the same time:

TEXT
token IDs and metadata in host memory
-> model scheduler and device commands
-> weights, KV cache, activations, and workspaces in GPU memory
-> kernels on NVIDIA streaming multiprocessors or AMD compute units
-> cache, memory-controller, HBM-channel, bank, row, and cell activity
-> links to another GPU, host, storage tier, or network when state is not local
-> electrical energy, heat, cooling work, and facility allocation

A GPU is not one large arithmetic box, and HBM is not a bucket inside it. The accelerator package contains compute blocks, registers, caches, schedulers, memory controllers, and link endpoints. HBM packages sit beside the accelerator on a dense package fabric. The controller turns requests into commands across channels, banks, rows, and columns. The physical details vary by vendor and generation, so diagrams in this article show functional boundaries unless a source says more.

Six words you need now, then a complete reference

You do not need to memorize the full stack before continuing. Keep these six words in view:

  • A GPU is a parallel processor with many compute blocks.
  • HBM is the GPU's nearby high-bandwidth working memory.
  • A tensor is an array of numbers with a declared shape and data type.
  • A weight is a learned number the model reads while producing an output.
  • An activation is temporary working data created while the model runs.
  • A KV cache is saved attention state that lets a language model reuse work from earlier tokens.

The full glossary below is a reference. Return to it when a term appears in a code, platform, or measurement section.

Term Plain-language meaning Where it lives in the path
AI inference Running a trained model to produce an answer, action, patch, image, or video The complete request-to-output process, not one chip operation
Model A structured program whose learned numeric parameters transform inputs into outputs Model code plus weight artifacts loaded by the runtime
Parameter or weight A learned number used by the model Stored in a model file, then placed in HBM or another memory tier for execution
Tensor A typed multidimensional array The software object that becomes physical allocations, loads, stores, and transfers
Shape The size of every tensor dimension Used by model code, compilers, kernels, and memory planners
Data type or precision The numeric encoding and number of bits used for a value BF16, FP8, FP4, integers, scales, and metadata affect capacity, traffic, math, and error
Token An integer unit consumed or produced by a language model Created by a tokenizer, stored in request state, and processed by model layers
Prompt and context The instructions and state supplied to the model System policy, user text, tools, retrieved data, repository files, and prior turns
Context window The model's configured limit or supported span of token positions A model/runtime contract; it is not a promise that every large request fits useful HBM capacity or latency
Layer A repeated model stage that applies attention, routing, normalization, or feed-forward work Model architecture and the sequence of operations executed by kernels
Attention A mechanism that lets token positions use information from other positions Q, K, V tensors, attention scores or equivalent algorithms, KV state, and attention kernels
Dense model Most or all parameters in a layer participate for each token Weight and compute demand is broadly active on each pass
Mixture of Experts, or MoE A router selects a subset of expert parameter blocks for each token Adds routing state, expert placement, grouped computation, imbalance, and communication decisions
Prefill Processing the input sequence to establish model state before incremental generation Often parallel across prompt positions; creates KV cache and can be compute intensive
Decode Producing subsequent token positions using prior state Often memory and latency sensitive; appends KV state and repeats many small steps
KV cache Saved attention key/value state from earlier token positions A mutable runtime allocation, commonly in HBM and sometimes restored or offloaded to lower tiers
Activation A temporary intermediate value produced while the model runs Registers, on-chip memory, HBM, or recomputed state depending on operation and schedule
GPU or accelerator A parallel processor built from many compute and data-movement resources The package containing compute blocks, caches, controllers, and high-speed link endpoints
Kernel A compiled device function that performs a bounded parallel operation Launched by the runtime and executed across GPU work groups, warps, or wavefronts
Cache A faster structure that keeps recently or strategically useful state near compute GPU registers, shared memory, L1/L2, software prefix caches, and KV tiers are different mechanisms
HBM Stacked DRAM beside an accelerator with a very wide local interface The hot external-memory tier on the accelerator package
Memory controller Logic that schedules and translates memory requests into device service Between GPU requesters and HBM channels, banks, rows, and timing rules
Interconnect or fabric Hardware and protocol that move data between components NVLink/NVSwitch, Infinity Fabric, PCIe, CXL, NICs, and scale-out networks have different scopes
Bandwidth Bytes that can be moved per unit time under a declared boundary A rate, not a latency or accepted-task result
Latency, p95, and p99 Latency is elapsed time between two named events. If 100 comparable requests are ordered from fastest to slowest, p95 is roughly where the 95th request finishes and p99 exposes the rarer slow tail. Formally, both are quantile estimates from a finite sample. Can describe a load, transfer, first token, token gap, tool call, or whole task, so the endpoints, sample population, sample count, interval, and percentile estimator must be named
Throughput Useful work completed per unit time Tokens, requests, clips, or verified tasks per second or hour
Concurrency and batch How many requests are active and which work is grouped together A scheduler decision that changes HBM occupancy, queueing, kernel shape, and tail latency
Offload and prefetch Moving state to a lower tier and bringing it back before it is needed HBM to host DRAM, CXL-attached memory, device memory, NVMe, or another node through a named path
Verifier and accepted task The rule that decides whether an output is actually useful Tests, review, quality checks, policy gates, and the denominator for cost and energy
Receipt A joined record of what ran, where, with which inputs, measurements, and outcome The evidence object connecting software, hardware, facility, and business claims

The terms are related but not interchangeable. A model is not a serving engine. A serving engine is not a kernel. A kernel is not an instruction. HBM capacity is not KV capacity. Link bandwidth is not memory bandwidth. A generated token is not a verified task.

Follow one line of code into the machine

Consider one PyTorch scaled dot-product attention operation:

PYTHON
out = torch.nn.functional.scaled_dot_product_attention(
    q,
    k,
    v,
    is_causal=True,
)

The line expresses a mathematical operation. It does not uniquely name the kernel, instruction sequence, cache behavior, or HBM command stream. The actual path depends on tensor shapes, dtype, layout, device, framework revision, serving engine, compiler state, backend selection, batching, and hardware.

Boundary What happens Evidence needed before making a physical claim
Python/model code The model requests scaled dot-product attention over q, k, and v Pinned source, input shapes, dtype, layout, and model revision
Serving engine vLLM, SGLang, or another engine chooses request grouping, pages, cache state, and model-runner inputs Engine config, scheduler trace, cache events, and request identity
Framework dispatch PyTorch and the selected backend choose an implementation compatible with the inputs and hardware Framework/backend versions and dispatch or profiler evidence
Compiler and executable A selected backend lowers the operation through one or more intermediate forms into a target-specific device executable Saved source, IR, binary, compile flags, target architecture, and disassembly where permitted
Kernel execution Threads cooperate to load tiles, compute attention work, synchronize, and store results Kernel name, launch geometry, device timeline, counters, and correctness result
On-chip data path Registers, shared memory or local data share, and caches serve some accesses Profiler counters or a bounded architecture model; not source-code inference alone
HBM path Misses and explicit transfers reach controllers, channels, banks, rows, and I/O Named hardware counters, simulator events, or vendor-scoped evidence
Fabric path Sharded state may cross a scale-up or scale-out link Topology, collective/transfer trace, endpoints, bytes, and timing
Product outcome The operation contributes to a token, tool step, patch, or clip Joined task trace and verifier result

Do not read compiler terms as peers. They occupy different levels:

TEXT
model operation
-> framework dispatch
-> backend or compiler
-> intermediate representation
-> device executable
-> hardware instruction
Term Category Exact role
CUDA NVIDIA platform, programming model, and toolkit CUDA includes runtime APIs, compilers, libraries, debugging tools, and optimization tools. It is not one instruction set or one binary format.
PTX NVIDIA virtual ISA and compiler intermediate form PTX is a low-level virtual instruction set that is translated offline or just in time into instructions for a target NVIDIA GPU. PTX is not the final target-specific machine binary.
cubin and SASS NVIDIA target binary and architecture-specific assembly A cubin is a binary file for a specific SM target. cuobjdump -sass disassembles cubin code into the architecture-specific CUDA assembly instruction stream commonly called SASS. The cubin file and its disassembled instructions are related but are not the same artifact.
HIP C++ runtime API and kernel language HIP is a source and runtime programming surface. On AMD ROCm, HIP-Clang or amdclang++ compiles device code through the AMDGPU backend. HIP is not an ISA.
LLVM IR and AMDGPU backend Compiler representation and target backend LLVM IR carries compiler-visible operations. The AMDGPU backend lowers compatible IR into AMDGPU ISA and packages executable machine code and metadata into an AMD GPU code object.
Triton Kernel language and compiler Triton is a language and compiler for custom GPU primitives. Its compiler uses MLIR and LLVM-based backend stages to produce target artifacts such as PTX or AMDGPU code. Triton is not CUDA, HIP, PTX, or a hardware ISA.

Primary references: CUDA Toolkit, PTX ISA, CUDA platform compilation, CUDA binary utilities, AMD HIP, AMD HIP compilers, LLVM AMDGPU backend, and the Triton compiler repository.

One source line can trigger many kernels. One kernel can execute the same instruction across many threads. One tensor read can hit a cache or reach HBM. A controller can reorder physical requests while preserving the programming contract. That is why the left visual uses color to show the active boundary but keeps unknown values marked unknown.

How to read the colors and diagrams

  • Blue active marks show the boundary currently being explained. They do not claim that every byte followed that path in a measured run.
  • Yellow text marks the accepted-task or executive decision boundary.
  • Gray text or panels expose engineering conditions, assumptions, and lower-level mechanisms.
  • Dashed marks and ? labels mean unknown, unmeasured, configuration-dependent, not public, or not applicable until a receipt resolves the field.
  • Side-by-side NVIDIA and AMD slots compare the same functional question. They remain blank or qualified when the public evidence scopes differ.
  • Animated particles, fills, or heat paths are schematic state transitions. A visual becomes a measured trace only when it is bound to a named replay receipt, counter source, clock interval, and topology.

The visual can simplify geometry. It cannot simplify evidence. If a diagram shows a GPU block lighting up, read it as this is the functional area under discussion, not we measured these exact transistors switching.

How to read the evidence labels

These labels stop a product announcement, a source-code path, a model, and a measured run from being treated as the same proof.

  • Measured or captured evidence includes [PAPER, MEASURED] and a joined workload receipt when one exists.
  • Shipping or vendor-reported evidence includes [SHIPPING VENDOR ARTIFACT] and [VENDOR TARGET / BENCHMARK]; the vendor and setup remain part of the claim.
  • Modeled or derived evidence includes [PAPER, MODELED] and [TOUCHDOWN DERIVATION]; the assumptions remain visible.
  • Proposed evidence includes [PATENT APPLICATION], [FIXTURE BACKED], and [PROPOSED]; these show a design, code path, or testable hypothesis without proving live workload execution.
  • Unknown evidence includes [NOT FOUND]; a missing receipt stays missing.

The exact definitions used throughout the article are:

  • [STANDARD] means an unrestricted public standards-body page or announcement defines a capability or term. It does not prove a product or workload result.
  • [SHIPPING VENDOR ARTIFACT] means a vendor says a named commercial artifact ships or is in production. Its performance claims remain vendor-reported unless independently measured.
  • [VENDOR TARGET / BENCHMARK] means a roadmap, internal test, or vendor benchmark. It applies only to the named setup.
  • [PATENT APPLICATION] means an architecture appears in a published application. It does not prove working silicon, a product, yield, cost, performance, or schedule.
  • [PAPER, MEASURED] and [PAPER, MODELED] retain the paper's exact apparatus. A modeled chip is not fabricated silicon.
  • [AUTHOR WORKLOG] is a public engineering report without independent reproduction.
  • [TOUCHDOWN DERIVATION] is transparent arithmetic from cited inputs, not a measurement.
  • [FIXTURE BACKED] means code or a schema passes saved tests. It is not live hardware proof.
  • [PROPOSED] means a hypothesis or experiment that still needs evidence.
  • [NOT FOUND] means a required public receipt was not located.

That distinction matters because HBM is shipping technology, Intel XBM is a patent application, Sandisk HBF is a proposed direction with vendor simulation, and a Touchdown state contract is a research hypothesis. Putting all four in one table without those labels would manufacture certainty that does not exist.

What breaks when an AI system cannot keep the right state close to compute?

In plain English. An AI system can own a fast GPU and still feel slow or expensive because the GPU keeps waiting for data, recreating data it already had, or moving data through the wrong tier. We begin with the result a person will accept, then work backward to the exact state that missed its deadline.

Start with the user.

A designer requests an 81-frame video. The clip is useful only if it follows the prompt, looks coherent, arrives before the product deadline, has the expected format, and costs an amount the product can support.

A developer asks a coding agent to inspect a repository, change a function, run the tests, repair a failed patch, and return a reviewable diff. The answer is useful only if the change is correct, the tests pass, the patch does not quietly break another path, and the developer does not spend the saved time debugging the agent.

Neither user buys terabytes per second. Neither user buys FLOPs. They buy an accepted clip or an accepted code change.

Memory enters because both workflows create state that has to remain available across time. The video request has prompt embeddings, model weights, a noisy latent, Q/K/V tensors, normalization and feed-forward intermediates, attention metadata, communication buffers, and decoded frames. The coding request has system instructions, tool schemas, a repository map, selected files, model weights, active KV cache, reusable prefixes, tool results, patches, test logs, and prior decisions.

Some of that state is tiny and hot. Some is large and read-only. Some is append-only. Some changes every step. Some can be regenerated. Some must survive a process crash. Some needs to arrive within microseconds or milliseconds. Some can wait for a storage read. The memory decision begins there, not with a product name.

The state object comes before the tier

Start with one concrete object: the KV cache is the language model's working record of earlier tokens. Before deciding where it belongs, the system needs to know whose request it belongs to, how large it is, whether it will grow, when it will be read again, how quickly it must arrive, and whether it can be rebuilt. The same questions apply to a video latent, a model-weight tile, a tool result, or a checkpoint.

Before saying HBM, CXL, host memory, or NVMe, write down the object:

Field Actual question
Identity Is this a weight tile, latent, Q/K/V tile, KV block, stable prefix, tool result, repository map, checkpoint, or output?
Shape and bytes What is its real shape, dtype, layout, padding, allocation size, and replication?
Access Is access sequential, strided, gathered, sparse, random, or broadcast?
Mutability Is it read-only, append-only, read-write, disposable, or durable?
Reuse How soon will it be read again, how often, and by how many consumers?
Deadline What first-byte and full-object p95/p99 deadline applies?
Correctness Must it be exact, can it be reconstructed, or is bounded approximation allowed?
Movement What source, target, interconnect, granularity, and fallback are permitted?

That table prevents three common errors.

First, model memory is not one thing. Read-only weights, mutable KV, temporary attention tiles, optimizer state, and checkpoints have different contracts.

Second, more capacity is not automatically useful. A lower tier can hold an object and still lose if the object arrives after the compute window, if page overfetch dominates, or if the miss path creates a p99 spike.

Third, faster memory is not automatically more valuable. HBM space occupied by cold repository artifacts can reduce concurrency for active inference state. A carefully selected lower tier can free the hot tier, but only if transfer and reuse are measured.

Where the leak appears

For video diffusion, the failure can appear as a model offload between steps, a latent spill, repeated materialization of intermediates, a gather pattern that does not use the memory channels well, or a collective that leaves compute waiting. The result is a longer denoising loop, less concurrency, more GPU-seconds, or a quality-losing optimization.

For a coding agent, the failure can appear as repeated prefill over the same repository prefix, evicted KV blocks, cold tool results, a blocked CPU sandbox, a failed patch, or another model call needed to recover state that was already known. The result is longer time to first useful token, more tool-loop latency, more retries, and more human review.

The accounting boundary is the accepted task:

TEXT
total_cost_per_accepted_task_USD =
  (accelerator_cost_USD
   + CPU_and_tool_cost_USD
   + memory_storage_network_cost_USD
   + facility_energy_and_water_cost_USD
   + retry_and_failure_cost_USD
   + human_rework_cost_USD)
  / accepted_tasks

[PROPOSED LEDGER] This is a cost-allocation shape, not a claim that every term is easy to observe. Every numerator field is denominated in USD before addition. The raw receipt must separately preserve accelerator and CPU/tool seconds; allocated and transferred bytes by tier and link; joules or kWh; water liters by declared boundary; retry and failure counts; accepted-output count; and verifier result. Named prices and allocation rules convert those physical quantities into cost. Seconds, bytes, joules, liters, and counts must not be added directly. If accepted_tasks is zero, the ratio is undefined and that state must be reported explicitly.

The technical path in this article will fill in that ledger one boundary at a time:

TEXT
user request
-> software state
-> runtime and kernel
-> cache and controller
-> HBM channel, bank, row, and cell
-> package power and heat
-> server and facility
-> accepted output
FIGURE 1 1600 × 900
Two timelines compare a Wan2.2 video request and a coding-agent task. The video path carries prompt embeddings, weights, a mutable latent, attention intermediates, and decoded frames. The coding path carries a stable prefix, active KV blocks, tool results, patches, tests, and an accepted diff.
The memory decision begins with the lifetime and access pattern of state, not the name of a memory product. Source state. Official workload paths plus Touchdown classification. No performance result. Geometry and timing are illustrative.

What this means by role. The CEO defines accepted output. The CFO asks where repeated work is paid twice. The CTO identifies which state limits concurrency and p99. The software engineer traces the object through allocation and movement APIs. The kernel engineer identifies the load, store, gather, or synchronization that creates the bytes. The hardware engineer asks what address distribution and service deadline reach the controller.

Now we can descend the stack. What does it mean to store one bit in the first place?

What is one bit of memory, physically?

In plain English. A DRAM bit begins as a tiny amount of electrical charge. A useful bounded analogy is a leaky cup connected to a shared measuring wire by a switch. The capacitor is the cup, the access transistor is the switch, the bitline is the measuring wire, and the sense amplifier reads the weak signal and restores what the read disturbed. The analogy stops there: the real device is an electrical circuit governed by voltage, capacitance, noise, timing, temperature, and fabrication variation.

A bit is not a microscopic printed 0 or 1. It is a physical state that a circuit agrees to interpret as one of two logical values.

In a dynamic random-access memory cell, that state is charge associated with a tiny capacitor. In a static random-access memory cell, it is the stable state of a feedback circuit. In NAND flash, it is related to the threshold behavior of a storage device. The software abstraction is the same bit. The physical promises are different.

Micron's Introduction to Memory is the foundational public teaching source for the DRAM cell, sensing, row and bank organization, read and restore, and refresh path summarized here. [OFFICIAL VENDOR EDUCATION]

Charge, voltage, and a DRAM cell

The first useful relation is:

TEXT
Q = C * V

Q is stored charge, C is capacitance, and V is voltage. The equation does not describe an entire memory system. It explains why a very small capacitance and small stored charge produce a small signal that must be sensed carefully.

A conventional DRAM cell is often described as 1T1C: one access transistor and one capacitor. The access transistor behaves like a controlled switch. Its gate connects to a wordline. One side of its channel connects the cell capacitor to a bitline.

Before a read, the bitline is prepared to a reference condition. The controller selects a row. A row decoder raises the corresponding wordline, turning on many access transistors along that row. Each selected cell shares a small amount of charge with its bitline. The bitline voltage moves slightly above or below its prepared level depending on the stored state.

That tiny movement is not yet a robust digital value. A sense amplifier compares and amplifies it to a full logic level. The row's resolved values become available in the row buffer, which is built from the active sense-amplifier state.

The read is destructive in the practical sense that charge sharing disturbs the cell. The sense amplifier therefore also restores the value into the capacitor while the row remains active. Later, the array precharges its bitlines so another row can be opened safely.

The complete local path looks more like this than a software load:

TEXT
precharge bitline
-> raise wordline
-> cell and bitline share charge
-> sense amplifier resolves small delta
-> row buffer holds full logic state
-> column path selects requested data
-> sense amplifier restores cell
-> close row and precharge for later access

A write reverses part of the intent. The write circuitry drives the desired bitline state, opens the access device, and charges or discharges the cell toward the new value. The array still has timing requirements for activation, write recovery, precharge, and interference control.

Why DRAM needs refresh

The cell is not a perfect bucket. Charge leaks through device and dielectric paths. Temperature, process variation, cell history, coupling, and disturbance affect how quickly a usable sensing margin shrinks.

DRAM therefore refreshes stored rows. A refresh operation restores charge even if software did not request the data. Refresh consumes command opportunities, internal activity, and energy. It can interact with tail latency because a demand request may arrive while some internal resource is unavailable.

It is still misleading to say refresh makes DRAM slow as if one number explains the system. Modern devices organize refresh and parallelism in implementation-specific ways. The observed penalty depends on generation, temperature, bank activity, controller policy, workload, and measurement boundary. The safe statement is narrower: refresh is mandatory work that a useful-bandwidth and energy model must include.

SRAM, eDRAM, and NAND make different bargains

SRAM usually stores a bit in a multi-transistor bistable circuit. As long as power remains and operating conditions hold, feedback maintains the state without periodic DRAM-style refresh. SRAM can provide fast, fine-grained access and sits close to compute in registers, caches, and scratchpads.

The cost is area and leakage. A multi-transistor cell takes much more silicon per bit than a dense DRAM cell. That is why an accelerator can have a valuable on-chip cache but not replace all HBM capacity with the same kind of SRAM at the same die size and cost.

Embedded DRAM, or eDRAM, puts dynamic cells closer to logic in a process designed to support them. It can offer more capacity than an SRAM structure in a given role while retaining dynamic-memory behavior such as refresh. It is not SRAM but denser, and it is not HBM on the logic die. Process integration, capacitor construction, thermal behavior, peripheral circuits, refresh policy, test, and yield all change.

NAND flash makes a different trade. It stores nonvolatile state and reaches very high density, but its useful access and update units are much coarser. Reads are page-oriented. Programming and erasing have their own granularity, latency, endurance, and controller requirements. A fast sequential flash stream is not a fine-grained DRAM load. That distinction will matter when we reach HBF and Carmack's deterministic-page argument.

Cell or medium Retention behavior Read/write character Density direction Typical role here
SRAM Volatile while powered, no periodic DRAM refresh Fast, fine-grained read and write Lowest density in this comparison Registers, caches, scratchpads
DRAM family Volatile, refresh required Fine-grained external interface backed by row operations High Working memory: HBM, DDR, LPDDR, GDDR
eDRAM Volatile, refresh required Near-compute dynamic working memory Between SRAM and external DRAM in role, process-dependent Workload-specific integrated memory
NAND Nonvolatile Page reads, programmed pages, erased blocks Very high Deep capacity, storage, read-mostly streams

Energy is a chain of events, not one formula

Two physical approximations are useful:

TEXT
dynamic_switching_energy scales with alpha * C * V^2
wire_delay scales with R * C

alpha is activity, C is switched capacitance, V is voltage, R is resistance, and RC is a first-order way to understand why longer, narrower, or more heavily loaded wires are harder to drive quickly.

Lower voltage can reduce switching energy, but it also changes noise margin, sensing, timing, and circuit design. Shorter wires can reduce capacitance and delay, but dense packaging adds routing, coupling, thermal, mechanical, and manufacturing constraints. Narrow interconnects can raise resistance. High simultaneous current can create voltage droop and I^2R conduction loss.

Real memory energy includes wordline drivers, sense amplifiers, row activation, restore, precharge, refresh, local and global buses, clocks, I/O, PHYs, controller work, error correction, and leakage. At the system level, it also includes link endpoints, voltage conversion, fans or pumps, and facility overhead. A C V^2 estimate is physical intuition. It is not a measured joule-per-token result.

FIGURE 2 1600 × 900
Four panels show a precharged DRAM bitline, wordline activation and charge sharing, sense amplification into a full logic level, and restoration of the cell before precharge. Small side diagrams contrast an SRAM feedback cell and page-oriented NAND storage.
A software bit can be stored under different physical contracts. A DRAM read senses a small charge difference, amplifies it, and restores the disturbed cell. Source state. Source-backed educational synthesis. Geometry is not to scale and does not represent a specific vendor process.

What this means by role. For software, one load hides an array-level sequence. For a memory engineer, sensing margin, refresh, timing, and variation are first-order constraints. For a CFO or investor, cell area, retention work, yield, and process complexity determine how much usable capacity can be built and powered. The logical bit is the beginning, not the product.

How does a DRAM bit become an HBM stack beside a GPU?

In plain English. HBM does not use a magical new kind of bit. It makes ordinary DRAM more useful to an accelerator by stacking dies, dividing work across many banks and channels, connecting the stack vertically, and placing it beside the accelerator on dense package wiring. The result is a much wider local path. The cost is harder packaging, power delivery, cooling, testing, repair, and yield.

High Bandwidth Memory is not a different logical kind of memory. It is a tightly organized and packaged form of DRAM designed to expose a very wide aggregate interface near compute.

That sentence contains the main engineering idea. HBM does not get its value from one magical fast cell. It gets value from hierarchy, parallelism, proximity, a wide interface, vertical die connections, a logic/interface layer, and a package that connects memory to an accelerator at high density.

Samsung's public HBM overview illustrates the stacked-DRAM and TSV direction. SK hynix and TSMC's base-die announcement establishes the public logic-base-die design surface for named HBM development. [OFFICIAL VENDOR EDUCATION / ROADMAP]

From one cell to many banks and channels

Cells sit in arrays. Arrays are divided into structures that let the device activate some storage while other storage can serve or prepare different work. Terms vary by generation, but the useful hierarchy is:

TEXT
cell
-> local bitline and sense amplifier
-> row buffer
-> subarray
-> bank
-> bank group and pseudo-channel organization
-> channel
-> die interface
-> vertical stack connection
-> base or interface logic
-> package route
-> accelerator memory controller

A bank is an independently managed array region with local sensing and row state. When a row is activated, the bank's row buffer holds the sensed row. A read or write command then selects columns from that open row.

If the next access targets the same open row, the device can use the row-buffer state. That is a row hit. If it targets a different row in the same bank, the current row may need to close and the new row to activate. That is a row conflict. If independent requests map across different banks or channels, the controller may overlap more of their work.

This is why an address is also a scheduling decision. The physical address bits, memory-controller mapping, tensor layout, stride, alignment, and request order influence which channels, banks, rows, and columns receive traffic. A large contiguous tensor can still use the hardware badly if its stride or mapping concentrates requests. An irregular gather can expose latency even when the device's headline bandwidth is enormous.

The command path includes operations commonly abbreviated as:

  • ACT: activate a row into sensing/row-buffer state.
  • RD: select and return data from an active row.
  • WR: select and update data in an active row.
  • PRE: precharge or prepare the bank to open another row.
  • REF: perform required refresh work.

Each has timing relationships. Commands cannot be issued arbitrarily close together. Power limits can also constrain how many activations occur in a window. Read-to-write and write-to-read direction changes consume bus and device time. Error correction, repair, and reliability mechanisms add logic and sometimes traffic.

The important result is not that one operation is always expensive. It is that useful service depends on the distribution and scheduling of many operations.

HBM makes the path wide and local

Conventional host DRAM modules favor capacity, serviceability, and a CPU-oriented memory system. Graphics memory favors high per-pin rate across discrete memory devices. HBM takes a different physical route: stack multiple DRAM dies, create vertical connections through the stack, expose many narrower channels or pseudo-channels, and place the stacks beside the accelerator on a dense package fabric.

A wide local interface can move many bits per transfer without requiring every pin to run at the highest possible signaling rate. More independent channels also give the controller more places to schedule outstanding work. The short package route reduces the electrical distance relative to board-level memory interfaces.

The theoretical interface calculation is simple:

TEXT
peak_payload_bandwidth =
  transfers_per_second * interface_width_bits / 8

The word payload matters. A vendor may report an interface data rate, a per-stack bandwidth, a device bandwidth, or a whole-accelerator aggregate. The number may be based on a maximum rate rather than a sustained workload. The comparison is valid only when the scope, encoding, direction, and count of stacks match.

Useful bandwidth is a different quantity:

TEXT
useful_bandwidth = useful_bytes_completed / wall_clock_time

utilization_of_peak = useful_bandwidth / peak_payload_bandwidth

Useful bytes must also be defined. If a kernel reads the same weight three times because tiling is poor, physical traffic can rise while useful tensor bytes stay unchanged. If a cache serves a request, HBM traffic can fall while the kernel still completes. If compression reduces link bytes but adds computation, the user may win or lose depending on the deadline.

Through-silicon vias are vertical wires, not free bandwidth

A through-silicon via, or TSV, is a conductor that passes vertically through thinned silicon. At a functional level, the structure needs a conductive path, electrical isolation from the silicon where required, reliable contacts to routing layers, and a geometry that can survive fabrication, thinning, bonding, thermal cycling, and current stress.

TSVs let signals and power travel through a stack rather than escaping every die to a wide board perimeter. They also consume area, create keep-out and stress concerns, add capacitance and resistance, and require testable connections. The exact formation sequence, metal system, dimensions, liner, barrier, and reveal process depend on vendor and generation. A research result for one narrow interconnect material is not evidence that an HBM vendor uses it throughout a stack.

The stacked dies need electrical and mechanical bonds. Shipping generations may use microbump-based connections, and industry development is moving toward finer-pitch direct or hybrid-bonding directions in advanced packages. It is unsafe to write that every HBM4 product uses one specific bond flow. The article therefore names the function and cites a named product when discussing its implementation.

The base die is where memory and logic meet

At the bottom of an HBM stack sits a base or interface die. Its job is not identical in every design. Broadly, it connects the vertical stack to the package interface and hosts interface, control, test, repair, RAS, and physical-layer functions according to the generation and vendor partition.

This is becoming technically important because advanced logic on the base die can do more than terminate wires. [VENDOR DESIGN SURFACE / PROPOSED] In a named custom design, it could host selected address translation, movement control, security, telemetry, reliability, or carefully bounded near-memory operations. Those functions do not exist in every HBM4 base die. That is the reason custom HBM and SPHBM4 matter later in this article.

The opportunity has a physical cost. Logic consumes power and area. It creates heat below a stack of temperature-sensitive DRAM. It needs a qualified process, known-good-die strategy, clock and power delivery, test access, repair behavior, and software contract. Moving a function to the base die is not automatically better than running it in the accelerator, controller, or software.

[SHIPPING VENDOR ARTIFACT] Samsung says its commercial HBM4 uses a logic base die manufactured on its 4 nm process, and Micron says it has begun volume production of a named HBM4 product for a named customer program. Those statements establish vendor-reported commercial artifacts as of their announcements. They do not make every performance or power number independent, and they do not establish a standard partition for all HBM4.

The interposer or package fabric closes the path

The accelerator and HBM stacks need dense horizontal wiring. A silicon interposer is one established way to provide that routing. Bridge and redistribution-layer package approaches create other topologies. The dense layer connects to an organic substrate, which connects the package to the board and the rest of the system.

The package therefore carries more than data:

  • Power arrives from board-level voltage regulation through planes, vias, package structures, bumps, and on-die grids.
  • Return current needs controlled paths.
  • Clocks and commands need timing margin.
  • Simultaneous switching creates noise and voltage droop.
  • Dense neighboring wires create crosstalk.
  • Logic and memory heat must cross thermal interfaces into a heat spreader, cold plate, air path, or liquid path.
  • Mechanical expansion, stack height, die thickness, underfill or mold, and warpage affect assembly and reliability.

HBM puts memory near compute electrically, but also thermally and mechanically. That coupling is part of the value and part of the constraint.

FIGURE 3A 1600 × 900
A GPU memory request moves through address translation and cache checks, a memory-controller queue, channel and bank mapping, row activation, cell sensing, column selection, correction, and return to the requesting kernel.
One logical load becomes a scheduled sequence across caches, a controller, an HBM channel, a bank, a row buffer, and a cell array. Source state. Architecture synthesis. The sequence is illustrative and omits implementation-specific details.
FIGURE 3B 1600 × 900
An illustrative package cross-section shows HBM core dies connected by TSVs, a base die, die bonds, a dense package route to an accelerator, an organic substrate, board power delivery, and heat flowing to a cooling structure.
HBM bandwidth is a cell, stack, interface, package, power, thermal, test, and controller result. The drawing is not to scale and does not represent a vendor cross-section. Source state. Source-backed functional synthesis. No proprietary geometry.

Test, repair, and yield decide whether the stack is usable

A stacked package multiplies dependencies. Each memory die must work well enough. Vertical connections must work. Bonds must work. The base die must work. The accelerator must work. The package must route and power the assembly. The complete unit must pass test and reliability screens.

A naive intuition is:

TEXT
naive_stack_yield =
  product(individual_die_yield) * assembly_yield

[TOUCHDOWN DERIVATION] That is not a vendor yield model. It simply shows why a multi-die assembly makes known-good-die screening, redundancy, repair, process correlation, and assembly control valuable.

Known-good-die testing tries to avoid stacking obvious failures. Built-in self-test can exercise internal paths. Error-correcting codes detect or correct some data errors. Redundant rows, columns, lanes, or other resources can replace defective structures at a design-specific granularity. Package test and burn-in or reliability screens can catch failures that wafer probing did not expose.

Repair is not free capacity. It consumes redundant resources, logic, fuses or state, test time, routing, validation, and diagnostic coverage. A repair scheme that fixes an array defect may not fix a TSV, bond, base-die, or package fault. Field RAS also asks whether errors are detected, isolated, logged, corrected, and recovered without corrupting the workload.

That is why more HBM supply is not only a wafer-capacity problem. It depends on qualified DRAM wafers, TSV and thinning steps, bonding, base dies, interposers or package fabrics, substrates, assembly capacity, test equipment, controller IP, firmware, cooling, and complete-system yield.

What this means by role. An investor should see a multi-stage supply chain, not one commodity number. A CFO pays for usable packaged capacity and spare/reliability policy, not gross wafer bits. A CTO cannot choose HBM independently from accelerator, package, controller, runtime, and workload. A memory or package engineer should reject any comparison that does not name the die, stack, package, and test boundary.

How does software turn peak HBM bandwidth into useful bandwidth?

In plain English. Peak bandwidth is the interface's advertised ceiling. Useful bandwidth is the part that delivers the right bytes to the right operation in time to help an accepted task. Layout, batching, tiling, caching, fusion, address patterns, controller scheduling, and heat can leave part of the advertised path unused.

The memory wall is the combined limit created by capacity, physical bandwidth, useful bandwidth, latency and queueing, movement energy, activation and refresh work, controller behavior, package wiring, power delivery, cooling, yield, repair, supply, and delivered cost. HBM attacks an important part of that wall by putting a wide DRAM interface close to accelerator compute. It does not remove every term. A larger HBM stack can relieve capacity while package yield becomes harder. A wider interface can raise peak bandwidth while a poor kernel still rereads the same bytes. A faster GPU can finish arithmetic sooner while the CPU tool loop, collective, network, or storage restore becomes the next delay.

That is the thought process used throughout this article:

TEXT
name the accepted task
-> name the state object and its deadline
-> identify which term of the memory wall blocks it
-> choose the smallest software, memory, fabric, package, or system change
-> measure where the bottleneck moved
-> accept or reject the change using the same task

Keep this traffic ladder intact:

TEXT
logical tensor bytes requested by the program
-> bytes allocated and resident in a memory tier
-> bytes served by cache or reaching the HBM controller
-> bytes carried across a fabric when the object moves
-> bytes that contribute to an output the verifier accepts

Each arrow can increase, decrease, or bypass bytes. A tensor can be padded in allocation, reread from HBM, served from cache, fused away, compressed in transit, retried, or computed for an output that is later rejected. That is why one byte count cannot stand in for the whole path.

The software engineer usually sees an allocation and an address. The hardware sees requests, queues, cache lines, page translations, channels, banks, row state, commands, errors, and returned data.

That translation is where headline bandwidth becomes useful bandwidth or disappears into avoidable traffic and stalls.

The NVIDIA CUDA Programming Guide and CUDA Best Practices Guide are the primary public references for device/global memory, transactions, alignment, coalescing, and the programmed GPU memory hierarchy used here. PyTorch CUDA semantics supplies the allocation-versus-caching-allocator-reservation boundary. [OFFICIAL DOCUMENTATION]

From a tensor allocation to an HBM command

Consider a tensor created by PyTorch on a CUDA device. The exact implementation changes by framework and runtime, but the functional path is:

  1. The framework requests memory for a shape and dtype.
  2. The allocator may reuse a pooled region, round the request, split a block, or reserve a larger segment.
  3. The runtime maps device virtual addresses to physical memory resources.
  4. A kernel launches threads that issue loads and stores through the GPU memory hierarchy.
  5. Address translation and caches determine whether a request needs HBM.
  6. Threads' accesses are coalesced or split into memory transactions according to layout, alignment, and hardware rules.
  7. The memory controller maps transactions to HBM channels, pseudo-channels, banks, rows, and columns.
  8. The controller schedules demand reads, writes, refresh, turnarounds, and fairness under timing and power constraints.
  9. Returned data enters a cache, register, shared-memory tile, tensor operation, DMA path, or communication buffer.

CUDA global memory is a programming and address-space concept. On a GPU with HBM, much of the device-memory backing is physically HBM. Those terms are related, not interchangeable. A global-memory load can hit in cache and never reach HBM. A DMA or peer transfer can create HBM traffic without looking like an ordinary scalar load in source code.

The allocator adds another distinction. A tensor's logical bytes equal the product of its shape and element size. Resident bytes can be larger because of alignment, block rounding, fragmentation, caching, replication, workspace reservation, or padding. Reserved memory is not necessarily active traffic. Allocated memory is not necessarily useful data.

Why peak bandwidth stays on the box

Peak bandwidth assumes a favorable interface pattern. Real kernels lose utilization for several different reasons:

  • Requests are too small, scattered, misaligned, strided, serialized, or dependent.
  • Too few warps or independent requests exist to hide service latency.
  • Address mapping concentrates traffic on too few channels or banks.
  • Row conflicts require additional activate and precharge work.
  • Read/write direction changes, refresh, timing constraints, or controller policy consume cycles.
  • Cache or TLB misses create extra work.
  • Temporary tensors are written to HBM and read again because operations were not fused.
  • A poor tile rereads weights or activations that could have remained on chip.
  • A collective, CPU stage, kernel launch, synchronization, or another GPU is actually the bottleneck.
  • Thermal or power management reduces sustained clocks or interface behavior.

A memory copy can report high sustained bandwidth and still say little about a sparse gather. A dense GEMM can have high arithmetic intensity and still wait on communication. An HBM counter can look busy while the user receives a rejected output.

The roofline model remains useful when the boundary is named:

TEXT
arithmetic_intensity = useful_operations / bytes_moved_at_named_boundary

bandwidth_ceiling =
  sustained_bandwidth_at_same_boundary * arithmetic_intensity

The denominator can be HBM bytes, L2 bytes, on-chip scratchpad bytes, PCIe bytes, NVLink bytes, or network bytes. Changing the boundary changes the value. A credible roofline plot states the counter, operation definition, dtype, shape, system, and time interval.

Locality, coalescing, tiling, and fusion

Locality means that data used now is likely to be used again nearby in time or address. Caches exploit it. Tiling changes an operation so a small working set remains in registers or shared memory while many arithmetic operations reuse it. Coalescing combines neighboring thread accesses into fewer efficient transactions. Fusion keeps an intermediate inside a kernel rather than writing it to HBM and launching another kernel to read it.

These are not just software tricks. They change physical demand.

If three separate kernels write and read an approximately 320 MiB intermediate, fusion can avoid some external traffic. If a sparse-attention index causes uncoalesced gathers, fewer logical elements may still create inefficient transactions. If a KV allocator rounds small sequences to large blocks, unused capacity can lower concurrency even before any bytes move.

The cheapest byte is often the byte that does not cross a boundary. That can happen through reuse, aliasing, zero-copy handoff, fusion, compression, quantization, or recomputation. Every mechanism has a cost: extra compute, precision risk, implementation complexity, larger on-chip state, or a stricter schedule.

Movement APIs and what profilers can see

Pinned host memory can support efficient DMA because pages stay resident and can be addressed predictably by the transfer machinery. Page migration can move backing between host and device under a managed-memory policy. Peer access can move or access data across accelerators through supported fabrics. NIXL and similar software can orchestrate movement across heterogeneous memory and storage paths.

Those mechanisms do not collapse all tiers into one latency. The route can include cache coherence, page faults, DMA setup, PCIe or a proprietary accelerator fabric, switches, NUMA placement, CPU memory controllers, CXL devices, storage controllers, and filesystems.

Profilers also have boundaries:

  • An allocation trace shows logical or reserved memory but not DRAM row commands.
  • HBM throughput counters show activity at a hardware boundary but not whether bytes were useful.
  • A kernel timeline shows duration and overlap but not automatically the customer acceptance result.
  • CPU offload counters prove movement to host memory, not CXL, LMCache, or disk unless the specific path is instrumented.
  • A cache-hit metric needs a defined key, block, version, and correctness rule.

This small calculator makes the distinction explicit:

PYTHON
# [TOUCHDOWN DERIVATION] No hardware result is embedded.
def bandwidth_receipt(
    transfer_rate_per_second,
    interface_width_bits,
    elapsed_seconds,
    useful_bytes,
    measured_interface_bytes=None,
):
    peak_bytes_per_second = (
        transfer_rate_per_second * interface_width_bits / 8
    )
    useful_bytes_per_second = useful_bytes / elapsed_seconds
    out = {
        "peak_bytes_per_second": peak_bytes_per_second,
        "useful_bytes_per_second": useful_bytes_per_second,
        "useful_over_peak": useful_bytes_per_second / peak_bytes_per_second,
    }
    if measured_interface_bytes is not None:
        achieved = measured_interface_bytes / elapsed_seconds
        out["measured_interface_bytes_per_second"] = achieved
        out["measured_over_peak"] = achieved / peak_bytes_per_second
        out["useful_over_measured"] = useful_bytes / measured_interface_bytes
    return out

The function contains no default rate, width, bytes, or time. A real receipt supplies them from a named device, direction, counter, kernel interval, and workload.

For the broader path from profiler evidence through compiler and kernel changes to useful GPU capacity, see Touchdown's Automated CUDA, Revenue per GPU, and Capability per GPU. That article supplies the optimization loop; this article supplies the memory-specific physical boundary.

FIGURE 4 1600 × 900
A waterfall descends from theoretical HBM interface bandwidth to sustained measured interface traffic, kernel-requested useful bytes, and bytes that contribute to an accepted task. Empty measurement boxes show that each stage needs a named receipt.
Peak bandwidth is an interface ceiling. Useful bandwidth survives allocation, caching, layout, controller scheduling, synchronization, and the workload acceptance gate. Source state. Touchdown systems model. No percentages are populated until measured.

What this means by role. The CEO should not expect a higher box number to guarantee more output. The CFO should compare sustained accepted-task throughput and reliability. The software engineer controls layout, tiling, reuse, allocation, batching, and movement policy. The kernel engineer measures traffic at named boundaries. The controller engineer sees address distribution and scheduling pressure that source code alone cannot reveal.

What memory path does Wan2.2 video diffusion create?

In plain English. Video diffusion begins with noise and repeatedly edits a compressed video representation until it matches the prompt closely enough to decode. That compressed representation is called a latent. The model rereads weights and creates large temporary tensors during many denoising steps, so its memory path is repeated, state-heavy, and different from a language model's growing KV cache.

Now the physical path needs a real workload.

Wan2.2 is useful because video diffusion creates large multi-dimensional state, repeated denoising, dense projections, attention, normalization, feed-forward work, classifier-free guidance, communication, and output decoding. It is not just one matrix multiplication, and it is not the same memory path as autoregressive text generation.

The source boundary matters. This walkthrough uses the official Wan-Video/Wan2.2 repository pinned to commit 42bf4cfaa384bc21833865abc2f9e6c0e67233dc. The configuration and code path are source-backed. The tensor arithmetic below is Touchdown-derived. It is not a hardware trace.

From prompt to a video someone accepts

For text-to-video generation, the functional path is:

TEXT
text prompt and negative prompt
-> text encoder creates conditioning tensors
-> runtime creates an initial noisy latent
-> scheduler selects a sequence of timesteps
-> one of the high-noise or low-noise denoisers runs for the current region
-> transformer blocks perform self-attention, cross-attention,
   normalization, feed-forward work, residual updates, and modulation
-> conditional and unconditional predictions are combined by guidance
-> scheduler updates the latent
-> repeat for the requested number of sampling steps
-> VAE decodes the final latent into video frames
-> product checks quality, format, latency, and safety

The A14B configuration does not use token-level top-k MoE routing in the usual language-model sense. It names separate high-noise and low-noise checkpoints and switches the model used across the timestep boundary. That distinction matters because memory planning can load, retain, or offload two full model states differently from a token-routed expert layer.

The pinned configuration includes:

PYTHON
# [OFFICIAL REPOSITORY SOURCE]
# Wan2.2 commit 42bf4cfaa384bc21833865abc2f9e6c0e67233dc
t2v_A14B.vae_stride = (4, 8, 8)
t2v_A14B.patch_size = (1, 2, 2)
t2v_A14B.dim = 5120
t2v_A14B.ffn_dim = 13824
t2v_A14B.num_heads = 40
t2v_A14B.num_layers = 40
t2v_A14B.sample_steps = 40
t2v_A14B.boundary = 0.875

This block proves configuration values at the pinned commit. It does not prove allocated memory, physical HBM traffic, or runtime.

One concrete tensor shape

Take an illustrative request of 81 frames at 832 by 480 pixels. These dimensions are chosen because they are compatible with the pinned temporal and spatial strides. The generation code constructs a latent target shape using the VAE stride, then calculates the sequence length from the spatial patch.

The derived latent grid is:

TEXT
temporal latent = (81 - 1) / 4 + 1 = 21
latent height   = 480 / 8 = 60
latent width    = 832 / 8 = 104

token grid = 21 * (60 / 2) * (104 / 2)
           = 21 * 30 * 52
           = 32,760 tokens

[TOUCHDOWN DERIVATION] That is a shape calculation from the pinned source, not an observed runtime tensor dump. A real distributed run may pad sequence length according to its sequence-parallel degree. The article therefore reports the unpadded logical grid and requires the runtime to report padding separately.

At model width 5,120, one BF16 activation of shape [32,760, 5,120] contains:

TEXT
32,760 * 5,120 * 2 bytes
= 335,462,400 bytes
= 319.921875 MiB
approximately 320 MiB

LOGICAL TENSOR BYTES, NOT MEASURED HBM TRAFFIC. If Q, K, and V outputs at that logical shape were all fully materialized, their logical writes would total approximately 959.77 MiB for one block evaluation. Multiplying that value by 40 blocks, 40 default steps, and two classifier-free-guidance passes gives approximately 2,999.27 GiB, or 2.929 TiB, of shape-derived Q/K/V output bytes across a request.

That number is deliberately labeled an unfused logical boundary. It is not measured HBM traffic. It excludes reads, weights, feed-forward tensors, attention results, communication, VAE work, padding, allocator overhead, caching, and output. It can also overstate physical writes when fusion or kernel design keeps intermediates on chip. Its value is to show why materialization choices matter.

For a second production-shaped reference, use 81 frames at 1280 by 720. From the same VAE stride and patch size, the logical latent is 16 x 21 x 90 x 160, or 4,838,400 latent elements, and the transformer grid contains 75,600 visual tokens. One BF16 hidden-state tensor at width 5,120 is 774,144,000 logical bytes, about 738.28 MiB. This is again shape arithmetic, not peak allocation or measured HBM traffic.

The default 40-step loop performs one conditional and one unconditional DiT evaluation per step. That is 80 transformer forwards. With 40 blocks, the request crosses 3,200 block evaluations before the VAE decodes the final latent. Prompt expansion and the UMT5 text encoder are separate workload phases. The stock repository uses BF16 mixed precision for the model path while retaining selected numerically sensitive work in FP32. An FP8 or NVFP4 result therefore needs its own named artifact, kernel path, quality gate, and receipt.

This is also where LLM vocabulary can mislead. Wan2.2's per-block Q, K, and V are transient diffusion-attention intermediates. They are not an autoregressive LLM KV cache that grows once per generated token and can be restored for a later turn. Offloading DiT weights, moving a mutable latent, sharding attention state, and restoring a coding-agent prefix are four different movement problems.

What the attention code actually does

The official WanSelfAttention.forward path projects Q, K, and V, applies normalization to Q and K, reshapes by head, applies rotary position handling, calls a FlashAttention-style function, flattens the result, and runs an output projection:

PYTHON
# [OFFICIAL REPOSITORY SOURCE, SHORTENED]
def qkv_fn(x):
    q = self.norm_q(self.q(x)).view(b, s, n, d)
    k = self.norm_k(self.k(x)).view(b, s, n, d)
    v = self.v(x).view(b, s, n, d)
    return q, k, v

q, k, v = qkv_fn(x)
x = flash_attention(
    q=rope_apply(q, grid_sizes, freqs),
    k=rope_apply(k, grid_sizes, freqs),
    v=v,
    k_lens=seq_lens,
    window_size=self.window_size)
x = self.o(x.flatten(2))

Source: wan/modules/model.py. Boilerplate is omitted. The excerpt is short enough to expose the operation boundary without copying the entire implementation.

For each block, ask three questions:

  1. What state is read or created? Input activations, projection weights, Q/K/V tensors, positional frequencies, sequence metadata, attention output, and output-projection weights.
  2. Which boundary is crossed? Depending on tiling and cache behavior: HBM to L2, L2 to on-chip storage and registers, on-chip accumulation, possible HBM materialization, and inter-GPU communication in a distributed layout.
  3. What receipt is missing? Kernel names, shapes, dtypes, HBM read/write counters, L2 traffic, duration, occupancy, synchronization, numerical parity, and clip-level acceptance.

The FlashAttention-style call is important. A naive attention explanation says attention creates an S by S matrix. A fused tiled implementation can avoid materializing that full score matrix in HBM. The algorithm still reads Q, K, and V and performs substantial work, but physical traffic differs from the naive graph.

The denoising loop shows another repeated path:

PYTHON
# [OFFICIAL REPOSITORY SOURCE, SHORTENED]
for _, t in enumerate(tqdm(timesteps)):
    latent_model_input = latents
    timestep = torch.stack([t])
    model = self._prepare_model_for_timestep(t, boundary, offload_model)
    sample_guide_scale = (
        guide_scale[1] if t.item() >= boundary else guide_scale[0]
    )
    noise_pred_cond = model(latent_model_input, t=timestep, **arg_c)[0]
    noise_pred_uncond = model(latent_model_input, t=timestep, **arg_null)[0]
    noise_pred = noise_pred_uncond + sample_guide_scale * (
        noise_pred_cond - noise_pred_uncond
    )
    temp_x0 = sample_scheduler.step(
        noise_pred.unsqueeze(0), t, latents[0].unsqueeze(0),
        return_dict=False, generator=seed_g)[0]
    latents = [temp_x0.squeeze(0)]

Source: wan/text2video.py. The shortened excerpt preserves the decision and data dependencies while omitting scheduler setup and cleanup.

This loop proves that conditional and unconditional model evaluations feed a guidance combination and latent update at each timestep. It also proves that the implementation has a model-offload control path. It does not tell us whether offload was enabled in a given benchmark, how much state moved, or whether the transfer overlapped compute.

FSDP and Ulysses turn the video sequence into a fabric problem

The single-device path is only one execution shape. Wan2.2's official repository also documents an eight-GPU path that combines FSDP for the diffusion transformer and text encoder with DeepSpeed Ulysses sequence parallelism:

BASH
# [OFFICIAL REPOSITORY COMMAND SHAPE]
# Pin the repository, weights, image, prompt, runtime, and seed before replay.
torchrun --nproc_per_node=8 generate.py \
  --task ti2v-5B \
  --size 1280*704 \
  --ckpt_dir ./Wan2.2-TI2V-5B \
  --dit_fsdp \
  --t5_fsdp \
  --ulysses_size 8 \
  --image examples/i2v_input.JPG \
  --prompt "Summer beach vacation style, a white cat wearing sunglasses sits on a surfboard."

Source: Wan2.2 generation instructions. The command proves that the project exposes the distributed path. It does not prove that eight GPUs beat one GPU at the same resolution, frame count, denoising steps, dtype, quality gate, latency target, or cost.

DeepSpeed-Ulysses partitions the sequence dimension across workers. At the attention boundary, an all-to-all redistributes Q, K, and V so each worker can compute attention for a subset of heads over the full sequence. Another all-to-all returns the output to the sequence-partitioned layout. [OFFICIAL PROJECT DESIGN / PAPER] The mechanism is also described in the DeepSpeed-Ulysses paper.

TEXT
local projection
-> Q/K/V all-to-all
-> local attention over assigned heads
-> output all-to-all
-> local output projection and feed-forward work
-> next transformer block
-> next denoising step

That repeated collective can matter because the video token sequence spans time, height, and width, and the cost recurs across transformer blocks and denoising steps. But a larger scale-up domain creates value only when the trace shows exposed collective wait on the critical path. The required receipt includes collective type, message bytes, rank map, topology, algorithm, overlap, p50/p95/p99 wait, DiT time, VAE time, peak allocated HBM, GPU-hours, retries, and accepted-clip rate.

The current A14B configuration has 40 attention heads. The official Ulysses implementation requires the sequence-parallel group size to divide the head count and, in the documented Wan path, match the participating world size. An unmodified 144-way Ulysses group on a Kyber NVL144 rack is therefore not a valid configuration. A future rack could host multiple bounded groups or combine other parallel dimensions, but that is an experiment to design, not a performance claim.

The product metric also continues after denoising. VAE decode, safety and format checks, quality rejection, and retries consume time and energy. The financial denominator is an accepted clip. A faster 40-step kernel path can still lose if quality forces more retries or if VAE and output handling dominate the job.

[NOT FOUND] No public Wan2.2-on-Kyber benchmark was verified as of July 10, 2026. NVIDIA platform bandwidth and topology specifications cannot be converted into a Wan2.2 latency, quality, energy, or cost result without a controlled replay.

The state-lifetime map

State Access shape Mutability Reuse horizon First placement question
High/low-noise weights Repeated layer-ordered reads Read-only in inference Across relevant denoising steps Keep resident, switch/offload by region, shard, or stream predictable groups?
Text conditioning Reused cross-attention input Read-only Request lifetime Replicate, cache, or recompute?
Latent Dense iterative update Read-write Every step Keep in hot writable memory by default
Q/K/V and FFN intermediates Dense or tile-local Ephemeral Inside block/kernel Fuse, tile, spill, or rematerialize?
Positional and sparse metadata Read-mostly indexes/frequencies Mostly read-only Blocks or steps Can gathers use memory efficiently?
Communication buffers Collective-dependent Read-write Per block or step Does sharding reduce compute but add movement stalls?
Decoded frames Sequential output Write-mostly then read End of request Move to host/storage without blocking the next request?

The table explains why one memory technology does not own the workload. Mutable latents and active intermediates need low-latency writable service. Read-only weights may offer prefetch or streaming opportunities if access order is predictable. Decoded frames can leave HBM. Sparse metadata can be small but latency-sensitive. Communication buffers couple memory placement to the accelerator fabric.

What the public optimization records do and do not prove

[VENDOR TARGET / BENCHMARK] Baseten's January 2026 article reports median Wan2.2 results for its named stack, with 40 sampling steps, 1280 by 720 output, and 81 frames. It reports 2.6 times speedup on H100 and 3.2 times on B200 relative to its stated reference. The article discusses GEMM, RoPE, LayerNorm, RMSNorm, communication overlap, and disabling CPU model offload. Those are workload-specific vendor results, not proof that one memory technology wins.

[AUTHOR WORKLOG] Ali's later public engineering record reports a broader composite improvement that includes reducing denoising steps, trained sparse attention, kernel work, fusion, scheduling, and NVFP4. The reported 54 times is an end-to-end composite claim, not Baseten's official result and not one kernel speedup. The worklog separately reports a VSA kernel change from 7.444 ms to 4.719 ms, about 1.58 times by arithmetic. Touchdown has not independently reproduced those numbers in this article.

[ACCEPTED PAPER, AUTHOR-REPORTED MEASUREMENTS] The Visual Sparse Attention paper, accepted by NeurIPS 2025 according to its current arXiv record, and the public FastVideo code provide a separate research boundary. Sparsity can reduce attention work, but indexes, gather patterns, load balance, communication, and quality tests become part of the system. Fewer selected tokens do not automatically mean proportional HBM or end-to-end savings.

FIGURE 5 1600 × 900
One Wan2.2 denoising step reads high-noise or low-noise model weights and text conditioning, updates a mutable latent, creates Q/K/V and feed-forward intermediates, consumes positional or sparse metadata, communicates when distributed, and feeds conditional and unconditional predictions into a scheduler update.
Wan2.2 creates read-only, mutable, ephemeral, sparse, and communication state with different movement requirements. Shape arithmetic comes from the pinned official source. Physical traffic is still unmeasured. Source state. Official repository plus Touchdown-derived tensor sizes; vendor and author benchmarks remain separate.

What this means by role. The product owner cares about an accepted clip. The CFO cares about GPU-seconds, retries, and concurrency. The software engineer needs allocated bytes and offload events. The kernel engineer needs HBM/L2/link traffic and quality parity. The hardware engineer needs the actual access distribution, not the phrase video is memory-bound.

What memory path does a coding agent create?

In plain English. A coding agent alternates between GPU model work and CPU-side tools. The model reads instructions and repository context, generates a proposed action, calls search, edit, test, or compilation tools, receives new evidence, and continues. Reusable prefixes and KV cache can avoid repeated model work, but tool delays, cache misses, retries, and failed patches can still dominate the cost of the accepted change.

A coding agent uses many of the same physical memories, but it creates a different state graph.

Its model still reads weights. Its transformer still performs prefill and decode. Its KV cache still occupies device memory. But the request repeatedly leaves the accelerator to search files, read code, edit text, run a compiler, execute tests, inspect failures, and return new context. The workflow branches. A tool result can invalidate a plan. A failed patch can trigger another prefill and decode cycle.

The accepted output is not a plausible paragraph. It is a correct, reviewable change with passing evidence.

From user request to accepted patch

TEXT
user request
-> system instructions and tool schemas
-> repository map, selected files, prior turns, and current diff
-> model prefill
-> active KV cache
-> token-by-token decode
-> tool call
-> CPU-side search, file I/O, edit, test, or lint
-> tool result appended to context
-> reused prefix or another prefill
-> more decode
-> candidate patch
-> test and review gate
-> accepted change or retry

There are three distinct kinds of state.

Model state includes weights and runtime workspaces shared across requests. Inference state includes active KV, reusable prefix KV, scheduler metadata, logits/sampling state, and communication buffers. Workflow state includes files, repository indexes, tool results, patch attempts, logs, tests, user preferences, and durable receipts.

On an HBM accelerator, model weights, active KV, and hot runtime buffers commonly occupy HBM. That does not mean HBM should hold every repository file or old test log. CPU DRAM, optional CXL-attached capacity, NVMe, and object storage can hold larger or more durable state. They help only when software knows what to retrieve, when it will be reused, and how long movement takes.

KV cache, one term at a time

The KV cache is the language model's working notebook for tokens it has already read. Without it, the model would have to recompute the same attention history before producing each next token. The notebook grows as generation continues, and every concurrent request needs its own correctly identified state. The formula below explains the logical size of that notebook before allocator padding, paging, replication, fragmentation, and runtime overhead.

During autoregressive inference, each transformer layer produces key and value vectors for prior tokens. Retaining them lets a new token attend to prior context without recomputing the full prefix at every decode step.

For a conventional attention layout, a useful logical-size formula is:

TEXT
logical_KV_bytes =
  2
  * layers
  * kv_heads
  * head_dim
  * bytes_per_element
  * sum(tokens_i for each live sequence i)

equal-length special case =
  2
  * layers
  * tokens_per_sequence
  * kv_heads
  * head_dim
  * bytes_per_element
  * concurrent_sequences

One bounded public coding-model example

Use the official Qwen/Qwen2.5-Coder-32B-Instruct configuration pinned to revision 381fc969f78efac66bc87ff7ddeadb7e73c218a7. [OFFICIAL MODEL CONFIGURATION] It declares 64 hidden layers, 40 attention heads, 8 key/value heads, hidden width 5,120, BF16 dtype, and 32,768 maximum position embeddings. The head dimension is derived as 5,120 / 40 = 128.

For one illustrative 32,768-token coding-agent context, split as 24,576 stable system/repository-prefix tokens plus 8,192 current-turn and tool-result tokens, the logical BF16 KV size is:

TEXT
2 for K and V
* 64 layers
* 32,768 tokens
* 8 KV heads
* 128 values per head
* 2 bytes per BF16 value
= 8,589,934,592 bytes
= 8 GiB per sequence

8 equal-length concurrent sequences
= 64 GiB of logical KV

[TOUCHDOWN DERIVATION] The token split and concurrency are an explicit experiment contract, not claims about typical Qwen usage. Eight GiB and 64 GiB are logical tensor capacities. A real serving receipt still needs a pinned engine and connector version, cache dtype, block size, allocation and fragmentation, sharding or replication, prefix sharing, device and lower-tier residency, transfer bytes, restore p50/p95/p99, prefill avoided, accepted patch, and failed-test or retry record.

The movement hypothesis is now concrete: the 24,576-token stable prefix might be worth preserving below HBM between turns if version-correct lookup and restore beat recomputing that prefix under the p99 budget. The active 8,192-token suffix remains the hotter mutable or append-only region. That is a hypothesis the receipt can reject.

The leading 2 is key plus value. kv_heads is not always the query-head count. Grouped-query attention and multi-query attention share K/V across more query heads, reducing KV state. head_dim and dtype come from the named model configuration.

Logical bytes are not resident allocator bytes. Paged KV systems divide state into blocks. The final block can be partly empty. Allocation can round, fragment, replicate, or attach metadata. Tensor or pipeline parallelism changes where blocks live. Prefix sharing can let multiple requests reference one physical block set. Quantization can reduce element size but adds compatibility and quality questions.

A complete KV receipt therefore reports:

TEXT
logical bytes
allocated blocks and bytes
used tokens per block
fragmentation and metadata
replication / sharding
shared-prefix references
evictions
device-to-host and host-to-device bytes
hit rate with key/version semantics
restore p50/p95/p99
prefill avoided
accepted task result

Touchdown's KV Cache Is Becoming the Memory Hierarchy of Inference follows one repeated agent workflow through KV placement in more operational detail. The current section adds the cell, package, alternative-tier, and hardware-admission path.

This schema shows what an agent trace should capture:

PYTHON
# [PROPOSED] Touchdown illustration, not framework source.
StateObject(
    name="repo_prefix_kv",
    bytes=measured_allocated_bytes,
    access="read_mostly",
    reuse_distance_ms=observed_reuse_distribution,
    deadline_ms=restore_p99_budget,
    mutable=False,
    recompute_ms=measured_prefill_time,
    correctness="exact",
    current_tier="gpu_hbm",
)

The crucial comparison is not HBM versus CPU memory in the abstract. It is measured restore plus miss cost versus measured prefill recomputation under the actual concurrency and p99 budget.

A real documented offload path

[OFFICIAL VENDOR DOCUMENTATION] NVIDIA Dynamo's KVBM guide, current v1.2.1 when reviewed, documents a write-through hierarchy and configuration for GPU, pinned CPU memory, and disk. The guide includes this cache-tier example:

BASH
# Official documentation example, not Touchdown defaults or a benchmark.
export DYN_KVBM_CPU_CACHE_GB=4
export DYN_KVBM_DISK_CACHE_GB=8

It also documents the vLLM connector shape:

BASH
vllm serve \
  --kv-transfer-config \
  '{"kv_connector":"DynamoConnector","kv_role":"kv_both","kv_connector_module_path":"kvbm.vllm_integration.connector"}' \
  MODEL_ID

The placeholder MODEL_ID replaces the documentation's example model because this article is not prescribing a benchmark. The option names and connector path come from the current guide.

The vendor guide carries an important warning: capacities should increase down the write-through tiers. If the CPU cache is smaller than the device KV allocation, the system can churn by offloading after forward passes without creating useful retained capacity. The guide also says that insufficient prefix hits can eliminate time-to-first-token benefit or create degradation.

That is exactly the point of the state-first method. A lower tier is not valuable because it exists. It is valuable when retained state is reused enough to avoid more work than movement costs.

This source proves a documented software path. It does not prove that the path helps this coding-agent workload. It does not prove CXL, because CPU memory can be ordinary host DRAM. It does not prove LMCache, because a different connector can own that path. It does not prove disk is healthy for active KV, because SSD latency, endurance, filesystem, reuse, and tail behavior still need measurement.

One open coding-agent loop, followed into memory

Hermes Agent is a useful public code example because its run_agent.py exposes the tool-calling loop instead of hiding it behind a hosted product. [OFFICIAL PROJECT SOURCE, REVIEWED 2026-07-10] The repository describes automatic tool calling, conversation history, error recovery, tool-result handling, and repeated model calls. That lets us trace the mechanism without claiming that every coding agent uses the same implementation.

The physical and software path is:

  1. The client assembles system instructions, tool schemas, prior messages, repository context, and the current request in host memory.
  2. The serving engine tokenizes that context and schedules a prefill. Model weights and the hot working set are read through accelerator HBM.
  3. Each transformer layer writes K and V tensors for the prompt into paged device allocations. Those allocations are logical model state represented by physical HBM pages, allocator metadata, and engine block tables.
  4. Decode reads weights and prior KV, produces a token, and repeats. When the model emits a tool call, generation pauses at the application boundary.
  5. Hermes executes the search, file read, patch, compiler, test, or lint operation on the CPU and storage path. The GPU may serve another request or wait, depending on scheduling and concurrency.
  6. The tool result is appended to conversation state. If the unchanged prefix is recognized under the same model, tokenizer, prompt-template, repository revision, and cache-key contract, reusable KV can avoid part of another prefill. If any identity field changes, the safe outcome is a miss.
  7. The loop continues until the agent returns a candidate patch. Tests, review, and acceptance determine whether the occupied compute, memory, and tool time produced useful work.

The word cache hides several different mechanisms. A provider prompt cache can reduce billed or processed input. An engine prefix cache can share device KV blocks. LMCache or HiCache can retain KV below the device tier. A repository index can cache file retrieval. A tool-result cache can avoid another external action. None proves another, and each needs its own key, version, hit, latency, and correctness receipt.

One enterprise task across Claude Code, Codex, and self-hosted GLM-5.2

Here is the CEO and CTO example this article needs.

An engineer asks an agent to update an authentication dependency across a private repository, follow the company's internal security standard, migrate every affected call site, run the named tests, and return a reviewable patch with citations to the policy it followed.

That one sentence creates this path:

TEXT
private repository and security documents
-> parse, version, permission, chunk, embed, and index
-> authenticate the engineer and retrieve permitted code and policy spans
-> rerank, deduplicate, and assemble the prompt
-> prefill the model and create active inference state
-> decode a search, read, edit, or test tool call
-> run CPU, filesystem, network, compiler, and sandbox work
-> append the tool result
-> reuse a valid prefix, restore KV, recompute prefill, or compact
-> repeat until a candidate patch exists
-> run tests, policy checks, citation checks, and human or CI review
-> accept the patch or pay for another attempt

The repository and security corpus normally live in object storage, databases, text indexes, vector indexes, CPU DRAM, and SSD. They do not all move into HBM. Retrieval selects a smaller set of spans. Those spans become prompt tokens. Prefill transforms the prompt into model state. Active KV or another architecture-specific cache then competes with model weights, workspaces, communication buffers, and concurrent requests for accelerator memory.

Microsoft's production RAG guidance describes query translation, parallel query execution, hybrid retrieval, reranking, and prompt assembly as separate operations. Its hybrid-search documentation combines text and vector search. [OFFICIAL PLATFORM DOCUMENTATION] These sources establish the retrieval path, not a Touchdown latency result.

Adding more chunks can improve recall and still make the product worse. More retrieved tokens increase prefill work, active cache state, time to first token, and possibly provider charges. Bad retrieval can therefore cost twice: first as wasted inference, then as a rejected patch.

The same task can enter three different agent paths:

Path What the public source exposes What the company can measure What remains unknown
Claude Code Repository tools, CLAUDE.md, skills, MCP, hooks, subagents, context loading, and compaction behavior Input and output usage where exposed, tool calls, tool duration, wall time, retries, tests, accepted patch, provider bill Exact serving GPU, weight precision, KV precision, HBM bytes, offload tier, NVLink topology, and per-task energy unless Anthropic discloses them
OpenAI Codex Repository instructions, local tools, file changes, tests, thread events, and context-compaction events Tool and terminal evidence, wall time, retries, accepted patch, usage or cost fields exposed by the selected product Exact serving GPU, precision, KV geometry, offload tier, fabric, and HBM traffic unless OpenAI discloses them
Hermes or OpenClaw with self-hosted GLM-5.2 Open harness loop plus a separately operated model endpoint All application fields plus engine, model revision, hardware, topology, weight precision, KV precision, cache policy, profiler counters, and power telemetry Nothing becomes measured automatically. The team still has to instrument and replay it.

Claude Code's extension documentation says CLAUDE.md loads in full at session start, tool schemas can be deferred, skills load on demand, and subagents receive isolated context. Codex's app-server protocol exposes thread, turn, item, and contextCompaction events. OpenClaw's memory documentation separates injected long-term memory from indexed daily records and flushes durable state before compaction. [OFFICIAL PRODUCT OR PROJECT DOCUMENTATION] These are context-management facts. They do not reveal provider HBM behavior.

That separation is the first executive lesson: hosted coding tools expose the business task and client-visible state, while the provider owns the lower inference stack. Self-hosting exposes more control and more operational responsibility.

GLM-5.2 makes the precision and memory decision concrete

GLM-5.2 is useful as the self-hosted control case because official artifacts expose the model architecture and current serving commands.

Z.ai's official repository lists GLM-5.2 as a 744-billion-parameter, 40-billion-active-parameter MoE and publishes BF16 and FP8 artifacts. The current FP8 configuration declares 78 hidden layers, 256 routed experts, eight selected experts per token, a 1,048,576-token configured maximum, FP8 E4M3 block quantization for the checkpoint, and architecture-specific low-rank KV fields. OFFICIAL Z.AI REPOSITORY AND MODEL CONFIGURATION

NVIDIA separately publishes nvidia/GLM-5.2-NVFP4. Its model card calls the artifact 753 billion total parameters and 40 billion activated, targets Blackwell, and documents SGLang and vLLM paths. It says linear operators inside MoE experts use NVFP4 while the shared expert remains unquantized. [OFFICIAL NVIDIA MODEL CARD] The 744B and 753B totals belong to different official artifact descriptions. This article preserves that discrepancy instead of silently choosing one number.

The first-order weight-capacity floor for the 744B Z.ai artifact is:

TEXT
ideal packed bytes = parameter count * bits per stored parameter / 8

BF16: 744B * 16 / 8 = 1.488 TB decimal
FP8:  744B *  8 / 8 = 0.744 TB decimal
4-bit ideal floor:  744B * 4 / 8 = 0.372 TB decimal

[TOUCHDOWN DERIVATION] These are arithmetic floors, not checkpoint sizes or runtime allocations. Quantization scales, unquantized modules, padding, duplicate tensors, expert placement, CUDA graphs, KV state, communication buffers, and allocator headroom all consume additional memory. Forty billion active parameters per token does not mean only forty billion parameters need to be resident. The runtime still has to place or fetch the experts that future tokens may select.

The precision decision is also not one switch:

Field BF16 baseline FP8 artifact NVIDIA NVFP4 artifact
Weight storage Larger Smaller, with FP8-specific scales and kernels Smaller for the quantized MoE linear operators, not every module
Activation/accumulation Runtime-specific mixed precision Not implied solely by the weight filename NVIDIA card states weight and activation quantization for selected operators; accumulation remains operator-specific
KV-cache dtype Separate engine decision Not automatically FP8 NVIDIA's vLLM command explicitly requests fp8_e4m3 KV; that command is artifact-specific
Hardware path Broad, but expensive at this size Runtime and accelerator support required NVIDIA card lists Blackwell compatibility
Quality evidence Base artifact and workload eval required Compare on the same agent harness NVIDIA publishes selected benchmark comparisons against its FP8 baseline; customer patch acceptance still requires replay

The official NVIDIA vLLM command shape is:

BASH
# [OFFICIAL NVIDIA MODEL-CARD COMMAND, VERSION-SENSITIVE]
vllm serve nvidia/GLM-5.2-NVFP4 \
  --tensor-parallel-size 8 \
  --enable-expert-parallel \
  --trust-remote-code \
  --reasoning-parser glm45 \
  --tool-call-parser glm47 \
  --enable-auto-tool-choice \
  --kv-cache-dtype fp8_e4m3 \
  --host 0.0.0.0 --port 8000

This proves that NVIDIA documents one Blackwell, vLLM, NVFP4, FP8-KV, TP8, expert-parallel serving shape. It does not prove a specific GPU count beyond the command's tensor-parallel requirement, a context length, concurrency, tokens per second, cache hit rate, accepted-patch rate, or cost.

[NOT FOUND] No public GLM-5.2-on-H200, GLM-5.2-on-B200, GLM-5.2-on-GB200, or GLM-5.2-on-Kyber serving row with context length, concurrency, tokens per second per user, tokens per second per GPU, accepted-patch rate, and cost was verified for this article as of July 11, 2026. Public GLM-5 benchmark rows are a different model version and must not be relabeled GLM-5.2.

That missing row is why the paid implementation starts from a run receipt, not a marketing comparison. Every run needs:

TEXT
task and repository revision
harness and version
model and artifact revision
hosted or self-hosted boundary
engine, image, hardware, and topology when self-hosted
weight, activation, accumulation, and KV precision
input, retrieved, tool-schema, tool-result, cached-prefix, and output tokens
context window requested and actually admitted
queue, prefill, decode, KV restore, tool, test, and wall time
cache hit, bytes moved, eviction, miss, and recompute
tool calls, failed calls, retries, compactions, and test result
accepted patch, review result, cost, and measured energy when available

Tokens per second answers only one line. A coding agent that decodes twice as fast but spends most of its wall time compiling, retries more often, or produces fewer accepted patches can have worse economics.

The same enterprise task makes prefix reuse visible

Use an explicit context sweep instead of pretending every enterprise agent fills one million tokens:

Run Stable policy and repository prefix Retrieved task-specific spans Prior tool results New tool result Question answered
A 30,000 tokens 5,000 2,000 0 Baseline prefill and first tool call
B same 30,000 same 5,000 2,000 4,000 Does exact-prefix reuse avoid rebuilding A?
C same 30,000 same 5,000 6,000 4,000 Does another appended result preserve the earlier prefix?
D policy or repository revision changed retrieval rerun prior history 4,000 Does invalidation correctly force a miss?

[ILLUSTRATIVE EXPERIMENT CONTRACT] These token counts are not production averages. They define a reproducible sweep. The important test is semantic validity: model revision, tokenizer, chat template, system policy, tool schema, tenant, repository commit, document ACL, and prefix hash must match before reuse.

For a conventional transformer, logical KV can be estimated from layers, KV heads, head dimension, tokens, dtype, and concurrency. Do not blindly apply that formula to GLM-5.2. Its sparse-attention and low-rank KV configuration changes the cached representation. The receipt must use the exact SGLang or vLLM cache geometry and measured allocation.

The break-even condition remains simple:

TEXT
prefix or offload path wins when

lookup + transfer + restore + miss risk
<
recomputed prefill under the required p99

This is also where privacy becomes a memory-system property. KV is derived from private code, retrieved policy, tool output, and prior reasoning. Tenant isolation, ACL-aware cache keys, deletion, invalidation, encryption, TTL, and audit logs belong in the cache design. A fast cross-tenant hit is a security failure.

Where Kyber could matter, and where the example stops

GLM-5.2 is a large MoE. Expert parallelism can create token dispatch, combine, and all-to-all traffic. A larger fast scale-up domain could change expert placement and reduce how often that traffic crosses a slower scale-out boundary.

That is the legitimate Kyber thought experiment:

TEXT
larger model and more experts
-> more pressure on weight placement and expert communication
-> larger fast-connected domain may keep more traffic inside scale-up fabric
-> runtime can test wider expert parallelism and different replica shapes
-> accepted-task replay decides whether the added domain helps

Kyber is NVIDIA's preliminary 144-GPU MGX NVL rack design for Rubin Ultra, not a current GLM-5.2 benchmark result. Exact Rubin Ultra Kyber HBM capacity, bandwidth, application topology, tokens per second, power, schedule, and accepted-agent economics remain not_public in the reviewed primary material.

There is a similar constraint on Wan2.2. Its current A14B reference uses 40 attention heads, and the official Ulysses path requires the group size to divide the head count. An unmodified 144-way Ulysses group is therefore not a valid conclusion. A future Kyber rack might run multiple bounded video groups or a different mixture of parallelism, but the stock repository does not prove it.

What this means by role. The CEO should compare accepted patches and accepted clips, not raw model calls. The CFO should count retrieval, provider or GPU cost, tool sandboxes, retries, review, energy, and failed outputs. The CTO should separate the hosted control boundary from the self-hosted engine, cache, precision, topology, and security decisions. The software engineer should stabilize prompt and repository identity. The kernel engineer should profile prefill, sparse attention, MoE dispatch, dequantization, KV copy, collectives, and the tool-idle gaps before claiming HBM is the bottleneck.

The serving libraries expose different memory-control surfaces

Layer What it controls Documented control surface What it does not prove
vLLM Device KV allocation, PagedAttention, prefix reuse, scheduling Engine flags and a versioned KV connector configuration That a hit occurred, that lower-tier storage was used, or that the result passed
LMCache MP Reusable KV outside the vLLM worker process lmcache server plus LMCacheMPConnector and optional L2 adapter That MP, L2, GDS, CXL, or remote storage is healthy until that exact mode is observed
SGLang HiCache Hierarchical cache policy inside the SGLang serving path --enable-hierarchical-cache, capacity, write policy, I/O backend, layout, and storage backend flags That a configured backend moved bytes or improved accepted-task cost
Dynamo KVBM GPU, pinned CPU, and disk KV hierarchy DYN_KVBM_* capacities and DynamoConnector That host memory is CXL or that disk belongs on the latency-critical path
NIXL Transfer orchestration across available memory and network transports Connector or runtime configuration plus UCX/libfabric transport logs A new physical link. NIXL selects and drives transports; it is not NVLink itself
NVLink and NVSwitch High-bandwidth GPU scale-up fabric inside the supported topology nvidia-smi topo, fabric-manager state, counters, and platform topology Cache identity, reuse, placement policy, or application correctness
BlueField DPU Infrastructure processing for networking and storage paths DOCA, RDMA, NVMe-oF, security, and storage services for the named system That KV bypassed the CPU or reached a particular storage tier without a trace
CMX NVIDIA's BlueField-4 STX context-memory storage direction Vendor-described Spectrum-X, NIXL, DOCA Memos, and KV prestaging path Shipping customer economics or the vendor's reported performance and power gains on this workload

The current Dynamo LMCache documentation shows why version and role matter. In aggregated mode, a sidecar can be launched as:

BASH
# Official documentation shape. Capacity is illustrative, not a Touchdown default.
lmcache server --l1-size-gb 100 --eviction-policy LRU &

python -m dynamo.vllm \
  --model MODEL_ID \
  --disable-hybrid-kv-cache-manager \
  --kv-transfer-config \
  '{"kv_connector":"LMCacheMPConnector","kv_role":"kv_both"}'

[OFFICIAL DOCUMENTATION, VERSION-SENSITIVE] In the reviewed documentation, disaggregated decode uses a NIXL connector for the prefill-to-decode transfer, while prefill can use a MultiConnector that combines LMCache retention with NIXL movement. The documentation also carries a compatibility warning for the MP connector with the then-pinned vLLM release. Copying a command without pinning versions is not a reproducible experiment.

SGLang exposes a different policy surface:

BASH
# Official flag names. Backend availability and exact semantics are version-sensitive.
python -m sglang.launch_server \
  --model-path MODEL_ID \
  --enable-hierarchical-cache \
  --hicache-ratio RATIO \
  --hicache-write-policy write_through \
  --hicache-io-backend BACKEND \
  --hicache-storage-backend STORAGE_BACKEND

That command requests a hierarchy. The receipt must still identify which blocks were admitted, bytes written and restored, hit and miss counts, eviction cause, p99 restore time, and whether repository and prompt versions matched.

NIXL is not NVLink, and offload is not one hop

NIXL is a software data-movement layer. NVLink is a physical scale-up interconnect. UCX and libfabric are communication frameworks that can expose RDMA-capable transports. InfiniBand, RoCE, Spectrum-X Ethernet, AWS EFA, PCIe, and storage paths have different controllers, topologies, latencies, and failure modes. Saying NIXL transfer without the selected backend is like saying file copy without naming whether the bytes crossed DRAM, NVMe, Ethernet, or object storage.

NVIDIA Dynamo documents the disaggregated path as:

TEXT
router selects prefill worker
-> prefill reads weights from HBM and creates KV in GPU memory
-> transfer metadata identifies blocks and endpoints
-> NIXL selects an available transport
-> same-domain transfer may use an applicable GPU path
-> cross-node transfer commonly uses RDMA through UCX or libfabric
-> decode worker installs the received KV blocks
-> decode reads local HBM KV and emits tokens

The minimum transfer proof is not a launch flag. It is the worker log showing the instantiated NIXL backend, topology capture, transferred bytes, first-byte and completion latency, overlap with compute, retry or fallback path, and a comparison against the same aggregated workload. A silent fallback to TCP can erase the reason for splitting prefill and decode.

BlueField and CMX add an infrastructure memory tier

BlueField is a DPU, not HBM. A DPU combines programmable processing with network and storage I/O so infrastructure work can be isolated from the host CPU. NVIDIA's BlueField-4 STX and CMX announcement describes a future context-memory storage tier built from Ethernet-attached flash, BlueField processing, Spectrum-X RDMA, NIXL, and DOCA Memos. [OFFICIAL VENDOR DIRECTION] The proposed read path is:

TEXT
request or agent session identifies reusable context
-> metadata lookup resolves an exact KV object and version
-> CMX storage locates flash-resident KV pages
-> BlueField handles storage/network protocol and integrity work
-> Spectrum-X RDMA and NIXL move or pre-stage pages
-> host or GPU destination receives the blocks
-> serving engine validates layout and installs them
-> decode consumes the restored KV from local accelerator memory

NVIDIA reports up to 5 times tokens per second and up to 5 times power efficiency relative to traditional storage approaches for its named CMX context. Those remain [VENDOR CLAIMS], not Touchdown measurements and not a Hermes, Qwen, LMCache, SGLang, or Wan2.2 result. The engineering questions are still exact: Which bytes moved? From which flash and network path? At what granularity? Under what reuse distribution? What was the miss path? What failed? Did restore beat recompute at p99? Did the final patch pass?

The CPU tool loop is part of inference economics

Suppose the model emits a search command. The GPU can become idle while a CPU process searches the repository. The result returns as tokens. The model prefills or reuses a prefix and decodes the next action. A test command can run for seconds or minutes. A failed test returns a long log that expands context and KV state.

The cost can show up outside the model call:

  • Repository indexing or retrieval returns too much low-value context.
  • Tokenization and prompt construction run on the critical path.
  • A stable prefix misses the cache because keys do not include the right version semantics.
  • Tool results are appended repeatedly instead of summarized with a verifiable pointer.
  • A CPU sandbox is underprovisioned, so the accelerator waits.
  • The agent retries a patch because the test contract was vague.
  • The cache restores stale state after the repository changed.

The unit is cost per accepted code change, not tokens per second alone.

What a complete coding-agent receipt contains

  1. Model ID, revision, attention configuration, dtype, serving engine, connector architecture, GPU, host, interconnect, and software versions.
  2. Input tokens split into stable prefix, repository context, current request, prior turns, and tool results.
  3. Prefill duration, verified prefix-cache key, hit/miss, tokens reused, and correctness policy.
  4. Active and reusable KV logical bytes, allocated bytes, block use, fragmentation, eviction, and transfer.
  5. Decode queue time, time to first token, inter-token latency, output tokens, and batching state.
  6. CPU search, file, edit, test, and lint time with accelerator overlap or idle time.
  7. Patch attempts, failed tests, retries, accepted diff, human corrections, and rollback.
  8. HBM, host, optional CXL, and disk residency by timestamp, including first-byte and full-restore p50/p95/p99.

If an LMCache/vLLM result is used, it also names the exact connector generation and mode. [OFFICIAL DOCUMENTATION, AS OF 2026-07-10] The LMCache quickstart recommends MP mode with a standalone lmcache server and LMCacheMPConnector, and documents connector-resolution differences below and above vLLM 0.20.0. The dynamic-connector documentation describes LMCacheConnectorV1 and LMCacheConnectorV1Dynamic while marking the in-process mode deprecated. These names are version-sensitive. Parser support is not live validation, and vLLM CPU-offload metrics are not automatically LMCache evidence.

FIGURE 6 1600 × 900
A coding-agent request moves through prompt construction, prefill, active KV in GPU HBM, decode, CPU tools and tests, optional reusable KV in host or CXL-attached memory, durable repository and logs on NVMe, retries, and an accepted patch.
A coding agent crosses accelerator, host, and storage boundaries. HBM holds hot inference state; lower tiers help only when reuse, versioning, transfer, and fallback are proven. Source state. Illustrative workflow plus official KVBM documentation. No Touchdown performance claim.

Same HBM, different workload

Property Wan2.2 video diffusion Coding-agent inference
Repeated read state Weights and conditioning across denoising steps Weights, stable system/tool prefix, reusable prefix KV
Hot mutable state Latent and current intermediates Active KV, scheduler state, current token stream
Irregular state Sparse-attention metadata and gathers Repository retrieval, tool output, cache lookup
CPU path Encoding, output, orchestration, possible offload Search, file I/O, edits, tests, lint, prompt rebuild
Acceptance gate Requested clip quality, format, latency Correct patch, passing tests, reviewability
Failure cost Wasted denoising run or rejected output Repeated prefill/decode/tools, failed patch, human rework
Predictable stream opportunity Layer/weight order can be predictable Stable prefixes and read-mostly artifacts can repeat, but tools branch

That symmetry is the reason to resist one universal memory slogan. Both workloads use HBM. They do not create the same state contract.

How does the NVIDIA software stack turn one GLM-5.2 coding task into HBM traffic?

In plain English. Hermes receives the coding task, a serving engine decides when and where model work runs, PyTorch expresses tensor operations, libraries or compilers choose implementations, CUDA launches GPU kernels, and the hardware reads and writes HBM. Those are different layers. A real receipt must name which branch was selected instead of listing every available library as if all of them ran.

The user does not buy CUDA, HBM bandwidth, or tokens per second. The user asks Hermes to inspect a repository, change code, run the tests, repair a failure, and return a patch that can be accepted. The unit that matters is one verified or accepted code change. Everything below that outcome is a path the system pays for: context construction, prefix lookup, prefill, MoE routing, matrix multiplication, attention, KV-cache writes, interconnect traffic, tool waits, retries, power, heat, cooling, and review.

This section follows that one task through NVIDIA's software and hardware path. It is a reference architecture, not a claim that every named component ran in one production request. A component is marked as observed only when a pinned run receipt contains its version, configuration, event, and artifact. Until then, the page says ARCHITECTURE ONLY / RUN NOT CAPTURED.

Code proof level: The companion examples are original Touchdown reference examples built against public APIs. They remain fixture_backed or not_started until the declared toolchain compiles them and a named GPU run produces correctness and profiler artifacts. The article never upgrades a documentation example into a measured GLM-5.2 result.

The same thirteen events stay visible from request to bill

The left learning journey and the article use the same event identity. Changing the audience view changes the explanation, not the workload, platform, precision, or event.

Event User and business boundary Software and GPU boundary Physics and economics boundary
1. Request One repository change with explicit tests and acceptance criteria Hermes creates the tool-capable conversation Admission starts the occupied-time boundary
2. Context Repository map, selected files, instructions, and tool schemas Tokenizer and prompt builder create input IDs Context bytes determine prefill work and initial KV demand
3. Route Quality, latency, and budget policy selects a path Model, engine, precision, parallelism, and fallback are pinned The route commits HBM capacity and infrastructure scope
4. Prefix lookup Reused context can avoid repeated work vLLM prefix caching, SGLang RadixAttention/HiCache, or LMCache performs its own exact lookup A hit has value only if restored state arrives before recomputation would finish
5. Prefill The model reads the complete admitted prompt Batched matrix multiplications and attention create KV state Weight reads, activation traffic, and KV writes pressure HBM bandwidth
6. Attention and MoE The model selects information and routed experts Attention, top-k routing, permutation, grouped GEMM, and expert combine execute Tensor Core work and irregular expert traffic compete for local and fabric bandwidth
7. Lowering No direct user-visible step PyTorch eager, Inductor/Triton, FlashInfer, cuDNN, cuBLASLt, CUTLASS, CuTe, or custom CUDA selects executable work Tile shape, data layout, registers, shared memory, TMEM, and launch shape decide useful traffic
8. GPU and HBM Work is still pending CUDA Runtime/Driver launches kernels and manages streams, events, graphs, and allocations SASS issues loads, stores, tensor operations, synchronization, and memory-controller requests
9. KV placement Conversation state must survive the next turn Engine allocator appends KV or restores/offloads it through LMCache, HiCache, KVBM, or another declared path HBM, pinned host memory, SSD, and remote memory have different latency and energy boundaries
10. Decode The next tool call or answer appears token by token Repeated small decode steps reuse weights and growing KV state Decode can become bandwidth and launch sensitive even when theoretical FLOPs are abundant
11. Tool call Search, edit, compiler, test, or lint leaves the model CPU process, filesystem, sandbox, and storage execute the tool GPU idle or overlap time still consumes reserved capacity and facility power
12. Retry and verifier Tests pass, fail, or force another turn Logs re-enter context; the model may prefill, decode, and call tools again Rejected work still consumed time, energy, cooling capacity, and money
13. Outcome and bill Patch is verified, accepted, rejected, or pending Run artifacts join to the explicit verifier Cost and energy per accepted patch are valid only when numerator and acceptance boundary align

Precision is four decisions, not one label

FP8 does not mean every number in the system uses eight bits. Weights, temporary activations, arithmetic accumulation, and KV cache can use different formats in the same request. Each choice changes capacity, traffic, accuracy, and kernel support differently.

Saying a request "uses FP8" is incomplete. The receipt must separate weight storage, activation storage, accumulation, and KV-cache precision. They can differ within one layer.

Path Weight storage Activations Accumulation KV cache Evidence boundary
BF16 reference BF16 BF16 FP32 or declared implementation BF16 unless configured otherwise Correctness reference, not automatically the cheapest path
FP8 path FP8 or mixed by module FP8/BF16 by operation Higher precision declared by kernel/library Independent setting Requires scales, amax policy, excluded modules, and parity results
NVFP4 preparation Block-scaled FP4 where supported Mixed FP4/FP8/BF16 Higher precision declared by kernel Independent setting A B200/GB200 reference experiment until export, compile, parity, and accepted-task replay pass
H200 comparison BF16 or FP8 BF16 or FP8 Declared implementation Independent setting HBM3E/Hopper path; do not imply Blackwell NVFP4 execution
MI355X comparison ROCm-supported BF16/FP8/FP4 path Declared by AMD runtime/kernel Declared implementation Independent setting Separate AMD receipt; never inferred from an NVIDIA result

The GLM-5.2 model revision, tokenizer, configuration digest, quantization artifact, calibration data hash, excluded modules, scale granularity, and output-quality result belong in the same receipt. A smaller weight file is not a deployment win if quality falls, KV pressure forces a smaller batch, or the runtime cannot execute the artifact.

PyTorch starts the operation, but it does not name the final kernel

At the Python level, a transformer layer expresses tensor operations. The same operation can remain in eager ATen, enter torch.compile, lower through TorchInductor and Triton, call a library backend, or hit a graph break and return to eager execution.

PYTHON
# TOUCHDOWN REFERENCE SHAPE. Not a captured GLM-5.2 production layer.
import torch

@torch.compile(dynamic=True)
def attention_slice(q, k, v):
    return torch.nn.functional.scaled_dot_product_attention(
        q, k, v, is_causal=True
    )

The code receipt must record the active source line, graph-break report, generated graph, selected backend, input shapes, strides, dtypes, and output parity. A Python function name is not proof that FlashInfer, cuDNN, or a particular fused kernel executed.

For the routed expert path, the useful symbolic shape is:

TEXT
X_e [tokens_for_expert, hidden]
  x W_up,e [hidden, intermediate]
  -> U_e [tokens_for_expert, intermediate]
  -> activation and gate
  x W_down,e [intermediate, hidden]
  -> Y_e [tokens_for_expert, hidden]

Before the grouped GEMM, routing selects experts and permutes tokens into expert-local groups. After it, outputs are combined into original token order. If expert parallelism crosses GPUs, token dispatch and combine add all-to-all traffic around the local matrix multiply. That is why a fast standalone GEMM does not prove a faster coding-agent turn.

The serving engine owns scheduling and KV allocation

Use four questions to keep the names straight:

  1. Who owns the state? The model server and its KV or prefix-cache manager identify and allocate it.
  2. Who decides placement? The scheduler and cache policy choose GPU, host, disk, or another declared tier.
  3. Who moves the bytes? A runtime, movement layer, or storage backend issues and tracks the transfer.
  4. What physical path carries them? NVLink, PCIe, RDMA, Ethernet, storage, or another named link supplies the actual route.

vLLM and SGLang are independent serving projects. LMCache is an independent KV-cache layer. NVIDIA Dynamo can orchestrate engines and separate prefill from decode. TensorRT-LLM is NVIDIA's optimized LLM runtime path. These are alternatives and integrations, not evidence that all of them executed together.

The reference packet contains separate recipes:

TEXT
Recipe A: Hermes -> GLM-5.2 -> vLLM -> engine-native prefix cache -> optional LMCache
Recipe B: Hermes -> GLM-5.2 -> SGLang -> RadixAttention -> optional HiCache/NIXL
Recipe C: supported model artifact -> TensorRT-LLM -> NVIDIA-native comparison
Recipe D: Dynamo router -> prefill worker -> NIXL KV transfer -> decode worker

Only one recipe can be the observed primary path for a single run. A prefix hit also needs the engine's real identity rule, requested block count, found block count, hit tokens, restore bytes, restore time, and fallback. A conceptual similarity score is not an engine-level hit.

The integrated C-001 harness joins the whole path without inventing a run.

The public code packet includes one integrated C-001 reference harness. It is deliberately different from a benchmark launcher. Its default mode runs on a CPU, makes no network request, does not load GLM-5.2, and does not claim that CUDA or HBM executed. It assembles the complete typed receipt shape so an engineer, executive, or investor can see exactly which evidence a real replay still owes.

PYTHON
EVENT_ORDER = (
    "request", "context", "route", "prefix_lookup", "prefill",
    "attention_moe", "lowering", "gpu_hbm", "kv_placement",
    "decode", "tool_call", "retry_verifier", "outcome_bill",
)

config = load_reference_config("reference-config.json")
receipt = architecture_receipt(config)

# Default truth boundary:
assert receipt["run_id"] is None
assert receipt["capture_kind"] == "architecture_only"
assert receipt["banner"] == "ARCHITECTURE ONLY / RUN NOT CAPTURED"
assert all(row["value"] is None for row in receipt["resource_ledger"])

Run the inspectable path with no GPU:

BASH
python3 examples/hbm-learning-journey/nvidia/08-c001-integrated/c001_reference_harness.py --pretty
python3 -m unittest discover \
  -s examples/hbm-learning-journey/nvidia/08-c001-integrated \
  -p 'test_*.py' -v

Capture mode does not start Hermes, vLLM, SGLang, CUDA, Nsight, or a facility meter. It joins artifacts a separate real workload already produced. The join requires fifteen hashed artifact kinds: request, context, route, prefix-cache identity, engine trace, operator trace, compiler artifacts, kernel profile, memory placement, fabric, tool, verifier, power, facility, and cost. Every artifact must carry the same non-placeholder run_id and synchronized clock boundary.

The harness rejects partial capture arguments, path traversal, hash mismatch, mixed run IDs, cache claims without model/tokenizer/engine/prefix/KV identity, accepted outcomes without a measured live verifier, duplicated energy ledgers, and coolant circulation mislabeled as consumed water. This is the actual code boundary between a useful architecture lesson and a false production claim.

Operator libraries make different promises

The operator layer is where the high-level tensor operation becomes a concrete implementation choice.

  • cuBLAS and cuBLASLt provide matrix multiplication. cuBLASLt adds descriptors, layouts, heuristic selection, workspace, fused epilogues, and reduced-precision paths. The receipt needs matrix dimensions, layouts, input and compute types, selected algorithm, workspace, and duration.
  • CUTLASS exposes source-visible hierarchical GEMM, tiling, pipelines, copies, Tensor Core operations, and epilogues. The packet includes BF16, FP8, block-scaled NVFP4, grouped MoE, decode-shape GEMV, and profiler examples tied to a pinned release.
  • CuTe and CuTe DSL expose layouts, thread-value mapping, tiled copies, tiled MMA, TMA, shared-memory layout, and architecture-specific narrow-precision movement. A layout drawing is not a measured kernel.
  • Transformer Engine manages transformer-oriented low-precision execution, recipes, scaling, and numerical comparison. Weight precision, activation precision, accumulation, and KV precision remain separate.
  • NVIDIA Model Optimizer prepares quantized or otherwise transformed artifacts. Its output must retain the original revision, calibration hash, quantization configuration, excluded modules, export format, and evaluation delta.
  • cuDNN Backend/Frontend can select operation graphs, heuristics, execution plans, attention paths, and workspaces. It stays an alternative until a trace names its plan.
  • FlashInfer is an independent serving-kernel project with paged-KV, prefill, decode, sampling, GEMM, and MoE surfaces. Its exact API must match the pinned release.
  • Triton is an independent kernel language and compiler used by PyTorch and inference projects. Generated Triton IR, LLVM IR, PTX, and cubin are artifacts, not assumptions.
  • Custom CUDA C++ gives direct control and also direct responsibility for bounds, synchronization, numerics, portability, and maintenance.

CUDA compilation makes the executable boundary visible

The offline path and runtime-compiled path must remain distinct:

TEXT
CUDA C++ source
  -> nvcc or NVRTC
  -> PTX virtual ISA
  -> ptxas or driver JIT
  -> cubin
  -> architecture-specific SASS
  -> scheduler, Tensor Core/CUDA Core, load-store unit
BASH
# REFERENCE COMMAND SHAPE. Select the architecture for the detected GPU and pinned toolkit.
nvcc -O3 --ptx kernel.cu -o kernel.ptx
nvcc -O3 --cubin kernel.cu -o kernel.cubin
cuobjdump --dump-sass kernel.cubin
nvdisasm kernel.cubin > kernel.sass

PTX is not final GPU machine code. SASS is the architecture-specific disassembly. NVRTC compiles device code at runtime; NVJitLink can join device modules. Compile latency must be reported separately from execution latency, and cached compilation must be distinguished from cold compilation.

Runtime allocation is not the same as engine KV allocation

The CUDA Runtime and Driver APIs expose devices, contexts, modules, streams, events, copies, launches, memory pools, peer access, and graphs. cudaMallocAsync can reuse storage from a stream-ordered pool. vLLM or SGLang can manage logical KV blocks above that allocation layer. The two allocators answer different questions.

CPP
// TOUCHDOWN REFERENCE SHAPE. Error checks omitted here; full example is in the repo.
cudaMemPool_t pool;
cudaDeviceGetDefaultMemPool(&pool, device_id);
cudaMallocAsync(&buffer, bytes, stream);
kernel<<<grid, block, shared_bytes, stream>>>(buffer);
cudaEventRecord(done, stream);
cudaFreeAsync(buffer, stream);

CUDA Graphs can reduce repeated CPU launch overhead by capturing a compatible sequence and replaying it. Batch shape, expert routing shape, addresses, or unsupported dynamic work can prevent reuse. The receipt therefore records graph eligibility, capture result, replay count, update failures, and fallback.

NCCL, NVLink, NVSwitch, and NIXL occupy different levels

NCCL asks for communication. NIXL coordinates state movement. NVLink carries bits over a physical scale-up link. NVSwitch routes those links across a supported domain. The same task can use more than one, but the names are not interchangeable.

NCCL implements collectives such as all-reduce, all-gather, reduce-scatter, all-to-all, and send/receive. NVLink and NVSwitch are physical scale-up fabric, not function calls. NIXL orchestrates point-to-point state movement and can select transports or storage backends. Dynamo KVBM manages KV blocks across declared memory tiers and uses NIXL for movement.

TEXT
MoE token dispatch
  -> NCCL or declared communication backend
  -> NVLink/NVSwitch inside a supported scale-up domain
  -> InfiniBand/RoCE and GPUDirect RDMA across nodes when qualified

Disaggregated KV movement
  -> Dynamo/NIXL request
  -> endpoint metadata and selected backend
  -> GPU, pinned host, SSD, or remote source
  -> physical transport
  -> destination and completion event

NCCL Tests qualify a topology. They do not become GLM-5.2 throughput. nvidia-smi topo -m, NVLink status, Fabric Manager, DCGM, UCX information, InfiniBand verbs tools, and NIXL logs must agree about the route. A fallback to host bounce or TCP remains visible.

The trace must join code, hardware, power, cooling, and acceptance

NVTX gives business phases names such as C-001/prefill, C-001/moe_grouped_gemm, C-001/tool/pytest, and C-001/verifier. CUPTI supplies lower-level activity and correlation data used by NVIDIA profiling tools. Nsight Systems shows the end-to-end CPU, CUDA, memory-copy, NCCL, and idle timeline. Nsight Compute inspects selected kernels. Compute Sanitizer checks memory, race, initialization, and synchronization failures.

BASH
nsys profile --trace=cuda,nvtx,osrt,cudnn,cublas \
  --output=glm52-hermes python run_hermes_fixture.py

ncu --set full --target-processes all \
  --kernel-name 'regex:.*moe.*gemm.*' \
  --output=glm52-moe-kernel python run_single_layer_fixture.py

NVML and DCGM can expose device identity, power, cumulative energy where supported, temperature, clocks, utilization, memory, health, and fabric state. They do not automatically provide facility energy or task attribution. A child GPU energy counter explains part of a rack or facility meter; it must not be added again when the parent meter already includes it.

The aligned accounting formulas remain:

TEXT
IT_kWh = IT_kW * occupied_seconds / 3600
facility_kWh = IT_kWh * PUE
site_water_L = IT_kWh * WUE
cost_per_accepted_patch = total_aligned_path_cost / accepted_patches

PUE and WUE require a declared facility and time boundary. Closed-loop coolant circulation is not site water consumption. Electricity-generation water is another boundary again. If those inputs are absent, the result stays unknown.

Search the complete NVIDIA and adjacent-library packet

The matrix below is generated from the repository registry. It includes every component named in the archived research input. Required, Trace, Alternative, and Later describe its role in the teaching packet. Observed in run is a separate field and remains false until a receipt proves execution.

Component and ownerCategoryPacket roleWorkload hookTrace stateCoverage and evidenceReceipts
CUDA cubin artifact
NVIDIA
binary_artifactRequiredcompiler_artifact
preserve target-specific GPU code object
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NCCL
NVIDIA
collective_libraryRequireddistributed_execution
collective traffic for tensor or expert parallelism
available_not_on_traceparser_only
source_backed_architecture
code · source
nvcc
NVIDIA
compilerRequiredoffline_compilation
compile CUDA C++ to PTX, cubin and executable artifacts
available_not_on_tracefixture_backed
source_backed_architecture
code · source
torch.compile and TorchInductor
PyTorch Foundation
compilerRequiredoperator_lowering
capture FX regions and lower them through Inductor
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVIDIA Container Toolkit
NVIDIA
container_runtimeRequiredenvironment
expose GPUs and driver libraries to OCI containers
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NIXL
NVIDIA / ai-dynamo open-source project
data_movementRequiredkv_transfer
orchestrate transfer through selected memory, network or storage plugins
available_not_on_traceparser_only
source_backed_architecture
code · source
GPUDirect RDMA
NVIDIA
data_pathRequiredscale_out_transfer
enable supported NIC DMA to registered GPU memory
available_not_on_traceparser_only
source_backed_architecture
code · source
NVIDIA Dynamo
NVIDIA / ai-dynamo open-source project
distributed_servingRequiredrouting_and_disaggregation
route requests and separate prefill from decode workers
available_not_on_traceparser_only
source_backed_architecture
code · source
NGC container
NVIDIA
environmentRequiredenvironment
freeze a compatible user-space CUDA and inference stack
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVLink
NVIDIA
hardware_interconnectRequiredscale_up_fabric
carry supported GPU-to-GPU scale-up traffic
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVSwitch
NVIDIA
hardware_switchRequiredscale_up_fabric
switch NVLink traffic inside the supported domain
available_not_on_tracefixture_backed
source_backed_architecture
code · source
vLLM
vLLM Project
inference_engineRequiredserving
paged KV allocation, scheduling, prefix caching, prefill and decode
available_not_on_traceparser_only
source_backed_architecture
code · source
FlashInfer
FlashInfer Project
inference_kernel_libraryRequiredprefill_and_decode_attention
single-request prefill and decode attention against KV state
available_not_on_tracefixture_backed
source_backed_architecture
code · source
PTX ISA artifact
NVIDIA
intermediate_representationRequiredcompiler_artifact
preserve virtual ISA before target-specific assembly
available_not_on_tracefixture_backed
source_backed_architecture
code · source
Custom CUDA C++ kernel
Touchdown Labs example using NVIDIA CUDA
kernelRequiredkernel_execution
launch a minimal checked kernel through the runtime and driver paths
available_not_on_tracefixture_backed
illustrative
code · source
CuTe and CuTe DSL
NVIDIA
kernel_dslRequiredkernel_authoring
compose layouts, copies and MMA operations
available_not_on_traceparser_only
source_backed_architecture
code · source
Triton language and compiler
Triton Project
kernel_dslRequiredoperator_lowering
author and JIT compile a GPU kernel
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CUTLASS
NVIDIA
kernel_libraryRequiredgemm_and_moe
source-visible BF16, FP8, NVFP4 and grouped GEMM kernels
available_not_on_traceparser_only
source_backed_architecture
code · source
Dynamo KVBM
NVIDIA / ai-dynamo open-source project
kv_cacheRequiredkv_offload_and_restore
manage KV blocks across GPU, pinned host, SSD and remote tiers
available_not_on_traceparser_only
source_backed_architecture
code · source
SASS disassembly
NVIDIA
machine_instruction_artifactRequiredcompiler_artifact
disassemble the target-specific instruction stream
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuBLAS
NVIDIA
math_libraryRequiredlinear_layer
ordinary GEMM baseline
available_not_on_traceparser_only
source_backed_architecture
code · source
cuBLASLt
NVIDIA
math_libraryRequiredlinear_layer
descriptor-based GEMM with heuristic tactic selection and workspace
available_not_on_traceparser_only
source_backed_architecture
code · source
CUDA stream-ordered memory pools
NVIDIA
memory_managementRequiredallocation
allocate and reuse device memory with cudaMallocAsync
available_not_on_traceparser_only
source_backed_architecture
code · source
PyTorch eager and ATen
PyTorch Foundation
model_frameworkRequiredmodel_execution
reference eager operator path
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVIDIA Model Optimizer
NVIDIA
model_optimizationRequiredquantization_preparation
calibrate and prepare a quantized model artifact
available_not_on_traceparser_only
source_backed_architecture
code · source
Transformer Engine
NVIDIA
precision_libraryRequiredtransformer_linear
FP8 execution with explicit recipe context
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CUDA Graphs
NVIDIA
runtimeRequireddecode_replay
capture and replay a stable launch graph
available_not_on_traceparser_only
source_backed_architecture
code · source
CUDA Runtime API
NVIDIA
runtimeRequiredkernel_dispatch
device allocation, asynchronous launch and error handling
available_not_on_traceparser_only
source_backed_architecture
code · source
CUDA streams and events
NVIDIA
runtimeRequiredkernel_dispatch
order asynchronous work and measure GPU elapsed time
available_not_on_traceparser_only
source_backed_architecture
code · source
CUDA Toolkit
NVIDIA
toolchainRequiredenvironment
provide compiler, runtime, libraries, binary tools and profilers
available_not_on_tracefixture_backed
source_backed_architecture
code · source
Compute Sanitizer
NVIDIA
correctness_toolTracekernel_correctness
check memory and synchronization errors
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NCCL Tests
NVIDIA
fabric_benchmarkTracefabric_qualification
measure collective bandwidth on the exact allocated topology
available_not_on_tracefixture_backed
source_backed_architecture
code · source
DCGM
NVIDIA
fleet_telemetryTracedevice_and_fabric_telemetry
collect health, power, energy, thermal, memory and link fields
available_not_on_traceparser_only
source_backed_architecture
code · source
CUPTI
NVIDIA
instrumentationTracegpu_activity_collection
provide callback and activity data beneath profiler tools
available_not_on_traceparser_only
source_backed_architecture
code · source
NVTX
NVIDIA
instrumentationTraceall_business_and_runtime_phases
name C-001 phases for cross-layer timeline joins
available_not_on_traceparser_only
source_backed_architecture
code · source
Nsight Compute
NVIDIA
kernel_profilerTraceselected_kernel_analysis
collect kernel-level memory, occupancy and instruction metrics
available_not_on_tracefixture_backed
source_backed_architecture
code · source
InfiniBand verbs and performance tools
linux-rdma community with vendor contributions
network_qualificationTracefabric_qualification
inspect RDMA devices and measure bandwidth or latency
available_not_on_tracefixture_backed
source_backed_architecture
code · source
nvidia-smi
NVIDIA
operator_cliTraceenvironment_and_topology
report device identity, topology, link and coarse telemetry
available_not_on_tracefixture_backed
source_backed_architecture
code · source
Nsight Systems
NVIDIA
profilerTraceend_to_end_timeline
capture CPU, CUDA, NVTX and selected system activity
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CUDA Driver API
NVIDIA
runtimeTracemodule_load_and_dispatch
load a target cubin, resolve a symbol and launch explicitly
available_not_on_traceparser_only
source_backed_architecture
code · source
NVML
NVIDIA
telemetry_apiTracedevice_power_and_health
sample timestamped power, temperature, clocks, utilization and memory
available_not_on_tracefixture_backed
source_backed_architecture
code · source
UCX
OpenUCX Project
communication_frameworkAlternativetransport_selection
provide selected shared-memory, TCP, RDMA or GPU-aware transports
available_not_on_tracefixture_backed
source_backed_architecture
code · source
nvJitLink
NVIDIA
compilerAlternativeruntime_linking
link device code into a loadable cubin
available_not_on_traceparser_only
source_backed_architecture
code · source
NVRTC
NVIDIA
compilerAlternativeruntime_compilation
compile an in-memory CUDA C++ device program to PTX
available_not_on_traceparser_only
source_backed_architecture
code · source
SGLang
SGLang Project
inference_engineAlternativeserving
RadixAttention scheduling, prefill, decode and tool-serving path
available_not_on_traceparser_only
source_backed_architecture
code · source
TensorRT-LLM
NVIDIA
inference_engineAlternativeserving
NVIDIA-native LLM runtime comparison
unsupportedparser_only
source_backed_architecture
code · source
TensorRT
NVIDIA
inference_runtimeAlternativesupporting_model_execution
build and run engines for encoders, rerankers or auxiliary models
available_not_on_traceparser_only
source_backed_architecture
code · source
cuDNN backend and frontend graph APIs
NVIDIA
kernel_libraryAlternativeattention_or_fusion
build and execute supported operation graphs
available_not_on_traceparser_only
source_backed_architecture
code · source
LMCache
LMCache Project
kv_cacheAlternativeprefix_lookup_and_offload
lookup, inject, store and multi-tier KV movement through an engine connector
available_not_on_traceparser_only
source_backed_architecture
code · source
SGLang HiCache
SGLang Project
kv_cacheAlternativeprefix_lookup_and_offload
hierarchical GPU, host and storage KV caching
available_not_on_traceparser_only
source_backed_architecture
code · source
CUDA Unified Memory
NVIDIA
memory_managementAlternativememory_residency_teaching
demonstrate managed virtual memory and migration
available_not_on_tracenot_started
source_backed_architecture
code · source
NVSHMEM
NVIDIA
one_sided_communicationAlternativedistributed_execution
GPU-initiated one-sided communication and synchronization
available_not_on_traceparser_only
source_backed_architecture
code · source
nvCOMP
NVIDIA
compressionLaterstate_storage
compress or decompress candidate state on GPU
available_not_on_traceparser_only
source_backed_architecture
code · source
CUB
NVIDIA / CCCL
cuda_cpp_core_libraryLateroptional_device_primitive
provide block and device scan, reduce, sort and selection primitives
available_not_on_traceparser_only
illustrative
code · source
libcu++
NVIDIA / CCCL
cuda_cpp_core_libraryLaterkernel_authoring
use CUDA-aware C++ atomics and standard-library facilities
available_not_on_traceparser_only
illustrative
code · source
Thrust
NVIDIA / CCCL
cuda_cpp_core_libraryLateroptional_parallel_algorithm
invoke CUDA C++ parallel algorithms
available_not_on_traceparser_only
illustrative
code · source
cuFFT
NVIDIA
cuda_library_catalogLaterscientific_or_media_operator
fast Fourier transforms when invoked by a supporting workload
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuRAND
NVIDIA
cuda_library_catalogLatersampling_or_diffusion_noise
generate random values for sampling or diffusion paths
available_not_on_traceparser_only
illustrative
code · source
cuSOLVER
NVIDIA
cuda_library_catalogLateroptional_solver
dense or sparse factorizations for supporting workloads
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuSPARSE
NVIDIA
cuda_library_catalogLateroptional_sparse_operator
sparse matrix operations when the selected workload invokes them
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuSPARSELt
NVIDIA
cuda_library_catalogLateroptional_structured_sparse_gemm
structured sparse matrix multiplication
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuTENSOR
NVIDIA
cuda_library_catalogLateroptional_tensor_contraction
general tensor contraction when selected by a workload
available_not_on_tracefixture_backed
source_backed_architecture
code · source
cuTensorNet
NVIDIA
cuda_library_catalogLatertensor_network_workload
tensor-network contraction outside the central GLM path
available_not_on_tracefixture_backed
source_backed_architecture
code · source
Cooperative Groups
NVIDIA
cuda_programming_modelLaterkernel_authoring
coordinate explicit thread groups
available_not_on_traceparser_only
illustrative
code · source
DALI
NVIDIA
data_pipelineLatermedia_loading_and_preprocessing
build accelerated input pipelines for declared media workloads
available_not_on_tracefixture_backed
source_backed_architecture
code · source
DOCA
NVIDIA
dpu_sdkLaternetwork_or_storage_offload
program supported BlueField networking, storage or DPA services
available_not_on_tracenot_started
source_backed_architecture
code · source
Multi-Instance GPU (MIG)
NVIDIA
gpu_partitioningLaterdeployment_and_isolation
partition supported GPUs into isolated instances
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CUDA Multi-Process Service (MPS)
NVIDIA
gpu_sharingLaterdeployment_and_concurrency
share GPU execution resources across CUDA processes
available_not_on_tracefixture_backed
source_backed_architecture
code · source
BlueField DPU
NVIDIA
hardware_dpuLaternetwork_or_storage_offload
host supported DOCA services and offload selected infrastructure work
available_not_on_tracenot_started
source_backed_architecture
code · source
NVIDIA GPU Operator
NVIDIA
kubernetes_operatorLatercluster_deployment
reconcile GPU drivers, device plugin, toolkit, telemetry and MIG components
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVIDIA Network Operator
NVIDIA
kubernetes_operatorLatercluster_deployment
reconcile RDMA and accelerated networking components
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVDEC
NVIDIA
media_hardware_apiLatervideo_decode
decode source or validation video for Wan2.2 workflows
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NVENC
NVIDIA
media_hardware_apiLatervideo_encode
encode accepted Wan2.2 output clips
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CV-CUDA
NVIDIA / CV-CUDA Project
media_libraryLatermedia_preprocessing
GPU image preprocessing for a declared media workload
available_not_on_tracefixture_backed
source_backed_architecture
code · source
NPP
NVIDIA
media_libraryLaterimage_or_signal_processing
run image and signal primitives when selected by a media path
available_not_on_tracefixture_backed
source_backed_architecture
code · source
nvImageCodec
NVIDIA
media_libraryLaterimage_codec_dispatch
select image codec implementations for a media workload
available_not_on_tracefixture_backed
source_backed_architecture
code · source
nvJPEG
NVIDIA
media_libraryLaterimage_decode
decode JPEG inputs for media workloads
available_not_on_tracefixture_backed
source_backed_architecture
code · source
nvJPEG2000
NVIDIA
media_libraryLaterimage_decode
decode JPEG 2000 inputs for media workloads
available_not_on_tracefixture_backed
source_backed_architecture
code · source
CUDA Virtual Memory Management
NVIDIA
memory_managementLateraddress_space_management
reserve, map and remap virtual address ranges
available_not_on_tracenot_started
source_backed_architecture
code · source
NVIDIA Triton Inference Server
NVIDIA
model_serverLaterdeployment
serve model repositories through backend adapters
available_not_on_traceparser_only
source_backed_architecture
code · source
cuDSS
NVIDIA
solver_libraryLateroptional_sparse_direct_solver
solve sparse linear systems outside the central LLM path
available_not_on_tracefixture_backed
source_backed_architecture
code · source
GPUDirect Storage and cuFile
NVIDIA
storage_data_pathLaterweight_or_state_storage
move supported file data to or from GPU memory with explicit fallback reporting
available_not_on_traceparser_only
source_backed_architecture
code · source

B200, GB200, H200, MI355X, Rubin, and Kyber are not interchangeable labels

The default reference profile is B200/GB200 with an NVFP4 preparation path. That is not a public GLM-5.2 NVFP4 workload receipt. H200 remains a Hopper/HBM3E BF16 or FP8 comparison. MI355X remains a separate AMD ROCm receipt path. Equal evidence means the same task, model revision or explicitly bounded substitute, quality gate, SLO, concurrency, meter boundary, and acceptance rule.

Rubin is a GPU generation. Kyber is a rack architecture for larger future scale-up domains. Neither name is evidence that this C-001 run executed there. Roadmap and preliminary system specifications stay source-backed architecture until hardware, software, topology, and workload receipts exist.

What the run receipt must prove

One real C-001 execution has one run_id. Wan2.2 V-001 has a different run_id; the two workloads join only through a declared comparison identifier.

The public page emits a demo manifest, not a fake run receipt:

YAML
schema_version: touchdown.hbm-demo-manifest.v1
fixture_id: C-001
comparison_id: CMP-C001-V001
publication_hold: false
trace_contract:
  run_id: null
  capture_kind: architecture_only
  banner: ARCHITECTURE ONLY / RUN NOT CAPTURED
run_receipt_emitted: false

A real receipt is a different object. It requires one non-placeholder run ID and a captured environment, typed events, artifacts, verifier, and resource boundaries:

YAML
schema_version: touchdown.hbm-run-receipt.v1
fixture_id: C-001
run_id: required
capture_kind: measured_run
environment:
  agent: Hermes
  model_revision: required
  engine_and_version: required
  weight_precision: required
  activation_precision: required
  accumulation_precision: required
  kv_precision: required
  gpu_uuid_and_topology: required
events:
  - request
  - context_snapshot
  - route_decision
  - cache_lookup
  - engine_phase
  - model_operation
  - kernel_dispatch
  - memory_activity
  - fabric_activity
  - tool_call
  - verifier_result
  - power_interval
  - thermal_cooling_interval
  - water_allocation
  - cost_allocation
outcome:
  state: pending

verified_patch requires the declared automated verifier. accepted_patch additionally requires explicit acceptance. No accepted output means energy and cost per accepted output are null even though total time, energy, and cost remain visible.

Primary documentation for this code path

How do Vera, Rubin, HBM4, CMX, and Kyber form one accepted-task system?

In plain English. Vera is the CPU side that handles control-heavy work and tools. Rubin is the GPU side that repeatedly transforms model tensors. HBM4 is Rubin's local working memory. NVLink connects selected CPU and GPU or GPU and GPU domains. CMX is a lower context-storage tier. Kyber is a future rack design. Together they describe possible homes and paths for one task; they are not one chip and they do not prove that GLM-5.2 ran on the platform.

The earlier sections explain each layer. This section keeps all of them on one page.

Start with the user. A developer does not buy an HBM stack, an NVLink domain, or a rack. The developer asks Hermes, backed by GLM-5.2, to inspect a repository, change code, call tools, run tests, repair failures, and return one accepted patch. The system earns its cost only when that verifier closes.

TEXT
C-001 request
-> repository, policy, tools, and prior turns
-> CPU orchestration and tokenization
-> cache identity and placement decision
-> GPU prefill, attention, MoE, and KV creation
-> decode a tool call
-> CPU sandbox, compiler, tests, or retrieval
-> restore a valid prefix or recompute it
-> more GPU decode
-> candidate patch
-> verifier
-> accepted patch or another paid attempt

NVIDIA's Vera Rubin architecture gives the path named physical homes. Vera is the CPU and CPU-memory side. Rubin is the GPU and HBM4 side. NVLink-C2C connects those two memory domains coherently. NVLink 6 and NVSwitch connect GPU domains for scale-up. ConnectX and Spectrum-X or InfiniBand handle other network boundaries. BlueField-4 handles infrastructure work. CMX is a flash-backed shared context tier below HBM and host memory. MGX packages compute, switching, power, cooling, mechanics, and serviceability into rack systems. Kyber is a later MGX rack architecture, not a GPU or memory generation.

Evidence boundary. NVIDIA's 2026 material is official product and preliminary roadmap documentation. It establishes named architecture, published specifications, and vendor targets. It does not establish a Touchdown GLM-5.2 run, accepted-patch rate, customer price, sustained HBM traffic, rack energy per task, CMX restore p99, or Kyber application result. Those fields remain unknown until a synchronized run receipt exists.

FIGURE 13 1600 × 900
A three-lane map follows one Hermes and GLM-5.2 coding task through Vera CPU orchestration, Rubin GPU and HBM4 execution, NVLink scale-up, BlueField-4 CMX context storage, verification, rack power and cooling, and supply qualification.
One accepted patch crosses software, memory, fabric, package, facility, and manufacturing boundaries. The arrows show functional order, not a measured GLM-5.2-on-Rubin trace. Source state. Official NVIDIA architecture plus Touchdown systems synthesis. Rubin and Kyber workload values remain unmeasured.

One request crosses six different boundaries

The word memory hides six different contracts in this example.

Boundary What crosses it Physical path Failure the user feels
Application to model Instructions, schemas, retrieved code, tool results CPU memory into the serving engine Wrong or bloated context, bad cache identity, longer prefill
Vera to Rubin Tokens, metadata, state, commands, and possibly reusable tensors Coherent NVLink-C2C between CPU LPDDR5X and GPU HBM4 domains Migration or remote access misses the latency deadline
Rubin to local HBM4 Weights, active KV, activations, expert buffers, workspaces GPU cache and controller into multiple HBM4 stacks Capacity pressure, low useful bandwidth, stalls, throttling
Rubin to Rubin Tensor, expert, sequence, or collective buffers NVLink 6 through NVSwitch inside a named scale-up domain Exposed collective wait, imbalance, retries, degraded links
Compute to context storage Reusable or evicted KV blocks plus identity and integrity metadata HBM or host memory through DPU/NIC, Spectrum-X, CMX flash, and back Restore arrives too late, misses, is stale, or reduces batch capacity
Silicon to accepted result Electricity, heat, telemetry, retries, and verifier state Package, rack power, liquid loop, facility boundary, and software receipt More rack time, energy, water allocation, or money per accepted patch

These are not one flat address space or one flat bandwidth number. Coherence can make CPU and GPU memory easier to program without making remote CPU memory equal to local HBM. A large NVLink domain can reduce a scale-out boundary without making all HBM bytes equally local. CMX can retain reusable context without becoming a hot-tier replacement.

Vera CPU owns the control-heavy half of the agent loop

The CPU is not a background detail. It prepares prompts, schedules GPU work, runs the operating system, searches repositories, launches tests, handles storage and networking, and joins the final evidence. A faster GPU cannot remove a serial tool or verification bottleneck on the CPU.

NVIDIA describes Vera as a data engine rather than a passive boot processor. Its official architecture pairs up to 1.5 TB of LPDDR5X with up to 1.2 TB/s of CPU-memory bandwidth. Second-generation NVLink-C2C provides a published 1.8 TB/s coherent CPU-GPU path. [OFFICIAL PRODUCT DOCUMENTATION, REVIEWED 2026-07-11] Those are vendor specifications, not measured C-001 values.

For the fixed coding task, the CPU side can own:

TEXT
Linux and containers
Hermes agent state
request parsing and authentication
repository and policy retrieval
tokenization and prompt construction
tool schemas and cache identity
scheduler, admission, block tables, and placement policy
CUDA Runtime and Driver calls
NCCL and NIXL setup
sandboxed search, editing, compilation, tests, and lint
retries, verifier, and telemetry joins

Rubin does not run git grep, a Python test suite, or the human acceptance rule merely because it is the accelerator. During a tool call, the GPU may serve another request, hold state, spill reusable state, or sit idle. That interval belongs in the economics. Faster decode does not fix a serial test suite; more HBM does not fix a broken sandbox.

The published coherent-memory model can make placement and migration easier, but four distinctions remain:

  1. A unified virtual address is not one uniform latency domain.
  2. Coherence is not proof that a tensor lived in Vera memory or moved over NVLink-C2C.
  3. CPU LPDDR5X bandwidth is not Rubin HBM4 bandwidth.
  4. Unified Memory is a mechanism, not a complete KV-cache policy. The engine still needs identity, deadlines, admission, prefetch, eviction, security, and fallback.

Rubin GPU and HBM4 own the hot repeated path

NVIDIA's current Rubin documentation publishes 288 GB of HBM4 per GPU, up to 22 TB/s of aggregate HBM4 bandwidth per GPU, and 3.6 TB/s of bidirectional NVLink 6 bandwidth per GPU. It also describes 224 streaming multiprocessors and fifth-generation Tensor Cores. [OFFICIAL PRODUCT DOCUMENTATION, REVIEWED 2026-07-11] The word up to and the physical scope matter. These figures do not say how much GLM-5.2 traffic is useful, how many HBM stacks are exposed to software, or whether one kernel reaches the device peak.

For C-001, Rubin's hot tier can contain:

Object Why it wants HBM4 Why capacity alone is insufficient
Model weights and experts Decode repeatedly reads weights; MoE needs candidate experts resident or fetchable Forty billion active parameters per token does not mean only forty billion parameters need placement
Active KV The next token reads prior attention state and appends new state Exact geometry depends on GLM-5.2 architecture, engine, dtype, block layout, and concurrency
Prefill activations Long prompts expose large attention and GEMM work Temporary peaks, tiling, fusion, and checkpointing change live bytes
MoE routing buffers Tokens are permuted, dispatched, combined, and sometimes moved across GPUs Expert imbalance and all-to-all wait can dominate even with free HBM
Kernel workspaces Libraries reserve temporary buffers for selected algorithms Reserved allocation is not logical tensor size or measured traffic
Communication buffers Collectives and transfers need staging and synchronization Aggregate link bandwidth is not application progress

The execution path remains:

TEXT
PyTorch operation
-> vLLM, SGLang, or another pinned engine
-> eager dispatch, torch.compile, or engine-specific lowering
-> cuBLASLt, CUTLASS, cuDNN, FlashInfer, Triton, or custom CUDA
-> PTX
-> target cubin and SASS
-> Rubin SM and Tensor Core execution
-> registers, shared memory or TMEM, L2, memory controller, HBM4

One Python line can fuse into a generated kernel, call an external library, trigger several kernels, or graph-break into fallback. The article's NVIDIA example registry therefore treats framework support, configured execution, dispatched artifacts, profiler evidence, and accepted output as separate states.

HBM4 begins as a DRAM bit and ends as qualified package capacity

HBM4 is still dynamic random-access memory. Each stored bit ultimately depends on charge, capacitance, voltage, sensing margin, leakage, restoration, and refresh. HBM becomes high-bandwidth by stacking dies, connecting them vertically, exposing a wide 2,048-bit-class interface, creating parallel channels, shortening package routes, and scheduling many requests, not by inventing a refresh-free cell.

The complete manufacturing path is:

TEXT
DRAM design and wafer process
-> known-good memory dies
-> TSV formation, thinning, reveal, and test
-> die bonding and stack assembly
-> base or interface logic die
-> dense package fabric, substrate, and accelerator die
-> electrical, memory, link, power, and thermal test
-> repair, burn-in, reliability, and customer qualification
-> cooled server and rack that can sustain the workload

In 2026, named vendor evidence has advanced beyond a generic roadmap. Micron reports 36 GB 12-high HBM4 in high-volume production for Vera Rubin and 48 GB 16-high samples. Samsung reports commercial HBM4 shipments using a 4 nm logic base die. Samsung and SK hynix report shipping 12-layer HBM4E samples, not HBM4E volume production. [VENDOR PRODUCT AND SAMPLE ANNOUNCEMENTS, REVIEWED 2026-07-11] Supplier announcements remain vendor evidence. A sample is not customer qualification; gross stack output is not delivered accelerator capacity.

The system supply bound is therefore:

TEXT
usable accelerator output = min(
  good accelerator dies,
  qualified HBM stacks,
  base-die and advanced-package capacity,
  substrates and assembly,
  test and repair capacity,
  firmware and system qualification,
  power- and cooling-ready racks
)

An HBM shortage reaches software because scarce hot-tier bytes become more valuable. Prefix stability, paged allocation, FP8 KV, NVFP4 weight paths, expert placement, compression, offload, batching, and reduced fragmentation can increase accepted tasks per constrained HBM GB. None is free: every saved byte must survive quality, latency, energy, failure, and engineering-cost gates.

HBM4E and custom HBM expand the logic boundary

HBM4E is not merely HBM4 at a higher clock. Public 2026 samples add higher vendor-reported pin rates and approximately 4 TB/s-class per-stack targets, while the base die, packaging, signal integrity, power density, thermals, test, yield, and customer qualification become harder. Those are sample-stage vendor targets, not a common shipping configuration.

Custom HBM changes another variable: customer-specific logic, interface behavior, or adjacent functions can move into the base-die and package design. Marvell, Samsung, Micron, and SK hynix have publicly described custom-HBM directions. Possible design surfaces include movement, telemetry, RAS, security, repair, remapping, layout conversion, compression, or carefully bounded near-memory operations. The public material does not prove that every function moves, that the full memory controller relocates, or that a startup can order it on useful terms.

Touchdown uses modular HBM more narrowly as a research term: expose composable memory-movement primitives such as place, prefetch, gather, scatter, compress, quantize, restore, repair, and remap, then choose whether each primitive belongs in software, a GPU kernel, the accelerator controller, a DPU, a package chiplet, or a custom base die.

That is a hypothesis, not a product claim. For each primitive, the receipt needs:

TEXT
object identity and security domain
preconditions and correctness rule
software API and lowering
source and destination
latency and useful/physical bytes
power, heat, and error counters
fallback and rollback
accepted-task replay
manufacturing, yield, NRE, and lock-in cost if silicon changes

The base die earns the function only when avoided movement exceeds its area, power, heat, verification, yield, non-recurring engineering, software, and supply-chain costs.

NVLink-C2C, NVLink 6, NVSwitch, and the network are different paths

Use the names at their physical scopes:

TEXT
Vera CPU <-> paired Rubin GPUs
  second-generation NVLink-C2C, coherent CPU-GPU path

Rubin GPU <-> Rubin GPU inside a named scale-up domain
  NVLink 6 through NVSwitch

compute node or rack <-> another network endpoint
  ConnectX or BlueField through InfiniBand or Spectrum-X Ethernet

HBM or host state <-> CMX context tier
  engine/KV manager + NIXL + transport + BlueField/Spectrum-X + flash

NIXL is software for data movement. NVLink is a physical scale-up link. NCCL is a collective-communication library. NVSwitch is switching silicon. BlueField is an infrastructure processor. CMX is a context-storage platform. Calling all of them NVLink erases the exact boundary an engineer must measure.

For GLM-5.2 MoE, the runtime may shard or replicate experts and use all-to-all exchange for token dispatch and combine. A larger scale-up domain offers more placement choices, but can also enlarge the participant set, synchronization surface, failure domain, and idle-capacity bill. The required artifact is not a topology diagram. It is a rank map joined to collective message sizes, p50/p95/p99, link state, HBM traffic, accepted patches, and cost.

CMX is a restore tier, not more local HBM

NVIDIA describes BlueField-4-powered CMX as a flash-backed, Ethernet-attached context tier for reusable KV at pod scale. Dynamo and KV block managers decide placement; NIXL orchestrates movement; Spectrum-X supplies the RDMA network; BlueField-4 handles the KV I/O plane; DOCA Memos supplies context communication and storage functions. NVIDIA publishes vendor targets of up to 5x sustained tokens per second and 5x power efficiency relative to traditional storage for its described use case. [OFFICIAL VENDOR ARCHITECTURE AND TARGET, REVIEWED 2026-07-11] This article does not convert those targets into C-001 results.

This supports the memory-wall argument across both accelerator generations without collapsing them. A Blackwell GPU keeps its active working set in its own HBM3E. A Rubin GPU keeps its active working set in HBM4. Vera's LPDDR5X is a separate CPU-memory domain. BlueField-4 is a DPU that handles infrastructure processing. CMX is the lower flash-backed context platform built around that DPU, Spectrum-X, NIXL, DOCA Memos, and the serving layer. None of those lower tiers turns into local GPU HBM. They reduce HBM pressure only when a reusable KV object can leave the hot tier and return before recomputation would have finished.

For the full software and memory-hierarchy explanation, read KV Cache Is Becoming the Memory Hierarchy of Inference. This HBM article uses that earlier work as the systems context, then follows the additional package, fabric, power, cooling, supply, and cost boundaries.

The safe hierarchy is:

TEXT
Rubin HBM4
  active weights, active KV, activations, workspaces

Vera LPDDR5X
  CPU work, warm state, metadata, possible offload or staging

local SSD or CMX
  reusable context that can be prefetched before its deadline

general storage
  durable repositories, indexes, logs, artifacts, and cold state

A CMX restore wins only when:

TEXT
identity lookup
+ queueing
+ network transfer
+ flash service
+ integrity check
+ HBM placement
+ miss and tail-risk cost
<
recomputed prefill under the same p99 and quality requirement

The cache key must include the model and tokenizer revision, template, tool schema, repository commit, policy version, tenant, ACL, KV format, and engine compatibility. A fast stale or cross-tenant restore is a correctness and security failure.

NVL72, Kyber NVL144, NVL576, and Feynman Kyber NVL1152 are different systems

HBM belongs to individual accelerator packages. When a model spans many accelerators, some data must cross package, rack, or multi-rack boundaries. The names below identify nested system scopes: one GPU, one rack, or several racks connected into a larger domain. A larger number is not automatically a faster task.

The latest official NVIDIA topology keeps four scopes separate:

Name Rack composition Published scale-up scope Evidence state
Vera Rubin NVL72 One MGX NVL rack with 72 Rubin GPUs and 36 Vera CPUs 72-GPU NVLink domain Official product architecture; NVIDIA says full production with shipment targeted for the second half of 2026
Vera Rubin Ultra Kyber NVL144 One future Kyber rack 144 GPUs in one rack-scale NVLink domain Official preliminary roadmap
Vera Rubin Ultra NVL576 Eight separate 72-GPU MGX NVL racks 576 GPUs in one multirack NVLink domain Official preliminary roadmap
Feynman Kyber NVL1152 Eight future 144-GPU Kyber racks 1,152 GPUs in one multirack NVLink domain Official preliminary roadmap

Kyber is the future 144-GPU MGX NVL rack architecture. It is not Rubin, Rubin Ultra, Feynman, HBM4E, NVL576, or a generic word for every NVIDIA rack. NVIDIA's older 800 VDC material used different illustrative Kyber language; the March 2026 POD topology above is the current naming authority used here.

NVIDIA also describes third-generation MGX features including modular cable cartridges, a PCB midplane, power steering, rack-level energy storage, 45 C liquid-cooling operation, and serviceable NVLink switch trays. Those are architecture and vendor claims. Final Kyber PCB stackup, board dimensions, impedance distributions, manufacturing yield, supplier allocation, qualified power envelope, cooling flow, customer price, and GLM-5.2 result are not public in the reviewed primary sources.

A larger domain matters only if the workload uses it. A low-concurrency coding service can pay for idle GPUs. A model-parallel MoE can gain from local expert placement and still lose to collective imbalance, tool-idle time, or retries. Wan2.2's current Ulysses constraints do not become a valid 144-way configuration merely because a future rack contains 144 GPUs.

Power, cooling, supply, and money close the same receipt

The rack does not stop at GPU telemetry:

TEXT
utility and switchgear
-> UPS, rectification, distribution, busbar, and voltage conversion
-> Vera, Rubin, HBM4, NVSwitch, ConnectX, BlueField, storage, pumps, and controls
-> package heat through TIM and cold plate
-> rack manifold and technology-cooling loop
-> CDU and facility-water loop
-> dry cooler, chiller, cooling tower, or heat-reuse sink

Power is a rate. Energy is power integrated over time. Heat transported by liquid is not electrical energy. Closed-loop coolant circulation is not site-water consumption. Site withdrawal, discharge, evaporation, drift, blowdown, and source-energy water are separate ledgers. PUE and WUE are useful only when their intervals and boundaries align with the task allocation.

For the same C-001 run, the CFO-facing denominator is:

TEXT
cost_per_verified_accepted_patch =
  accelerator_or_provider_time
  + CPU, sandbox, storage, and network
  + allocated facility electricity and cooling
  + allocated site-water cost where measured
  + retries and rejected attempts
  + engineering, operations, and review allocation
  ------------------------------------------------
  verified accepted patches

No value is inferred from a roadmap. Unknown remains visible. A vendor rack envelope sizes infrastructure; only synchronized meters and an allocation rule turn it into task energy. A supplier shipment statement informs diligence; only qualified delivered capacity turns it into available service.

The alternatives change the boundary, not the proof rule

The article's later sections examine each alternative in detail. This is the compact system map:

Architecture What changes Current evidence boundary C-001 question
Standard HBM4 Wider stacked-DRAM hot tier and more capable base/interface logic Named vendor products are shipping or ramping; workload-specific traffic remains unmeasured Do useful HBM bytes and accepted throughput justify the package and supply cost?
HBM4E Higher vendor targets within the next HBM generation Named vendor samples and development targets; not one common volume product Does the higher interface target survive power, heat, signal, package, yield, qualification, and useful-bandwidth measurement?
Custom HBM or cHBM Customer-specific base/interface-die logic, controller, security, telemetry, or bounded compute surfaces Vendor design direction; access, function partition, NRE, ownership, and production qualification remain customer-specific Which function belongs in the base die, and does replay pay back NRE, heat, yield, and lock-in?
SPHBM4 JEDEC standard-package HBM4-stack integration direction using a different buffer/interface boundary Public standard record; a shipping product, access terms, and workload value must still be named Does the standard package path improve access or integration without losing the state deadline, bandwidth, or qualification contract?
Memory-on-logic DRAM moves above logic through much denser vertical connections Research and modeled systems unless a named product says otherwise Does shorter movement survive heat, PDN, test, repair, yield, and serviceability?
Qualcomm HBC Near-memory compute and 3D-stacked LPDDR-based architecture Official forward-looking Dragonfly roadmap; HBC Gen 1 commercial samples expected in 2027 Do vendor effective bandwidth targets become lower cost per accepted task at matched quality and p99?
Intel XBM Patent-described backend-transistor DRAM, fine subchannels, repair, and UCIe-facing serialization Published patent application, not shipping silicon or a committed product Can the cell, stack, base-die arbitration, link, package, and repair path be built and measured?
Huawei HiBL / HiZQ Phase-specific proprietary HBM direction for prefill/recommendation versus decode/training Official Huawei roadmap and vendor specifications; supplier, process, yield, and independent results remain undisclosed Does workload phase specialization beat a complete system under the same verifier and facility boundary?
CXMT reported HBM path Potential DRAM and packaging integration path CXMT officially lists DDR and LPDDR; HBM product, qualification, bandwidth, yield, and volume were not found on its product site Which reported supply-chain elements become qualified shipping artifacts?
NVIDIA CMX Shared flash-backed context tier below hot memory Official vendor architecture and performance targets, not a local-HBM replacement Does verified restore beat recomputation without harming batch capacity, p99, security, or energy?

The memory wall is not one wall

The useful part of the argument in the Wafer post is physical: moving a bit costs energy, and a shorter, denser connection can reduce that cost. HBM already applies that idea. It moves stacked DRAM close to the accelerator and replaces a narrow off-package memory channel with a very wide package-local interface.

The argument becomes too simple when distance is treated as the whole memory wall. The complete problem has at least eight coupled limits:

TEXT
capacity
+ physical bandwidth
+ useful bandwidth after access-pattern and protocol losses
+ latency and queueing
+ data-movement energy
+ DRAM activation, refresh, retention, and controller energy
+ package, power-delivery, and cooling limits
+ yield, repair, qualification, supply, and cost
= the memory wall seen by a real system

A design can shorten one wire and still lose on capacity, bank conflicts, refresh, heat, yield, software coverage, or total cost. It can publish a very large effective bandwidth number without exposing physical pin bandwidth or sustained application throughput. It can also help memory-bound decode while helping a compute-bound phase much less.

The numerical claims need their original boundaries:

Repeated claim What the strongest reviewed source actually supports What remains unresolved
HBM movement costs 3 to 4 pJ/bit A 2017 NVIDIA Research paper estimates about 3.97 pJ/bit for a complete HBM2 access in its modeled scope This is not a universal HBM link constant and should not be transferred unchanged to HBM3E or HBM4
Dense vertical I/O costs about 0.5 pJ/bit A separate research interface reports 0.65 pJ/bit in its own measured test-chip scope It does not prove total DRAM-access energy, package energy, or a commercial memory-on-logic product
DRAM starts losing bits above 85°C Higher temperature can reduce retention margin and increase refresh or reliability pressure Micron's current HBM3E brief specifies an operating range through 105°C; 85°C is not a universal failure cliff
Qualcomm has 18x and 54x more bandwidth Qualcomm publishes 133 TB/s effective memory bandwidth per card for AI250 and describes 18x and 54x generation comparisons against its AI200 baseline These are future vendor targets for effective bandwidth; HBC Gen 1 commercial sampling is expected in mid-2027
Qualcomm proves HBM is obsolete Qualcomm positions HBC for high-capacity, memory-bound inference and decode No shipping HBC system, independent benchmark, physical-bandwidth disclosure, yield result, or matched HBM replacement receipt is public
d-Matrix already uses stacked DRAM under compute d-Matrix's current Corsair architecture is SRAM-based digital in-memory compute; its later Raptor program targets 3D DRAM, and the company says Pavehawk test silicon was validated in its labs The future commercial product, manufacturing yield, full software path, independent performance, and HBM comparison remain unproven
Samsung has already bonded HBM directly onto a GPU Samsung publicly documents HBM4, logic base dies, custom-HBM direction, and hybrid-bonding research The exact primary artifact for a shipping Samsung HBM stack bonded directly over a GPU compute die was not found in the reviewed sources

This produces a more accurate replacement map:

TEXT
standard HBM
-> wider HBM plus more capable logic base dies
-> custom HBM with bounded customer-specific functions
-> near-memory compute beside or beneath high-capacity memory
-> memory-on-logic with much denser vertical connections
-> in-memory compute that removes selected data movement

That is not a countdown to HBM's disappearance. These architectures can coexist. Standard HBM can remain the best hot tier while custom HBM removes a bounded movement, near-memory compute owns a narrow phase, and a capacity tier holds state that does not deserve expensive local bandwidth.

Any proposed HBM replacement must disclose the same comparison packet:

TEXT
same model, weights, precision, batch, sequence, and quality rule
+ physical capacity and physical pin bandwidth
+ useful and effective bandwidth with the derivation shown
+ bytes read, written, avoided, and moved across each boundary
+ p50, p99, sustained throughput, and failure rate
+ card and rack power, temperature, refresh, throttling, and cooling scope
+ package topology, repair, yield, qualification, and software coverage
+ delivered-system cost and constrained supply
= a replacement claim that can be evaluated

Without that packet, closer memory, effective bandwidth, and bandwidth per watt are research or vendor signals. They are not proof that a technology replaces HBM.

The Touchdown view is not that HBM has failed. HBM is exceptional at the role it was built for. The research question is whether explicit, composable movement and state primitives can put each byte at the cheapest physical boundary that still meets correctness and latency. Software, kernels, controllers, DPUs, packages, custom base dies, and new memory devices are all candidate implementation levels.

The proof rule never changes:

TEXT
same model and revision
same Hermes task and repository
same verifier and quality rule
same latency and reliability target
same accounting boundary
different placement, precision, primitive, or hardware
-> compare accepted tasks, useful/physical bytes, tail latency,
   energy, heat, failures, constrained capacity, and total cost

What each reader should take into the next section

  • Student or intern: Picture one state object moving through a hierarchy. Ask where it lives, who can read it, how long it remains useful, and what physical boundary a read crosses.
  • Software engineer: Stabilize prompt, repository, model, tokenizer, tenant, and tool identity before calling reuse a cache hit. Instrument the CPU tool loop as carefully as GPU decode.
  • Kernel engineer: Join source operation, compiler path, executable artifact, launch, cache/HBM/fabric counters, correctness, and accepted output. A peak number is not a bottleneck diagnosis.
  • Infrastructure engineer: Keep NVLink-C2C, NVLink/NVSwitch, PCIe/CXL, RDMA, and scale-out traffic separate. Capture topology, degraded state, queueing, and restore deadlines.
  • Memory or package engineer: Follow the bit through array, TSV, bond, base die, package, PDN, thermal path, test, repair, and qualification. Do not let a system aggregate hide the component boundary.
  • CTO: Decide which experiment is reversible before committing to a model, engine, topology, cache tier, or custom-silicon path.
  • CFO: Price occupied capacity and failed work, not theoretical accelerator throughput. Separate capital, energy, cooling, water, support, and engineering assumptions.
  • CEO or investor: Ask what receipt connects the technical advantage to accepted product output, and which supplier, qualification, software, power, and cooling dependencies can block it.

This section is the bridge. The remaining article keeps descending into the memory technologies, movement policies, facility physics, manufacturing constraints, and falsification tests that make the bridge real.

Primary sources for this system bridge

How do the AMD MI400 Series, MI455X, HBM4, Venice, UALink, and Helios form one rack-scale system?

In plain English. Venice is the CPU side of the rack. MI455X is the accelerator. HBM4 is its local working memory. ROCm is the AMD software path that turns framework operations into GPU work. UALink or UALink over Ethernet connects GPUs for scale-up, while Vulcano and Ultra Ethernet serve scale-out roles. Salina handles bounded infrastructure work. Helios is the reference rack that combines these layers. A production announcement and peak specification still do not prove one Touchdown workload ran.

Publication state: integrated additive architecture-only lane. run_id is null. The release candidate is not deployed.

The previous NVIDIA section follows one accepted coding task across Vera, Rubin, HBM4, NVLink, BlueField, CMX, power, cooling, and verification. AMD needs the same treatment. A developer does not buy an MI455X, an HBM4 stack, or a UALink cartridge. The developer asks an agent to inspect a repository, change code, run tools and tests, repair failures, and return one accepted patch.

If you searched for MI4100, stop at the name before trusting any number. No reviewed official AMD source establishes an MI4100 product. AMD uses MI400 Series for the family and MI455X for the launched CDNA 5 accelerator inside Helios. Older MI450 Series, MI450 architecture, and MI450X labels remain attached to their exact sources. Their relationship is not public.

TEXT
request and repository state
-> Venice CPU orchestration, tokenization, tools, and scheduling
-> route, cache, and placement decision
-> MI455X prefill, attention, MoE, and KV creation
-> HBM4 reads and writes
-> UALink-over-Ethernet scale-up when state or experts cross GPUs
-> Vulcano / Ultra Ethernet scale-out when work crosses rack boundaries
-> Salina DPU for bounded network, storage, security, and state-I/O work
-> CPU sandbox and verifier
-> accepted patch or another paid attempt

July 23, 2026 launch correction. AMD's official Advancing AI keynote displayed AMD HELIOS above the status IN PRODUCTION TODAY, and Lisa Su said Helios is “in full production.” AMD described the MI455X enhanced accelerator module as a production module, said shipments are on track to start at the end of Q3, and said the ramp continues into Q4. AMD's same-day MI400 launch page and Helios launch page replace the older pre-launch product-state and bandwidth fields. MI455X product identity and Helios production state are now official current. That does not mean all customer shipments were delivered today.

Evidence boundary. Helios is still an AMD reference design, not a product for sale directly by AMD: OEM and ODM partners build the branded systems. Production status, shipment timing, product form, launch specifications, ROCm qualification, and workload execution are separate claims. The same-day launch sources establish a launched product, a production reference-design system, and published peak ceilings. They do not establish a universal orderable AMD rack SKU, a Touchdown GLM-5.2 or Wan2.2 run, customer price, MI455X TBP, sustained HBM traffic, UALoE p99, rack energy per task, production yield, or an accepted Touchdown result. The older MI450-001 engineering-projection footnote remains attached to the older 19.6 TB/s brochure snapshot and its performance comparisons; it does not override the July 23 launch specification.

AMD terms used in this section

Term Plain-language meaning
OAM OCP Accelerator Module, a standardized accelerator-module form used by products such as MI355X
EAM AMD's enhanced accelerator module label for the MI455X form described in the Helios launch
XCD Accelerator Complex Die, one compute chiplet in an AMD Instinct package
CU Compute Unit, AMD's repeated GPU execution block
LDS Local Data Share, fast on-chip memory shared by work-items in a workgroup
MFMA Matrix Fused Multiply-Add, an AMDGPU instruction family used for matrix work
ROCr and HSA ROCr is ROCm's low-level runtime; HSA supplies the heterogeneous queue, memory, and execution model it exposes
COMGR and HSACO COMGR manages AMD GPU compilation artifacts; an HSACO is an HSA code object containing target-specific executable code and metadata
UALink and UALoE UALink is the scale-up interconnect contract; UALoE is AMD's UALink-over-Ethernet path for the Helios rack
Vulcano 800 AI NIC AMD Pensando's named 800 Gb/s-per-NIC path for scale-out and scale-across networking; AMD describes up to three NICs per GPU and therefore up to 2.4 Tb/s per GPU as a vendor-published ceiling
Pollara 400 AI NIC AMD Pensando's 400 Gb/s AI NIC for extending current Ethernet clusters; it is an AI NIC, not the Salina DPU
Salina DPU AMD Pensando's front-end DPU for bounded network, storage, security, encryption, integrity, telemetry, and DPU-managed NVMe work
VAST AI Operating System VAST Data's storage, data, and event-driven software platform; in the cited AMD test, a VAST flash partition is one remote home for reusable KV blocks
OEM and ODM Partners that turn AMD's reference design into branded, orderable systems with their own final bill of materials, firmware, service, and qualification

CDNA 4 is the concrete baseline

MI355X is a current CDNA 4 OAM accelerator. AMD publishes 256 compute units, 1,024 matrix cores, 16,384 stream processors, 160 KB LDS per CU, 256 MB last-level cache, 185 billion transistors, 288 GB HBM3E, 8 TB/s of peak HBM bandwidth, 1,400 W TBP, direct-liquid cooling, and 10.1 PFLOPS MXFP4. Its compiler target is gfx950.

The public physical path is:

TEXT
PyTorch / vLLM / SGLang
-> ROCm and HIP
-> AITER, Composable Kernel, Triton, hipBLASLt, or rocBLAS
-> CDNA 4 wavefronts and matrix instructions
-> registers, LDS, cache, and XCDs
-> package Infinity Fabric and I/O dies
-> HBM controller
-> eight 36 GB HBM3E stacks

That does not mean each kernel receives 8 TB/s. Useful bandwidth still depends on layout, concurrency, channel balance, bank behavior, cache reuse, fusion, communication, power, and the actual operation.

ROCm is the memory software stack, not a PyTorch checkbox

Seeing torch.version.hip proves that a PyTorch build knows about AMD's HIP platform. It does not tell you which operator library ran, which GPU executable was dispatched, which HBM counters moved, or whether the final patch or clip passed. This section follows those layers separately.

PyTorch is one framework entry point. It does not identify the allocator, queue, selected operator, generated device object, collective, physical movement, profiler interval, or management receipt. A real AMD walkthrough must keep the complete path visible:

TEXT
framework and serving:
  PyTorch / vLLM / SGLang / workload repository

operator and kernel selection:
  AITER / Composable Kernel / Triton AMD
  rocBLAS / hipBLASLt / MIOpen / sparse / FFT / solver / random / tensor libraries

compiler and executable identity:
  hipcc or clang -> LLVM AMDGPU -> COMGR -> AMDGPU code object / HSACO

dispatch and memory:
  HIP runtime -> ROCr / HSA queues -> amdgpu driver
  allocation, copies, streams, events, page state, kernels, registers, LDS, cache, HBM

collectives and movement:
  RCCL / rocSHMEM / MoRI / UCX where the pinned workload actually selects them
  package or node links -> UALink or UALoE scale-up -> Vulcano / Ultra Ethernet scale-out

evidence and operations:
  rocprofiler SDK / rocprofv3 / ROCm Systems Profiler / ROCtx
  AMD SMI / RAS / topology / firmware / clocks / power / temperature

media:
  rocDecode / rocJPEG only when ingest or output processing selects them

Each name answers a different question. torch.version.hip can establish a framework build identity. It cannot establish that AITER replaced an ATen fallback, that a particular HSACO was dispatched, that RCCL moved bytes across a named link, that HBM controllers sustained useful traffic, or that the accepted result consumed a measured amount of energy.

The matrix below is the source-bounded AMD memory-software packet used by the interactive trace. Available means the component and its documented role exist. Walkthrough role says where C-001 or V-001 could select it. Observed remains false until one pinned run joins the source, executable artifact, dispatch, counters, output, verifier, power interval, and run_id.

Component and branchLayerMemory roleWalkthrough rolePlatform scopeCoverage and evidenceExact next proof
ROCm platform and compatibility matrix
rocm_hip
platform release, driver, firmware, operating-system, and user-space compatibilityDefines the version envelope in which HBM allocations and device execution can be attributed to an AMD runtime.C-001: identity, prefill, decode, kv_state
V-001: identity, latent_prepare, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: support_must_be_queriedparser_only
official_current · not_observed
bash -lc 'rocminfo > rocminfo.txt && hipconfig --full > hipconfig.txt && amd-smi version > amd-smi-version.txt'
official source
HIP runtime and programming model
rocm_hip
device runtime, allocations, streams, events, launches, graphs, and peer accessOwns device allocation and copy APIs that may place or move model state and intermediate tensors in HBM.C-001: identity, prefill, decode, kv_state
V-001: identity, latent_prepare, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocprofv3 --hip-trace --memory-copy-trace --kernel-trace --output-directory hip-trace -- ./fixed-workload'
official source
PyTorch on ROCm
rocm_hip
framework tensor, ATen, autograd, distributed, and compilation entryCreates the framework-visible tensors and caches later lowered into AMD runtime allocations and kernels.C-001: prompt_context_ingest, prefill, decode, kv_state, tool_loop
V-001: prompt_encode, latent_prepare, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
python3 -c "import json,torch; print(json.dumps({'torch':torch.__version__,'hip':torch.version.hip,'devices':[torch.cuda.get_device_name(i) for i in range(torch.cuda.device_count())]},indent=2))"
official source
ROCr and HSA runtime
rocm_hip
queue, signal, agent, executable loading, and low-level runtimeBinds executable queues and signals to device agents and memory pools below HIP.C-001: identity, prefill, decode
V-001: identity, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocprofv3 --hsa-trace --kernel-trace --memory-copy-trace --output-directory hsa-trace -- ./fixed-workload'
official source
AITER inference operator library
math_kernel
attention, MLA, paged attention, MoE, GEMM, normalization, RoPE, quantization, sampling, and communication-aware operatorsMay read weights and KV pages, write activations and KV state, and allocate operator workspaces in HBM.C-001: prefill, decode, kv_state
V-001: denoise_attention, denoise_mlp
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$AITER_SRC" rev-parse HEAD > aiter-revision.txt && rocprofv3 --kernel-trace --memory-copy-trace --output-directory aiter-trace -- "$WORKLOAD_RUNNER"'
official source
Composable Kernel and CK Tile
math_kernel
templated HIP kernels, tile programming, examples, instances, and profiler clientsDefines tiled reads, writes, LDS staging, and HBM-facing kernel layouts for covered operators.C-001: prefill, decode
V-001: denoise_attention, denoise_mlp, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$ROCM_LIBRARIES_SRC" rev-parse HEAD > rocm-libraries-revision.txt && "$CK_PROFILER" gemm "$CK_GEMM_ARGS" | tee ck-profile.txt'
official source
hipBLASLt
math_kernel
descriptor-driven GEMM, heuristics, epilogues, tuning, and solution selectionReads matrix operands and scale metadata from HBM, uses workspace, and writes GEMM outputs to HBM.C-001: prefill, decode
V-001: prompt_encode, denoise_mlp, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$HIPBLASLT_BENCH" $HIPBLASLT_ARGS 2>&1 | tee hipblaslt-bench.txt'
official source
rocBLAS
math_kernel
BLAS operations and reference GEMM baselineReads dense matrix and vector operands from HBM and writes BLAS results, with optional workspace.C-001: prefill, decode
V-001: prompt_encode, denoise_mlp, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocblas-bench -f gemm -r f32 --transposeA N --transposeB N -m 4096 -n 4096 -k 4096 --alpha 1 --beta 0 | tee rocblas-bench.txt'
official source
MIOpen
math_kernel
deep-learning primitives, convolutions, normalization, fusion, and solution selectionMay read feature maps and weights from HBM, allocate workspace, and write convolution or normalization outputs.C-001:
V-001: denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'MIOpenDriver conv -n 1 -c 64 -H 224 -W 224 -k 64 -y 3 -x 3 -p 1 -q 1 -F 1 -V 1 | tee miopen-driver.txt'
official source
Triton AMD backend
math_kernel
tile-language compilation, scheduling, and AMD target loweringGenerated kernels define global-memory tile loads and stores that may map to HBM traffic.C-001: prefill, decode
V-001: denoise_attention, denoise_mlp, vae_decode
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
public_code · not_observed
bash -lc 'git -C "$TRITON_SRC" rev-parse HEAD > triton-revision.txt && TRITON_ALWAYS_COMPILE=1 python3 "$PINNED_TRITON_FIXTURE" 2>&1 | tee triton-run.txt'
official source
rocWMMA
math_kernel
wave-level matrix-fragment API and matrix-core programmingLoads operand fragments from HBM through cache and LDS and writes matrix results back to device memory.C-001: prefill, decode
V-001: denoise_mlp
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$ROCM_LIBRARIES_SRC" rev-parse HEAD > rocm-libraries-revision.txt && "$ROCWMMA_SAMPLE" | tee rocwmma-sample.txt'
official source
hipcc and amdclang++
compiler_artifact
offline HIP compilation driver and compiler frontendLowers source memory operations and target flags into device code that later issues cache and HBM transactions.C-001: compile, prefill, decode
V-001: compile, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'hipcc --version > hipcc-version.txt && hipcc --offload-arch=gfx950 -O3 -save-temps "$HIP_SOURCE" -o fixed-kernel 2> compile.log'
official source
HIPRTC
compiler_artifact
runtime HIP source compilationProduces runtime device code whose kernels may later read and write HBM-resident objects.C-001: compile, prefill, decode
V-001: compile, denoise
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$HIPRTC_FIXTURE" --arch gfx950 --dump-code-object hiprtc-output.hsaco 2>&1 | tee hiprtc.log'
official source
LLVM AMDGPU backend
compiler_artifact
LLVM target lowering, code generation, metadata, and AMDGPU ISA emissionLowers address spaces, loads, stores, atomics, cache policy, and synchronization into target-specific device instructions.C-001: compile, prefill, decode
V-001: compile, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
public_code · not_observed
bash -lc 'llvm-objdump --mcpu=gfx950 --disassemble --source "$DEVICE_ARTIFACT" > amdgpu-disassembly.txt'
official source
AMD COMGR, ROCm Device Libraries, and HSACO code objects
compiler_artifact
device compilation support, linking, device libraries, metadata, and executable code-object packagingPackages the device instructions and metadata for kernels that later operate on HBM-resident objects.C-001: compile, prefill, decode
V-001: compile, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'readelf -h -n -s "$DEVICE_ARTIFACT" > hsaco-readelf.txt && sha256sum "$DEVICE_ARTIFACT" > hsaco-sha256.txt'
official source
RCCL
collective_fabric
collective communication library for all-reduce, all-gather, reduce-scatter, all-to-all, broadcast, and point-to-point operationsReads and writes collective buffers that may live in HBM and cross accelerator or node boundaries.C-001: prefill, decode, expert_parallel, tensor_parallel
V-001: denoise_attention, sequence_parallel, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$RCCL_TESTS_SRC" rev-parse HEAD > rccl-tests-revision.txt && "$RCCL_TESTS_SRC/build/all_reduce_perf" -b 8 -e 8G -f 2 -g 8 | tee rccl-all-reduce.txt'
official source
rocSHMEM
collective_fabric
partitioned global address-space communication and device-initiated memory operationsMay expose symmetric device buffers and remote operations involving HBM-resident data.C-001: expert_parallel, tensor_parallel
V-001: sequence_parallel
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$ROCSHMEM_SRC" rev-parse HEAD > rocshmem-revision.txt && "$ROCSHMEM_FIXTURE" 2>&1 | tee rocshmem-run.txt'
official source
MoRI, MoRI-IO, and MoRI expert-parallel communication
collective_fabric
modular RDMA interface, prefill/decode state transfer, and expert dispatch/combineMay move KV blocks and expert-routing buffers between GPU HBM and RDMA-visible communication paths.C-001: prefill, decode, kv_state_transfer, expert_parallel
V-001: sequence_parallel
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$MORI_SRC" rev-parse HEAD > mori-revision.txt && "$PINNED_MORI_RUNNER" 2>&1 | tee mori-run.txt'
official source
Infinity Fabric, UALink or UALoE, and Ultra Ethernet boundary
collective_fabric
physical scale-up and scale-out topology beneath communication softwareCarries communication derived from HBM-resident tensors across accelerator, tray, rack, or cluster boundaries.C-001: kv_state_transfer, expert_parallel, tensor_parallel
V-001: sequence_parallel, vae_decode
mi355x: source_backed_candidate, mi455x_helios: source_backed_candidateparser_only
official_preliminary · not_observed
bash -lc 'amd-smi topology --show-weight > topology-weight.txt && amd-smi topology --show-hops > topology-hops.txt && amd-smi topology --show-link-type > topology-links.txt'
official source
ROCprofiler-SDK and rocprofv3
profiler_telemetry
HIP, HSA, kernel, memory-copy, allocation, marker, counter, and communication tracingCan observe allocation, copy, dispatch, and counter surfaces needed to attribute candidate HBM activity.C-001: identity, prefill, decode, kv_state, tool_loop
V-001: identity, latent_prepare, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocprofv3-avail > rocprofv3-available.txt && rocprofv3 --runtime-trace --kernel-trace --memory-copy-trace --marker-trace --output-format csv --output-directory rocprofv3-out -- "$WORKLOAD_RUNNER"'
official source
ROCm Compute Profiler
profiler_telemetry
kernel counters, roofline, memory analysis, occupancy, and baseline comparisonCan expose counters related to cache, HBM bandwidth, LDS, occupancy, and instruction mix for a selected kernel.C-001: prefill, decode
V-001: denoise_attention, denoise_mlp, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocprof-compute profile -n "$RUN_NAME" --format-rocprof-output csv -- "$WORKLOAD_RUNNER" && rocprof-compute analyze -p "./workloads/$RUN_NAME/$GPU_TARGET"'
official source
ROCm Systems Profiler
profiler_telemetry
CPU, thread, runtime, system, accelerator, and communication timelineCan place HBM-facing kernels and copies in the same timeline as CPU scheduling, tools, storage, and network work.C-001: prompt_context_ingest, prefill, decode, tool_loop, accepted_task
V-001: prompt_encode, latent_prepare, denoise, vae_decode, accepted_clip
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocprof-sys-sample --output-path rocprof-sys-out -- "$WORKLOAD_RUNNER"'
official source
AMD SMI telemetry
profiler_telemetry
device inventory, clocks, power, temperature, memory use, topology, links, and health telemetryCan report device-level memory allocation and health signals, but not which workload object produced the bytes.C-001: identity, prefill, decode, tool_loop
V-001: identity, latent_prepare, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'amd-smi version > amd-smi-version.txt && amd-smi metric --csv > amd-smi-metric.csv'
official source
AMD SMI management and RAS controls
management_ras
inventory, configuration, reset, process, firmware, error, topology, and reliability administrationExposes device and memory health boundaries needed before treating an HBM result as valid.C-001: preflight, identity, postflight
V-001: preflight, identity, postflight
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'amd-smi static --asic --board --vbios --driver > amd-smi-static.txt && amd-smi metric --ecc --pcie --xgmi > amd-smi-ras.txt'
official source
ROCm Validation Suite
management_ras
system qualification, stress, memory, PCIe, peer, and GPU validationCan exercise memory and platform health before workload-specific HBM claims are accepted.C-001: preflight, postflight
V-001: preflight, postflight
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rvs -g > rvs-gpu-list.txt && rvs -c "$RVS_CONFIG" -l rvs-run.log'
official source
AMD Instinct MI355X system acceptance guide
management_ras
platform bring-up and acceptance tests for inventory, memory, fabric, power, cooling, and sustained operationDefines the acceptance boundary that must pass before a workload-specific HBM receipt is trusted.C-001: preflight, postflight
V-001: preflight, postflight
mi355x: source_backed_candidate, mi455x_helios: not_applicableparser_only
official_current · not_observed
bash -lc 'sudo lspci -d 1002:75a3 > mi355x-pcie.txt && test "$(wc -l < mi355x-pcie.txt)" -eq 8'
official source
rocDecode
media
video decode API and hardware-assisted media ingestMay create decoded frame surfaces and staging buffers before preprocessing or model input.C-001:
V-001: input_decode, preprocess
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$ROCDECODE_SRC" rev-parse HEAD > rocdecode-revision.txt && "$ROCDECODE_SAMPLE" "$PINNED_VIDEO" 2>&1 | tee rocdecode-run.txt'
official source
rocJPEG
media
JPEG decode and image ingestMay create decoded image surfaces and staging buffers before multimodal or video preprocessing.C-001: multimodal_input
V-001: input_decode, preprocess
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'git -C "$ROCJPEG_SRC" rev-parse HEAD > rocjpeg-revision.txt && "$ROCJPEG_SAMPLE" "$PINNED_IMAGE" 2>&1 | tee rocjpeg-run.txt'
official source
rocAL
media
accelerated data loading, decode, augmentation, and preprocessing pipelineMay allocate input batches, decoded surfaces, augmentation intermediates, and model-ready tensors.C-001: multimodal_input
V-001: input_decode, preprocess
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCAL_FIXTURE" --input "$PINNED_MEDIA_MANIFEST" --output rocal-output 2>&1 | tee rocal-run.txt'
official source
ROCm Performance Primitives
media
image and tensor preprocessing primitivesMay read decoded image or tensor batches, apply preprocessing, and write model-ready outputs.C-001: multimodal_input
V-001: preprocess, output_postprocess
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$RPP_FIXTURE" --manifest "$PINNED_IMAGE_BATCH" 2>&1 | tee rpp-run.txt'
official source
ROCm 7.14 TheRock gfx1250 source-bring-up registry
compiler_artifact
ROCm build and distribution target registry at tag therock-7.14A source target can make future HBM4-capable device code buildable, but it does not allocate HBM or establish a supported runtime.C-001: source_bringup, compile_candidate
V-001: source_bringup, compile_candidate
mi355x: source_backed_candidate, mi455x_helios: source_bringup_not_qualifiedparser_only
public_code · not_observed
bash -lc 'git -C "$THEROCK_SRC" checkout therock-7.14 && git -C "$THEROCK_SRC" rev-parse HEAD > therock-revision.txt && rg -n "gfx1250.*MI450/MI450X/MI455X" "$THEROCK_SRC/cmake/therock_amdgpu_targets.cmake" > therock-gfx1250.txt'
official source
amdgpu, KFD, and GPU firmware boundary
rocm_hip
kernel driver, compute-device interface, firmware loading, memory mapping, queues, and reset boundaryOwns the kernel-visible GPU memory, process, queue, page-mapping, fault, reset, and firmware boundary beneath ROCr and HIP.C-001: preflight, allocation, dispatch, fault_recovery
V-001: preflight, allocation, dispatch, fault_recovery
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'uname -a > kernel.txt && modinfo amdgpu > amdgpu-modinfo.txt && dmesg --level=err,warn | rg -i "amdgpu|kfd|firmware" > amdgpu-kfd-dmesg.txt || true'
official source
rocminfo HSA agent and memory-pool inventory
profiler_telemetry
HSA system, agent, cache, ISA, queue, and memory-pool enumerationReports the runtime-visible agents and memory pools required to distinguish a detected accelerator from an assumed one.C-001: identity, preflight
V-001: identity, preflight
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rocminfo > rocminfo.txt && sha256sum rocminfo.txt > rocminfo.sha256'
official source
hipBLAS portability wrapper
math_kernel
BLAS portability interface and backend dispatch wrapperAccepts matrix and vector operands, strides, layouts, and workspaces that may reside in HBM before dispatching to a backend.C-001: prefill, decode, expert_compute
V-001: prompt_encode, denoise, vae_decode
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$HIPBLAS_BENCH" $HIPBLAS_ARGS 2>&1 | tee hipblas-bench.txt'
official source
hipTensor
math_kernel
tensor contraction, permutation, reduction, and plan selectionReads multidimensional tensors and metadata, may allocate workspace, and writes contracted or transformed tensors.C-001: candidate_tensor_contraction
V-001: candidate_tensor_contraction, denoise, vae_decode
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$HIPTENSOR_FIXTURE" --manifest "$HIPTENSOR_CASE" 2>&1 | tee hiptensor-run.txt'
official source
hipSPARSE and rocSPARSE
math_kernel
sparse matrix and vector portability interface plus AMD backendReads sparse values, indices, pointers, dense operands, and workspaces and writes sparse or dense outputs.C-001: candidate_sparse_operator, expert_compute
V-001: candidate_sparse_operator
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCSPARSE_FIXTURE" --manifest "$SPARSE_CASE" 2>&1 | tee rocsparse-run.txt'
official source
hipSPARSELt
math_kernel
structured-sparse matrix multiplication, pruning, compression, plan, and algorithm selectionMay store structured-sparse weights, compressed metadata, dense activations, workspace, and outputs in HBM.C-001: candidate_structured_sparse_gemm, expert_compute
V-001: candidate_structured_sparse_gemm
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$HIPSPARSELT_BENCH" $HIPSPARSELT_ARGS 2>&1 | tee hipsparselt-run.txt'
official source
hipFFT and rocFFT
math_kernel
FFT portability interface, plan generation, kernels, work buffers, and transformsReads signal tensors, twiddle or plan data, and workspace and writes frequency-domain or inverse-transform outputs.C-001: not_selected_by_current_source
V-001: candidate_frequency_transform
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCFFT_RIDER" $ROCFFT_ARGS 2>&1 | tee rocfft-run.txt'
official source
hipRAND and rocRAND
math_kernel
random-number portability API, generators, distributions, states, and output buffersMay create random states and output tensors for sampling or latent initialization in device memory.C-001: sampling_candidate
V-001: latent_prepare_candidate
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCRAND_FIXTURE" --manifest "$RNG_CASE" 2>&1 | tee rocrand-run.txt'
official source
hipSOLVER and rocSOLVER
math_kernel
dense and sparse linear-system and factorization portability API plus AMD backendReads matrices and solver metadata, uses workspaces and pivot buffers, and writes factors or solutions.C-001: not_selected_by_current_source
V-001: not_selected_by_current_source
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCSOLVER_FIXTURE" --manifest "$SOLVER_CASE" 2>&1 | tee rocsolver-run.txt'
official source
rocALUTION
math_kernel
iterative sparse solvers, preconditioners, matrix formats, and host-device movementMay place sparse matrices, vectors, preconditioners, staging buffers, and solver state across host memory and HBM.C-001: not_selected_by_current_source
V-001: not_selected_by_current_source
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCALUTION_FIXTURE" --manifest "$ROCALUTION_CASE" 2>&1 | tee rocalution-run.txt'
official source
hipCUB and rocPRIM
math_kernel
device and block primitives for scan, reduce, sort, select, partition, and memory operationsReads and writes device ranges and temporary storage used by routing, sampling, indexing, sorting, and reduction paths.C-001: routing_candidate, sampling_candidate, indexing_candidate
V-001: reduction_candidate, indexing_candidate
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCPRIM_FIXTURE" --manifest "$PRIMITIVE_CASE" 2>&1 | tee rocprim-run.txt'
official source
rocThrust
math_kernel
parallel algorithms, iterators, containers, scans, reductions, sorts, and transformsMay own device containers and temporary ranges for higher-level indexing, sorting, selection, transform, and reduction work.C-001: routing_candidate, sampling_candidate, indexing_candidate
V-001: indexing_candidate, postprocess_candidate
mi355x: source_backed_candidate, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$ROCTHRUST_FIXTURE" --manifest "$ROCTHRUST_CASE" 2>&1 | tee rocthrust-run.txt'
official source
UCX transport boundary
collective_fabric
communication framework and transport selection above network and memory registrationCan register and move buffers between local or remote memory domains, but it is not an HBM allocator, cache policy, or physical fabric.C-001: kv_state_transfer_candidate, expert_parallel_candidate
V-001: distributed_tensor_transfer_candidate
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
public_code · not_observed
bash -lc 'ucx_info -v > ucx-version.txt && ucx_info -d > ucx-devices.txt && "$UCX_FIXTURE" 2>&1 | tee ucx-run.txt'
official source
hipFile and AMD Infinity Storage
storage_kv
early-access direct-to-GPU storage I/O with synchronous, asynchronous, batch, and POSIX fallback pathsCan move file-backed state between storage and device memory while preserving a distinct fallback path through host I/O.C-001: kv_offload_candidate, model_state_load_candidate
V-001: model_state_load_candidate, latent_or_output_io_candidate
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_preliminary · not_observed
bash -lc 'ais-check --output ais-check.json && "$HIPFILE_FIXTURE" --manifest "$HIPFILE_CASE" 2>&1 | tee hipfile-run.txt'
official source
ROCm AMD Infinity Context
storage_kv
early-access disaggregated KV-cache inference stack across GPU memory, CPU DRAM, local NVMe, and NFS over RDMADefines a tiered KV-state architecture and test harness; it does not prove that C-001 selected it or that any tier moved bytes.C-001: kv_lookup_candidate, kv_offload_candidate, prefill_decode_disaggregation_candidate
V-001: not_applicable_to_diffusion_kv
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_preliminary · not_observed
bash -lc 'git -C "$AIC_SRC" rev-parse HEAD > aic-revision.txt && docker compose -f "$AIC_COMPOSE" config > aic-compose.txt && "$AIC_BENCH" --manifest "$C001_MANIFEST" 2>&1 | tee aic-run.txt'
official source
vLLM, LMCache, and NIXL ROCm integration
storage_kv
version-pinned serving, KV block management, and transfer integration inside the AIC technology previewSeparates the serving engine, cache identity and policy, and movement backend across GPU, CPU, local storage, and network-storage tiers.C-001: prefill, decode, kv_lookup_candidate, kv_transfer_candidate
V-001: not_applicable_to_diffusion_kv
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_preliminary · not_observed
bash -lc 'docker compose -f "$AIC_COMPOSE" config > integration-compose.txt && "$AIC_BENCH" --manifest "$C001_MANIFEST" --capture-kv-trace 2>&1 | tee integration-run.txt'
official source
ROCm Data Center Tool
management_ras
data-center GPU discovery, groups, monitoring, diagnostics, policy, health, and telemetryCan collect device health and memory-related telemetry around a workload interval without proving object-level HBM traffic.C-001: preflight, runtime_monitoring, postflight
V-001: preflight, runtime_monitoring, postflight
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc 'rdci discovery -l > rdc-discovery.txt && rdci dmon -l > rdc-fields.txt && "$RDC_CAPTURE" --manifest "$WORKLOAD_INTERVAL" > rdc-capture.json'
official source
TransferBench and ROCm Validation Suite boundary
management_ras
simultaneous transfer benchmark versus broader platform validation and diagnosticsTransferBench measures configured copy paths; RVS checks system health. Neither proves a workload selected the path or that logical bytes equal HBM or wire bytes.C-001: preflight_transfer_baseline, postflight
V-001: preflight_transfer_baseline, postflight
mi355x: support_must_be_queried, mi455x_helios: not_publicparser_only
official_current · not_observed
bash -lc '"$TRANSFERBENCH" "$TRANSFER_CONFIG" 2>&1 | tee transferbench-run.txt && rvs -c "$RVS_CONFIG" -l rvs-run.log'
official source
Legacy ROCm-SMI, profiler, and RBT negative guard
management_ras
fail-closed routing away from legacy ROCm-SMI, old profiler entry points, and the ROCm Bandwidth TestPrevents stale tool output from being treated as current AMD SMI, rocprofv3, ROCm profiler, or TransferBench evidence.C-001: preflight_negative_guard
V-001: preflight_negative_guard
mi355x: not_applicable, mi455x_helios: not_applicableparser_only
official_current · not_observed
bash -lc 'for tool in rocm-smi rocprof rocprofv2 rocm-bandwidth-test amd-smi rocprofv3 TransferBench; do command -v "$tool" || true; done > tool-routing.txt'
official source

For C-001 Hermes plus GLM-5.2, the expected inspection path is:

TEXT
model and engine identity
-> HIP allocation and stream identity
-> selected attention, GEMM, MoE, normalization, or sampling operator
-> AMDGPU target and HSACO digest
-> dispatch geometry and kernel name
-> HBM allocation, read/write counters, cache behavior, and stalls
-> RCCL / rocSHMEM / MoRI only if the placement plan communicates
-> CPU tool window and verifier
-> accepted patch, latency, energy, and cost under one run_id

For V-001 Wan2.2, the expected inspection path is different:

TEXT
prompt, model revision, latent contract, and dtype
-> HIP allocations for weights, latent state, transient Q/K/V, activations, and workspaces
-> selected attention, convolution, normalization, scheduler, and VAE operations
-> AMDGPU target, HSACO digests, and repeated denoising dispatches
-> HBM traffic and allocation high-water mark across the denoising timeline
-> RCCL collective and rank map only if the declared parallel plan communicates
-> rocDecode or rocJPEG only if the declared media path actually uses it
-> accepted clip, quality rule, latency, energy, and cost under a different run_id

The catalog deliberately includes libraries that may be available but absent from both traces. Package presence is not execution. Import success is not dispatch. A microbenchmark is not an accepted coding task or video. A profiler capture without synchronized workload and verifier identity is not a business receipt.

The current MI455X software evidence has one important split:

TEXT
public source bring-up:
  TheRock maps gfx1250 to MI450 / MI450X / MI455X
  several ROCm 7.14 components contain gfx1250 source or changelog entries
  RCCL contains an MI455X performance-test configuration

published qualification:
  the ROCm 7.14 hardware-support table still stops at MI350 / MI355X gfx950
  TheRock's supported-GPU table does not list gfx1250
  AITER's published hardware table still stops at MI355X
  MoRI marks the MI450X family path as work in progress
  ROCm Infinity Context labels Helios / MI455X as upcoming

local observation:
  no selected MI455X operator
  no captured gfx1250 HSACO or disassembly
  no dispatch, HBM counter, fabric trace, power interval, output, or run_id

AMD's Helios page also uses Day-0 support and production-ready AI performance language for the software ecosystem. That is a vendor platform direction, not a component-level release receipt. The trace therefore labels MI455X software source bring-up, not qualified, not selected, and not observed until the release table, exact package set, hardware, and workload artifacts agree.

MI455X changes the memory and system boundary

The headline change is straightforward: one MI455X exposes more local HBM capacity and a much higher published peak HBM bandwidth than the MI355X baseline. That creates more room for weights, active state, batching, and workspaces. It does not predict application speed because kernels, collectives, software support, power, and the workload can still become the bottleneck.

AMD currently publishes these MI455X and Helios launch specifications. Product, production, and launch-specification states are current. Peak compute, memory-bandwidth, scale-up, and scale-out fields are still vendor-published ceilings, not measured workload results:

Boundary Published value Evidence state
Product identity MI455X / CDNA 5 official current launch
MI455X module state production enhanced accelerator module official current launch statement
Helios production state in full production on July 23, 2026 official current launch statement
Shipment state on track to start at the end of Q3 and ramp into Q4 2026 official forward-looking schedule
HBM4 per GPU 432 GB official current launch specification
Peak HBM bandwidth per GPU 23.3 TB/s official current launch specification; peak ceiling
Transistors per GPU 320 billion official current launch specification
Peak OCP MXFP4 per GPU 40 PFLOPS official current launch specification; theoretical peak
Peak OCP MXFP8 per GPU 20 PFLOPS official current launch specification; theoretical peak
GPUs per Helios rack 72 across 18 four-GPU compute trays official current launch specification
Rack HBM4 31 TB official current launch specification
Aggregate peak rack HBM bandwidth 1.7 PB/s official current launch specification; peak ceiling
Peak rack OCP MXFP4 2.9 EFLOPS official current launch specification; theoretical peak
Peak rack OCP MXFP8 1.4 EFLOPS official current launch specification; theoretical peak
Aggregate scale-up bandwidth 260 TB/s official current launch specification; peak ceiling
Aggregate scale-out bandwidth 43 TB/s official current launch specification; peak ceiling

The June brochure and the product page captured before the launch used 19.6 TB/s per MI455X. Multiplying that older field by 72 produced the old 1.4112 PB/s Touchdown arithmetic. Both values remain useful only as a dated pre-launch snapshot. The July 23 launch values that govern this article are 23.3 TB/s per MI455X and 1.7 PB/s aggregate peak per Helios rack.

The direct CDNA 4 to CDNA 5 arithmetic is:

TEXT
HBM capacity:
  432 / 288 = 1.50x

HBM bandwidth:
  23.3 / 8.0 = 2.91x

FP4 peak:
  40 / 10.1 = 3.96x

primary GPU domain:
  72 / 8 = 9x as many GPUs

rack HBM peak reconciliation:
  432 GB * 72 = 31,104 GB = 31.104 decimal TB
  AMD publishes the rounded rack capacity headline as 31 TB

  72 * 23.3 TB/s = 1,677.6 TB/s = 1.6776 PB/s
  AMD publishes the rounded rack headline as 1.7 PB/s

The capacity is a vendor-published product total, and the bandwidth is a vendor-published peak ceiling. The arithmetic does not reveal HBM stack count, stack height, supplier, reserved capacity, sustained traffic, or application speedup. CDNA 5 compute grows faster than memory capacity and bandwidth. A weak kernel or exposed collective can strand more peak arithmetic than before.

AMD also prints a 10x performance increase versus MI355X headline. The brochure footnote bounds that claim to engineering projections: a future MI400 Series design versus MI355X using an MoE inference model with 2K and 16K prefill under TP8 and EP8, plus projected training improvements for GEMM and attention. A separate footnote compares the theoretical dense MXFP4 peak of a 72-MI455X Helios rack with an eight-MI355X platform. Neither is a universal 10x application result.

What is inside the Helios reference rack?

Think of the rack as six connected responsibilities: CPU control, GPU compute, local HBM, GPU-to-GPU scale-up, rack-to-rack scale-out, and power/cooling/service. The partner that delivers the final system must turn AMD's reference design into one qualified configuration with exact parts, firmware, topology, support, and facility requirements.

AMD describes an Open Rack Wide system built from MI455X GPUs, 6th Gen EPYC Venice CPUs, Vulcano 800G AI NICs, Salina DPUs, four UALink-over-Ethernet scale-up cartridges, central power shelves, a vertical busbar, and direct-liquid cooling.

TEXT
host and control:
  6th Gen EPYC Venice
  up to 256 CPU cores and 1.6 TB/s CPU-memory bandwidth

accelerator hot tier:
  72 MI455X GPUs across 18 four-GPU compute trays
  31 TB HBM4 and 1.7 PB/s aggregate peak HBM bandwidth per rack

scale-up:
  UALink or UALink over Ethernet
  four scale-up cartridges in the reference architecture

scale-out and offload:
  Pensando Vulcano 800 Gb/s AI NIC path
  Pensando Salina DPU role named in the Helios brochure

mechanical, power, and cooling:
  double-wide Open Rack Wide
  central power shelves and vertical busbar
  liquid manifold and quick-disconnect service boundaries

The product page makes the scope explicit: this is a partner blueprint. OEMs and ODMs can turn it into branded systems, so the final tray population, NIC and DPU count, firmware bundle, rack power, cooling inputs, service contract, price, and qualification result belong to the delivered partner configuration.

The published network, offload, security, and serviceability claims also need separate scopes:

Boundary Published statement What remains unproven
Scale-up cartridges four cartridges; UALink over Ethernet; up to 72 GPUs; 260 TB/s aggregate message latency, switch path, collective mapping, retries, congestion, useful application bandwidth
Vulcano AI NIC 800 Gb/s Ethernet per named NIC; brochure also says up to 2.4 Tb/s of scale-out bandwidth per GPU exact NIC count and topology, how the per-NIC, per-GPU, and 43 TB/s rack figures reconcile, delivered throughput
Helios scale-out 43 TB/s aggregate peak endpoint count, traffic class, oversubscription, p95/p99, failure recovery
Salina DPU network, security, storage, encryption, integrity, and telemetry offload final rack population, workload-selected function, bytes, latency, power, accepted-work effect
Salina projected limits 400 Gb/s bandwidth, 10M connections/s, 100M packets/s, 400 Gb/s encryption offload, 4M storage IOPS in the brochure footnote final silicon, Helios integration, customer configuration, matched-workload result
Security hardware root of trust, continuous attestation, isolation, device identity, encrypted memory and interconnect claims exact trust chain, key ownership, cipher boundary, attestation evidence, overhead, tenant policy
Serviceability integrated power, cooling, and networking connections intended to avoid recabling during sled replacement measured replacement time, failure isolation, spare policy, leak handling, uptime

The important point is not tray count alone. Helios exposes several physically different paths:

Path Intended role What it is not
MI455X HBM4 active weights, active KV, activations, workspaces pooled cold storage
Venice DDR CPU work, metadata, warm state, staging local HBM4
Infinity Fabric CPU-GPU or package/node connectivity where supported rack-wide persistent storage
UALink over Ethernet GPU scale-up and collective traffic NVMe or object storage
Vulcano / Ultra Ethernet AI scale-out local GPU memory
Salina DPU network, storage, security, encryption, integrity, telemetry, bounded transfer execution dense tensor compute
FIGURE 12D 1600 × 900
A left-to-right architecture map follows one request through Venice CPU control, the ROCm software path, MI455X compute and HBM4, UALink or UALoE scale-up, Vulcano and Ultra Ethernet scale-out, the Salina DPU role, vertical busbar power, liquid cooling, and an unfilled accepted-task receipt.
Helios is a complete CPU, GPU, HBM4, fabric, networking, power, cooling, and verification path. Product identity, production state, and July 23 launch specifications are official current; peak values remain vendor-published ceilings. No Touchdown GLM-5.2 or Wan2.2 Helios run has been captured. Source state. AMD's Helios and MI400 launch pages plus the official Advancing AI keynote were reviewed July 23, 2026; Helios full production announced, shipments scheduled, ROCm component qualification separate, Touchdown run not captured. The older 19.6 TB/s and 1.4112 PB/s fields remain historical.

The HBM4 package is a pressure point, not a disclosed stack map

AMD publishes up to 432 GB HBM4 for MI455X. That total does not establish the exact stack count, stack height, package organization, supplier mix, qualification, yield, or customer allocation. Those fields remain not public until an AMD package disclosure, qualified product document, or physical analysis proves them.

HBM4 doubles the external stack interface from the 1,024-bit HBM3-class boundary to 2,048 bits. More wires buy bandwidth, but increase the pressure on:

TEXT
base-die logic
microbump or bonding density
package routing
power delivery
signal integrity
warpage
thermal extraction
known-good-die test
stack and package yield

The base die becomes more strategic because advanced logic can support PHY, repair, telemetry, RAS, security, power management, address mapping, remapping, and possible customer-specific movement functions. It also adds area, heat, verification, process, yield, and lock-in cost.

AMD and Samsung announced an HBM4 supply direction for MI455X. That is evidence of collaboration and intent. It is not evidence of final volume allocation, price, yield, or customer qualification.

UALink bandwidth is not MoE latency

Peak bandwidth describes how much data a link can move under a declared scope. MoE latency asks how quickly small, uneven expert-routing messages finish, especially at the slow tail. Those are different measurements.

Helios publishes 260 TB/s of rack scale-up bandwidth. That does not answer the most important inference questions:

TEXT
64-byte to 4-KB message latency
all-to-all p99
expert-dispatch jitter
barrier cost
incast behavior
congestion control
retransmission
collective setup
failure recovery

MoE inference can send small, bursty, skewed messages. Decode is latency sensitive. A fabric can deliver excellent large-transfer bandwidth and still expose bubbles on expert routing or token-by-token work.

The required receipt joins:

TEXT
rank and expert map
message-size histogram
RCCL operation
p50 / p95 / p99
link and retry state
GPU idle time
HBM traffic
accepted tasks
power and cost

CXL is a conditional warm tier, not Helios HBM

CXL can add CPU-attached memory capacity, but the GPU does not experience it as local HBM. A restore must cross a real topology and still beat recomputation or another placement choice before the next operation needs the data.

No reviewed AMD product artifact establishes a production Helios CXL KV-cache path. CXL can still be tested as CPU-attached capacity behind Venice or a separate memory server.

Good candidates:

TEXT
stable repeated prefixes
paused-session KV
recommendation embeddings
retrieval indexes
model and checkpoint staging
predictably prefetched cold experts

Bad default candidates:

TEXT
active weights used every token
active per-token KV
hot activations
GPU workspaces
MoE all-to-all

The decision is not HBM or CXL in the abstract:

TEXT
CXL fetch and restore
vs
recompute prefix
vs
wait for state owner
vs
route request to state owner
vs
use host DDR or storage

Use CXL only if the complete path improves accepted tasks per rack-hour under the same p99 and quality rule.

Penguin Solutions' MemoryAI is a useful comparison. It is a separate 4U memory server built around dual EPYC CPUs, 3 TB DDR5, and up to eight 1 TB CXL cards for up to 11 TB of system memory. It targets KV and context capacity. It does not make CXL-attached DDR behave like local MI455X or Rubin HBM.

Astera Labs sells CXL controllers, retimers, switches, cable modules, and management software to hyperscalers, cloud platforms, OEMs, and memory-module vendors. Penguin sells complete systems, integration, deployment, and operations. These are different points in the stack.

MI400 is a family, not one universal SKU

Keep source labels exact:

Name Public role
MI430X high-precision HPC, AI for science, and sovereign systems; 432 GB HBM4 and 19.6 TB/s disclosed
MI440X enterprise/on-prem eight-GPU role; final memory, bandwidth, power, and package not public
MI450 Series label used in major deployment agreements
MI455X named flagship Helios accelerator
Meta custom MI450-based GPU customer-specific accelerator; exact memory and package not publicly specified

Do not silently convert every MI450-family customer into a stock MI455X customer.

Customer claims need exact labels

Customer Confirmed public object Unsafe rewrite
OpenAI MI450 Series, up to 6 GW, first 1 GW in 2H 2026 all stock MI455X
Meta custom GPU based on MI450 architecture, Helios, up to 6 GW confirmed 144 GB / six-stack product
Microsoft Azure ND MI455X v7 / Helios direction measured production economics
Oracle 50,000 MI450-Series GPUs in Helios deployment all stock MI455X
Anthropic no primary confirmation found in this audit confirmed customer

How Helios compares with NVIDIA's HBM progression

The comparison must preserve scope and precision.

Platform Per-GPU memory Per-GPU HBM bandwidth Rack GPU memory Rack HBM bandwidth Scale-up
NVIDIA GB300 NVL72 288 GB HBM3E architecture about 8 TB/s 20 TB 576 TB/s 130 TB/s NVLink 5
AMD Helios 432 GB HBM4 23.3 TB/s peak 31 TB 1.7 PB/s aggregate peak 260 TB/s UALoE
NVIDIA Vera Rubin NVL72 288 GB HBM4 22 TB/s 20.7 TB 1.58 PB/s 260 TB/s NVLink 6

Helios publishes about 1.5x Rubin's HBM capacity per GPU and per rack, about 1.06x Rubin's per-GPU peak HBM bandwidth, and about 1.08x its rack aggregate peak HBM bandwidth. Both publish a 260 TB/s scale-up headline. Their latency, collective behavior, precision formats, runtime maturity, and accepted-work performance are not equivalent, and no same-workload receipt here establishes a winner.

The NVIDIA upgrade itself is revealing:

TEXT
GB300 -> Rubin:
  HBM capacity per GPU stays at 288 GB
  rack GPU memory stays near 20 TB
  per-GPU HBM bandwidth rises from about 8 to 22 TB/s
  rack HBM bandwidth rises from 576 TB/s to 1.58 PB/s
  scale-up bandwidth doubles from 130 to 260 TB/s

Rubin uses HBM4 mainly to increase bandwidth and system integration, not to produce a large local-capacity jump.

Where NVIDIA places reusable context

NVIDIA's public Rubin context answer is CMX, not raw CXL. CMX combines Dynamo, NIXL, Spectrum-X Ethernet, BlueField-4, DOCA Memos, and flash/NVMe-backed context capacity. KV is restored into host or GPU memory before its deadline.

TEXT
Rubin HBM4:
  active weights, active KV, activations, workspaces

Vera LPDDR5X:
  CPU work, metadata, warm state, staging

CMX / local flash:
  reusable context that can be restored before deadline

storage:
  durable repositories, logs, indexes, and artifacts

CMX is not HBM. CXL is not HBM. A DPU is not HBM. The system wins only when the placement and restore policy make useful work cheaper without breaking p99, correctness, security, or quality.

How C-001 Hermes plus GLM-5.2 would traverse Helios

This is the fixed coding-agent workload from the rest of the article. The model is not a generic chatbot. It must inspect a repository, use tools, edit files, run tests, repair failures, and return one accepted patch.

TEXT
1. Venice CPU boundary
   hold system instructions, permissions, repository metadata, tool schemas,
   tokenizer work, scheduling, sandbox state, test processes, and verification

2. Engine decision
   select and pin PyTorch plus vLLM, SGLang, or another declared engine
   capture container, ROCm, driver, firmware, model, tokenizer, and precision

3. Native AMD operation path
   ROCm / HIP
   -> AITER, Triton, Composable Kernel, hipBLASLt, or rocBLAS
   -> RCCL or rocSHMEM when the chosen parallel plan communicates
   -> AMDGPU compilation and HSACO only after the exact target is captured

4. MI455X and HBM4
   read FP8 or BF16 model state according to the exact artifact
   create and reuse prefix and active KV state
   hold activations, MoE expert shards, routing buffers, and workspaces
   never infer useful bandwidth from the 23.3 TB/s peak

5. Rack fabrics
   use UALink or UALoE only when tensor, pipeline, expert, or other state
   crosses the selected GPU placement
   use Vulcano and Ultra Ethernet only when traffic crosses the rack boundary

6. Tool and verifier return
   pause or continue GPU work under an explicit tool-window policy
   run repository tools on the declared CPU sandbox
   join the patch, tests, retries, latency, power, and cost under one run ID

ARCHITECTURE ONLY / RUN NOT CAPTURED. The current article has no selected MI455X compiler target, device executable, instruction stream, dispatch trace, HBM counter, UALoE trace, power interval, or accepted-patch receipt. NVIDIA code examples cannot fill those AMD fields.

One VAST Data and AMD test makes the reusable-KV path concrete. It is evidence for a memory hierarchy, not evidence that storage replaced HBM.

VAST Data published a July 23, 2026 benchmark using one AMD MI355X node, GPT-oss 120B, ROCm, vLLM 0.22.1, LMCache, tensor parallelism of one, concurrency from one through 300, and an explicitly described 128K-token setup. The remote path used the LMCache filesystem plugin, NFSv3 over RDMA, NFS multipath, an 800 Gb/s network, and a VAST storage partition. The source calls the adapters AMD Infinity Fabric 800Gb NICs; it does not identify them as Vulcano, so this article does not silently rename the test hardware. It also describes local host RAM as a comparison tier without identifying that memory as DDR5 or LPDDR5X.

TEXT
active request
-> vLLM schedules GPT-oss 120B on MI355X
-> ROCm kernels create and consume active KV in MI355X HBM3E
-> LMCache identifies reusable KV chunks
-> filesystem plugin writes or restores through NFSv3 over RDMA
-> 800 Gb/s NIC and switch carry the transfer
-> VAST flash retains the reusable object
-> restored KV returns to the serving path before decode can use it

Step by step: from one long-context request to the physical memory system.

Step What moves or changes Physical layer Why the CEO or CFO should care Receipt
1. Request identity Prompt, policy, model revision, tokenizer, tenant, and session identity enter the serving system CPU registers, caches, and host memory A wrong cache key can return stale or cross-tenant state even if the transfer is fast Request ID, exact revisions, tenant and ACL policy
2. Prefill GPT-oss 120B transforms the 128K-token context into attention state MI355X compute units read weights and write active KV through cache and HBM controllers This is the expensive computation the remote cache is trying not to repeat Kernel timeline, input tokens, TTFT, HBM bytes, output check
3. HBM cell write Each KV bit is encoded as charge in a DRAM cell and recovered through wordlines, bitlines, sense amplifiers, and restore HBM3E DRAM arrays inside vertically stacked memory dies HBM is fast because many channels operate in parallel close to the accelerator, but every access still consumes time and energy Memory-controller counters and a synchronized board-power interval
4. Vertical and package movement Signals cross on-die metal, dense vertical TSVs, die-to-die bonds or bumps, the stack interface or base layer, and the accelerator package Copper-dominated wiring, dielectrics, silicon, solder or hybrid-bond structures, redistribution layers, substrate, power delivery, and thermal interfaces The memory wall is physical: wire length, capacitance, resistance, heat, yield, and package area turn bytes into cost Vendor stack map, package cross-section, channel counters, thermal map, yield and repair data
5. Chunk and policy decision LMCache groups reusable KV into 8,192-token chunks and decides whether the object is eligible to leave local memory Software metadata in host memory; active source blocks still reside in accelerator HBM Chunk size trades metadata and transfer overhead against wasted bytes and late restores Cache key, chunk IDs, admitted bytes, rejected bytes, eviction reason
6. Read out of the hot tier Eligible KV is copied out of accelerator memory toward the host or I/O path HBM read, accelerator I/O, PCIe or platform path, DMA engines, host staging only if the selected path requires it Copying state frees scarce HBM capacity but uses bandwidth and can interrupt useful work Source and destination addresses, bytes, DMA path, overlap, p50/p95/p99
7. Remote transfer NFSv3 over RDMA carries the filesystem I/O through the named 800 Gb/s adapters and Mellanox switch NIC SerDes, copper or optical links, switch buffers, packet headers, flow control, and retransmission or fallback behavior A nominal 800 Gb/s link does not guarantee useful KV goodput or a stable tail Link topology, backend, packet and retry counters, transfer latency, NIC and switch power
8. VAST flash retention The KV object is written to or read from the VAST storage partition Storage controllers, DRAM metadata and caches, flash channels, NAND media, power supplies, and cooling Remote flash is cheaper capacity than accelerator HBM only if reuse, latency, durability, and operational cost work together Object bytes, media and array identity, queue depth, service time, array energy, failure state
9. Restore A later matching request retrieves the exact chunks and reinstalls them into the serving path Flash to network to accelerator I/O, followed by HBM writes before the GPU consumes the KV The restore must arrive before recomputing prefill would have finished and must not crowd out active work Hit, identity validation, bytes restored, full-restore p99, recompute baseline, correctness
10. Accepted result Decode continues and the application decides whether the output is useful GPU decode, CPU tools, application verifier, logs, and durable output Faster tokens are not business value if the result fails, retries, or needs human repair Accepted output, retries, total latency, energy, cost, and failure ledger under one run ID

The deepest physical point is easy to lose in a software diagram. A logical KV block is not a box floating inside the GPU. Its bits are represented by electrical state in DRAM cells. Access activates rows, perturbs bitlines, uses sense amplifiers to resolve very small voltage differences, restores cell charge, and moves data through on-die wiring and the HBM stack interface. Dense TSVs shorten the vertical path relative to board-level memory, while the wide package connection creates parallelism. The benefit comes with more difficult wafer processing, die thinning, alignment, bonding, thermal management, package routing, test, repair, and known-good-die economics. Exact MI355X stack materials, TSV dimensions, bonding process, and per-boundary energy are not public in the cited benchmark, so the article does not substitute a generic HBM diagram for a product teardown.

The offload path does not make those HBM writes disappear. It changes when they happen:

TEXT
recompute path:
  read weights from HBM
  execute prefill again
  write new KV to HBM

restore path:
  read saved KV from remote flash
  move it through storage, switch, NIC, and accelerator I/O
  write restored KV to HBM

restore wins only when:
  transfer time + queueing + validation + HBM reinstall + miss risk
  is lower than
  repeated prefill time under the same correctness and p99 requirement

The power and materials follow every arrow. HBM power includes array activation, sensing, restore, refresh, I/O, controller, and package losses. Compute power includes matrix, vector, cache, register, clock, and control activity during prefill. The remote path adds storage controllers and media, NIC SerDes, switch buffers, CPUs or DPUs, fans or liquid loops, voltage conversion, and facility overhead. Copper reduces resistance but does not remove capacitance. Shorter TSV and package paths can reduce movement energy relative to longer board or network paths, but manufacturing and cooling costs do not vanish.

For this exact benchmark, the honest ledger is:

TEXT
known:
  MI355X node
  GPT-oss 120B
  vLLM 0.22.1 plus LMCache plus ROCm
  128K detailed context setting
  tensor parallelism 1
  concurrency 1 through 300
  8,192-token chunks
  NFSv3 over RDMA
  800 Gb/s network
  vendor-reported 9x TTFT and 9.7x token throughput

not published:
  exact HBM bytes read and written
  exact KV bytes stored and restored
  cache-hit distribution
  GPU, CPU, NIC, switch, and storage watts over the same interval
  joules per request
  electricity tariff
  facility PUE
  facility WUE and water allocation
  hardware and support price
  accepted-task result

The equations a real receipt must fill are:

TEXT
task energy, kWh =
  integral of
  accelerator + CPU + DPU + NIC + switch + storage + allocated cooling power
  over the same task interval
  / 3,600,000

electricity cost =
  task energy in kWh * contracted tariff

allocated facility water, liters =
  measured IT energy in kWh * declared WUE in liters per kWh

cost per accepted task =
  electricity
  + allocated accelerator, network, and storage capacity
  + software, support, and engineering
  + failed attempts and human rework
  divided by accepted tasks

Those formulas make the decision inspectable; they do not manufacture missing inputs. The older modeled HBM2 literature cited elsewhere in this article reports an approximately 3.97 pJ/bit access value for its own configuration. It is useful for understanding why distance and movement matter. It is not an MI355X, HBM3E, VAST, BlueField, CMX, or facility-water measurement.

The headline comparison is the VAST-optimized remote path against no offload, where the context was recomputed at the start of each decode turn. VAST reports a 9x improvement in time to first token and 9.7x greater token throughput for its named setup. These are [VENDOR-REPORTED BENCHMARK RESULTS]. VAST's page says AMD reviewed the claims but did not independently verify them, and that the result is specific to VAST and may not be typical. The page introduction says 120K tokens while the detailed setup and results say 128K; this article uses the explicit 128K setup and preserves the discrepancy instead of averaging it away.

The software configuration matters as much as the storage brand. The published path used asynchronous loading, direct I/O, 8,192-token LMCache chunks, a 16 MiB read-ahead value, and a filesystem base path on the shared VAST mount. That proves a named configuration was reported. It does not prove a cache hit, byte transfer, tail-latency distribution, power result, or accepted coding task on Touchdown infrastructure.

The new AMD networking parts occupy different boundaries.

Part Boundary Memory-wall role What it does not prove
Salina DPU Front-end networking, security, software-defined networking, storage access, and DPU-managed NVMe Can move infrastructure and storage work away from host CPU cores and make an NVMe-backed KV tier easier to operate That an application used the DPU, that KV bypassed the host, or that restore beat recompute
UALink over Ethernet Scale-up inside Helios Connects 72 MI455X accelerators so distributed model state and communication can cross GPU packages Cache identity, storage access, scale-out behavior, or application goodput
Vulcano 800 AI NIC Scale-out and scale-across AMD publishes 800 Gb/s per NIC and up to three NICs per GPU, or a 2.4 Tb/s-per-GPU ceiling Sustained useful bandwidth, p99, switch behavior, power, or KV hit rate
Pollara 400 AI NIC Existing Ethernet clusters Brings AMD Pensando's AI-networking path to current deployments The Salina storage role or the Vulcano Helios topology
VAST AI Operating System plus flash Remote reusable-data and KV tier Gives LMCache a remote filesystem target so reusable context can outlive local HBM pressure Local-HBM latency, cache correctness, universal 9x performance, or lower total energy
LMCache KV identity, chunking, admission, movement, and restore policy Decides what reusable KV object is stored and requested A physical wire, DPU, NIC, switch, or storage array

The equivalent NVIDIA path uses different products but the same memory question:

TEXT
Blackwell HBM3E or Rubin HBM4
  active weights, active KV, activations, and workspaces

Grace or Vera CPU memory
  CPU tools, metadata, warm state, and staging

BlueField-4 plus CMX
  DPU-managed, Spectrum-X-connected flash context tier

Dynamo / KV block manager plus NIXL and DOCA Memos
  identity, placement, and movement policy

BlueField is the DPU. CMX is the context-memory storage platform. NIXL is movement software. Dynamo and the KV manager make serving decisions. Blackwell HBM3E and Rubin HBM4 remain the hot tiers. For the complete KV-cache software hierarchy, read KV Cache Is Becoming the Memory Hierarchy of Inference. The point here is narrower: these systems exist because HBM capacity, bandwidth, energy, and cost make repeated recomputation expensive, not because flash has become HBM.

Memory, power, and money require one joined comparison.

TEXT
candidate A: keep reusable KV in accelerator HBM
candidate B: spill to host DDR or LPDDR5X
candidate C: restore from VAST or CMX flash
candidate D: miss and recompute prefill

for every candidate:
  same model, precision, prompt, context, concurrency, and output check
  measure hit rate, bytes, queueing, restore p50/p95/p99, and recompute time
  integrate accelerator, CPU, DPU, NIC, switch, storage, and cooling power
  apply the actual electricity tariff and infrastructure allocation
  divide by accepted tasks, not cache requests or generated tokens

The remote path can save money when avoided GPU-seconds and freed HBM capacity are worth more than the DPU, NIC, switch, flash, CPU, software, support, and tail-latency cost. It can save energy when avoided recompute exceeds the energy used to store, find, move, validate, and reinstall the KV blocks. It can lose on both when reuse is low, objects are stale, transfers arrive late, or the network and storage tier stay powered for little accepted work.

The VAST benchmark does not publish a synchronized power trace, electricity rate, storage-array energy, facility PUE, facility WUE, water allocation, purchase price, or accepted-task result. Those fields remain unknown. Coolant flow is not water consumption. A DPU or NIC TDP is not task energy. Water belongs in the facility ledger only when the same measured IT-energy interval is joined to a declared WUE and allocation method.

What this means by reader. The CEO asks whether more long-context sessions complete on time. The CFO compares avoided GPU-seconds and deferred accelerator capacity against storage, network, power, support, and engineering cost. The CTO owns cache identity, tenant isolation, failure fallback, topology, and rollback. The serving engineer captures the exact LMCache, vLLM, ROCm, filesystem, and transport versions. The infrastructure engineer measures DPU, NIC, switch, and storage queues. The kernel and GPU engineer decides when restore is slower than recompute and which active blocks must remain in HBM.

How V-001 Wan2.2 would traverse Helios

Wan2.2 creates a different state lifecycle. It does not grow an autoregressive LLM KV cache token by token. Its attention Q/K/V tensors are transient inside repeated denoising work.

TEXT
1. Venice CPU boundary
   accept the prompt, prepare metadata, schedule the text encoder,
   allocate the latent contract, and manage the output and quality gate

2. ROCm candidate runtime
   pin PyTorch, the Wan2.2 repository revision, ROCm, libraries,
   model weights, dtype, resolution, frame count, and sampling steps

3. MI455X HBM4
   retain model weights, latent state, transient Q/K/V,
   activations, communication buffers, scheduler state, and VAE workspace

4. Distributed denoising
   cross UALink or UALoE only if the pinned tensor, sequence, FSDP,
   Ulysses, expert, or pipeline plan actually communicates
   capture RCCL operation, message size, rank map, wait time, and retries

5. Decode and acceptance
   decode through the selected VAE path
   verify frame count, dimensions, codec, quality rule, and failure policy
   divide total time, energy, and cost by accepted clips, not attempts

ARCHITECTURE ONLY / RUN NOT CAPTURED. No reviewed source proves Wan2.2 executed on Helios. The MI455X model path, kernels, transient HBM traffic, collective behavior, output quality, power, energy, water, and cost remain unknown.

What Touchdown should measure

For a coding-agent workload:

TEXT
model and revision
precision and quality parity
engine and kernel path
context and concurrency
TTFT, TPOT, end-to-end p50/p95/p99
HBM occupancy and fragmentation
prefix-hit rate
KV bytes
KV offload tier and object identity
restore bytes and p50/p95/p99
model and expert shards
RCCL/NCCL wait
fabric message histogram
DPU/NIC activity
accelerator, CPU, DPU, NIC, switch, and storage power
facility PUE, electricity tariff, and separately measured WUE when available
cost and joules per accepted patch

For Wan2.2-style video:

TEXT
latent dimensions and token count
denoising phases and steps
attention, FFN/MoE, and VAE time
QKV and activation traffic
HBM reads and writes
collective wait
power
quality gate
cost and joules per accepted clip

What each reader should take from the Helios lane

Reader Decision this architecture can inform now Receipt still required
Investor AMD has launched MI455X and moved the Helios reference-design system into full production shipment ramp, partner configurations, supply, yield, customer acceptance, utilization, and economics
CEO Helios could become another infrastructure option for the same accepted business workflow accepted-task throughput, reliability, deployment time, and operator ownership
CFO More HBM and rack bandwidth may change how many replicas, shards, or racks are needed price, rack power, facility work, support, failures, utilization, and cost per accepted task
CTO / infrastructure The port crosses runtime, kernels, collectives, topology, failure, and service boundaries pinned version matrix, degraded-mode plan, topology tests, p95/p99, and rollback
Software engineer The framework call is only the top of a different AMD runtime path exact engine flags, fallbacks, artifact identity, logs, and correctness
Kernel engineer Useful bandwidth depends on shapes, layout, fusion, occupancy, MFMA path, and communication selected target, IR/HSACO, kernel name, counters, numerical parity, and profiler interval
Hardware / facility engineer The rack combines HBM4 packages, fabrics, power shelves, busbar, liquid manifold, and service points final BOM, rack power, flow, pressure, coolant, CDU, thermal map, repair, and yield

Final AMD conclusion

Helios is serious competition because AMD now combines:

TEXT
CDNA 5
HBM4
Venice
Vulcano
Salina
UALink
Ultra Ethernet
ROCm
semi-custom silicon
large customer commitments

The remaining question is not answered by a specification sheet:

Can AMD convert its published peak capacity and bandwidth advantage at the named MI455X and Helios scopes, plus its open customization path, into stable, quality-matched accepted-work throughput before NVIDIA's fabric, software, context, and deployment integration closes the system-level gap?

The answer needs synchronized workload, HBM, fabric, power, failure, quality, and cost receipts.

Primary sources

Which memory technologies fit which state roles?

In plain English. No memory technology wins every job. Hot, frequently changing state needs to stay close. Warm, reusable state can sometimes tolerate a transfer. Cold, durable state belongs in cheaper storage. Some state is cheaper to recompute than to preserve. Start with the object's lifetime, access pattern, deadline, and correctness rule, then choose a tier.

There is no useful answer to What memory is best? without a state object, boundary, and workload.

The comparison needs the same questions for every technology:

TEXT
What physical state stores the bit?
How far is it from compute?
What interface and controller expose it?
What is the useful access granularity?
How do first-byte latency and sustained service behave?
How do reads differ from writes?
Does it need refresh? What is its retention and endurance?
How large can the usable tier be in the named system?
What happens under contention, errors, and failure?
What ships, what is modeled, and what is only proposed?

The objective matrix

Technology Physical and interface distinction Plausible role What breaks first Evidence state
SRAM, cache, scratchpad Multi-transistor volatile cells close to logic Smallest, hottest, repeatedly reused state Area and leakage limit capacity Shipping, implementation-specific
eDRAM Dynamic cells integrated closer to logic Larger near-compute working memory where process and refresh fit Process integration, refresh, temperature, capacity, design/test complexity Shipping precedents plus workload-specific papers
HBM3E Stacked DRAM with a very wide local package interface Hot weights, active KV, activations, latents, communication buffers Capacity, heat, package/PDN, yield, supply, software utilization Named shipping vendor artifacts; metrics remain source-scoped
HBM4 Stacked DRAM with a wider local package interface and expanded base-logic design surface Hot weights, active KV, activations, latents, communication buffers Capacity, heat, package/PDN, yield, supply, software utilization Public standard plus named shipping vendor artifacts; metrics remain source-scoped
Custom HBM and custom base logic HBM-compatible stack with more configurable logic/control at the base/interface layer Workload-specific movement, security, RAS, telemetry, or bounded near-memory functions Non-recurring engineering, access, integration, qualification, power, heat, and yield Standards and vendor roadmap/design-surface evidence
Standard Package High Bandwidth Memory (SPHBM4) JEDEC JESD330-4 direction for HBM4-stack integration through a different buffer/interface die and standard-package path Capacity or bandwidth integration under the public standard's scope Ecosystem access, implementation, validation, and workload value Public standard record exists; product and access must be named separately
Intel XBM application Backend-transistor DRAM organization with fine subchannels, repair/control, and serialized-package embodiments Architecture simulation and diligence candidate Unmeasured latency, energy, retention, yield, cost, thermals, software, product status Patent application only
GDDR7 Discrete graphics DRAM, high per-pin signaling over a narrower aggregate interface than HBM stacks High-bandwidth boards where package, cost, serviceability, or capacity trade differs Signal/power complexity and aggregate board design Shipping/vendor product evidence
DDR5 Commodity system DRAM, typically DIMM-oriented and CPU-controlled Large host-visible working memory Lower accelerator-local bandwidth, longer path, transport overhead Standard and shipping products
LPDDR5/LPDDR5X Low-power DRAM interface optimized for energy and capacity-density trade-offs Client, edge, and some data-center designs where power/capacity matter Different bandwidth, package, controller, serviceability, and ecosystem Shipping/vendor product evidence
CXL-attached memory Coherent, load/store-capable expansion over a CXL link and topology Capacity expansion, pooling, workload-defined warm state Link latency/bandwidth, topology, contention, coherency, software policy, failure domain Public consortium capability plus named vendor products; shipment state is checked per part
NVMe SSD and NAND Block storage protocol over nonvolatile flash Durable state, checkpoints, repositories, cold cache, predictable pages First-byte latency, page/block behavior, writes, endurance, filesystem/controller path Shipping standard and products
Sandisk HBF direction Proposed high-bandwidth flash stack/interface Very large read-mostly parameter or predictable-page state First-byte latency, writes, page usefulness, software/controller maturity Vendor roadmap and internal simulation
Kioxia high-bandwidth flash prototype Separate flash-module prototype Physical research evidence for higher-bandwidth flash modules Prototype scope, workload proof, productization, software Physical prototype announcement, not Sandisk HBF
Emerging device families PCM, ReRAM/RRAM, MRAM, FeFET, oxide/gain-cell and related candidates Specific retention, endurance, density, or compute roles if evidence survives Variability, write cost, endurance, temperature, process, yield, control, ecosystem Research or named vendor state only
Compression and rematerialization Software reduces retained or moved bytes Avoid memory demand when compute/quality permit Extra compute, latency, precision risk, complexity Shipping techniques; workload result required

HBM remains the hot-tier baseline. The alternatives change capacity, distance, granularity, retention, energy, or programmability. None removes the need to measure the state path.

SRAM and eDRAM: closer is expensive in different ways

SRAM is the natural home for register files, caches, queues, and software-managed scratchpads because it can deliver fine-grained low-latency access close to compute. Its area cost prevents it from becoming a capacity-equivalent substitute for stacked DRAM.

That area trade affects kernel design. Tiling exists because a small on-chip SRAM working set can reuse data many times before HBM sees another request. Larger tiles can improve reuse but consume registers or shared memory, reduce occupancy, or constrain scheduling. The right tile is an interaction between the operation, data layout, on-chip capacity, and HBM path.

eDRAM offers another point. The RANA paper from ISCA 2018 explored refresh-aware eDRAM for CNN acceleration using synthesized RTL, cycle-accurate simulation, and modeled energy. [PAPER, MODELED] Its useful idea is that workload error tolerance, layer scheduling, bank allocation, and refresh control can be co-designed. Its limit is equally important: it is not fabricated RANA silicon, HBM, an LLM, video diffusion, robotics proof, or a current-node benchmark. It shows that retention and refresh can become workload-visible design variables under a bounded experiment.

HBM3E and HBM4: the hot-tier tournament baseline

HBM's strength is aggregate local bandwidth and channel parallelism under a tight package. HBM4 continues that direction while expanding the logic base-die and interface design space.

[VENDOR PRODUCT ARTIFACT] Micron's dated HBM3E product brief establishes a named HBM3E product family and its vendor-scoped specifications. It does not establish a universal HBM3E system result.

[SHIPPING VENDOR ARTIFACT] Samsung's dated announcement says it shipped a commercial HBM4 product and describes a 4 nm logic base die. It reports 11.7 Gb/s operation and capability up to 13 Gb/s for its named product context. [SHIPPING VENDOR ARTIFACT] Micron's dated announcement says it began volume production of a 36 GB 12-high HBM4 product designed for NVIDIA Vera Rubin and reports more than 2.8 TB/s for that named part. These are vendor statements, not Touchdown measurements. The scope is a part or stack, not every accelerator.

HBM4 does not solve everything by becoming wider or faster. The accelerator needs enough outstanding work, useful layout, controller balance, on-chip reuse, power, and thermal headroom. Capacity can still limit model size or concurrency. More interface pins and logic increase package and validation demands.

The accelerator roadmap is also a memory roadmap

Comparing generations is easy to get wrong because vendors publish at different scopes. A single-GPU capacity cannot be placed beside a 72-GPU rack total without naming the boundary. Peak theoretical bandwidth is not sustained useful bandwidth. A roadmap projection is not a shipping receipt.

Use this five-number decoder before reading the table:

Number What it answers What it does not answer
Memory capacity, GB or TB How much state can fit at the named device or rack scope How fast the state arrives or how much is usable after runtime reservation
Local HBM bandwidth, GB/s or TB/s The peak or measured byte rate between an accelerator and its local HBM Fabric speed, latency, or accepted-task throughput
Compute rate, FLOPS A theoretical or measured rate for eligible arithmetic at a named precision Quality equivalence, memory utilization, or application speed
Fabric bandwidth, GB/s or TB/s The peak or measured rate between named components or domains Small-message p99, collective efficiency, or local HBM bandwidth
Latency, seconds or milliseconds Time between two named events Throughput, capacity, or the reason for the delay

None of these five numbers proves how many patches, clips, or other tasks the system accepts.

Vendor platform Evidence state on July 23, 2026 Published memory scope Published memory bandwidth scope Interconnect and system boundary
NVIDIA H100 SXM Official current product 80 GB HBM per GPU 3.35 TB/s per GPU NVLink 4, 900 GB/s per GPU; system topology varies
NVIDIA H200 SXM Official current product 141 GB HBM3E per GPU 4.8 TB/s per GPU Same 900 GB/s NVLink-class per-GPU figure; more local capacity and bandwidth than H100
AMD Instinct MI350P Official current product 144 GB HBM3E per PCIe card 4 TB/s peak theoretical per card PCIe 5.0 x16; passive full-height double-slot card; 450 W configurable and 600 W maximum TBP
AMD Instinct MI355X Official current product 288 GB HBM3E per GPU 8 TB/s peak theoretical per GPU Seven scale-up Infinity Fabric links; 1,400 W published TBP for the OAM
NVIDIA GB200 NVL72 Official current rack system 13.4 TB HBM3E across 72 GPUs 576 TB/s aggregate HBM 130 TB/s aggregate NVLink domain, plus 36 Grace CPUs and LPDDR5X
NVIDIA GB300 NVL72 Official current rack system 20 TB HBM3E across 72 GPUs Up to 576 TB/s aggregate HBM 130 TB/s aggregate NVLink domain; more HBM capacity than GB200
NVIDIA Vera Rubin NVL72 Official preliminary 20.7 TB HBM4 across 72 GPUs 1,580 TB/s aggregate HBM NVLink 6, 260 TB/s aggregate scale-up bandwidth; specifications subject to change
AMD MI455X / Helios Official current launch and production reference design; shipments scheduled to begin near end-Q3 and ramp in Q4 432 GB HBM4 per MI455X; 31 TB across 72 GPUs 23.3 TB/s peak per MI455X; 1.7 PB/s aggregate peak per rack 72-GPU UALink/UALoE scale-up domain; OEM/ODM reference design rather than one direct AMD rack SKU
NVIDIA Vera Rubin Ultra NVL72 / Kyber NVL144 / NVL576 Official preliminary topology Final public per-system HBM specification is not treated as fixed here Final delivered useful bandwidth is not public Three different 72-, 144-, and 576-GPU domain options; Kyber is the 144-GPU rack
AMD Instinct MI500 Series Official roadmap HBM4E named; capacity not public in the reviewed primary source Not public CDNA 6 and 2 nm named; planned for 2027; AMD performance statements are projections
NVIDIA Feynman Kyber NVL1152 Official preliminary topology Final HBM implementation and delivered capacity not public Not public Eight 144-GPU Kyber racks in one planned scale-up domain

[OFFICIAL PRODUCT / ROADMAP SOURCES] H100, H200, GB200, GB300, and Vera Rubin figures above use NVIDIA's named product pages and current platform documentation. MI350P uses AMD's current product page and enterprise deployment article. MI355X uses AMD's product page. The July 23 MI400 launch and Helios launch govern the current MI455X and Helios fields. The older 19.6 TB/s MI400-family projection and 1.4112 PB/s arithmetic rack total are historical, not current launch specifications. MI500 uses AMD's CES 2026 roadmap statement naming CDNA 6, 2 nm, HBM4E, and a planned 2027 launch. Future rows are not purchase specifications, and current peak specifications are not measured workload results.

How should a buyer compare RTX 6000, B200, B300, GB200, GB300, MI350P, MI355X, and Helios?

Start with scope. These products do not belong in one flat leaderboard. An RTX workstation card, a passive MI350P server card, one MI355X OAM, an eight-GPU DGX system, a Grace Blackwell Superchip, and a 72-GPU rack expose different memory domains, CPU boundaries, fabrics, cooling requirements, software contracts, and failure domains. The first question is not which number is largest. It is which state must stay local for which workload, and at what physical boundary.

Compare at the same boundary NVIDIA path AMD path Memory fact that matters Design decision and limitation
Workstation or enterprise PCIe card RTX 6000 Ada has 48 GB GDDR6. RTX PRO 6000 Blackwell Workstation, Max-Q, and Server each have 96 GB GDDR7, with different power and deployment roles documented below. MI350P has 144 GB HBM3E and 4 TB/s peak local bandwidth in a passive PCIe server card. This is local accelerator memory on one card. HBM3E and GDDR7 make different capacity, bandwidth, packaging, power, graphics, media, and serviceability bargains. Choose from the actual job. Local graphics, media, CAD, visualization, and model development are not the same path as dense server inference, private RAG, or an eight-card agent service. No peak-memory ratio proves accepted-task throughput.
Accelerator and eight-GPU server NVIDIA publishes 1,440 GB total GPU memory and 64 TB/s aggregate HBM3E bandwidth for DGX B200. NVIDIA publishes 2.1 TB total GPU memory for DGX B300. AMD publishes 288 GB HBM3E and 8 TB/s peak per MI355X OAM. The buyer must name the exact eight-accelerator UBB or OEM server before multiplying or comparing a node. A product page may report per accelerator, per baseboard, or per complete server. Aggregate capacity is the sum of separate local domains unless the runtime and fabric prove a usable distributed placement. B200 and B300 package a qualified NVIDIA system and software boundary. MI355X gives more local HBM per named accelerator, but server topology, ROCm and kernel qualification, cooling, collectives, and support remain configuration-specific.
CPU-GPU superchip and 72-GPU rack GB200 NVL72 publishes 13.4 TB HBM3E at 576 TB/s aggregate, 17 TB Grace LPDDR5X at 14 TB/s, and a 130 TB/s NVLink domain. GB300 NVL72 publishes 20 TB GPU memory, up to 576 TB/s aggregate memory bandwidth, the same published CPU-memory scope, and a 130 TB/s NVLink domain. Helios publishes 31 TB HBM4 and 1.7 PB/s aggregate peak memory bandwidth across 72 MI455X accelerators, with Venice CPUs, UALink over Ethernet scale-up, Vulcano 800 scale-out NICs, and Salina DPUs. The OEM or ODM owns the exact host-memory configuration. CPU memory, accelerator-local HBM, scale-up bandwidth, scale-out networking, and storage are five different boundaries. Coherent or routable access does not make them one pool with one latency. Compare the whole model placement, KV and activation lifetime, collective graph, p99 latency, rack power, cooling, failure behavior, and accepted output. Published rack ceilings do not establish useful bandwidth, tokens per watt, or cost per accepted task.
Context and data outside local accelerator memory NIXL coordinates movement. BlueField handles named infrastructure work. CMX is an announced context-memory storage direction. NVMe or remote storage retains colder objects. ROCm and the serving engine own placement. Infinity Fabric or UALink handles named GPU domains. Vulcano and Salina own different network and infrastructure roles. The VAST plus LMCache example uses remote flash over a named RDMA path. A DPU, NIC, CXL device, host-memory tier, or flash system does not become HBM. It changes where an object waits and how it returns. Offload wins only when lookup, identity, transfer, restore, miss handling, energy, and failure behavior beat recomputation before the reuse deadline.

The design sequence is the same for every row:

  1. Write down the exact weight, KV, activation, latent, workspace, communication-buffer, and tool-state objects.
  2. Prove which objects fit in each local memory domain after runtime reservation and fragmentation.
  3. Trace every byte that leaves local memory through the selected CPU, PCIe or coherent link, scale-up fabric, NIC or DPU, network, and storage path.
  4. Name the existing framework, compiler, kernel library, collective library, serving engine, placement layer, driver, firmware, and observability tools that support that exact hardware revision.
  5. Join output quality and accepted work to latency, measured power, integrated energy, cooling and water allocation, retries, failures, and money.

That sequence prevents three common mistakes: treating aggregate rack memory as one flat address space, treating a software package import as workload support, and treating a roadmap component as a delivered system.

In plain English. MI350P is the CDNA 4 card for an enterprise that wants local AI inference inside familiar PCIe servers without moving straight to an eight-OAM baseboard or a 72-GPU Helios rack. It has 144 GB of local HBM3E and 4 TB/s of published peak memory bandwidth. It is not the 288 GB MI350X or MI355X in another shell.

[OFFICIAL CURRENT PRODUCT SPECIFICATION, REVIEWED 2026-07-23] AMD publishes these product-scoped fields:

Compute field MI350P official value
GPU architecture and process CDNA 4; TSMC 3 nm and 6 nm FinFET
Stream processors / matrix cores / compute units 8,192 / 512 / 128
Peak engine clock / transistors 2,200 MHz / 73 billion
Peak MXFP4 / MXFP6 / MXFP8 matrix 4.6 / 4.6 / 2.3 PFLOPS
Peak OCP FP8 matrix 2.3 PFLOPS dense; 4.6 PFLOPS with structured sparsity
Peak FP16 and BF16 matrix 1.15 PFLOPS dense; 2.3 PFLOPS with structured sparsity
Peak INT8 matrix 2.3 POPS dense; 4.6 POPS with structured sparsity
Peak scalar FP16 72 TFLOPS
Peak FP32 matrix / scalar 72 / 72 TFLOPS
Peak FP64 matrix / scalar 36 / 36 TFLOPS
Memory, board, and software field MI350P official value
Local memory 144 GB HBM3E
Peak local HBM bandwidth / interface 4 TB/s / 4,096 bit
Last-level cache / Infinity Cache 128 MB / yes
Data protection Full-chip ECC
Board power 450 W configurable; 600 W maximum TBP
Host and external power PCIe 5.0 x16; 12V-2x6
Physical card PCIe add-in card; full height; double slot; 10.5 in / 267 mm
Cooling Passive
Operating system Linux x86-64
RAS and virtualization RAS, page retirement, page avoidance; current AMD web specification says SR-IOV is supported, while the May 2026 brochure describes SR-IOV as future support for up to four partitions
Supported technologies CDNA 4, 4th Gen AMD Infinity Architecture, ROCm
APIs OpenMP, OpenCL, HIP, ROCm
Frameworks listed by AMD TensorFlow, PyTorch, ONYX-RT, SGLang, JAX, Triton, Kokkos, RAJA

AMD's page renders the framework label ONYX-RT; it is preserved here rather than silently corrected. AMD also says some technologies need third-party enablement, features vary by operating system, and the buyer should confirm support with the system manufacturer. Every compute and bandwidth rate above is a vendor-published theoretical peak, not a Touchdown benchmark.

AMD's current web specification and product brochure conflict on SR-IOV timing. Production virtualization therefore needs the exact ROCm, hypervisor, firmware, OEM-server, and support-matrix receipt. The article does not turn either source into an unqualified deployment claim.

[OFFICIAL AMD DEPLOYMENT POSITIONING, REVIEWED 2026-07-23] AMD describes MI350P as a dual-slot card for standard air-cooled partner servers, including configurations with up to eight cards. AMD names on-premises inference, RAG, generative AI, agentic AI, and small, medium, and large model inference. It also names ROCm, PyTorch, AMD Inference Microservices, Kubernetes GPU Operator, and partner bare-metal, virtualized, Kubernetes, and hybrid-cloud paths.

The useful enterprise role is specific:

TEXT
private RAG and document assistants
agent and tool-calling inference
departmental or tenant-isolated model serving
embeddings, reranking, and generation pipelines
fine-tuning that fits the measured memory budget
regulated on-premises inference
incremental one-card to qualified eight-card deployment

That list is workload fit to test, not a claim that every workload wins.

Memory topology. One MI350P has one local 144 GB HBM3E domain. AMD publishes PCIe 5.0 x16 for the card. AMD does not publish the seven external scale-up Infinity Fabric links that it publishes for MI350X and MI355X OAM products. In an eight-card server, do not call the aggregate memory one flat HBM pool.

TEXT
one card:
  144 GB local HBM3E
  4 TB/s local peak HBM bandwidth

eight cards, arithmetic only:
  1,152 GB aggregate local HBM3E
  32 TB/s aggregate local peak HBM bandwidth
  3.6 kW configurable accelerator-board sum
  4.8 kW maximum accelerator-board TBP sum

[TOUCHDOWN DERIVATION] The eight-card values are multiplication across eight separate cards. They are not one 1.152 TB address space, one shared 32 TB/s interface, complete server power, or measured workload performance. Independent replicas and tenant placement can use separate local memory domains. Model, tensor, expert, or sequence shards that span cards must cross the qualified server topology, and communication-heavy scaling needs its own PCIe and peer-transfer receipt.

Compatibility with an existing data center. The card uses a familiar PCIe, full-height, double-slot, passive-air-cooled server form. That reduces the mechanical and facility jump relative to an OAM baseboard or liquid-cooled rack. It does not make installation automatic.

The operator still has to qualify:

  • the exact OEM server, BIOS, BMC, and GPU firmware;
  • slot spacing, retention, service access, and a front-to-back fan curve for a passive 600 W card;
  • 12V-2x6 cabling, branch power, PSU capacity, redundancy, and conversion loss;
  • PCIe 5.0 x16 lanes, root complexes, switches, NUMA locality, and peer behavior;
  • CPU, host DRAM, NIC, storage, and virtualization balance;
  • Linux, ROCm, framework, container, Kubernetes or OpenShift, and GPU Operator versions;
  • page retirement, page avoidance, SR-IOV, monitoring, failure isolation, and recovery;
  • sustained inlet temperature, component temperature, throttling, noise, and error behavior under the real workload.

Eight cards can occupy 3.6 to 4.8 kW of accelerator-board power before CPUs, DRAM, NICs, storage, fans, PSUs, and losses. Passive cooling means the server fans move the heat. It does not mean cooling power is zero.

[ARCHITECTURE ONLY / RUN NOT CAPTURED] Touchdown has not captured an MI350P run_id, useful HBM counters, PCIe traffic, peer-transfer p99, node power, thermal behavior, accepted-task throughput, or cost per accepted task. The first receipt must join the exact server and card topology to firmware, ROCm, engine, model, precision, HBM allocation, transfer bytes, latency, power, temperature, throttling, retries, and the accepted output.

MI350P versus the RTX 6000 family: trace one workload before comparing ratios

WORKLOAD COMPARISON · INPUT-DRIVEN · NO WINNER WITHOUT A MATCHED RECEIPT

Start with the user. A developer asks an agent to inspect a repository, change code, run tests, repair failures, and return a reviewable patch. The model path creates weights, prompt state, active KV or architecture-specific attention state, workspaces, and communication buffers. The tool path leaves the accelerator for CPU search, file I/O, editing, compilation, tests, and review. The accepted patch, not a peak specification, is the final denominator.

  1. Hermes assembles system instructions, tool schemas, repository context, and the user's request.
  2. GLM-5.2 FP8 weights are sharded across local accelerator-memory domains. Its 744-billion-parameter FP8 weight floor is 744 GB decimal before scales, unquantized modules, padding, workspaces, graph buffers, attention state, and fragmentation.
  3. The serving engine performs prefill, sparse-MoE routing, decode, and the cross-card communication required by the exact placement.
  4. The workflow returns to the CPU for search, edits, tests, and logs, then adds new context for another model turn.
  5. A verifier accepts or rejects the patch. Retries, failed tests, tool time, GPU time, and energy remain part of the cost.

[TOUCHDOWN CAPACITY DERIVATION, NOT A RUN] The ideal 744 GB FP8 weight floor requires at least six 144 GB MI350P cards or eight 96 GB RTX PRO 6000 Server cards by raw capacity arithmetic. Eight MI350P cards provide 1,152 GB across eight separate HBM3E domains, leaving a 408 GB arithmetic remainder over that floor. Eight RTX PRO 6000 Server cards provide 768 GB across eight separate GDDR7 domains, leaving 24 GB before every runtime allocation. Neither aggregate is one flat pool. The calculation does not prove engine support, kernel correctness, communication performance, attention-state residency, or an accepted result.

Exact product Local memory Published peak local bandwidth Board power envelope Host interface and role
AMD Instinct MI350P 144 GB HBM3E; 128 MB last-level cache 4,000 GB/s 450 W configurable; 600 W maximum TBP PCIe 5.0 x16; passive dual-slot enterprise card
NVIDIA RTX PRO 6000 Blackwell Server Edition 96 GB GDDR7 ECC 1,597 GB/s Configurable up to 600 W PCIe Gen 5; air or liquid server form; MIG up to four instances
NVIDIA RTX PRO 6000 Blackwell Workstation Edition 96 GB GDDR7 ECC 1,792 GB/s 600 W maximum PCIe Gen 5; workstation graphics, media, AI, and local model development
NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition 96 GB GDDR7 ECC 1,792 GB/s 300 W maximum PCIe Gen 5; lower-power workstation configuration
NVIDIA RTX 6000 Ada Generation 48 GB GDDR6 ECC 960 GB/s 300 W maximum PCIe 4.0 x16; earlier Ada workstation product

What the specification sheet can say. MI350P has 48 GB more local memory than RTX PRO 6000 Server and 2.50 times its published peak local-memory bandwidth. Dividing peak bandwidth by a maximum board-power envelope produces 6.67 versus 2.66 GB/s per watt at 600 W. That is only a specification-envelope quotient. It is not measured useful bandwidth, tokens per watt, joules per token, node energy, or application efficiency.

HBM3E, GDDR7, DDR5, and LPDDR5X are different tiers. MI350P HBM3E and RTX PRO 6000 GDDR7 are accelerator-local memory. A conventional PCIe server reaches host DDR through its CPU, root complex, switches, and PCIe path. LPDDR5X belongs only when the named host actually contains it, such as a Grace-based system. Neither compared card includes LPDDR5X. Moving a prefix, weights, or attention state to a host tier changes capacity, but it also adds lookup, serialization, transfer, NUMA, page-pinning, latency, and CPU-memory power boundaries.

One reproducible RTX server result, and why it still cannot answer tokens per watt

[REPRODUCIBLE EXTERNAL RECEIPT, DIFFERENT WORKLOAD] MLPerf Inference v5.1 contains an HPE closed-division submission for one HPE ProLiant Compute DL380a Gen12 with eight RTX PRO 6000 Blackwell Server Edition cards, two Intel Xeon 6787P CPUs, 2,048 GB host memory, TensorRT 10.11.0.33, CUDA 12.9, and driver 575.57.08. Its Llama 2 70B 99% Server result is 26,005.90 tokens/s at the eight-GPU system boundary. It is not a single-card result, not C-001, and the submission does not include measured system power, price, or an accepted-patch denominator. The system JSON also labels accelerator memory as HBM2e, which conflicts with NVIDIA's official GDDR7 specification. The official NVIDIA product page controls the memory type.

No equivalent MI350P result was found that joins the exact model revision, precision, runtime, prompt and output lengths, concurrency, latency target, quality gate, board and node power, and price. A cross-vendor tokens-per-watt claim therefore remains unknown.

YOUR RECEIPT

Calculate capacity, energy, and cost without borrowing a benchmark

1. Capacity-only workload inputs

The defaults reproduce the article's bounded Qwen2.5-Coder-32B BF16 example. They do not describe GLM-5.2, whose low-rank and sparse-attention state must come from the exact runtime allocation.

2. Same-run measured inputs for each platform

Blank means unknown. Use average measured power over the same interval as the token count. Do not enter TBP, TDP, or a maximum-power limit as measured workload power.

AMD Instinct MI350P
RTX PRO 6000 Blackwell Server Edition

Capacity math is ready. Power, energy, transfer, and cost ratios remain unknown until matching measured inputs are entered.

WORKLOAD LOWER BOUND

Weight floor
Calculating
Logical KV
Calculating
Working set plus reserve
Calculating

MI350P

Capacity fit
Calculating
Tokens per measured watt
Unknown
kWh per 1M output tokens
Unknown
Cost per 1M output tokens
Unknown
Accepted tasks/hour
Unknown
Accepted tasks/kWh
Unknown
Cost per accepted task
Unknown

RTX PRO 6000 SERVER

Capacity fit
Calculating
Tokens per measured watt
Unknown
kWh per 1M output tokens
Unknown
Cost per 1M output tokens
Unknown
Accepted tasks/hour
Unknown
Accepted tasks/kWh
Unknown
Cost per accepted task
Unknown

HOST MOVEMENT

One-way transfer time
Unknown
Required proof
Payload bytes, useful GiB/s, direction, NUMA route, p99, CPU and memory power

Primary and reproducible sources: AMD MI350P, MI350P brochure, RTX PRO 6000 Server, RTX PRO 6000 Workstation, RTX PRO 6000 Max-Q, RTX 6000 Ada, and MLPerf Inference v5.1 HPE result.

This table describes installed capability. It does not say which platform wins a coding agent, video generator, RAG service, or training job. The workload still decides how much state remains local, how much crosses a fabric, whether kernels use the available bandwidth, how much power is occupied, and whether the output is accepted.

The two maintained internal vendor registries preserve the full claim history, source links, unknown fields, command paths, and update rules: 2026-07-10-amd-instinct-platform-engineering-ssot.md, 2026-07-10-nvidia-accelerator-platform-engineering-ssot.md, and 2026-07-10-amd-nvidia-platform-ssot-index.md. These internal draft files are evidence routers. The cited primary vendor page still controls if a specification changes.

FIGURE 12A 1600 × 900
Two evidence-labeled paths compare AMD Instinct and NVIDIA accelerator generations, their HBM generations, and whether each number applies to one accelerator, one rack, or a multirack domain.
Current products and preliminary roadmaps are separate evidence classes. Per-accelerator, per-rack, and multirack totals are not interchangeable. Source state. Official current product pages, official preliminary roadmaps, and Touchdown scope normalization. This is an architecture map, not a workload performance ranking.

AMD and NVIDIA expose two native software and fabric paths

The hardware is not interchangeable below the framework. A PyTorch model can look portable while the actual fast path changes at the operator library, compiler target, collective library, topology, network endpoint, and offload layer.

Before tracing those software branches, freeze the CPU-to-accelerator topology. Coherent addressing can make two memory domains easier to program without making their physical latency, bandwidth, capacity, or power identical.

Named system boundary CPU-side memory Accelerator-local memory CPU-to-accelerator path What remains topology-specific
Grace plus Blackwell in a GB200 Superchip Grace LPDDR5X Blackwell HBM3E Coherent NVLink-C2C Actual page placement, remote access, cache behavior, transfer counters, and task latency
Vera plus Rubin Vera LPDDR5X Rubin HBM4 NVIDIA describes coherent second-generation NVLink-C2C Delivered-system topology, placement policy, sustained behavior, and matched workload evidence
Conventional MI350P PCIe server The qualified host's DDR and NUMA domains 144 GB card-local HBM3E PCIe 5.0 x16 OEM slot wiring, NUMA route, peer path, useful transfer rate, airflow, power, and runtime qualification
MI355X OAM or UBB platform The exact qualified host-memory tier 288 GB local HBM3E per accelerator Platform-specific host and Infinity Fabric topology The orderable server, firmware, ROCm, peer, collective, cooling, and failure-domain contract
MI455X and Helios Venice is the CPU and control boundary; the exact installed host-memory configuration belongs to the OEM or ODM system 432 GB HBM4 per MI455X The published platform separates CPU control, GPU scale-up, scale-out networking, and storage boundaries Do not infer a universal CPU-to-MI455X link, coherent placement rule, or useful transfer rate from the rack name alone

The current AMD path is:

TEXT
request, training step, or diffusion step
  -> PyTorch, JAX, vLLM, or SGLang
    -> ROCm and HIP
      -> AITER, Triton, Composable Kernel, hipBLASLt, or rocBLAS
        -> RCCL or rocSHMEM
          -> XCD cache and local HBM
            -> Infinity Fabric scale-up
              -> NIC and scale-out network
                -> CPU DDR, NVMe, object storage, or a version-supported offload tier

The current NVIDIA path is:

TEXT
request, training step, or diffusion step
  -> PyTorch, JAX, TensorRT-LLM, vLLM, or SGLang
    -> serving engine, Dynamo, KVBM, LMCache, or HiCache policy where configured
      -> CUDA, CUTLASS, cuBLASLt, cuDNN, and engine kernels
        -> local HBM for the active working set
          -> NCCL collectives where the operation spans GPUs
            -> NIXL orchestration where a tensor or KV object must move
              -> selected backend such as UCX or libfabric
                -> NVLink, PCIe, or an RDMA path
                  -> ConnectX or BlueField endpoint, switch, and network
                    -> destination HBM, CPU memory, storage, or CMX

[OFFICIAL CURRENT SOFTWARE PATHS, VERSION-SENSITIVE] AMD documents current ROCm architecture targets, AITER, RCCL, vLLM, and SGLang support surfaces for Instinct products. NVIDIA and the named open-source projects document CUDA, NCCL, TensorRT-LLM, vLLM, SGLang, Dynamo, NIXL, LMCache, and the CMX direction. A feature in documentation is not proof that the pinned model, kernel, connector, and hardware combination works. The receipt needs package versions, container digest, GPU architecture, launch command, topology, and logs.

That NVIDIA path is a set of conditional branches, not one mandatory sequence. A local single-GPU operation may touch none of NCCL, NIXL, a NIC, a DPU, or storage. A distributed or offloaded request must name which branch executed.

NIXL is software that coordinates transfer. NVLink is a physical/link scale-up fabric. NCCL is collective software. ConnectX is a NIC family. BlueField is a DPU and infrastructure-processing boundary that incorporates networking functions. LMCache is a KV-cache layer. CMX is an announced context-memory storage platform. On AMD, Infinity Fabric is the current scale-up path for MI300 and MI350 systems. UALink and UALoE belong to the launched Helios production reference design; the exact partner topology, ROCm package set, collective path, and workload behavior still require qualification. Do not label an MI355X UBB as UALink unless the exact orderable platform proves that topology.

Network or infrastructure component Role in the memory path Platform or generation scope Claim boundary
ConnectX-7 NVIDIA NIC generation used in Hopper and early Blackwell-era system paths Exact server and network configuration required A NIC specification does not prove KV transfer, collective behavior, or accepted-task benefit
ConnectX-8 NVIDIA 800 Gb/s-class NIC generation associated with GB300-era systems Exact port, switch, rail, and transport configuration required Peak link rate is not payload goodput or p99
ConnectX-9 NVIDIA NIC generation named for the Vera Rubin platform path Roadmap or delivered-system state must remain explicit Availability does not prove a selected backend or workload result
BlueField-4 DPU and infrastructure processor with ConnectX-class networking functions in the Vera Rubin and CMX direction Named platform and software stack required It is not merely a NIC and does not become more local HBM
Pollara 400 AMD Pensando AI NIC for existing Ethernet-cluster paths Exact switch, ROCm, transport, and server configuration required The product role does not prove that a workload used it
Vulcano 800 AMD Pensando scale-out and scale-across NIC direction for Helios Helios partner topology and delivery state required Vendor link ceilings are not collective latency or application goodput
Salina AMD Pensando DPU for infrastructure services Selected system and software role required A DPU is not an AI NIC, scale-up link, or local memory tier
FIGURE 12B 1600 × 900
Parallel AMD and NVIDIA software-to-hardware paths show different runtimes, kernel libraries, collectives, HBM, scale-up links, scale-out networks, and offload layers.
One framework API can hide different libraries, fabrics, failure boundaries, and validation work. Source state. Official project and vendor documentation, normalized by layer. Every arrow remains a versioned compatibility and measurement boundary.

Keep current products separate from roadmap systems

Current orderable or production-described products and future systems answer different questions.

Current product boundary Memory and fabric fact What still requires a workload receipt
AMD MI300X 192 GB HBM3 and 5.3 TB/s per OAM; eight-GPU Infinity Fabric platform direction sustained kernel bandwidth, RCCL topology, engine support, accepted output
AMD MI325X 256 GB HBM3E and 6.0 TB/s per OAM; same gfx942 architecture family whether capacity avoids another shard or only increases idle memory
AMD MI350P 144 GB HBM3E and 4.0 TB/s per PCIe card; 450 W configurable and 600 W maximum TBP; passive double-slot server form OEM qualification, slot and NUMA topology, peer path, airflow, node power, useful HBM traffic, accepted output
AMD MI350X / MI355X 288 GB HBM3E and 8.0 TB/s per OAM; gfx950; different 1,000 W and 1,400 W product envelopes cooling boundary, native kernel coverage, delivered throughput per watt
AMD MI455X / Helios 432 GB HBM4 and 23.3 TB/s peak per accelerator; 31 TB and 1.7 PB/s aggregate peak across the 72-GPU production reference design released gfx1250 qualification, selected operator and HSACO, useful HBM traffic, partner configuration, power, accepted output, and run_id
NVIDIA H100 / H200 Hopper with 80 GB HBM3 or 141 GB HBM3e; NVLink 4 whether added H200 capacity changes parallelism, KV residency, and cost
NVIDIA GH200 Hopper HBM plus coherent Grace LPDDR5X through NVLink-C2C page residency, remote access, offload overlap, TTFT and TPOT
NVIDIA B200 GPU Blackwell GPU product boundary with HBM3E and NVLink 5 installed memory configuration, selected kernel, useful bandwidth, power, and accepted result
NVIDIA HGX B200 Eight B200 GPUs in an HGX baseboard and server-platform boundary exact OEM server, CPU, host memory, NICs, cooling, topology, collective path, and task result
NVIDIA DGX B200 Complete eight-B200 NVIDIA system boundary delivered software image, operating configuration, node power, workload throughput, and accepted-work economics
NVIDIA GB200 Superchip One Grace CPU plus two Blackwell GPUs connected through NVLink-C2C placement across separate LPDDR5X and HBM3E domains, remote-access behavior, and matched workload result
NVIDIA GB200 NVL72 36 Grace CPUs and 72 Blackwell GPUs in a rack-scale NVLink domain useful HBM and NVLink bandwidth, collective wait, rack utilization, facility power, and accepted-work economics
NVIDIA B300 / Blackwell Ultra GPU-architecture and product boundary with up to 288 GB HBM3E installed product configuration, sustained behavior, power, and accepted result
NVIDIA GB300 NVL72 Blackwell Ultra rack boundary with approximately 20 TB aggregate GPU memory reasoning or video result for the exact model, precision, quality gate, topology, and power boundary

[TOUCHDOWN STRATEGY EXTRAPOLATION, NOT A BENCHMARK] The design signals become useful only when their pros and cons stay attached to the workload:

Platform choice Plausible use case to test Design advantage Trade-off or failure mode Strategy signal, not a forecast
RTX 6000 Ada or RTX PRO 6000 Blackwell workstation variants Local model development, graphics, visualization, media, CAD, and bounded inference One card combines graphics, media, and AI functions with a mature workstation software path 48 or 96 GB local memory can force smaller models, more quantization, more sharding, or more host movement A mixed graphics and AI workload can value functions that a pure HBM capacity comparison ignores.
RTX PRO 6000 Blackwell Server Edition Enterprise PCIe inference where 96 GB per card fits the selected shard and NVIDIA software is important Familiar server form, GDDR7, MIG, and NVIDIA's deployment and inference stack Less local capacity and lower published local bandwidth than MI350P; distributed state can amplify movement and communication Software and deployment integration can compensate for a weaker paper memory ratio only if accepted-task throughput and operations prove it.
AMD MI350P Air-cooled PCIe inference, private RAG, agent services, and incremental enterprise deployment where 144 GB local HBM changes the shard plan More local memory and published bandwidth than the compared RTX PRO server card in a conventional PCIe form Passive 600 W maximum card, OEM airflow, PCIe topology, ROCm versioning, kernel coverage, and peer behavior all require qualification AMD can compete by making larger local HBM available without forcing the buyer directly into an OAM baseboard or liquid-cooled rack.
NVIDIA B200 or B300 systems Large training, inference, and reasoning workloads that benefit from an integrated NVIDIA server and software boundary Qualified system architecture, NVLink, mature libraries, management, and broad deployment support High system and facility commitment, vendor-specific integration, and no guarantee that the workload uses peak HBM or compute Integration reduces time and qualification risk, which can be worth more than an isolated component ratio.
AMD MI355X systems Memory-heavy training or inference where 288 GB and 8 TB/s per OAM reduce shards or preserve more hot state locally Large HBM3E capacity and bandwidth per accelerator with an eight-GPU platform path 1,400 W product envelope, liquid cooling, exact server topology, ROCm and operator coverage, collectives, and support are workload-specific More state per accelerator can reduce communication or increase batching, but only if the software uses it and the facility can sustain it.
NVIDIA GB200 or GB300 NVL72 Rack-scale training, inference, and reasoning that benefit from Grace CPUs, LPDDR5X, coherent CPU-GPU access, NVLink, and one integrated system CPU memory, HBM, scale-up, networking, software, and operations are designed together The rack is a large power, cooling, networking, and operational boundary. Coherence does not erase placement cost. NVIDIA is selling a memory and data-movement system, not only GPUs. The value must still close at the accepted workload.
AMD MI455X and Helios Frontier training and inference where 432 GB HBM4 per accelerator and a 72-GPU open reference design change model or state placement Very large published local and rack HBM capacity and bandwidth, plus an open OEM or ODM system direction Exact partner configuration, released software qualification, useful HBM and UALoE behavior, rack power, cooling, reliability, and cost remain unproved here AMD is attacking the rack-scale memory boundary directly. The strategic question is whether open system choice and HBM scale convert into faster qualification and accepted-work economics.

The roadmap stays separate:

Roadmap system Source-backed statement Unknown that must remain unknown
AMD MI430X AMD names 432 GB HBM4 and 19.6 TB/s in the product context final package, power, broad availability, delivered application result
AMD MI440X AMD names an eight-GPU enterprise deployment role final HBM, bandwidth, fabric, power, price
AMD MI450 Series / MI450X and other MI400 members AMD uses source-specific family and architecture labels alongside the launched MI455X whether those labels are equivalent to MI455X, final member taxonomy, member-specific configuration, qualification, and delivered performance
AMD MI500 family AMD names CDNA 6, 2 nm, HBM4E, and a planned 2027 generation capacity, bandwidth, package, power, scale-up topology, shipping system
NVIDIA Vera Rubin NVL72 NVIDIA says the platform is in production while its published specification table remains explicitly preliminary final non-preliminary values and customer-specific availability
NVIDIA Rubin Ultra NVL576 eight 72-GPU MGX NVL racks in one announced scale-up domain final Rubin Ultra GPU, HBM, power, delivered multirack behavior
NVIDIA Rubin Ultra Kyber NVL144 one 144-GPU next-generation MGX NVL rack PCB qualification, yield, rack power, final GPU specification, customer schedule
NVIDIA Feynman Kyber NVL1152 eight 144-GPU Kyber racks in one announced domain Feynman GPU, HBM, process, power, availability, workload result

The MI450 naming boundary matters. AMD's own current public material uses MI450 Series, MI450 architecture, MI450X, and MI455X in different official contexts. The relationship is [NOT PUBLIC] until AMD publishes a final taxonomy. The safe move is to retain the exact source label, not normalize all four into one fictional SKU.

The Rubin boundary matters for the same reason. A current NVIDIA statement that the Vera Rubin platform is in production can coexist with a product table marked preliminary. Production status and specification finality are separate fields.

FIGURE 12C 1600 × 900
Four topology diagrams distinguish one 72-GPU Helios rack, an eight-rack Rubin Ultra NVL576 domain, one 144-GPU Kyber rack, and an eight-Kyber-rack Feynman NVL1152 domain.
A GPU count names a scale-up domain, not necessarily one rack. Helios is a production OEM/ODM reference design; the NVIDIA Rubin Ultra, Kyber, and Feynman configurations remain preliminary. Source state. Official preliminary vendor topology. Exact PCB geometry, yield, supplier allocation, customer schedule, and delivered workload behavior remain not public.

A native-support claim needs more than an import

The complaint that AMD software has not caught up is too vague to be useful. It mixes six separate questions: Can the framework express the workload? Does the compiler target the exact GPU? Does a tuned operator exist for the exact shape and precision? Do collectives match the installed topology? Can the serving and placement layers move state correctly? Can an operator deploy, observe, recover, and get support for the whole configuration?

AMD has made concrete progress. ROCm and HIP expose a broad framework and programming path. PyTorch, JAX, vLLM, and SGLang have AMD surfaces. AITER, Composable Kernel, Triton, hipBLASLt, rocBLAS, MIOpen, RCCL, rocSHMEM, and MoRI cover different operator, compiler, collective, and communication jobs. ROCm profiling and system-management tools expose additional diagnostics. AMD also publishes open source and source-level architecture signals that make the stack inspectable.

NVIDIA's current advantage is not captured by saying CUDA has more libraries. The stronger claim is integration depth across qualified hardware, CUDA, mature operator coverage, compiler and engine paths, NCCL, TensorRT-LLM, Dynamo, NIXL, observability, OEM systems, deployment tooling, and support. Even that advantage is workload-specific. A library existing in CUDA does not prove that the selected model, precision, kernel, topology, and memory policy are correct or economical.

Software and operations layer AMD current advancement NVIDIA current integration What is current, projected, or still unproved
Framework and serving surface ROCm, HIP, PyTorch, JAX, vLLM, SGLang, and AMD Inference Microservices expose supported paths on named current products CUDA frameworks, TensorRT-LLM, vLLM, SGLang, Dynamo, and NVIDIA AI Enterprise expose multiple supported paths on named current products Package presence is current architecture evidence. Exact model revision, engine version, precision, and accepted output still require a run.
Operators and kernels AITER, Composable Kernel, Triton, hipBLASLt, rocBLAS, MIOpen, and other domain libraries provide different tuned or lower-level paths CUTLASS, cuBLASLt, cuDNN, TensorRT kernels, Triton, and engine-specific kernels provide different tuned or lower-level paths Coverage is shape, datatype, architecture, and version specific. A fallback kernel can run correctly while leaving much of the memory or compute ceiling unused.
Compiler and device code HIP, LLVM, COMGR, HSACO code objects, ROCr/HSA, and amdgpu/KFD form the current AMD lowering and dispatch path CUDA compilation, PTX, SASS, cubins, the CUDA driver, and architecture-specific device code form the current NVIDIA path MI455X gfx1250 source bring-up signals are not a released qualification packet. A target string is not proof of selected device code or dispatch.
Collectives and scale-up RCCL, rocSHMEM, MoRI, Infinity Fabric, and the launched UALink or UALoE Helios direction cover different communication layers NCCL, NVLink, NVSwitch, and the named rack topology cover different communication layers A collective library is not a fabric. The receipt needs the exact graph, message sizes, topology, p50, p95, p99, retries, errors, and useful overlap.
Placement, KV, and external memory vLLM or SGLang policy, LMCache, host memory, VAST, NFS over RDMA, and storage can form a remote-state path when configured Dynamo, KVBM, LMCache or HiCache, NIXL, host memory, BlueField, and the announced CMX path can form different remote-state paths when configured VAST's MI355X result is a vendor-reported named benchmark. CMX is an announced direction. Neither proves lower total energy or cost for C-001 or V-001.
NICs, DPUs, and specialized infrastructure compute Pollara or Vulcano NICs and Salina DPUs can own named transport and infrastructure functions ConnectX NICs and BlueField DPUs can own named transport and infrastructure functions These components can remove work from the host CPU or change the movement boundary. They do not increase local HBM, become the scale-up fabric, or prove an application gain by existing.
Deployment and operations ROCm containers, operators, system management, profiling, RAS, OEM qualification, and support matrices are moving the stack toward repeatable deployment NVIDIA's software distribution, management, diagnostics, OEM qualification, and support integration are broader on many deployed configurations Operational maturity must be scored per configuration: install, upgrade, observability, failure isolation, rollback, replacement, security, support, and accepted workload.

Current versus projected software must stay separate. The current MI350P and MI355X rows can be evaluated against released ROCm packages and qualified OEM systems. The launched MI455X and Helios hardware direction has public source bring-up and vendor day-zero support language, but the article has no released, joined MI455X workload packet. Current Blackwell systems can be evaluated against released NVIDIA software and delivered configurations. Vera Rubin, BlueField-4 STX, and CMX contain current vendor architecture statements alongside preliminary, announced, or configuration-dependent fields. Neither vendor earns a future workload result from a current diagram.

The useful buyer question is therefore not whether AMD software is good or bad. It is: for this model, operation, precision, server, topology, and service-level objective, which exact path is qualified, what falls back, what can be observed, what fails, and who fixes it?

For either vendor, call a workload native_supported only when this packet exists:

YAML
model:
  id: null
  revision: null
  tokenizer_revision: null
runtime:
  framework: null
  engine: null
  engine_version: null
  container_digest: null
hardware:
  vendor: AMD_or_NVIDIA
  exact_product: null
  architecture_target: null
  hbm_capacity_gb: null
topology:
  scale_up: null
  scale_out: null
  gpu_ids: []
execution:
  precision: null
  tp: null
  pp: null
  ep: null
  cp_or_sp: null
receipts:
  launch_log: null
  topology_snapshot: null
  profiler_trace: null
  collective_test: null
  accepted_output: null
  power_log: null
metrics:
  throughput: null
  p95_latency: null
  p99_latency: null
  accepted_tasks_per_hour: null
  cost_per_accepted_task: null

An installed ROCm or CUDA package is parser_only evidence for the environment. A synthetic fixture can be fixture_backed. A real model that completes with captured topology and diagnostic artifacts is live_validated. Release readiness adds documentation, failure modes, supported versions, packaging, CI, and rollback. That state model prevents a software import from becoming a false claim of full AMD or NVIDIA coverage.

The generational path exposes several different engineering bets:

  1. H100 to H200 increases local memory capacity and bandwidth while keeping the Hopper compute generation. This is a clean example of memory changing the usable workload envelope without changing every compute block.
  2. GB200 makes the unit of discussion a 72-GPU liquid-cooled rack with Grace CPUs, LPDDR5X, NVSwitch, networking, power shelves, and facility requirements. HBM is now one tier inside a rack computer.
  3. GB300 increases HBM capacity and targets reasoning and attention-heavy use, but the application still needs scheduling and placement that turn capacity into goodput.
  4. Vera Rubin moves to HBM4 and doubles the published rack-level NVLink bandwidth relative to Blackwell NVL72. It also introduces a new CPU, networking, DPU, storage, power, and cooling context. A GPU-only comparison misses the system change.
  5. Kyber changes the physical rack and scale-up domain. Doubling GPUs per rack increases the number of local placement choices, collective participants, links, failure combinations, power-delivery constraints, and service operations. It does not make 144 HBM pools one uniform load/store address space for arbitrary code.
  6. AMD MI355X emphasizes very large HBM3E capacity per OAM and an eight-GPU UBB path. MI455X is launched, and Helios is an AMD production 72-GPU OEM/ODM reference design with HBM4 and UALink/UALoE. AMD says partner-system shipments are scheduled to start at the end of Q3 and ramp into Q4 2026. MI500 names the next memory generation but still lacks a public production specification.

For procurement or architecture, the row needs these additional fields before a decision:

TEXT
exact orderable SKU and availability
GPU and HBM lot / revision
usable capacity after runtime reservation
sustained bandwidth for the named kernel mix
scale-up and scale-out topology
collective p50 / p95 / p99 by message size
board, tray, rack, and facility power boundaries
coolant and water accounting boundaries
software versions and supported models
accepted-workload throughput, latency, quality, and failure rate
price, support, deployment labor, spares, and utilization

Custom HBM and SPHBM4: more control, more obligation

[STANDARD] Standard Package High Bandwidth Memory (SPHBM4) is JEDEC JESD330-4. Its public record describes an HBM4-stack direction with a different buffer/interface die for standard-package integration. The standard does not by itself prove a shipping product or workload result. This article links the unrestricted public record but does not reproduce standard text or tables.

[VENDOR ROADMAP / DESIGN SURFACE] Marvell publicly describes a custom-HBM compute architecture. The significance is not a fixed partition. It is that a system designer may be able to move selected controller, security, RAS, telemetry, data-movement, or near-memory functions into the base/interface layer.

That possibility creates a full engineering burden:

  • What exact function moves, and why is the base die better than the accelerator or software?
  • What traffic does it avoid at which boundary?
  • How does software invoke it and fall back?
  • What is the added logic area, power, thermal density, and verification cost?
  • What changes in test, repair, fault containment, and field update?
  • Who supplies and qualifies the design?
  • What non-recurring engineering and volume make it economic?

No public source reviewed for this article proves that Touchdown can order custom HBM, that a startup gets access on useful terms, or that one custom partition wins. Those remain diligence questions.

Intel XBM, claim by claim

[PATENT APPLICATION] The source calls it XBM, expanded in the description as cross-batch memory. It is not XBEM, and this document does not call it extreme bandwidth memory.

The public record is US 2026/0191095 A1, application 19/001,921, titled Ultra High Bandwidth Memory With Backend Transistors. The application was filed December 26, 2024 and published July 2, 2026.

Intel did not launch an XBM product in this publication. A patent application can still matter because it shows architectural territory the applicant considered worth describing and claiming. It does not show that the device works, a package yields, software exists, customers can buy it, or economics beat HBM4.

The filing describes several coupled ideas:

Ingredient What the application describes Why it is interesting What remains unknown
Backend-transistor 1T1C DRAM Memory dies with DRAM devices using backend transistors Changes cell and interconnect integration space Density, retention, speed, leakage, variability, thermal budget, yield
Fine subchannels and datablock regions Smaller independently useful organizational regions May expose request parallelism or reduce unnecessary activation Controller efficiency, queueing, real workload use
TSV gutters Via regions between datablock regions Couples vertical delivery to floorplanning Area, stress, routing, heat, yield
Base-die options Embodiments with base logic and different mounting organizations Opens PHY, routing, test, repair, package control Which embodiment, if any, becomes a product
UCIe-facing links Serialized die-to-die package links, including a 32 GT/s example Separates internal memory parallelism from package-facing serialization Lane count, payload efficiency, latency, power, generation, interoperability
BIST, redundancy, repair Per-die test and spare subchannel/datablock examples Makes manufacturing recovery part of architecture Coverage, repair granularity, area, performance, qualification
Package variants Interposer and other package organizations in figures Broadens integration choices Signal integrity, PDN, heat, assembly, cost

The key distinction is that serialized outside does not mean serial inside. The memory can retain substantial internal parallelism while a base die or distributed logic arbitrates traffic onto high-speed package-facing links.

That creates a raw bandwidth check. The original public HBM4 reference point used in the existing research is 8 Gb/s across 2,048 bits:

TEXT
8e9 transfers/s * 2,048 bits / 8
= 2.048 TB/s raw

one ideal 32 GT/s lane-equivalent
= 32 Gb/s
= 4 GB/s raw

2.048 TB/s / 4 GB/s
= 512 ideal raw lane-equivalents

[TOUCHDOWN DERIVATION] The 512 result assumes one raw bit per transfer, one direction, and 100 percent efficiency. It is not an Intel physical lane count, not a UCIe payload result, and not a product topology. The filing includes a 64 by 32 GT/s callout in one figure, while total bundles, duplex behavior, protocol overhead, retry, arbitration, and sustained application bandwidth remain unknown.

The 32 GT/s filing example is also not a current UCIe ceiling. [STANDARD] The UCIe Consortium's public specifications page lists later 48 and 64 GT/s capability in UCIe 3.0. The architecture must treat link rate as a versioned target parameter, not a frozen identity.

The real questions are harder than raw division:

  • How many physical lanes and bundles operate in each direction?
  • What fraction becomes payload after framing, flow control, retry, repair, and idle cycles?
  • What first-byte and tail latency come from serialization, arbitration, and turnarounds?
  • What is energy per useful transferred bit at the link and system boundaries?
  • Can base logic sustain mixed reads and writes without becoming the bottleneck?
  • How do backend transistors behave across retention, leakage, variation, temperature, aging, and test?
  • What package cost is removed or added in the chosen embodiment?
  • How well do BIST, redundancy, and repair improve usable yield?

Do not say XBM eliminates an interposer. One filing figure uses an interposer. The safe claim is that described embodiments create multiple package choices, some of which may reduce dependence on a conventional HBM-style interposer.

Do not publish the filing's illustrative capacity language as a committed product. Do not publish bandwidth, latency, energy, thermals, yield, cost, qualification, availability, or schedule as measured. Those fields are null until evidence exists.

YAML
# [PATENT APPLICATION] Claim boundary
artifact: US-2026-0191095-A1
described_architecture: source_backed
measured_silicon: false
product_commitment: false
bandwidth_latency_energy_yield_cost: null
touchdown_analysis:
  raw_lane_equivalence_only: true
  physical_lane_count: null

XBM is important because it makes memory cells, internal organization, controller behavior, repair, package links, and topology one architecture search space. It is not important because it has already beaten HBM. It has not established that.

GDDR7, DDR5, and LPDDR5X: DRAM is not one interface

GDDR7 uses discrete graphics-memory devices and high per-pin signaling to provide board-level graphics bandwidth. Vendor pages describe PAM3 signaling and named part data rates. [VENDOR PRODUCT] Those figures are useful for a named device and bus width. A system comparison still needs device count, aggregate interface width, board routing, power, capacity, controller, and workload.

GDDR can be attractive when a design values a different package and cost structure from HBM. Its narrower aggregate interface and higher per-pin rate change signal integrity, power, board area, and controller choices. HBM is faster or GDDR is cheaper without a complete system is too high level.

DDR5 is the host-memory baseline for many servers. DIMMs provide large, serviceable CPU-visible capacity through the host memory controller. The path from a GPU to host DDR can cross an accelerator link, PCIe or another fabric, CPU root complex, NUMA topology, and memory controller. The capacity can be useful while the movement deadline still fails.

LPDDR5X is designed around lower-power I/O and capacity-density trade-offs and is widely associated with mobile and client systems. It also appears in data-center design discussions because power and memory capacity matter there. LPDDR is not a drop-in HBM or DDR DIMM replacement. It changes packaging, soldered-down serviceability, controller support, channel organization, capacity, bandwidth, and system architecture.

The comparison is therefore role-specific:

  • HBM favors the widest, closest accelerator-local hot tier.
  • GDDR favors high board-level graphics bandwidth under a different package and per-pin trade.
  • DDR favors large, serviceable host working memory.
  • LPDDR favors power and capacity-density directions under a different integration model.

CXL memory: capacity with a fabric contract

[STANDARD] The CXL Consortium's unrestricted public overview describes a cache-coherent interconnect built on PCIe physical infrastructure and capabilities for processors, accelerators, and memory expansion. Different device types, generations, switches, and system topologies expose different functions.

CXL-attached memory can expand addressable capacity, support pooling or sharing in applicable systems, and create another software placement option. It does not become HBM by joining the address space.

[VENDOR PRODUCT ARTIFACT] Samsung's MD210 CMM-D product page lists a named CXL memory product with Mass Production status. That establishes a vendor product state, not a workload result or a universal CXL latency and bandwidth profile.

The path still has link serialization, switch or topology effects, controller queues, coherency behavior, media access, contention, RAS, and software policy. The first byte and full object can arrive later than local HBM. Whether that is acceptable depends on the object and lookahead.

Warm tier is a workload-placement description in this article. It is not a CXL standards category.

We deliberately do not process, quote, or summarize the CXL 4.0 evaluation-specification download because its agreement restricts AI processing and use. We use the consortium's unrestricted public overview and announcements. We also do not reproduce JEDEC standard tables.

NVMe, NAND, and HBF: capacity changes granularity

NVMe exposes block storage over an efficient host/controller protocol. The underlying NAND can hold large durable state without power. It is the natural home for repositories, checkpoints, model files, cold cache, logs, and outputs.

It is not a transparent HBM extension. The path can include a filesystem, page cache, block layer, NVMe queue, SSD controller, flash translation layer, NAND page read, error correction, DMA, and destination allocation. Direct-storage paths can remove some CPU copies, but they do not erase first-byte latency, page granularity, contention, or endurance.

[VENDOR TARGET / BENCHMARK] Sandisk's HBF fact sheet proposes a high-bandwidth flash direction and reports projected capacity/bandwidth plus internal Llama simulation. Those are vendor claims. [STANDARDIZATION WORKSTREAM] SK hynix and Sandisk separately announced an OCP standardization effort. Neither establishes shipping HBF.

[PHYSICAL PROTOTYPE] Kioxia announced a separate high-bandwidth flash module prototype with its own capacity and bandwidth boundary. That is evidence of a physical prototype. It is not the Sandisk HBF product, and it is not an AI workload result.

The right HBF question is not Can flash replace HBM? It is Can a predictable, read-mostly, page-useful state stream arrive ahead of demand while a smaller hot tier holds mutable working state?

Emerging memories: candidates need a state role

PCM, ReRAM/RRAM, MRAM, FeFET, gain-cell, oxide-semiconductor and related research families offer different combinations of density, retention, write cost, endurance, analog behavior, and process integration.

That design space is real. It is also easy to misuse. A cell paper can measure one device under one condition without proving an array, controller, ECC, package, compiler, yield, or product. A nonvolatile bit can still have an expensive write. A dense cell can still have variability. A compute-in-memory primitive can still lose on data conversion, routing, precision, programming, or utilization.

The article therefore does not choose a Touchdown material stack. A candidate enters only with a named state role, device/array evidence, controller and software contract, thermal/process path, and merchant baseline.

FIGURE 7 1600 × 900
Broad qualitative regions place SRAM, eDRAM, HBM, GDDR, DDR, LPDDR, CXL-attached memory, NVMe/NAND, HBF, and recomputation by hotness and access granularity on one axis and capacity and distance from compute on the other. Evidence badges distinguish shipping, prototype, patent, roadmap, and research states.
Memory technologies occupy different state roles. Position depends on a named implementation and workload; the regions are qualitative, not a universal performance ranking. Source state. Standards, vendor artifacts, patent, papers, and Touchdown synthesis as labeled.

What this means by role. The CEO chooses a user outcome, not a medium. The CFO compares complete system and migration cost. The CTO uses HBM as the tuned baseline and tests lower or custom tiers one state role at a time. The hardware engineer rejects performance claims without package, controller, and evidence state. The software engineer must make placement and fallback explicit.

When should state move through HBM, host memory, CXL, or NVMe?

In plain English. Moving data is never free. Offload wins only when the HBM capacity it releases is more valuable than the transfer time, energy, extra software, failure risk, and tail-latency penalty it introduces. The object must also return before the next operation needs it.

A larger lower tier helps only when the system can predict, move, reuse, and retire state within the deadline. Capacity without a movement schedule is not useful capacity.

The minimum transfer model

For an object of bytes, the physical lower bound is not just size divided by a headline link rate:

TEXT
transfer_time_lower_bound =
  first_byte_latency
  + bytes / sustained_link_bandwidth

That still omits queueing, page faults, DMA setup, source/destination allocation, topology, contention, coherency, decompression, synchronization, and tail behavior. A useful receipt reports first byte and full completion at p50, p95, and p99 for the relevant direction and object size.

The placement decision compares complete paths:

TEXT
placement_benefit =
  avoided_recompute_time
  + avoided_hot_tier_occupancy
  + queueing_relief
  - transfer_time
  - movement_energy
  - miss_and_retry_cost
  - correctness_and_reliability_risk

[PROPOSED MODEL] Some terms are measured in milliseconds or bytes, others need a cost or constraint model. The equation is a decision checklist, not dimensionally valid addition until every term is normalized to a common objective.

Keep, prefetch, stream, spill, compress, or recompute

Keep state in HBM when reuse is near, mutation is active, or restore cannot meet the deadline. Active KV and a video latent are default examples, subject to a real trace.

Prefetch when access order is predictable enough to start movement before demand. The prefetch needs a confidence estimate, completion deadline, capacity reservation, cancellation path, and miss fallback.

Stream when large read-only regions are consumed in order. Double buffering can let compute consume one window while the next arrives. The schedule fails if lookahead shrinks, page usefulness falls, or tail latency exceeds the compute window.

Spill and restore when an inactive object is costly to recompute and likely to return. This can free HBM but adds write and read traffic.

Compress when fewer bytes justify encode/decode work and any precision loss. Rematerialize when recompute is cheaper than storage plus restore. Replicate when fan-out or locality beats capacity cost. Evict when expected reuse value falls below occupancy cost.

The fallback must remain correct:

PYTHON
# [PROPOSED] Illustrative policy, not a production scheduler.
if obj.mutable and obj.next_use_ms < restore_p99_ms:
    keep("gpu_hbm")
elif obj.read_mostly and schedule_confidence >= required_confidence:
    prefetch(target="gpu_hbm", source=best_measured_lower_tier)
elif obj.recompute_ms < restore_p99_ms:
    rematerialize()
else:
    run_correct_fallback_and_record_miss()

A real scheduler also accounts for capacity pressure, concurrent transfers, topology, bandwidth sharing, energy, fan-out, reliability, versioning, and quality.

CXL placement is still software policy

CXL can make additional memory coherently addressable or shareable under supported system designs. That reduces some software friction. It does not remove physical distance.

For a stable agent prefix, CXL-attached capacity might avoid recomputing prefill if the prefix is reused and restore meets time-to-first-token. For a mutable Wan2.2 latent used every step, the same tier may be a poor default because the object is read and written too frequently. For cold model weights, it may help model load or staging but still fail steady-state layer deadlines.

The evaluation must record the exact CPU, root complex, switch, CXL device, NUMA placement, firmware, kernel, allocator, link state, concurrency, and error behavior. CXL enabled is not an experiment description.

Page-streaming needs useful pages

For flash-class storage, define page usefulness:

TEXT
useful_fraction =
  bytes_consumed_before_eviction / bytes_fetched

effective_useful_bandwidth =
  physical_read_bandwidth * useful_fraction

[TOUCHDOWN DERIVATION] This ignores queue and compute overlap but exposes overfetch. If software reads 1 MiB and consumes 64 KiB before discarding the page, the physical interface can be busy while useful bandwidth is only a fraction of the headline result.

Page streaming becomes plausible when access is predictable, pages are mostly consumed, queue depth is sufficient, first-byte and tail latency fit the lookahead window, double buffers fit in the hot tier, and writes are rare. It fails for frequent fine-grained mutation, unpredictable branches, sparse low-use pages, or a missing correct fallback.

FIGURE 8 1600 × 900
HBM, host DRAM, optional CXL-attached memory, NVMe/NAND or HBF, and remote storage are aligned with a timeline showing prefetch start, first byte, object completion, compute consumption, eviction, and a p99 deadline.
A lower tier is useful only when state arrives before consumption and the avoided work or hot-tier occupancy exceeds movement and miss cost. Source state. Touchdown-derived movement model. No unmeasured tier is assigned a latency.

What do energy, heat, materials, yield, and supply do to the economics?

In plain English. Every useful byte begins in a manufactured device, consumes electrical work when it moves, adds heat that must leave the package, and occupies infrastructure that someone pays for. This section keeps five ledgers separate: electricity entering the system, IT work performed, heat transported, cooling work and water, and money allocated to an accepted task.

HBM is often described with an interface number and a stack capacity. A data center has to buy, power, cool, test, replace, schedule, and keep the whole system useful.

That means following three connected paths:

TEXT
materials and qualified processes
-> dies, stack, package, board, server, rack

electrical events
-> heat, cooling, throttling, facility energy

usable capacity and throughput
-> accepted tasks, failures, margin, trust

Material science by function, not by buzzword

The exact bill of materials is vendor, generation, node, and package specific. The safe way to teach it is by function.

Silicon and device structures. The DRAM array begins on a semiconductor wafer with devices, isolation, contacts, and memory capacitors formed through repeated deposition, lithography, etch, implantation or doping, cleaning, annealing, and planarization steps. The logic base die has its own transistor and interconnect process. A high-performance logic process and a dense DRAM process optimize different things, which is one reason separate dies are useful.

DRAM capacitor. A capacitor needs two conductive electrodes separated by a dielectric. As cell footprint shrinks, the structure must preserve enough effective capacitance and low enough leakage to maintain a sensing margin. High-permittivity dielectric materials, three-dimensional capacitor geometry, electrode interfaces, defects, contamination, and process uniformity matter. It would be inaccurate to assign one dielectric or electrode stack to every HBM vendor without a named process source.

Transistor and contacts. The access device must switch the cell onto the bitline while limiting leakage when off. Threshold variation, mobility, junction leakage, contact resistance, temperature, disturbance, and aging influence retention and timing. Intel's XBM application makes backend-transistor integration part of its described search space, but it does not publish measured arrays that resolve those questions.

On-die interconnect. Local and global wires use conductors, barrier or liner structures, and insulating dielectrics. As dimensions narrow, resistance and reliability become harder. Capacitance between wires affects delay and energy. Copper remains important in semiconductor and package wiring, while tungsten, cobalt, ruthenium, and other materials appear in specific process roles and research. An Imec 16 nm-pitch ruthenium test structure is evidence about that structure, not proof of an HBM vendor's complete wiring stack. [RESEARCH TEST STRUCTURE]

TSV. A TSV needs an etched opening, an insulating liner where required, barrier/seed and conductive fill choices, contacts, and a wafer-thinning or reveal process. Copper is a common functional example, but the article does not treat one material sequence as universal. Mechanical stress from the via and thermal expansion can affect neighboring devices and keep-out zones.

Die bond. Microbumps can use solder and under-bump metallization to connect die pads. Direct or hybrid bonding can connect fine-pitch metal/dielectric surfaces under another process flow. Bond pitch, surface planarity, contamination, voids, alignment, thermal cycling, and current density affect yield and reliability.

Interposer, redistribution, and substrate. A silicon interposer can provide very dense wiring. Redistribution layers and bridges offer other routes. The organic substrate uses dielectric layers, copper traces, vias, solder connections, and reinforcement to connect the dense package to the board. Glass substrates are an industry development direction, not a universal shipping HBM assumption.

Thermal materials. Thermal interface material, heat spreaders, cold plates, solders or attach layers, and coolant or airflow form the heat-removal path. Thermal conductivity is not the only variable. Interface quality, thickness, pressure, pump/fan power, corrosion, reliability, serviceability, and the location of heat sources matter.

This functional view explains why raw materials is only the first supply layer. A data center cannot buy copper, silicon, and dielectric powder and obtain HBM. It needs ultra-pure materials, qualified recipes, process tools, masks, wafers, design IP, inspection, metrology, TSV and bonding capacity, substrates, assembly, test, firmware, controller validation, cooling, and field support.

From wafer to a qualified HBM package

A simplified dependency chain is:

TEXT
memory and logic design
-> masks and qualified wafer processes
-> DRAM core dies and base logic dies
-> wafer probe and known-good-die screening
-> TSV, thinning, reveal, singulation, and preparation as applicable
-> die stacking and bonding
-> stack test, redundancy, and repair
-> accelerator plus memory package assembly
-> package test and reliability qualification
-> board, firmware, system, and cooling validation
-> production deployment and field RAS

SK hynix's public wafer-level package process overview is one source for the functional core/base-die, known-good-die, bump, stacking, and package sequence. It is not a universal vendor recipe. [OFFICIAL VENDOR EDUCATION]

Memory vendors, logic foundries, OSATs, substrate suppliers, equipment makers, materials companies, accelerator designers, board manufacturers, server companies, and cloud operators can each own different boundaries. Vertical integration changes the contractual map but does not eliminate the tasks.

Supply risk can enter at every point:

  • Good DRAM wafers exist, but advanced packaging capacity is constrained.
  • Stacks are available, but a logic die or accelerator limits packages.
  • Packages exist, but substrates, boards, power delivery, or liquid cooling limit servers.
  • Servers exist, but the software cannot use HBM efficiently.
  • Capacity is installed, but reliability or acceptance failures reduce useful output.

That is why an HBM forecast cannot be translated directly into AI revenue without the middle path.

Touchdown's Computex 2026 physical-layer field guide shows the same hot, warm, and cold state problem at rack, power, cooling, and service boundaries. Here, the path is narrowed to HBM and its alternatives.

Where the electricity goes

At the cell and array level, energy appears in activation, sensing, restoration, precharge, read/write movement, and refresh. At the die and stack level, energy appears in internal buses, TSVs, clocks, I/O, ECC, BIST, and control. At the package and system level, it appears in PHYs, links, voltage conversion, board loss, accelerator compute, fans, pumps, and facility overhead.

A useful event model is:

TEXT
memory_dynamic_energy_proxy =
  sum(event_count_i * energy_per_event_i)

[MODELED] This is only as good as its counters and per-event model. A simulator may estimate activate, read, write, and refresh events. A real device may expose aggregate counters but not every internal action. Vendor per-bit numbers may use different traffic patterns and boundaries.

The direct measurement path is:

TEXT
chip_energy = integral(chip_power_over_time)

That requires a power source with known sampling, calibration, and attribution. Board power includes more than memory. Device telemetry can be filtered or estimated. A rack power distribution unit includes host, networking, fans, and power conversion. The article must say which one it uses.

Facility overhead is often approximated with power usage effectiveness, using The Green Grid's public definition for the facility-to-IT boundary:

TEXT
facility_energy_proxy = IT_energy * PUE

PUE is total facility energy divided by IT equipment energy over a defined boundary and time. It does not identify HBM's share of chip power. It changes with facility, load, weather, cooling mode, and accounting. This formula is useful only when the PUE source and interval match the workload estimate.

Water needs three separate ledgers. Technology-loop inventory tracks coolant fill, circulation, makeup, drain, leak, and recovery without calling circulation consumption. Site water uses the exact numerator reported by the source, such as withdrawal, discharge, consumption, or another explicitly defined usage metric at the data-center boundary. Source-energy water tracks off-site water associated with producing the electricity used by the facility. Combining them without names, intervals, and meter boundaries can double count or hide the larger term.

The Green Grid's WUE convention uses site water consumption divided by IT equipment energy over the reporting period. A different source may publish withdrawal or another water numerator, so the receipt must preserve the source's exact definition instead of relabeling it:

TEXT
WUE_L_per_kWh = annual_site_water_L / annual_IT_energy_kWh

site_water_L_for_workload =
  workload_IT_energy_kWh * matched_WUE_L_per_kWh

A workload accounting chain can be written as:

TEXT
IT_energy_kWh =
  measured_average_IT_kW * occupied_seconds / 3600

facility_energy_kWh =
  IT_energy_kWh * matched_PUE

site_water_L =
  IT_energy_kWh * matched_WUE_L_per_kWh

source_water_L =
  facility_energy_kWh * matched_grid_water_intensity_L_per_kWh

Do not multiply a whole-facility energy reading by PUE again. Do not use a yearly average WUE as if it were an instantaneous meter. Do not infer water from liquid cooled: a closed equipment loop can reject heat through dry coolers, cooling towers, chillers, or a hybrid system, each with different water and power behavior. The Lawrence Berkeley National Laboratory review of data-center workload water use finds that workload water estimates can vary by more than four orders of magnitude across location, cooling, electricity, and efficiency assumptions. [RESEARCH SYNTHESIS] That is a warning to expose inputs, not permission to choose a convenient average.

This no-default calculator forces the boundary to stay visible:

PYTHON
from dataclasses import dataclass

@dataclass(frozen=True)
class FacilityReceipt:
    measured_it_kw: float
    occupied_seconds: float
    pue: float
    wue_l_per_it_kwh: float
    grid_water_l_per_facility_kwh: float
    electricity_usd_per_kwh: float

def facility_cost_and_water(r: FacilityReceipt) -> dict[str, float]:
    it_kwh = r.measured_it_kw * r.occupied_seconds / 3600.0
    facility_kwh = it_kwh * r.pue
    return {
        "it_kwh": it_kwh,
        "facility_kwh": facility_kwh,
        "site_water_l": it_kwh * r.wue_l_per_it_kwh,
        "source_water_l": facility_kwh * r.grid_water_l_per_facility_kwh,
        "electricity_usd": facility_kwh * r.electricity_usd_per_kwh,
    }

The inputs must come from the named facility, utility, interval, meter boundary, and allocation method. A CFO also needs demand charges, capacity reservation, networking, storage, software, support, labor, financing, depreciation, rejected work, and stranded time. A CEO needs useful tasks delivered under the product SLO. A CTO needs the architecture and failure envelope. The same HBM platform can look efficient at the chip boundary and expensive at the accepted-task boundary.

Power, cooling, and money belong to one accepted-task receipt

The bad baseline is simple: read a GPU power number, multiply it by time, add a generic PUE, and call the result the cost of AI. That shortcut hides the parts that actually decide whether the infrastructure made money: how much verified work crossed the acceptance gate, how many retries were paid for, how much power capacity sat stranded, which state was recomputed, which links and storage tiers moved bytes, and what the facility spent carrying the resulting heat away.

A 600 W card, a 14.5 kW server, an eight-module accelerator baseboard, and a 120 kW rack cooling design are not four points on one comparable chart. They are different physical boundaries. A coolant flow rate is not a water-consumption rate. A rack design envelope is not a task-energy measurement. A faster token stream is not a cheaper coding task if the patch fails tests.

The decision boundary is:

TEXT
minimum facility joules and dollars per verified successful task

subject to:
- output-quality parity
- TTFT, ITL/TPOT, and task-latency SLOs
- useful concurrency and throughput
- GPU, HBM, CPU, NIC, switch, and optic temperature limits
- power, flow, pressure, dew-point, and redundancy limits
- cache correctness and tenant isolation
- declared site-water constraints
- reliability and rollback requirements

The same run_id, operation_id, facility_interval_id, and accepted_output_id must join the application trace to GPU telemetry, rack power, cooling state, water accounting, and the final test or review result. Power, cooling, and water run beside the workload timeline. They are synchronized accounting boundaries, not fictional stages that begin only after the model finishes.

Interactive visual. The continuous machine on the left follows this same accepted-task boundary: request -> runtime and state -> GPU, HBM, and fabric -> measured electrical power -> integrated energy -> heat transport -> cooling mode -> site and source-water ledgers -> allocated total cost -> verified success. Each subsection activates the corresponding physical boundary and evidence state without replacing the full article on the right. Power is a rate; energy accumulates over time; transported heat, cooling electricity, closed-loop circulation, site-water consumption, and cost remain separate ledgers until the accepted-output receipt joins them.

The three-layer decision

Layer Question What must appear in the receipt
CEO / CFO Did this system produce more accepted work from the same constrained power, capital, and engineering budget? Verified successes, failures, retries, occupied time, capacity released, energy and total cost per success
CTO / serving engineer Which software decision changed the physical work? Prompt identity, cache hit or miss, prefill/decode phase, routing, batching, precision, placement, transfer path, fallback, p95/p99
Kernel / hardware / facility engineer Where did electricity go, where did heat move, and what limited sustained operation? Kernel and collective trace, HBM and fabric traffic, power counters, temperatures, throttle state, coolant flow and pressure, CDU/fan/chiller state, meter coverage and uncertainty

The executive decision is not “air or liquid?” It is “which complete architecture produces the required accepted work inside our power, cooling, latency, reliability, and cash constraints?”

Power is a rate. Energy is accumulated work.

Power describes the current operating rate:

TEXT
power_W = joules_per_second

For a simple 60-second teaching example, a complete server averaging 10,000 W for the aligned interval would use:

TEXT
10,000 W * 60 s = 600,000 J
600,000 J / 3,600,000 = 0.1667 kWh

That energy ultimately becomes heat inside the declared system boundary. Pumps, fans, a coolant distribution unit, chillers, or heat-rejection equipment can consume additional electricity outside the server meter. The example still says nothing about cost per useful result until the same interval records retries, shared utilization, prices, and how many outputs the verifier accepted.

Energy is power integrated over the same declared interval:

TEXT
energy_J = integral(power_W(t), time)

energy_kWh = energy_J / 3,600,000

For sampled power, use the trapezoidal rule:

TEXT
energy_J =
  sum(((power_i_W + power_i+1_W) / 2)
      * (timestamp_i+1_s - timestamp_i_s))

Prefer a cumulative hardware energy counter when its unit, timestamp, resolution, reset behavior, and wrap behavior are known. A task-level result is invalid when the counter resets, clock synchronization is unresolved, the sampling gap is material, or the facility interval does not contain the declared pre-task, task, post-task, and thermal-lag windows.

A published maximum is useful for electrical and cooling design. It is not a workload receipt:

TEXT
vendor-rated maximum power * task duration
!= measured task energy

The equality holds only when the complete named boundary actually operated at that power for the full interval and the allocation method proves that the task owned it.

Lock the physical scope before comparing platforms

Public platform boundary Source-backed power or cooling fact Correct use What it does not prove
NVIDIA DGX B300 complete server NVIDIA specifies 14.5 kW power consumption, 49,476 BTU/h maximum heat output, 1,500 CFM for the PSU configuration or 1,350 CFM for the busbar configuration at 70% PWM, and 10–30°C inlet air Size power, airflow, containment, and facility capacity for the complete 10RU server Energy for one coding task, fan energy for a workload, rack PUE, or accepted-task cost
NVIDIA GB200 NVL72 rack design NVIDIA's OCP contribution specifies direct liquid cooling sized for 120 kW of rack cooling capacity; the DGX guide says CPUs and GPUs use cold plates while networking and storage remain air cooled Explain a high-density hybrid or design-specific liquid path and the facility capacity it requires That the rack continuously consumes 120 kW, that all heat enters liquid, or that 120 kW can be multiplied by one request's duration
AMD MI350P PCIe card boundary AMD specifies passive air cooling, 450 W configurable and 600 W maximum TBP, a 12V-2x6 connector, and partner air-cooled systems with up to eight cards Size the qualified slot, card power, PSU, cabling, server airflow, and board-only 3.6 to 4.8 kW eight-card envelope Complete server or facility power, fan energy, inlet limit, workload energy, or universal eight-card compatibility
AMD MI355X UBB 2.0 accelerator-module boundary AMD specifies PG-25, 43°C maximum inlet, 2.1 L/min recommended flow per OAM, and 1,400 W maximum TBP per module Derive an eight-OAM module-only ceiling of 11.2 kW and a recommended circulation sum of 16.8 L/min Complete server/rack power, pressure drop, pump power, CDU power, or 16.8 L/min of water consumption
NVIDIA Vera Rubin NVL72 rack direction NVIDIA describes 100% liquid cooling with no system fans, up to 45°C rack inlet and roughly 55°C return in its reference narrative; the current MGX description says the actual rack fluid depends on the data center and may use deionized water or PG25 Explain warm-water, fanless rack design and why dry cooling can become practical in suitable climates Configuration-grade rack power, universal coolant chemistry, universal zero-water operation, or task-level savings
AMD Helios reference design AMD describes a 72-MI455X double-wide Open Rack Wide design with a vertical power busbar and liquid-cooling manifold feeding compute and switch trays through quick disconnects Show the AMD rack-scale liquid and power architecture direction Final orderable SKU, rack power, flow, pressure drop, coolant chemistry, CDU power, or accepted-workload result
RTX PRO 6000 Blackwell NVIDIA specifies 300 W maximum for Max-Q with active cooling; the server edition is configurable up to 600 W and is available in dual-slot air or single-slot liquid form factors Separate workstation/card experimentation from complete server and rack infrastructure Complete node power, OEM airflow or liquid conditions, facility energy, or cost per accepted task

Sources: NVIDIA DGX B300 User Guide, NVIDIA DGX GB rack guide, NVIDIA GB200 NVL72 OCP design, AMD MI350P product page, AMD MI350P enterprise deployment article, AMD MI355X platform brief, NVIDIA 45°C liquid-cooling architecture, NVIDIA Vera Rubin POD and MGX architecture, AMD Helios, RTX PRO 6000 Max-Q, and RTX PRO 6000 Server Edition.

Follow the electrical path from grid to silicon

TEXT
utility meter
-> transformer and switchgear
-> UPS or rack energy storage
-> facility distribution
-> rack PDU or power shelf
-> rectifier and DC busbar
-> board voltage regulators
-> GPU, HBM, CPU, DRAM, NIC, DPU, switch, retimer, storage, and fans

Each conversion can lose energy:

TEXT
conversion_loss_J =
  input_energy_J
  - output_energy_J
  - stored_energy_change_J

conversion_efficiency =
  output_energy_J / input_energy_J

If only a vendor efficiency curve exists, mark the result modelled and preserve the exact equipment revision, input voltage, load fraction, temperature, and source. If the operating point is missing, leave the value unknown.

Do not add component telemetry to a parent meter that already contains it. A rack PDU total already includes the node's GPUs, CPUs, memory, NICs, storage, local fans, and conversion losses downstream of that meter. Those component readings explain composition; they are not extra energy.

TEXT
reconciliation_residual_J =
  measured_parent_J
  - sum(mutually_exclusive_measured_children_J)

sum(task_allocations_J) + unallocated_J
  = measured_parent_J +/- uncertainty

The residual remains visible. It is where missing meters, clock error, sensor accuracy, storage changes, and unobserved components show up.

Follow the runtime mechanism into the power trace

The same accelerator can draw similar average power while producing very different useful work.

TEXT
C-001 coding-agent task
-> stable system, policy, tool, RAG, and repository prefix
-> cache identity validation
-> prefix hit, partial hit, restore, or miss
-> prefill for non-reused tokens
-> decode and MoE or dense model operations
-> KV append and possible collective traffic
-> tool call on CPU, filesystem, network, or sandbox
-> tool-result suffix prefill
-> more decode
-> patch, tests, review, and acceptance

A cache miss can repeat expensive prefill. A cache hit can avoid it. A route to a cooler GPU can destroy prefix locality and create more HBM and fabric traffic than it saves thermally. A lower core frequency can save GPU-device energy during memory-bound decode, but it can also extend runtime while HBM, NICs, CPUs, pumps, and static platform power continue to accumulate.

That is why the control objective is not minimum instantaneous watts:

TEXT
choose:
  frequency
  power cap
  route
  batch
  cache location
  state placement
  transfer path
  cooling setpoint where exposed

to minimize:
  facility_joules_per_verified_success
  and total_cost_per_verified_success

while preserving:
  quality
  TTFT / ITL / p99
  useful capacity
  reliability
  thermal and water limits

VoltanaLLM is useful research evidence for this mechanism. Its authors report up to 36.3% GPU energy savings relative to a static maximum-frequency baseline in their evaluated SGLang configurations while maintaining high SLO attainment. That is an author-reported GPU-device result for the paper's named setup. It is not a Touchdown reproduction, accepted-task result, complete-node result, rack result, facility result, or water result. The next step is not to copy its A100 clock table. It is to replay the control idea on each exact model, engine, kernel, topology, and power boundary. Source: VoltanaLLM.

Electricity becomes heat, but heat transport is not electrical energy

Within a declared control volume over a steady-state interval, nearly all IT electrical energy that is not exported as electrical or optical work or retained in stored state eventually appears as heat. During transients, stored thermal energy must remain an explicit term. The heat can leave through air, liquid, or both.

TEXT
GPU logic + HBM + CPU + DRAM + NIC + switch + VRM + storage
-> die and package
-> thermal interface
-> heat sink or cold plate
-> chassis air or technology cooling loop
-> CDU or room cooling equipment
-> facility loop
-> dry cooler, chiller, cooling tower, or verified heat-reuse sink

For air cooling:

TEXT
heat_transport_air_W =
  air_density_kg_m3
  * air_specific_heat_J_kgK
  * airflow_m3_s
  * (exhaust_C - inlet_C)

fan_wire_power_W =
  pressure_rise_Pa * airflow_m3_s / fan_efficiency

For a single-phase liquid:

TEXT
mass_flow_kg_s =
  coolant_density_kg_m3 * volumetric_flow_m3_s

heat_transport_liquid_W =
  mass_flow_kg_s
  * coolant_specific_heat_J_kgK
  * (return_C - supply_C)

pump_hydraulic_power_W =
  differential_pressure_Pa * volumetric_flow_m3_s

pump_wire_power_W =
  pump_hydraulic_power_W / wire_to_water_efficiency

For a phase-changing coolant:

TEXT
heat_transport_W =
  mass_flow_kg_s
  * (enthalpy_out_J_kg - enthalpy_in_J_kg)

Do not use a constant specific_heat * delta_T approximation across boiling or condensation. Do not model PG-25 with pure-water properties. Record the property source and temperature.

The aligned heat balance is:

TEXT
IT electrical power
+ pump and fan heat entering the observed control volume
=
liquid heat transport
+ air heat transport
+ exported useful heat
+ ambient loss
+ stored thermal energy increase
+ exported electrical or optical power
+ reconciliation residual

Heat carried by liquid is not GPU electrical energy. It is not cooling electrical energy. It is not automatically heat reused. These are different ledgers.

Cooling is a runtime state, not a product label

“Liquid cooled” is not a complete description. The facility can operate in different modes over one long video job or agent loop:

Operating mode Heat-rejection path Main electrical loads Direct site-water path
Dry or free cooling Facility loop to dry cooler to ambient air Pumps, dry-cooler fans, controls No routine evaporation; service loss and optional adiabatic assistance remain separate
Adiabatic assist Dry cooler plus spray or wetted media Pumps and fans Weather-dependent spray consumption
Mechanical chiller Facility loop through evaporator, compressor, and condenser Pumps, compressor, condenser fans or tower Depends on air-cooled versus water-cooled condenser
Cooling tower Condenser loop to evaporative tower Pumps, tower fans, possible chiller Makeup, evaporation, drift, blowdown, and leaks
Useful heat export Warm facility loop crosses a meter into a verified sink Pumps and possibly a heat pump Water effect remains separately measured

ASHRAE recommends matching cooling topology to AI rack densities around 60–120 kW/rack and above, using scalable manifolded liquid distribution, tracking PUE and water metrics, and considering warm-water heat recovery where a real sink exists. OCP guidance separates the technology cooling loop from the facility loop and emphasizes thermal ride-through, pump redundancy, and dry heat rejection where feasible. DOE's typical tower path carries heat from IT through room or liquid systems, chillers, condenser water, and finally evaporative rejection. Sources: ASHRAE AI Data Center Framework, OCP Advanced Cooling Facility reference design, and DOE cooling-water guidance.

Record every mode transition:

TEXT
mode_start and mode_end
ambient dry-bulb and wet-bulb temperature
humidity and dew point
coolant supply and return temperature
flow and differential pressure
pump, fan, CDU, chiller, tower, and adiabatic state
thermal storage change and lag
redundancy, leak, throttle, and fault state

Warm-water liquid cooling can free facility power for compute, but the result depends on climate, final heat rejection, pump and fan work, redundancy, and the accepted workload. NVIDIA's Rubin reference narrative describes 45°C supply, roughly 55°C return, a closed loop, and the potential to use dry coolers in suitable climates. Treat its zero-water and dollar-savings examples as vendor scenarios tied to geography and design, not universal task measurements.

Water needs three ledgers

TEXT
1. Technology-loop inventory and service loss
   initial fill + makeup - drain - leak - recovered coolant

2. On-site withdrawal and consumption
   potable + reclaimed + groundwater + surface water
   - returned or discharged water
   - positive storage change

3. Source-energy water
   off-site water associated with generating the electricity used

The critical distinction:

TEXT
closed_loop_circulation_L =
  integral(flow_rate_L_s, time_s)

closed_loop_circulation_L
!=
site_water_consumption_L

For the AMD example:

TEXT
8 OAMs * 2.1 L/min per OAM
= 16.8 L/min recommended circulation sum

not:
16.8 L/min of water consumed

For a cooling tower:

TEXT
tower_makeup_L =
  evaporation_L
  + drift_L
  + blowdown_L
  + leaks_L
  + inventory_change_L
  - internal_recovery_L

For the site:

TEXT
site_consumption_L =
  potable_input_L
  + reclaimed_input_L
  + groundwater_input_L
  + surface_water_input_L
  - returned_or_discharged_L
  - positive_storage_delta_L

PUE and WUE remain facility metrics over aligned boundaries and intervals:

TEXT
PUE_T =
  total_facility_energy_T / IT_energy_T

WUE_T =
  site_water_usage_L_T / IT_energy_kWh_T

site_water_usage_L_T preserves the numerator used by the named WUE source. The receipt must map that numerator explicitly to withdrawal, consumption, or another defined usage boundary; it cannot silently treat those terms as interchangeable.

Do not apply PUE when facility energy is already measured. Do not call the full PUE overhead cooling; it also includes electrical conversion, distribution, controls, lighting, and other non-IT loads. Do not present an annual WUE as an instantaneous task meter. A per-task result from PUE or WUE is an allocation model unless synchronized meters cover the same task and thermal-lag interval.

The full money path

Electricity matters. It is rarely the whole bill.

TEXT
total_relevant_cost =
  accelerator_and_memory_capital
  + host_CPU_and_DRAM_capital
  + networking_and_storage_capital
  + electrical_and_cooling_buildout
  + installation_and_financing
  + software_and_support
  + operations_and_maintenance
  + measured_energy_cost
  + allocated_peak_or_capacity_cost
  + spares_and_expected_failures
  + retries_and_rejected_work
  + human_rework
  - residual_value

Use separate ledgers:

TEXT
variable_energy_cost =
  facility_energy_kWh * blended_energy_rate_USD_per_kWh

monthly_peak_cost =
  billing_peak_kW * demand_or_capacity_rate_USD_per_kW

straight_line_annualized_capital_cost =
  (purchase
   + electrical_buildout
   + cooling_buildout
   + network
   + storage
   + installation
   + financing
   - residual_value)
  / useful_life_years

capital_recovery_factor =
  discount_rate * (1 + discount_rate)^useful_life_years
  / ((1 + discount_rate)^useful_life_years - 1)

finance_grade_annualized_capital_cost =
  net_present_capital_cost * capital_recovery_factor

allocated_capital_cost_per_success =
  allocated_annualized_capital_cost
  / annual_verified_successes

The first expression is straight-line allocation, not finance-grade annualization. Use one accounting horizon and present-value convention across every cost term. Tariffs, financing, hardware prices, service contracts, discount rate, useful life, and residual value are site- and contract-specific. Keep every unsourced input visible as an assumption or range.

The business denominator is:

TEXT
gross_cost_per_verified_success =
  all_allocated_costs_for_all_attempts
  / verified_successes

Retries and rejected outputs stay in the numerator. Only verified successes enter the denominator. If verified_successes == 0, the result is null, not zero and not an impressive-looking infinity.

The capacity view is often more important than the energy line item:

TEXT
verified_successes_per_constrained_kWh =
  verified_successes / facility_energy_kWh

contribution_margin_per_constrained_kWh =
  verified_successes_per_constrained_kWh
  * contribution_margin_per_success

realizable_capacity_value_released =
  min(additional_verified_successes, additional_demand)
  * contribution_margin_per_success

A cache improvement, kernel optimization, or thermal-aware route is valuable when it releases enough constrained GPU, rack, or facility capacity to serve more accepted work without violating p99 or quality. Released technical capacity becomes realizable financial value only when demand and positive contribution margin exist. The real prize is not merely a lower electricity bill. It is more trustworthy product capacity from the same expensive power and hardware envelope.

One fully labeled money example

This is an illustrative sensitivity, not a Touchdown benchmark.

NVIDIA publishes a 14.5 kW power-consumption figure for one complete DGX B300. Suppose an operator uses that published system figure as a conservative one-hour planning envelope, assumes PUE = 1.10, and tests three explicitly assumed electricity rates.

TEXT
IT energy proxy =
  14.5 kW * 1 hour
  = 14.5 kWh

facility energy proxy =
  14.5 kWh * 1.10
  = 15.95 kWh
Assumed electricity rate Energy-only cost for the hour If 100 tasks are verified If only 70 tasks are verified
$0.05/kWh $0.7975 $0.007975/success $0.011393/success
$0.10/kWh $1.5950 $0.015950/success $0.022786/success
$0.15/kWh $2.3925 $0.023925/success $0.034179/success

Every number after the 14.5 kW vendor specification is a Touchdown derivation from stated assumptions. This is energy-only. It excludes the server purchase, networking, storage, facility buildout, peak charges, support, labor, spares, failed attempts, and financing. It also assumes the system held that published figure for the full hour; a real task result must integrate measured power.

The table illustrates the denominator effect: identical facility energy looks 43% more expensive per verified success when 30 of 100 attempts fail the acceptance gate.

The controlled before/after experiment

Use four equal-duration energy controls:

TEXT
E00 = loaded-service idle
E10 = compute-only control
E01 = transfer-only control
E11 = actual combined compute + transfer

Then preserve the interaction term:

TEXT
delta_E_compute =
  E10 - E00

delta_E_transfer =
  E01 - E00

delta_E_combined =
  E11 - E00

E_interaction =
  E11 - E10 - E01 + E00

E_interaction captures overlap, synchronization, contention, DVFS, batching, queueing, fabric congestion, and thermal state. Do not report compute energy + transfer energy as the actual combined result without measuring or preserving that interaction.

Required experiment arms:

  1. Vendor default.
  2. Fixed-frequency sweep.
  3. Power-cap sweep.
  4. State-aware routing only.
  5. Phase-frequency control only.
  6. Joint frequency and state-aware routing.
  7. Joint frequency, routing, state placement, and fabric path.
  8. Joint compute control plus thermal-headroom control.

Freeze:

TEXT
workload and arrival trace
repository / prompt / retrieval identity
model and artifact revision
engine, container, precision, and KV format
quality and acceptance rubric
concurrency and retry budget
topology and background load
coolant and ambient conditions
warmup, steady-state, and lag windows

Report:

TEXT
attempts and verified successes
quality and test results
TTFT, ITL/TPOT, p50, p95, and p99
prefill, decode, tool, restore, and queue time
cache hit, miss, eviction, and bytes moved
GPU, node, rack, and facility joules by valid boundary
pump, fan, CDU, chiller, and tower energy
site and source-water liters
temperature, throttle, leak, and failure events
gross and incremental cost per verified success
useful capacity and successful tasks per kWh

A policy wins only when the accepted-task result improves without quality, latency, retry, capacity, thermal, or reliability regression.

The minimum facility receipt

YAML
FacilityBoundaryReceipt:
  schema_version: touchdown.facility-boundary-receipt.v1
  run_id: string
  workload_id: C-001 | V-001
  fixture_digest: string
  accepted_output_ids: [string]

  attempts: int
  verified_successes: int
  retries: int
  failures: int
  quality_parity: bool | null
  slo_parity: bool | null

  interval:
    utc_start: string
    utc_end: string
    monotonic_start_ns: int
    monotonic_end_ns: int
    clock_domain: string
    max_sync_error_ms: float
    sample_period_ms: float
    missing_sample_fraction: float
    warmup_s: float
    cooldown_s: float
    thermal_lag_window_s: float

  execution:
    model_revision: string
    engine_revision: string
    precision: string
    kv_format: string
    topology_id: string
    policy_decision_ids: [string]
    fabric_event_ids: [string]

  energy:
    boundary_ids: [string]
    parent_child_coverage: object
    measurement_kind: measured | allocated | modelled | unknown
    meter_ids: [string]
    reconciliation_residual_J: float | null

  thermal:
    boundary_ids: [string]
    coolant_and_property_source: string | null
    supply_return_flow_pressure: object
    air_side_remainder: float | null
    throttle_and_fault_events: [object]

  water:
    boundary_ids: [string]
    inventory_L: float | null
    circulation_L: float | null
    withdrawal_L: float | null
    site_consumption_L: float | null
    source_energy_water_L: float | null

  finance:
    electricity_rate_source: string | null
    demand_rate_source: string | null
    capital_allocation_method: string
    allocated_energy_cost_USD: float | null
    allocated_total_cost_USD: float | null

  derived:
    gross_J_per_verified_success: float | null
    incremental_J_per_verified_success: float | null
    successful_tasks_per_kWh: float | null
    site_L_per_verified_success: float | null
    total_cost_per_verified_success: float | null

  uncertainty: object

Every populated field carries a source, timestamp, physical boundary, unit, measurement kind, and uncertainty. Unknown remains visible and never becomes zero.

Copy-paste: executive decision rule

TEXT
Do not buy or reject a platform from TDP, peak bandwidth, or tokens per second.

Compare the same verified workload at the same quality and p99:

1. accepted tasks per hour;
2. gross and incremental facility joules per accepted task;
3. total cost per accepted task;
4. useful capacity released inside the constrained power envelope;
5. thermal, water, reliability, and rollback limits.

A vendor envelope sizes infrastructure. A synchronized accepted-task receipt makes the decision.

Copy-paste: measurement gate

TEXT
A power or cooling claim is publishable only when it names:
- the exact platform and physical scope;
- included and excluded components;
- workload, model, engine, precision, concurrency, and topology;
- meter or sensor, sampling semantics, calibration, and clock sync;
- air/liquid loop, coolant, temperatures, flow, pressure, and rejection mode;
- attempts, retries, failures, and verified successes;
- formula, allocation method, uncertainty, and source revision.

Closed-loop flow is not water consumption.
Heat transport is not electrical energy.
GPU energy is not facility energy.
A faster output is not a cheaper task until it passes the acceptance gate.

What this means by role. The CEO sees whether infrastructure created more accepted product capacity. The CFO sees capital, energy, peak power, cooling, failures, and margin on one denominator. The CTO sees which routing, cache, precision, placement, and control decision changed the result. The software engineer can join the request and state path to the physical meters. The kernel and hardware engineers can separate logical bytes, measured traffic, electrical work, heat transport, and cooling overhead without double counting them.

Heat changes sustained behavior

With the facility accounting boundary now closed, return to the package and ask how HBM temperature changes the sustained service that the software actually receives.

Electrical energy becomes heat. The temperature rise depends on heat flux, material conductivities, interfaces, airflow or coolant, and thermal resistance.

HBM sits close to a high-power accelerator. The logic base die can add activity under the memory stack. Heat can increase leakage and reduce retention margin. Controllers or firmware can throttle clocks or traffic to stay inside safe limits. A peak bandwidth specification can remain true while sustained workload service falls under a thermal or power cap.

Cooling also consumes energy and infrastructure. A liquid-cooled rack needs pumps, distribution, heat exchangers, controls, leak management, water or coolant treatment, and maintenance. More efficient chip-level movement can still lose at the facility level if it requires a much harder cooling path or lowers packaging yield.

The correct receipt joins performance and thermal data on the same timeline:

TEXT
timestamp
kernel / workload phase
HBM and link traffic
chip / board power source
memory and accelerator temperature
clock / throttle state
cooling operating point
accepted output

Kyber shows where the HBM package becomes a rack system

HBM ends at the accelerator package, but the workload does not. Once a model spans GPUs, memory placement couples to NVLink, NVSwitch, board routing, rack power, cooling, and the physical boundary of the scale-up domain.

NVIDIA now defines Kyber officially as its next-generation MGX NVL rack design. It is not a GPU generation, an HBM generation, or another name for every Rubin Ultra system. NVIDIA's current preliminary topology map is: NVIDIA Vera Rubin POD.

TEXT
Vera Rubin Ultra Kyber NVL144:
  one Kyber rack
  144 GPUs in one NVLink scale-up domain

Vera Rubin Ultra NVL576:
  eight separate 72-GPU MGX NVL racks
  576 GPUs in one multirack domain

Feynman Kyber NVL1152:
  eight 144-GPU Kyber racks
  1,152 GPUs in one multirack domain

Those are different physical products. NVL576 is not one 576-GPU rack. NVL1152 is not one 1,152-GPU rack. Kyber first appears with Vera Rubin Ultra as a standalone NVL144 system, while eight Kyber racks later form the Feynman NVL1152 system. [OFFICIAL PRELIMINARY]

This matters to an HBM article because aggregate HBM capacity and bandwidth do not automatically become one useful memory pool. Tensor, expert, sequence, pipeline, and data parallelism decide which bytes remain local, which bytes cross NVLink, and which bytes cross a scale-out network. A rack can contain enormous HBM bandwidth while a workload waits on collectives, synchronization, host orchestration, storage, or a rejected output.

NVIDIA also connects Kyber to its 800 VDC power architecture. That establishes an official architecture direction around higher-voltage facility-to-rack distribution. It does not establish delivered Kyber power efficiency or accepted-workload economics.

Named independent reports from Tom's Hardware and Data Center Dynamics describe vertical-tray and orthogonal-PCB-midplane implementation details and report possible schedule risk tied to manufacturability. Both also report NVIDIA's response that its roadmap is intact. [THIRD-PARTY REPORT]

The reviewed NVIDIA primary sources do not publish the final Kyber PCB stackup, layer count, trace geometry, impedance distribution, board yield, failure Pareto, named supplier qualification, supplier allocation, customer schedule, or Wan2.2 result. Those fields remain [NOT PUBLIC]. Repeated secondary numbers do not become a production datasheet.

What must physically work in a Kyber-class system

The official architecture is enough to identify the engineering categories, even when private implementation values stay unknown.

Signal path. A serializer launches bits into package and board channels. Vias, connectors, copper traces, cable cartridges, and any direct-optical boundary add insertion loss, reflection, crosstalk, skew, and jitter. Equalization and clock recovery must recover the eye at the receiver across process, voltage, temperature, aging, and assembly variation. A clean schematic or simulated nominal channel is not qualification.

Power path. Facility power passes through switchgear, conversion, busbars, shelves, voltage regulators, package delivery, and on-die rails. Current transients create droop and noise. NVIDIA's current MGX description adds rack-level energy storage, dynamic power steering, and power smoothing. Its 800 VDC direction attempts to change the facility-to-rack conversion boundary. Those mechanisms must be evaluated with workload transients, protection, service safety, fault isolation, and conversion loss.

Thermal path. GPU logic, HBM stacks, NVSwitch silicon, optics or copper interfaces, voltage conversion, and DPUs produce heat at different locations. Heat crosses dies, underfill, package lids, cold plates, coolant, manifolds, CDUs, facility loops, and heat rejection equipment. Flow imbalance, fouling, bubbles, pump failure, leaks, and high inlet temperature can change sustained clocks or availability even when peak specifications remain unchanged.

Mechanical path. A dense rack must survive board fabrication, component placement, reflow, tray assembly, connector mating, shipping shock, installation, repeated service, and thermal cycling. Tighter channel and power constraints can reduce the manufacturing margin. The relevant yield is not only a bare PCB yield or a good GPU package. It is the probability that the assembled tray, fabric, cooling loop, power path, firmware, and rack pass qualification and remain serviceable.

Test and repair path. Manufacturing needs boundary scan, link training, BIST, lane margining, power and thermal test, leak checks, burn-in, fault injection, and traceable component identity. Operations need a way to isolate a bad GPU, switch, link, DPU, power shelf, sensor, or cooling component without turning an entire scale-up domain into unusable inventory. NVIDIA describes continued rack operation during NVLink switch maintenance for Vera Rubin. The exact Kyber degradation and repair behavior remains a future product question.

What software sees instead of one giant memory

The runtime sees ranks, devices, address spaces, allocations, communication groups, and failures. A placement compiler or scheduler must map each tensor or state object onto those boundaries:

TEXT
weights and active KV
-> shard or replicate across GPU HBM

attention heads / sequence partitions / experts
-> assign to ranks in a topology-aware group

all-reduce / all-gather / reduce-scatter / all-to-all
-> choose collective algorithm and route over the scale-up fabric

checkpoint / reusable KV / dataset / tool state
-> place in host memory, CMX or storage only under an explicit policy

rank, switch, link, or rack failure
-> degrade, remap, retry, or reject under a named correctness rule

For Wan2.2 Ulysses sequence parallelism, the all-to-all exchange changes how attention state is distributed before and after attention. A 144-GPU domain creates more possible partitions, but it also enlarges the participant set and the consequences of imbalance or failure. For an MoE model, expert placement and token routing create a different all-to-all. For a coding agent, a small model or low-concurrency service may gain nothing from a giant domain and can pay more for idle capacity and operational complexity.

What the Kyber receipt must contain

TEXT
named platform and evidence state
GPU packages per rack and racks per scale-up domain
HBM capacity and bandwidth at per-GPU and aggregate scopes
rank map and physical topology capture
collective message-size histogram and p50 / p95 / p99
link training, retry, error, and degraded-mode counters
board, tray, rack, and facility power with timestamps
temperature, coolant, flow, throttle, and leak events
planned and unplanned maintenance time
software, firmware, compiler, engine, and model revisions
accepted-task throughput, quality, latency, retry, and cost

Until those fields exist for a delivered Kyber system, the honest result is an architecture and diligence model, not a workload benchmark.

The decision rule remains workload-first:

TEXT
HBM-local kernel time
+ exposed collective wait
+ synchronization and host time
+ VAE, storage, and review time
+ retries and rejected outputs
= occupied rack time per accepted result

The operator should compare cost_per_accepted_task or rack_seconds_per_accepted_clip, not only peak HBM or NVLink bandwidth.

Yield, repair, and failures reach the customer

Manufacturing yield determines how many good units emerge from a process. Test coverage determines how many bad units are caught before shipment. Repair determines which defects can be recovered. Field RAS determines how errors are detected and contained after deployment.

A lower package yield can raise effective cost and constrain supply. More redundancy can recover units but consumes area and may complicate routing. Aggressive repair can improve usable capacity while still leaving performance nonuniformity or untested failure modes. Strong ECC can correct a defined error pattern but cannot repair every controller, link, or package fault.

At runtime, an uncorrectable error can kill a request or a worker. A correctable-error storm can reduce performance. A link retry can add tail latency. A stale cache entry can return a logically wrong result even when every bit is physically correct. Reliability therefore crosses physics and software.

Turn it into business value without inventing prices

The complete system cost is broader than the accelerator invoice:

TEXT
usable_system_cost =
  accelerator_and_memory_package
  + host_CPU_and_memory
  + board_and_network
  + storage
  + power_delivery_and_cooling
  + expected_failures_and_spares
  + software_and_operations

The useful capacity side is:

TEXT
accepted_output_capacity =
  tasks_per_hour
  * acceptance_rate
  * availability

And the business boundary returns to:

TEXT
cost_per_accepted_task =
  total_relevant_cost / accepted_tasks

No public dollar appears without a dated source or labeled assumption. Instead, run sensitivity:

  • If extra HBM raises batch or concurrency, how many additional accepted tasks fit before p99 breaks?
  • If host or CXL offload frees HBM, how much restore latency and link contention enters?
  • If streaming weights from flash avoids a larger package, how often does first-byte or page overfetch stall compute?
  • If custom base logic removes traffic, how much non-recurring engineering, qualification, thermal, and yield risk is added?
  • If a sparse kernel reduces attention work, does quality remain acceptable and do index/gather/collective costs stay below the savings?

A scenario tool should preserve evidence state in its inputs:

PYTHON
# [TOUCHDOWN DERIVATION] Structure only. No default prices or power.
inputs = {
    "accepted_tasks": {"value": measured_tasks, "state": "measured"},
    "gpu_seconds": {"value": measured_gpu_s, "state": "measured"},
    "gpu_rate": {"value": contract_rate, "state": "assumption"},
    "IT_energy_kWh": {"value": metered_energy, "state": "measured"},
    "PUE": {"value": facility_PUE, "state": "measured_or_assumed"},
    "transfer_p99_ms": {"value": restore_p99, "state": "measured"},
}

The output is a range with sensitivity, not a precise answer built on hidden defaults.

FIGURE 9 1600 × 900
A chain runs from qualified materials and wafer processes through DRAM dies, TSV stacking, base logic, package, board, server, rack power and cooling, runtime placement, and accepted AI output. Loss markers show die yield, bond/test failures, thermal throttling, link contention, cache misses, retries, and rejected output.
Business value appears only after physical supply and workload acceptance survive the complete path. Source state. Systems synthesis. No vendor yield, price, or savings number.

What this means by role. The investor identifies which supply constraint is durable and which is substitutable. The CEO ties installed capacity to accepted output. The CFO separates package capex, energy, failure, and engineering cost. The CTO requires a workload replay before committing to a new tier. The hardware team shows test, repair, PDN, and thermal evidence. The software team shows placement, retry, and acceptance receipts.

What do the hardware lottery and Carmack actually change?

In plain English. Hardware makes some algorithms easy and others awkward. Software then evolves around the machines that are available. The practical question is not whether hardware determines every idea. It is whether a repeated, predictable workload pattern is important enough to justify a different software primitive, controller, memory placement, or physical design.

The evidence so far points in two directions at once.

HBM is a phenomenal commercial answer for hot AI state. The software and hardware stack around dense tensor operations is deeply optimized. At the same time, the two workloads contain state and operations that do not all look like dense GEMM.

The right response is not to reject the current stack. It is to ask which ideas were selected because the stack makes them cheap, and which memory schedules become possible when access is predictable.

Sara Hooker: available systems shape successful ideas

Sara Hooker's paper The Hardware Lottery describes a selection effect: research ideas are more likely to succeed and spread when available hardware and software make them easy to run.

The feedback loop is straightforward:

TEXT
available hardware and libraries
-> reliable, cheap implementations of favored operations
-> more experiments and deployment
-> more models designed around those operations
-> more investment in the same hardware path

This is not an argument that matrix multiplication is wrong. Dense matrix multiplication is extraordinarily useful. Accelerators, compilers, kernels, quantization schemes, and model architectures have made it efficient at enormous scale.

The question is where the fit is incomplete.

Wan2.2 adds normalization, positional operations, attention selection, sparse indexes, guidance, scheduler updates, VAE decode, communication, and repeated state movement. A coding agent adds prefix reuse, active KV growth, CPU tools, storage reads, test execution, retries, and durable state. Optimizing only the dense operation can leave a movement, synchronization, or workflow bottleneck untouched.

Ali's worklog is useful here because the reported end-to-end result combines algorithm, kernels, precision, step count, fusion, and scheduling. It does not prove hardware should abandon GEMM. It shows why an end-to-end workload can improve through several layers at once and why each layer needs its own receipt.

The hardware-lottery lesson is practical: expose and measure the operations and state paths that the current stack handles poorly. Then compare a tuned commercial baseline before asking for new hardware.

John Carmack: predictable reads create a scheduling opportunity

John Carmack's public post argues that inference weight access can be deterministic enough to schedule large pages ahead of use and keep a smaller active window in faster memory. We paraphrase the primary post because ordinary web extraction is incomplete. [PRIMARY SOCIAL / AUTHOR ARGUMENT] It is a proposal, not a benchmark.

The strongest version of the idea is not flash is as fast as HBM. It is:

TEXT
known access order
+ high useful bytes per fetched page
+ enough lookahead
+ parallel page reads
+ bounded first-byte and tail latency
+ double-buffer capacity
+ correct miss fallback
= a chance to hide slower, denser storage behind compute

Imagine that a layer consumes weight pages in a known order. Buffer A holds the pages currently feeding compute. Buffer B receives the next pages. When compute finishes A, the buffers swap. If transfer of B always completes before the swap and most fetched bytes are used, a large read-mostly tier can reduce how much state must remain in HBM.

The schedule breaks when a branch changes the access order, the page contains little useful data, queue contention stretches p99, writes become frequent, the hot window does not fit, or compute is too fast to hide transfer.

It also does not automatically apply to every AI state. Active KV is append-only and repeatedly read. A diffusion latent is mutable every step. Sparse attention can gather irregular regions. Coding-agent tools branch based on external results. Safety-critical state needs bounded worst-case behavior and fault containment.

Sandisk's HBF proposal gives the idea a vendor architecture direction. Its capacity, bandwidth, and Llama figures remain vendor targets or internal simulation. The SK hynix/Sandisk announcement establishes a standards workstream. Kioxia's separate module establishes a prototype. None is a shipping HBF workload result.

Unconventional AI is pursuing a first-principles, physics-led hardware and software co-design direction that treats parameter memory and working memory as separate design problems. We admire the attempt to connect algorithms, circuits, devices, and manufacturing, and leave the comparison there.

The shared lesson is not build a new memory. It is turn the access hypothesis into a trace, a schedule, a correct fallback, and a comparison against tuned HBM.

FIGURE 10 1600 × 900
Three loops converge on a named workload receipt. Hooker's loop links available hardware to research selection. Carmack's loop links predictable access to scheduled page streaming. The workload loop records state, movement, latency, correctness, and accepted output.
External arguments create hypotheses. A reproducible workload trace determines where they survive. Source state. Source-backed paper and primary-author argument plus Touchdown synthesis.

What is the narrow Touchdown hypothesis?

In plain English. This is a research hypothesis, not a shipping memory product. First measure a real accepted workload. Then exhaust competent software and available memory options. Only a repeated bottleneck that survives those tests earns a small, typed hardware or controller experiment.

Touchdown starts with HBM as the commercial baseline. Our hypothesis is that software should describe state precisely enough to compare placement, movement, retention, repair, and fallback across commercial tiers now and physical candidates later.

A public state contract can be small:

YAML
# [PROPOSED] Illustrative public contract
state_object:
  identity: named_tensor_or_workflow_object
  shape_dtype_bytes: measured_or_derived
  access: {pattern: named, mutability: named, predictability: measured}
  reuse: {distance_distribution: measured, fanout: measured}
  timing: {first_byte_p99_ms: named, completion_p99_ms: named}
  correctness: {error_budget: exact_or_bounded, recompute_ms: measured}
  movement: {sources: [named], targets: [named], granularity_bytes: named}
  fallback: named_correct_path
  evidence: {source: immutable_trace, version: immutable_id}

The contract does not make physical memories interchangeable. It gives a compiler or runtime enough information to test primitives such as place, prefetch, stream, gather, compress, rematerialize, checkpoint, evict, migrate, replicate, refresh, repair, scrub, and fallback.

Each primitive needs preconditions, a target-specific lowering, observable counters, a failure rule, and a correct fallback. stream could lower to CUDA/NIXL movement from host memory, an NVMe read, a future HBF controller, or a custom base-die function. Those are different implementations with different physics, not loose hardware Lego.

The direction is supply-chain-first. Each physical lowering must name the commercial parts, interfaces, packaging, test, repair, qualification path, and merchant fallback that exist before it earns a custom controller or memory program. The software contract can be composable. The manufacturing stack remains constrained by real materials, tools, suppliers, yields, and volumes.

HBM remains the tournament baseline. A candidate must beat tuned HBM plus caching, batching, fusion, compression, recomputation, and commercial offload for one named state role. It does not need to replace HBM universally.

The admission gates are intentionally hard:

Gate Evidence required
G0: real joined trace Join a real request, runtime events, state objects, accepted output, and business outcome without inventing physical bytes.
G1: durable bottleneck Show that the bounded movement or retention bottleneck recurs under controlled repetition and relevant variation.
G2: tuned commercial baseline Show that the opportunity survives competent HBM, DDR/LPDDR, NUMA/CXL, caching, offload, batching, kernel, and runtime tuning with parity.
G3: portable recipe value Show that a versioned recipe creates value across more than one supported backend under equivalent output contracts.
G4: simulator plus physical anchor Show that a calibrated simulator agrees closely enough with an FPGA or RISC-V controller prototype to justify physical continuation, with fallback, fault injection, and resource/timing receipts.
G5: joint-IC and manufacturing case Show value across three workloads with named memory, logic, package, and manufacturing partners, qualification risks, kill criteria, and a merchant fallback.

The existing public compiler work is fixture-backed where its tests pass. It is not a G4 physical anchor, custom HBM, a tape-out, or a manufacturing commitment.

FIGURE 11 1600 × 900
A versioned state contract lowers through placement and movement primitives to commercial memory targets, then advances from G0 reproducible trace through G5 repeated cross-workload and manufacturing evidence, with a merchant fallback at every physical stage.
A hardware idea earns progress through workload, commercial-baseline, portability, physical-anchor, and manufacturing evidence. Source state. Proposed Touchdown research process. No custom silicon result.

Updated July 27, 2026 · investor-to-engineer guide

CXMT’s HBM path starts with DRAM manufacturing, then adds stacking, packaging, qualification, and workload proof.

CXMT already makes conventional DRAM. HBM stacks thin DRAM dies beside an AI accelerator so far more data can move at once. This fourteen-step chapter separates shipping products, reported plans, patents, engineering requirements, supplier coverage, qualification, and the receipts still needed to prove a working HBM system.

Open the machine-readable CXMT evidence registry

What evidence would prove or disprove this architecture?

In plain English. A diagram, installed package, vendor specification, or successful import is not enough. The proof has to join one workload identity, pinned source, selected runtime and kernel, real hardware, measured movement and energy, the final output, and the verifier that accepted or rejected it. Missing fields remain unknown.

The hypothesis weakens if state descriptors fail on held-out workloads, transfer overhead erases capacity gains, compression or recomputation wins at lower complexity, or a tuned HBM-only baseline remains better on accepted-task cost, latency, energy, and reliability. It also fails if tail misses, stale state, thermal limits, test coverage, yield, repair, supplier access, or qualification make the physical path unsafe or uneconomic.

The next public receipt should use one schema and harness for both workloads:

TEXT
pinned Wan2.2 trace on a named GPU
+ coding-agent KV/tool trace on a named model and GPU
+ tuned HBM-only baseline
+ tuned host-offload baseline
+ CXL or NVMe lane only when the named hardware exists
+ placement bytes, p50/p95/p99, capacity, energy source, acceptance result
+ versions, raw artifacts, negative cases, and correct fallback

A simulator remains modeled. An absent CXL device remains NOT FOUND, not a synthetic win.

Reader Decision Receipt before spending
CEO Which workflow outcome is memory-limited? Accepted-task baseline and failure ledger
CFO Do savings exceed movement, engineering, energy, and supply cost? Range model tied to measured acceptance and throughput
Investor Is advantage in software, controller, package, device, supply, or workload data? Evidence gate, ownership, dependencies, time to proof
CTO Which reversible test comes first? Trace, tuned baseline, migration and rollback
Software/kernel engineer Which object and operation create the traffic? Source, allocation, counters, timeline, parity
Memory/package engineer Can it be powered, cooled, tested, repaired, yielded, and supplied? PDN, thermal, DFT/repair, process and supply diligence

The complete explanation stays public. The companion repository can carry schemas, calculators, synthetic fixtures, and source-linked examples. A paid release earns its price only when it adds runnable, pinned labs, traces, notebooks, failure cases, and maintained updates.

Frequently asked questions

What is HBM and why is it used for AI?

HBM is stacked DRAM connected to an accelerator through a very wide, short package interface. Its channel parallelism and proximity provide high local bandwidth for weights, KV cache, activations, latents, and communication buffers. Value still depends on useful traffic, capacity, power, heat, yield, and software.

How is HBM different from ordinary DRAM?

HBM uses DRAM cells, but organizes, stacks, and packages them differently from common DDR DIMMs. Vertical TSV connections, many channels, base/interface logic, and dense accelerator-side routing create a wider local path. The cell still needs sensing, restore, refresh, test, and repair.

What are TSVs and what does the HBM base die do?

TSVs are vertical conductors through thinned silicon. They connect dies in the stack. The base/interface die connects stack traffic to the package and can contain PHY, control, test, repair, RAS, or vendor-specific logic. Its exact partition is not universal.

Why is useful HBM bandwidth lower than peak bandwidth?

Peak bandwidth is an interface ceiling. Useful bandwidth loses to scattered access, insufficient concurrency, row conflicts, refresh, turnarounds, cache/TLB behavior, repeated materialization, synchronization, communication, power, or another bottleneck. Measure the named workload at the named boundary.

Is HBM always faster than GDDR7, DDR5, or LPDDR5X?

No universal comparison is valid. HBM favors aggregate local width, GDDR favors high per-pin board-level graphics bandwidth, DDR favors host capacity and serviceability, and LPDDR favors power/capacity-density integration. Device, interface width, package, controller, access pattern, and workload decide.

Can CXL memory replace HBM?

CXL can expand or pool coherent memory capacity, but it does not turn farther memory into HBM. It can hold workload-defined warm state when measured transfer and access meet the deadline. Hot mutable state usually remains a poor default candidate.

Can NVMe or HBF hold model weights?

Predictable, read-mostly pages may be streamable from flash-class capacity when lookahead, page usefulness, queue depth, buffering, and tail latency work. Active mutable state and random fine-grained access remain poor default fits. HBF is not yet shipping workload proof.

What is Intel XBM?

Intel XBM is cross-batch memory in US 2026/0191095 A1. The application describes backend-transistor DRAM, fine subchannels, TSV organization, test/repair, base-die options, serialized UCIe-facing links, and package variants.

Is Intel XBM a shipping product?

No. Intel XBM is a published patent application, not a confirmed shipping product. The filing does not prove working silicon, measured bandwidth, latency, energy, thermals, retention, yield, cost, software, schedule, qualification, or availability.

What is custom HBM or SPHBM4?

Custom HBM expands the logic and control that can be tailored around an HBM stack, especially at the base/interface die. Standard Package High Bandwidth Memory (SPHBM4) is JEDEC JESD330-4, an HBM4-stack direction with a different buffer/interface die for standard-package integration. A standard or vendor roadmap does not prove access, economics, or workload benefit.

How large is an LLM KV cache?

Logical size is 2 * layers * kv_heads * head_dim * bytes_per_element * sum(tokens_i across live sequences). For equal-length sequences, this becomes the familiar per-sequence token count times concurrency. Actual residency also includes paging, block rounding, fragmentation, metadata, sharing, replication, sharding, and quantization. Use the named model configuration and runtime allocation trace.

Why does Wan2.2 stress memory differently from a coding agent?

Wan2.2 repeatedly updates a large latent and runs conditional and unconditional denoising across many transformer blocks and steps. A coding agent grows KV, reuses prefixes, leaves the GPU for CPU tools and tests, branches, and retries. Their state lifetimes differ.

What is NVIDIA Kyber and how does it relate to HBM?

Kyber is NVIDIA's official next-generation MGX NVL rack design, not a GPU or HBM generation. One Kyber rack is planned as Vera Rubin Ultra NVL144, and eight Kyber racks later form Feynman NVL1152. HBM supplies local accelerator memory; Kyber changes the rack and scale-up boundary across which distributed workloads may move data. No public Wan2.2-on-Kyber benchmark was verified.

What does the hardware lottery mean for AI accelerators?

It means ideas that fit available hardware and software are easier to test, scale, and adopt. It does not mean GEMM is wrong. It tells engineers to measure operations and state paths that the current stack handles less well before proposing new hardware.

What would prove composable memory-movement primitives are useful?

They need a reproducible trace, durable bottleneck, tuned commercial baseline with parity, portable value across supported backends, a calibrated model plus bounded physical anchor, and repeated cross-workload evidence with a credible manufacturing plan and merchant fallback.

Secondary analyst research that helped us find the right questions. SemiAnalysis helped frame the history and economics of the memory wall, the HBM roadmap, Vera Rubin co-design, and the HBM4, custom-HBM, cooling, and packaging questions raised at ECTC 2026. SemiVision helped surface the AI memory-supply problem, the three HBM battlegrounds, 3D-stacked SRAM, and the question of whether ZAM could complement or replace HBM. These are secondary analyst synthesis, not the authority for product specifications or Touchdown measurements. Some links are paid. We do not reproduce their prose, tables, images, forecasts, or proprietary data. Primary vendor documents, standards, papers, patents, public code, and named receipts govern every public factual claim in this article.

Primary sources and evidence notes

The source list preserves evidence state. Vendor metrics apply only to the named artifact and date.


Additive implementation appendix: long-context decode datapaths, Blackwell qualification, and accepted-task resource receipts

In plain English. The main article explains the complete system. This appendix shows how to turn several important ideas into code and measurement packets: long-context attention, block-parallel softmax, low-precision Blackwell experiments, tool-window state, checkpointing, power, cooling, water, and cost per accepted patch.

Freeze notice. Everything above this line began as the canonical public-manuscript draft supplied on July 11, 2026. This appendix remains additive. The July 22 production pass installed three source-reviewed Touchdown figures. The July 23 update added the current AMD launch receipt, corrected the publication date, and preserved the earlier technical record. Neither pass replaced, deleted, reordered, compressed, or silently corrected the technical paragraphs, tables, equations, evidence labels, or source notes. Any later technical conflict must be recorded as a dated correction note rather than rewritten out of history.

Why long-context decode needs its own physical explanation

In plain English. During language-model decode, each new token uses attention state from earlier tokens. As the context grows, the model repeatedly reads a larger history while producing one next position. That makes data movement, placement, and tail latency different from the highly parallel first pass over the prompt.

The article has already separated prefill from decode. The next missing step is to show why long-context decode attention can stop looking like the large dense matrix work that accelerators handle best.

During prefill, many prompt positions are available together. Large projections and attention tiles can create substantial matrix-matrix work. During low-batch decode, one or a few new query positions attend over a large existing cache. The runtime repeatedly reads prior state, updates a running normalization, and reduces value vectors into one new output.

A simplified single-head decode step is:

TEXT
scores_i = dot(q, k_i) * scale

probability_i =
  exp(scores_i)
  / sum_j(exp(scores_j))

output =
  sum_i(probability_i * v_i)

The mathematical expression is compact. The physical work is not:

TEXT
read one new q
-> stream or gather many cached k vectors
-> calculate score fragments
-> maintain a numerically stable maximum and normalization sum
-> decide which v vectors must be read
-> accumulate the weighted output
-> write one result

At long context and low batch size, the query can be reused while the KV cache dominates bytes. This is a memory, reduction, and state-dependency problem even though dot products and multiply-accumulate operations remain inside it.

The KVStream post as an evidence-bounded design signal

[AUTHOR-REPORTED HACKATHON RESULT] Aarav Wattal reports that the KVStream team built a serial HLS/RTL streaming-attention tile, validated it against NumPy, passed RTL cosimulation, and synthesized it in Vivado at 200 MHz during an Anthropic, Etched, Cognition, and Mercor hackathon. The post reports:

TEXT
serial measured tile projection:
  approximately 1.3 times an H100 bandwidth baseline

modeled block-parallel plus Skip-Softmax:
  approximately 5.9 times with a streaming policy

modeled two-pass upper bound:
  approximately 7.7 times on peaked attention

These statements belong to the named 4K-context, attention-only projection and the team's stated model assumptions. They are not a GLM-5.2 benchmark, an end-to-end agent result, a B200 or GB200 result, a fabricated ASIC result, a complete power result, or independent reproduction.

The useful architectural claim is narrower:

Long-context decode attention can justify a datapath built around streamed KV access, local online-softmax state, block-level parallel reduction, and conditional suppression of value-path work.

That claim now becomes a Touchdown experiment lane.

Online softmax is a recurrence

In plain English. Softmax turns attention scores into normalized weights. An online version processes the score stream in pieces while carrying a small running summary: the largest score seen so far and the scaled sum needed for normalization. That avoids storing the full score matrix, but later pieces still depend on the summary produced by earlier pieces.

A stable streaming softmax can process scores without first materializing the complete score vector.

For scores processed in order, maintain:

TEXT
m_i:
  running maximum through position i

l_i:
  running normalized denominator through position i

o_i:
  running unnormalized weighted-value accumulator

For a new score s_i and value vector v_i:

TEXT
m_new = max(m_old, s_i)

old_scale = exp(m_old - m_new)
new_scale = exp(s_i - m_new)

l_new =
  old_scale * l_old
  + new_scale

o_new =
  old_scale * o_old
  + new_scale * v_i

At the end:

TEXT
output = o_final / l_final

This avoids a full score materialization, but it creates a dependency:

TEXT
state_i
depends on
state_i-1

A serial datapath must wait for the running maximum, denominator, and output state. KV bandwidth alone does not remove that recurrence.

Reference implementation

The following reference is intentionally simple. It is a correctness oracle, not an optimized kernel:

PYTHON
from __future__ import annotations

import math
from collections.abc import Sequence


def online_attention(
    scores: Sequence[float],
    values: Sequence[Sequence[float]],
) -> list[float]:
    if not scores or len(scores) != len(values):
        raise ValueError("scores and values must be non-empty and aligned")

    width = len(values[0])
    if any(len(vector) != width for vector in values):
        raise ValueError("all value vectors must have equal width")
    if any(math.isnan(score) or score == math.inf for score in scores):
        raise ValueError("scores may be finite or -inf for a masked position")

    running_max = -math.inf
    running_sum = 0.0
    accumulator = [0.0] * width

    for score, value in zip(scores, values):
        if score == -math.inf:
            continue
        new_max = max(running_max, score)
        old_scale = math.exp(running_max - new_max)
        new_scale = math.exp(score - new_max)

        running_sum = old_scale * running_sum + new_scale
        accumulator = [
            old_scale * old + new_scale * current
            for old, current in zip(accumulator, value)
        ]
        running_max = new_max

    # Explicit reference semantics for a fully masked row. A production kernel
    # must match its framework/backend contract rather than inheriting this choice.
    if running_sum == 0.0:
        return [0.0] * width

    return [component / running_sum for component in accumulator]

Required tests:

TEXT
compare against full softmax attention
random values
extreme score ranges
single element
all-equal scores
masked positions with -inf
all-masked row returns the declared zero vector
NaN and positive-infinity inputs fail visibly
peaked distributions
mixed positive and negative scores
FP32 reference
target dtype error envelope

Breaking the recurrence across blocks

In plain English. Independent blocks can compute partial softmax summaries, but the system must merge them with mathematically correct rescaling. Parallelism is useful only if the merge preserves the same answer within the declared numeric tolerance.

Divide the KV sequence into blocks. Each block independently computes a summary:

TEXT
block maximum:
  m_b

block denominator relative to m_b:
  l_b = sum_i_in_block(exp(s_i - m_b))

block weighted-value accumulator:
  o_b = sum_i_in_block(exp(s_i - m_b) * v_i)

Two block summaries can be merged exactly in real arithmetic:

TEXT
m = max(m_a, m_b)

l =
  exp(m_a - m) * l_a
  + exp(m_b - m) * l_b

o =
  exp(m_a - m) * o_a
  + exp(m_b - m) * o_b

This creates a reduction tree:

TEXT
KV blocks process in parallel
-> each emits m_b, l_b, o_b
-> hierarchical reducer merges block summaries
-> final normalization produces output

The recurrence has not disappeared. It moved from every KV position to a much smaller tree over block summaries.

Merge code

PYTHON
from __future__ import annotations

import math
from dataclasses import dataclass


@dataclass(frozen=True)
class SoftmaxBlock:
    maximum: float
    denominator: float
    weighted_value: tuple[float, ...]


def merge_blocks(left: SoftmaxBlock, right: SoftmaxBlock) -> SoftmaxBlock:
    if len(left.weighted_value) != len(right.weighted_value):
        raise ValueError("block widths must match")

    maximum = max(left.maximum, right.maximum)
    left_scale = math.exp(left.maximum - maximum)
    right_scale = math.exp(right.maximum - maximum)

    denominator = (
        left_scale * left.denominator
        + right_scale * right.denominator
    )

    weighted_value = tuple(
        left_scale * a + right_scale * b
        for a, b in zip(left.weighted_value, right.weighted_value)
    )

    return SoftmaxBlock(
        maximum=maximum,
        denominator=denominator,
        weighted_value=weighted_value,
    )

The hardware design must still answer:

TEXT
How many KV blocks execute concurrently?
Where does q remain resident?
Where do K and V stream from?
How large is each block?
How are exp and reduction units provisioned?
How wide is the block-summary network?
How much local SRAM is needed?
How does the design handle causal masking and page boundaries?
How does it handle GQA, MLA, DSA, quantized KV, and sparsity metadata?
How does numerical error change with tree order?

Skip-Softmax-style value-path gating

In plain English. A value-path gate asks whether some work can be skipped or approximated without changing the accepted result beyond its tolerance. The optimization is not valid because it runs faster. It is valid only when the quality gate, error bound, and workload receipt still pass.

A block can have a score maximum far below the global maximum. Its exponential contribution can become very small after rescaling.

A hardware policy may first bound the block's unnormalized exponential mass:

TEXT
unnormalized_mass_bound_b
=
number_of_items_b
* exp(block_max_b - current_global_max)

That expression does not by itself bound weighted-value or final-output error. To make an output claim, the experiment also needs a declared norm, component range, or other conservative bound for every value vector in the skipped block. If delta_b bounds the block's normalized probability mass and V_max bounds the chosen norm of its values, then a conservative bound for a kept-output that is renormalized after omission is:

TEXT
output_error_bound <= 2 * delta_b * V_max

The exact factor depends on the approximation and norm. The implementation must derive the bound it actually enforces. If the denominator, value bound, numeric state, or mask semantics are unavailable, the safe fallback is to read and compute the block.

Only when the declared mass and value bounds remain below the accepted error budget may the design avoid some work:

TEXT
skip or reduce:
  V reads
  exponential evaluations
  probability-value MACs

This is not automatically exact. The experiment must name:

TEXT
threshold
probability-mass bound
value-vector norm or range bound
derived output-error bound
dtype
accumulation order
error envelope
quality gate
fallback

For exact inference, gating must be proven not to alter the required output contract. For approximate inference, the accepted quality budget must be explicit.

Required labels:

TEXT
measured serial RTL result
modeled block-parallel result
modeled gated result
ideal upper bound
fabricated-silicon result
end-to-end accepted-task result

Never merge those categories into one speedup.

How this maps to GLM-5.2

In plain English. GLM-5.2 adds sparse expert routing to the attention and decode path. Total model capacity, active experts per token, numeric format, KV state, expert placement, and communication are separate decisions. A single label such as FP8 model does not settle all of them.

Do not assume that GLM-5.2 uses a conventional dense K/V layout.

The pinned model, engine, and backend can use architecture-specific state including:

TEXT
compressed latent KV
RoPE-related state
indexer state
indexer-K state
multiple cache groups
paged blocks
backend-specific scales and metadata

The exact GLM-5.2 decode experiment must begin by extracting the runtime cache geometry.

YAML
GLM52CacheGeometryReceipt:
  schema_version: touchdown.glm52-cache-geometry.v1
  run_id: string

  model:
    artifact_id: string
    revision: string
    architecture: string

  runtime:
    engine: sglang | vllm
    revision: string
    backend: string
    container_digest: string

  groups:
    - group_id: string
      semantic_role: string
      dtype: string
      block_size_tokens: int
      bytes_per_block: int
      scale_bytes_per_block: int | null
      metadata_bytes_per_block: int | null
      allocated_blocks: int
      used_blocks: int
      shared_blocks: int
      device_resident_blocks: int
      lower_tier_blocks: int

  evidence:
    startup_log: string
    allocator_dump: string
    profiler_trace: string | null

Only after this receipt exists may the design ask whether a KVStream-like tile can consume the actual state efficiently.

Blackwell lane: B200 before GB200 rack claims

In plain English. Prove the operation on one B200-class accelerator before using a GB200 rack label. A rack adds CPUs, multiple GPUs, scale-up fabric, networking, power, and cooling boundaries that a single-device run does not measure.

The qualification order is mandatory:

TEXT
1. One B200 GPU, synthetic attention fixture
2. One B200 GPU, pinned GLM-5.2 operation
3. Eight B200 GPUs, TP8 / EP configuration
4. One GB200 superchip or bounded tray allocation
5. Bounded GB200 NVL72 worker group
6. Shared rack under realistic concurrency

At every step record:

TEXT
exact GPU UUIDs
driver
CUDA
firmware
engine
model revision
quantization
KV dtype
parallelism
topology
background load
cooling boundary
power boundary
accepted output

No rack-level specification becomes per-request performance.

FP8 versus NVFP4 experiment

In plain English. Lower-precision formats can reduce weight capacity and traffic, but they can also change accuracy, kernel selection, scaling metadata, and hardware compatibility. Compare them on the same model, workload, quality gate, device scope, and software revision.

The comparison has four independent precision decisions:

TEXT
weight storage
operator input activations
accumulation
KV/cache representation

Required experiment arms:

TEXT
A:
  GLM-5.2 FP8 artifact
  runtime-resolved activation and accumulation path
  runtime-resolved KV dtype

B:
  NVIDIA GLM-5.2 NVFP4 artifact
  NVFP4 only for documented eligible expert linear operations
  shared expert and other modules preserved at their actual dtype
  runtime-resolved KV dtype

C:
  same as A with explicit FP8 E4M3 KV where supported

D:
  same as B with explicit FP8 E4M3 KV where supported

Freeze:

TEXT
repository revision
prompt and tools
retrieval spans
accepted-patch rubric
concurrency
parallelism
engine revision
hardware group
thermal starting state
retry budget

Report:

TEXT
checkpoint bytes
runtime allocated bytes
KV bytes
free HBM
largest free block
TTFT
inter-token latency
committed output rate
tool wall time
end-to-end task time
GPU and rack joules
temperature and throttle state
accepted patch

Real tool-calling timeline

In plain English. The GPU does not run continuously during an agent task. The model may pause while a CPU searches files, edits code, runs tests, waits for storage, or returns tool output. The receipt must show those gaps instead of charging every delay to HBM or GPU kernels.

The same task must remain alive across GPU and CPU phases:

TEXT
TURN-01
  prefix validation
  prefill
  reasoning decode
  read_file tool call

TOOL-01
  CPU / storage reads repository file
  GPU worker may serve other requests
  current request KV remains, moves, or is released by policy

TURN-02
  append tool result
  suffix prefill
  decode
  apply_patch tool call

TOOL-02
  patch writes
  lint and tests
  log capture

TURN-03
  append test result
  valid prefix reuse or safe miss
  repair decode if needed
  final answer

ACCEPTANCE
  diff review
  test verification
  accepted or rejected patch

Required phase instrumentation:

PYTHON
from contextlib import contextmanager
import time

import torch


@contextmanager
def trace_phase(name: str, event_writer):
    start_ns = time.monotonic_ns()
    torch.cuda.nvtx.range_push(name)
    event_writer({
        "event_type": "phase_start",
        "phase": name,
        "monotonic_ns": start_ns,
    })

    try:
        yield
    finally:
        end_ns = time.monotonic_ns()
        torch.cuda.nvtx.range_pop()
        event_writer({
            "event_type": "phase_end",
            "phase": name,
            "monotonic_ns": end_ns,
        })

The tool runner must never put raw secrets into the resource receipt. Store digests, counts, classifications, and approved excerpts only.

Checkpointing is a state-ownership experiment

In plain English. A checkpoint decides which state must survive interruption and who can restore it correctly. Saving everything is expensive. Saving too little can force recomputation or lose the exact context needed to continue safely.

The Doubleword checkpoint reverse engineering is architecture-specific to an RTX 4090 and driver 590.48.01. It proves that, in that fixture:

TEXT
checkpoint:
  GPU allocations and driver state moved into anonymous host memory
  NVIDIA VMAs and file descriptors disappeared
  the process disappeared from nvidia-smi

restore:
  CUDA resources were recreated
  fresh allocations were refilled
  device state survived

It also reports that, for an 8,578 MiB fixture:

TEXT
baseline full cycle:
  approximately 4.5 seconds

direct pipe protocol:
  approximately 3.9 seconds

transparent huge pages:
  approximately 1.6 seconds

preallocated staging and asynchronous unmap:
  approximately 1.2 seconds

Those measurements do not transfer automatically to Blackwell, Grace ARM, TP8, NCCL, SGLang, vLLM, or GLM-5.2.

The portable lesson is:

Checkpoint time can be dominated by host virtual-memory allocation, page faulting, and zeroing rather than the nominal accelerator link alone.

Blackwell checkpoint gates

Begin every row at NOT_RUN.

Gate B200 GB200
Single-process counter survives NOT_RUN NOT_RUN
GPU allocation disappears NOT_RUN NOT_RUN
Small one-GPU inference restores NOT_RUN NOT_RUN
CUDA graphs restore NOT_RUN NOT_RUN
HTTP service remains valid NOT_RUN NOT_RUN
Prefix cache is correct NOT_RUN NOT_RUN
KV state is correct NOT_RUN NOT_RUN
TP8 ranks quiesce together NOT_RUN NOT_RUN
NCCL state restores or rebuilds safely NOT_RUN NOT_RUN
CUDA IPC state restores or rebuilds safely NOT_RUN NOT_RUN
GLM-5.2 FP8 accepted patch NOT_RUN NOT_RUN
GLM-5.2 NVFP4 accepted patch NOT_RUN NOT_RUN
Restore p99 meets SLO NOT_RUN NOT_RUN
Total facility energy improves NOT_RUN NOT_RUN
Cost per accepted patch improves NOT_RUN NOT_RUN

Allowed values:

TEXT
PASS
FAIL
BLOCKED
NOT_OBSERVABLE
NOT_APPLICABLE
NOT_RUN

Quiescing a distributed worker

In plain English. To capture consistent distributed state, workers must reach a point where no hidden operation is still changing the data being saved. That controlled pause is quiescence.

A safe distributed checkpoint starts above CUDA:

TEXT
stop new admissions
-> mark worker DRAINING
-> finish or cancel active requests
-> finish tool-result prefills
-> finish or abort collectives under runtime rules
-> flush event and allocator receipts
-> distributed barrier
-> prove every rank has zero active requests and collectives
-> lock CUDA processes
-> checkpoint in a documented order
PYTHON
from __future__ import annotations

from dataclasses import dataclass
from enum import Enum


class RankState(str, Enum):
    SERVING = "serving"
    DRAINING = "draining"
    QUIESCED = "quiesced"
    CHECKPOINTED = "checkpointed"
    FAILED = "failed"


@dataclass(frozen=True)
class RankReceipt:
    rank: int
    pid: int
    active_requests: int
    active_collectives: int
    state: RankState


def group_is_quiesced(receipts: list[RankReceipt]) -> bool:
    return bool(receipts) and all(
        receipt.active_requests == 0
        and receipt.active_collectives == 0
        and receipt.state is RankState.QUIESCED
        for receipt in receipts
    )

Never checkpoint one rank in the middle of a collective while peers continue.

Full checkpoint versus selective reconstruction

In plain English. A full checkpoint stores more state and can simplify recovery. Selective reconstruction stores less and rebuilds some state later. The right choice depends on restore time, determinism, correctness, storage traffic, and the deadline for resuming useful work.

Compare four modes:

Mode Weights KV/prefix CUDA context/graphs Host/application state
Full CUDA image Copy Copy Copy where supported Retain
Selective session image Reload from host/storage Preserve only required sessions Copy bounded control state Retain
Runtime-native rebuild Reload Restore through engine/cache layer or recompute Recreate Retain receipts and identities
Cold restart Reload Miss/recompute Recreate Reconstruct from durable workflow state

The preferred design is not known in advance.

A full GLM-5.2 worker can contain hundreds of gigabytes of weights and runtime state. Copying all of it may lose to reloading from a staged host or storage path. Conversely, reconstructing graphs and runtime state can dominate a smaller image. Measure the complete transition.

State-aware tool-window policy

In plain English. Agent context should not grow without a rule. Preserve evidence that still changes the decision, summarize or move older material when provenance survives, and reject any compaction that causes the verifier or policy checks to fail.

A tool call creates a next-use prediction problem.

Inputs:

TEXT
expected tool duration distribution
current GPU pressure
other queued requests
request KV bytes
shared-prefix references
weight residency ownership
checkpoint bytes
measured checkpoint and restore p99
next-turn deadline
host-memory pressure
cooling and power state

Scheduling actions and state-lifecycle actions are separate. Serving another request does not say whether this request's KV stayed in HBM, moved, or was checkpointed.

TEXT
scheduling:
  HOLD_CAPACITY
  SERVE_OTHER_REQUESTS

state lifecycle:
  KEEP_WARM
  RELEASE_REQUEST_KV
  OFFLOAD_REQUEST_KV
  SELECTIVE_CHECKPOINT
  FULL_CHECKPOINT

supervisor-only failure or rollback outcome:
  TERMINATE_AND_REBUILD

TERMINATE_AND_REBUILD is deliberately not a StateAction below. A local tool-window helper must not destroy an active session because a cost heuristic prefers it. Termination requires a separate supervisor decision with explicit authorization, a receipt describing which state will be lost or reconstructed, and a verifier-backed recovery path.

PYTHON
from dataclasses import dataclass
from enum import Enum


class SchedulingAction(str, Enum):
    HOLD_CAPACITY = "hold_capacity"
    SERVE_OTHER_REQUESTS = "serve_other_requests"


class StateAction(str, Enum):
    KEEP_WARM = "keep_warm"
    RELEASE_REQUEST_KV = "release_request_kv"
    OFFLOAD_REQUEST_KV = "offload_request_kv"
    SELECTIVE_CHECKPOINT = "selective_checkpoint"
    FULL_CHECKPOINT = "full_checkpoint"


@dataclass(frozen=True)
class ToolWindow:
    predicted_ms: float
    next_turn_slack_ms: float
    restore_p99_ms: float
    full_checkpoint_restore_p99_ms: float
    gpu_pressure: float
    checkpoint_cost_usd: float
    keep_warm_cost_usd: float
    active_session_requires_exact_kv: bool
    selective_checkpoint_preserves_exact_kv: bool
    queued_request_count: int


@dataclass(frozen=True)
class ToolWindowDecision:
    scheduling: SchedulingAction
    state: StateAction


def choose_state_action(window: ToolWindow) -> StateAction:
    # Exact-KV correctness is checked before cost or utilization.
    if window.active_session_requires_exact_kv:
        if (
            window.selective_checkpoint_preserves_exact_kv
            and window.restore_p99_ms < window.next_turn_slack_ms
        ):
            return StateAction.SELECTIVE_CHECKPOINT
        if window.full_checkpoint_restore_p99_ms < window.next_turn_slack_ms:
            return StateAction.FULL_CHECKPOINT
        return StateAction.KEEP_WARM

    if (
        window.gpu_pressure >= 0.9
        and window.restore_p99_ms < window.next_turn_slack_ms
    ):
        return StateAction.OFFLOAD_REQUEST_KV

    if (
        window.checkpoint_cost_usd < window.keep_warm_cost_usd
        and window.restore_p99_ms < window.next_turn_slack_ms
    ):
        return StateAction.SELECTIVE_CHECKPOINT

    return StateAction.KEEP_WARM


def choose_tool_window_decision(window: ToolWindow) -> ToolWindowDecision:
    scheduling = (
        SchedulingAction.SERVE_OTHER_REQUESTS
        if window.queued_request_count > 0 and window.predicted_ms > 0
        else SchedulingAction.HOLD_CAPACITY
    )
    return ToolWindowDecision(
        scheduling=scheduling,
        state=choose_state_action(window),
    )

This is a teaching policy, not a production optimum. A real implementation still needs measured queueing, checkpoint and restore distributions, cache identity, failure handling, power state, and a verifier-backed rollback rule.

Power measurement boundaries

In plain English. Power is an instantaneous rate. Energy is power integrated over a named time interval. A task receipt must align the workload start and stop with the device, server, or facility meter instead of treating a rated maximum as measured task energy.

Collect at least:

TEXT
GPU device telemetry
complete server input meter
rack PDU
cooling equipment boundary when available
facility meter interval when available

Never add children to a parent meter that already includes them.

YAML
PowerBoundary:
  boundary_id: string
  parent_boundary_id: string | null
  includes:
    - gpu
    - cpu
    - host_memory
    - nvlink_switch
    - nic
    - storage
    - fans
    - pumps
    - conversion_losses
  measurement_kind: measured | allocated | modeled | unknown
  meter_id: string | null
  sampling_interval_ms: float | null

Energy integration:

PYTHON
from __future__ import annotations

from dataclasses import dataclass


@dataclass(frozen=True)
class PowerSample:
    monotonic_ns: int
    power_w: float


def integrate_joules(samples: list[PowerSample]) -> float | None:
    ordered = sorted(samples, key=lambda sample: sample.monotonic_ns)

    if len(ordered) < 2:
        return None

    total = 0.0

    for left, right in zip(ordered, ordered[1:]):
        if right.monotonic_ns <= left.monotonic_ns:
            raise ValueError("timestamps must increase")

        elapsed_s = (
            right.monotonic_ns - left.monotonic_ns
        ) / 1_000_000_000

        total += (
            (left.power_w + right.power_w)
            / 2.0
            * elapsed_s
        )

    return total

Unknown or invalid intervals return null, not zero.

Cooling lag and heat balance

In plain English. Electrical work becomes heat quickly, but sensors, coolant, and facility equipment respond on different time scales. The accounting window must include that lag without pretending coolant flow is electrical energy or water consumption.

Electrical power changes before temperature and coolant return temperature.

The aligned timeline must include:

TEXT
kernel or tool phase
GPU power
server/rack power
GPU and memory temperatures
clock/throttle state
coolant supply temperature
coolant return temperature
flow
differential pressure
pump/CDU power
ambient condition

For a single-phase coolant:

TEXT
heat_transport_W =
  mass_flow_kg_s
  * specific_heat_J_kgK
  * delta_T_K

This is the heat transported across that loop boundary. It is not automatically task electrical power.

Required lag windows:

TEXT
pre-task baseline
task electrical interval
post-task thermal response
return-to-baseline or declared truncation

Water remains an allocation ledger

In plain English. Water circulating in a closed cooling loop is not the same as water consumed at the site. Site withdrawal, discharge, consumption, and water associated with electricity generation need separate boundaries and records.

Record separately:

TEXT
closed-loop circulation
technology-loop makeup and loss
site withdrawal
site consumption
source-energy water
embodied manufacturing water

Never convert a coolant flow rate into water consumption.

YAML
WaterReceipt:
  run_id: string
  facility_interval_id: string

  circulation_L: float | null
  technology_loop_makeup_L: float | null
  site_withdrawal_L: float | null
  site_consumption_L: float | null
  source_energy_water_L: float | null
  embodied_water_estimate_L: float | null

  allocation_method: string
  evidence_state: measured | allocated | modeled | unknown

Accepted-patch economics

In plain English. A cheap model call is not a cheap coding result if retries, tool time, GPU idle time, failed tests, and human repair erase the savings. The denominator is an accepted patch, not a token.

The final denominator is the verified code change.

TEXT
total_cost_per_accepted_patch =
  (
    model_and_accelerator_cost
    + CPU_and_tool_cost
    + memory_storage_network_cost
    + checkpoint_restore_cost
    + facility_energy_cost
    + allocated_cooling_and_water_cost
    + retry_and_failure_cost
    + human_review_cost
    + allocated_capital_and_support
  )
  / accepted_patches

All attempted runs stay in the numerator. Only accepted patches enter the denominator.

Required comparison matrix:

Artifact Cache Tool result Lifecycle Acceptance
FP8 hit first pass warm measured
FP8 miss retry warm measured
FP8 hit long tool selective checkpoint measured
FP8 hit long tool full checkpoint measured
NVFP4 hit first pass warm measured
NVFP4 miss retry warm measured
NVFP4 hit long tool selective checkpoint measured
NVFP4 hit long tool full checkpoint measured

Every cell begins unpopulated.

Visual additions

In plain English. The visual system follows the same task identity through software, hardware, movement, facility, and cost. Animation explains a transition; it does not upgrade an architectural illustration into a measured run.

Append these scenes to the continuous left-side machine:

  1. Decode is not one GEMM

    • one query vector
    • many KV blocks
    • online maximum, denominator, and output accumulator
  2. Serial recurrence

    • each score updates the previous state
    • KV bandwidth and recurrence highlighted separately
  3. Block-parallel summaries

    • independent block summaries
    • reduction tree
    • exact merge equations
  4. Gated value path

    • block contribution bound
    • skipped V reads and MACs labeled modeled or measured
  5. GLM-5.2 actual cache groups

    • populated from runtime receipt
    • generic K/V graphics prohibited when the backend uses another representation
  6. B200 worker topology

    • exact eight-GPU group
    • HBM pools
    • NVLink/NVSwitch path
    • host memory and process ranks
  7. GB200 rack slice

    • Grace CPU
    • Blackwell GPUs
    • LPDDR and HBM
    • NVLink-C2C and NVSwitch
    • liquid loop and rack meter boundaries
  8. Tool-call window

    • request pauses
    • sandbox works
    • worker serves, offloads, checkpoints, or waits
  9. Checkpoint host boundary

    • page allocation
    • zeroing
    • copy out
    • context teardown
    • host residency
    • restore
  10. Electrical and thermal lag

    • GPU power responds first
    • junction temperature follows
    • coolant return follows
    • facility response follows
  11. Accepted-patch ledger

    • tokens
    • KV
    • GPU and server joules
    • cooling
    • water allocation
    • tool and human work
    • acceptance

Additive proof gates

In plain English. Every deeper claim needs a stronger receipt. Source code can prove that a path exists. A trace can prove that it ran. Counters can prove measured activity at a named boundary. Only the verifier can prove that the final task was accepted.

This appendix becomes publishable only when:

TEXT
the frozen manuscript remains byte-for-byte unchanged above the marker
the exact GLM-5.2 artifact and engine are pinned
actual cache geometry is captured
FP8 and NVFP4 mixed-precision scopes are explicit
tool calls are joined to model phases
B200 measurements are separated from GB200 measurements
checkpoint claims carry platform and driver receipts
power boundaries do not double count
cooling lag is included
water circulation is not labeled consumption
all author-reported KVStream numbers retain their evidence label
accepted-patch results close the trace
unknown remains unknown

Final additive doctrine

In plain English. Keep the real task, technical path, physical boundary, evidence state, and accepted result connected. A useful explanation can become simpler to navigate without deleting the detail required to verify it.

Long-context decode can be limited by streamed state, reduction recurrence, and value-path work, not merely dense matrix throughput.

A custom datapath earns attention only after it is compared with a tuned Blackwell software baseline on the same model state, quality contract, and accepted task.

B200 and GB200 are not interchangeable measurement scopes.

FP8, NVFP4, accumulation precision, and KV precision are separate choices.

Checkpointing does not save energy merely because GPU allocations disappear.

Coolant flow is not water consumption.

Tokens are not the final product.

The trace ends only when the patch passes.