Public interactive companion. This page is intentionally noindex so the complete article owns the canonical search result. It proves deterministic teaching arithmetic over pinned source text. It proves no model load, kernel dispatch, HBM or fabric traffic, latency, power, energy, cooling, water, cost, or accepted workload run. Back to the complete HBM article.
C-001 code-to-HBM micro-lab: one matrix multiply, one term at a time
The C-001 workload is one bounded coding-agent task: Hermes Agent, backed by pinned
GLM-5.2-FP8, must produce a patch that passes a named verifier. Inside that path, the
GLM-5.2 router decides which experts each token reaches, and that decision starts as one
matrix multiply. This micro-lab steps a small 3×4 · 4×3 teaching multiply
through all 46 deterministic states: one start state, 36
multiply-accumulate terms, and 9 output commits. The same reproducible state model is
used throughout this standalone companion.
The pinned source line this teaches
# huggingface/transformers @ b7f0101522ddf1fb7b49aef9aff85fa22ceff36b# src/transformers/models/glm_moe_dsa/modeling_glm_moe_dsa.py - Apache-2.0# GlmMoeDsaTopkRouter forward body - PINNED SOURCE / NOT EXECUTEDhidden_states = hidden_states.view(-1, self.hidden_dim)router_logits = F.linear(hidden_states.type(torch.float32), self.weight.type(torch.float32))scores = router_logits.sigmoid()
The highlighted line is a real GEMM: every router logit is a dot product between one token's hidden state and one expert row of the router weight. The multiply below uses small Touchdown-authored integer teaching values, not workload tensors.
Touch the math
Keyboard: arrow keys step, Space plays or pauses, Home and End jump. The state index
is kept in the URL (?step=N) so a link reproduces the exact state.
Values are shown at display scale; the canonical model is integer-exact
(Q10 inputs, Q100 accumulators), so every state is reproducible bit for bit.
Where this sits in the machine
- Pinned source - the router linear above (
PINNED SOURCE / NOT EXECUTED) - PyTorch operator -
F.linear; the executable kernel is selected at run time (UNKNOWNuntil a run is captured) - NVIDIA B200 SM / Tensor Core role - the compute unit class that would execute a real dispatch (
ARCHITECTURE ONLY) - HBM3E - where weights and activations would reside (
ARCHITECTURE ONLY / RUN NOT CAPTURED)
What is derived and what is not measured
- MAC terms (derived)
- 3 × 4 × 3 = 36; plus 9 commits and 1 start = 46 states
- Router weight payload (derived)
- 256 experts × 6,144 hidden × 1 byte (FP8 E4M3) = 1,572,864 bytes before scales
- Physical HBM read/write bytes
- NULL / RUN NOT CAPTURED
- Latency, throughput, power, energy
- NULL / RUN NOT CAPTURED
- Cooling, water, cost, accepted patch
- NULL / RUN NOT CAPTURED
- run_id
- null
Continue through the complete HBM path
Open the standalone compute learning atlas at the coding-agent HBM layer. The atlas keeps source, operation, physical hardware, data movement, business meaning, evidence state, and the exact next proof synchronized. Finishing this deterministic teaching exercise proves only that the authored arithmetic and navigation work. It is not course completion, a workload run, or an accepted customer outcome.