Merge remote-tracking branch 'build/master' into reference-apps

This commit is contained in:
igneum-labs 2026-10-08 15:06:22 +00:00
commit 9d9b46de99
4 changed files with 72 additions and 8 deletions

View file

@ -178,6 +178,18 @@ win, because on this hash the memory clock is the rate and the core clock is onl
climb may raise the memory clock (5 percent of the range per probe) where the vendor exposes it, and no measured row
yet shows a gain from it (the 5090's memory sits at 13,801 of 14,001 MHz).
Amendment, 15:5x UK (floor lane 1's memory-clock ladder on PC 1, the 5090 alone at the 1,300 MHz core lock, job
run-ca4-pc1-memclk-5090-20261008-b, every row bit-exact; `docs/analysis/class-v6/floor/sm-sparse.md` section 3.1 at
66bc6ca7): the memory clock is not a lever, now measured. The driver holds the 5090's memory in two states only (asks
of 8,001 MHz and above read 13,801; 6,001 and below read 7,001, the PHY's half-rate state). Class v3 at the lock:
13,801 MHz 134.1 to 134.5 MH/s at 222.7 to 223.7 W (1.66); 7,001 MHz 76.0 MH/s at 141 to 142 W (1.86 to 1.87).
Class v4: 134.3 to 134.5 at 308 to 309 W (2.29 to 2.30); 76.1 at 184 to 187 W (2.42 to 2.46). The half-rate state
loses 43 percent of the rate (the activate ceiling scales with the memory clock) for 37 percent of the watts, so the
energy per hash rises 12 percent on class v3 and 5 on class v4; the 82 W between the states at 7.4 G reads per second
lost is 11 nJ per read at the margin, equal to the DRAM chase probe's whole-path figure, so the memory-side term is 80
to 90 W of the 223 W at the lock and is bought one for one with the rate. The table's "never touched" line stands as
a measurement; the tiers table carries `mem_mhz` 0 on every row.
## 4. The LPDDR6 and unified-memory tier, five years out
The method. The M5 Max row is the only measured unified-memory point and it splits three ways at the meter: the DRAM

View file

@ -40,8 +40,9 @@ M5 Max's 0.78 is measured (6 October). "Edge" is the card's microjoules over the
On the H100 the persistent shape itself loses 4 percent of the class v4 rate and the self-tune installs nothing.
PC 1's class v3 at stock agrees with the rented ladder at every rung (43 SMs 136.5 of 136.7 MH/s at 10 W under
base; 170 SMs x 8 warps 137.1 at 310.8 W, 3.5 percent under base, the best shape on that board).
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is the one
knob left after the core lock.** The decomposition probes (section 3) put the 5090's draw over idle at the hash
4. **The residual is the core clock domain and the memory path on the GPU side, and the memory clock is not a
lever either (section 3.1: the 5090's PHY has two states, 13,801 and 7,001 MHz; the half-rate state loses 43
percent of the rate for 37 percent of the watts, so the energy per hash rises 12 percent).** The decomposition probes (section 3) put the 5090's draw over idle at the hash
into: the resident SMs issuing nothing (the SM clock domain awake), the L2 and crossbar path per read, and the DRAM
path per read on top of the memory's own modelled 2.0 nJ; the memory-clock ladder on PC 1 at the 1,300 lock reads
the PHY's share directly. The honest floor per tier is the base row at stock within 2 percent, the core lock
@ -536,7 +537,30 @@ The floorsm ladder
### PC 1, the memory-clock ladder at the 1,300 MHz lock, job `run-ca4-pc1-memclk-5090-20261008`
The memory-clock ladder: not landed.
The memory-clock ladder
| Pack | State | Shape | MH/s | W | Microjoules | SM MHz | Mem MHz | Temp max | Check |
|---|---|---|---|---|---|---|---|---|---|
| v4-devnet-epoch0 | unlocked | base (grid) | 136.552 | 464.3 | 3.400 | 2852.1 | 13801 | 70 | PASS |
| mx8-devnet-epoch0 | unlocked | base (grid) | 136.29 | 328.1 | 2.407 | 2850 | 13801 | 65 | PASS |
| v4-devnet-epoch0 | 1300 | base (grid) | 134.34 | 309 | 2.300 | 1290 | 13801 | 60 | PASS |
| mx8-devnet-epoch0 | 1300 | base (grid) | 134.089 | 222.9 | 1.662 | 1290 | 13801 | 57 | PASS |
| v4-devnet-epoch0 | 1300m14001 | base (grid) | 134.513 | 308.1 | 2.290 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m14001 | base (grid) | 134.342 | 222.7 | 1.658 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m12001 | base (grid) | 134.528 | 308.6 | 2.294 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m12001 | base (grid) | 134.252 | 223.1 | 1.662 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m10001 | base (grid) | 134.487 | 308.9 | 2.297 | 1290 | 13801 | 59 | PASS |
| mx8-devnet-epoch0 | 1300m10001 | base (grid) | 134.287 | 223.7 | 1.666 | 1290 | 13801 | 55 | PASS |
| v4-devnet-epoch0 | 1300m8001 | base (grid) | 134.541 | 308.7 | 2.294 | 1290 | 13801 | 60 | PASS |
| mx8-devnet-epoch0 | 1300m8001 | base (grid) | 134.33 | 223 | 1.660 | 1290 | 13801 | 56 | PASS |
| v4-devnet-epoch0 | 1300m6001 | base (grid) | 76.09 | 187 | 2.458 | 1290 | 7001 | 53 | PASS |
| mx8-devnet-epoch0 | 1300m6001 | base (grid) | 76.009 | 142.2 | 1.871 | 1290 | 7001 | 51 | PASS |
| v4-devnet-epoch0 | 1300m5001 | base (grid) | 76.091 | 184.4 | 2.423 | 1290 | 7001 | 50 | PASS |
| mx8-devnet-epoch0 | 1300m5001 | base (grid) | 76.007 | 141.1 | 1.856 | 1290 | 7001 | 48 | PASS |
| v4-devnet-epoch0 | 1300m3001 | base (grid) | 76.082 | 183.8 | 2.416 | 1290 | 7001 | 50 | PASS |
| mx8-devnet-epoch0 | 1300m3001 | base (grid) | 75.96 | 140.9 | 1.855 | 1290 | 7001 | 47 | PASS |
| v4-devnet-epoch0 | unlocked-end | base (grid) | 136.562 | 452.5 | 3.314 | 2866.6 | 13801 | 61 | PASS |
| mx8-devnet-epoch0 | unlocked-end | base (grid) | 136.298 | 318.6 | 2.338 | 2865 | 13801 | 60 | PASS |
### The decomposition probes (the worker's --microbench: sleep = full residency and no issue; the L2 chase; the dependent DRAM chase; the ALU probe; 30 s each with the sampler's watts and both clocks)
@ -626,6 +650,35 @@ clock for the PHY. The core lock takes the first three down with the clock (the
reads per second on PC 1, 20.3); the memory-clock lock is the only knob on the fourth, and no rented host allows it, so
its row is PC 1's (the memory-clock ladder table in section 1, the job `run-ca4-pc1-memclk-5090-20261008`).
### 3.1 The memory-clock ladder (PC 1, the 5090 alone, job `run-ca4-pc1-memclk-5090-20261008-b`, 15:0x to 15:4x UK, the helper's `lmc` verb at the 1,300 MHz core lock; the amendment of 15:50 UK)
The driver holds the 5090's memory clock at two points only: every ask at or above 8,001 MHz reads 13,801 on the row
and every ask at or below 6,001 reads 7,001 (the GDDR7 PHY's half-rate state). The rows, both bit-exact against the
Mac's fingerprints:
| Pack | Core lock, memory ask | Memory MHz on the row | MH/s | Watts | Microjoules per hash | nJ per dependent read, whole card | Label |
|---|---|---|---|---|---|---|---|
| class v3 (mx8) | 1,300, no memory lock | 13,801 | 134.09 | 222.9 | 1.662 | 13.0 | measured |
| class v3 | 1,300, lmc 14001 / 12001 / 10001 / 8001 | 13,801 | 134.3 to 134.5 | 222.7 to 223.7 | 1.658 to 1.666 | 13.0 | measured (no change) |
| class v3 | 1,300, lmc 6001 / 5001 / 3001 | 7,001 | 76.0 | 140.9 to 142.2 | 1.855 to 1.871 | 14.5 | measured |
| class v4 (sh256x27) | 1,300, no memory lock | 13,801 | 134.34 | 309.0 | 2.300 | | measured |
| class v4 | 1,300, lmc 8001 and above | 13,801 | 134.5 | 308.1 to 308.9 | 2.290 to 2.297 | | measured |
| class v4 | 1,300, lmc 6001 and below | 7,001 | 76.1 | 183.8 to 187.0 | 2.416 to 2.458 | | measured |
| both, unlocked (the pass's own base) | none | 13,801 | 136.3 / 136.6 | 318.6 to 328.1 / 452.5 to 464.3 | 2.34 to 2.41 / 3.31 to 3.40 | | measured (the heat drift across the pass) |
The reading: the memory clock is not a lever either. At the half-rate PHY state the rate falls 43 percent (134 to 76
MH/s: the hash is bound by the memory's activate ceiling, which scales with the memory clock) and the watts fall 37
percent (82 W on class v3, 125 W on class v4), so the energy per hash RISES 12 percent on class v3 (1.66 to 1.86) and
5 percent on class v4 (2.30 to 2.42). Read as a decomposition, the 82 W between the two memory states on class v3 is
the memory path's clock-scaled share at 7.4 G reads per second of rate lost: 11 nJ per read at the margin, the same
figure the DRAM chase probes give for the whole memory path on the GPU side (10.9 to 11.8 nJ). So the memory-side term
of the 5090's draw is about 80 to 90 W of its 223 W at the lock (the PHY, controllers and L2 at the memory clock plus
the devices' own 55 W modelled), and it is bought one for one with the rate: the honest floor at the 1,300 lock stays
1.66 microjoules on this board (1.58 on 20.3b's cooler pass), and no clock or occupancy knob on the card takes it lower
without taking the rate with it. The class v4 premium at the half-rate state is 43 W for 76 MH/s x 102,100 ops = 5.5
pJ per counted op (6.3 at the full memory clock on the same pass): the shadow's marginal is the core domain's, not the
memory's, as expected.
## 4. The honest floor per tier, and the fraction a chip cannot strip
The chip of the record pays the memory's own energy per read, its static power and a controller and PHY of its own
@ -702,10 +755,7 @@ AMD or Apple owner gets nothing from this lane (the shapes are CUDA's; Metal and
reading the sparse shapes lower than 20.3b did on the same board at the same lock (85 SMs 111.0 against 129.8 MH/s);
the pass ran across the founder's card swap on PC 1 (14:2x UK), so the class v3 lock row stays 20.3b's (43 SMs
126.3 MH/s at 205.4 W against the full grid's 134.0 at 211.4: 1.63 against 1.58 microjoules, nothing saved) and the
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder is queued behind the hash lane's two
jobs (run-ca4-pc1-memclk-5090-20261008-b) and lands in section 1 when its upload does.
- The decomposition probes on the 4090 and the 5090 host c: in section 1 when their runs end (queued behind the
sweeps on the same cards so no row shares the card).
ladder is re-run on a quiet PC 1 before it replaces it. The memory-clock ladder landed at 15:4x UK (section 3.1).
- The memory side's own share per tier is modelled (chip-model 5.3); the memory-clock ladder is the one measurement of
its PHY term, and only on PC 1.
- Class v5 on the rented 5090s was not run (the v5-kits worker carries the leaf upload and not the sparse shapes; the

View file

@ -4,6 +4,8 @@ Three texts and one ledger row, written by the Counter ASIC lane, which owns the
## 1. The home page's chip line (replaces the hero sentence served since 6 October 16:21Z)
SUPERSEDED on the served pages on 8 October 2026 (the class v6 floor close, docs/design/class-v6-rotating-family.md section 10, amended into three statements at the founder's word): the served chip text is now the audit lane's spec-accept-23 wording (the 64-register window, 2.2x to 2.4x modelled against the GPU tier, 2.0x on the GPU's own node; economic resistance stated as conditions, not a forecast; response capability per boundary), with the 2.1x/3.4x line, the 5x to 9x baseline, the k about 0.33 column and the lifetime claim struck. The text below is the 7 October wording, kept as the record of what was served between 7 October and the 8 October landing.
Built for graphics cards. At launch the strongest chip in our public model reaches 2.1x per joule against an RTX 5090 with a core as good as a GPU lane, 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. Class v5 then makes the dataset the chain's own state, so a chip that stores it or recomputes it is wrong on every item. Without class v4 the same chip would reach 5x to 9x. The model and every measurement are public.
## 2. The litepaper's chip section (replaces the paragraph that begins "The chip model: 5x to 9x per joule")

File diff suppressed because one or more lines are too long