Class v6: the sound per-load form's 5090 rows measured (shl4096x1 at class v4's count: 116 W over class v3 against the whole block's 152 unlocked, 59 against 84 at the knee; shl2304x3 19 W and 6.6 W over the whole block)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 11:44:21 +00:00
parent 024c5aecb4
commit 6cc3bfec87

View file

@ -251,7 +251,7 @@ The wording this lane proposes for main's word, if the SRAM-store reading stands
| Layer | What it is | The chip rows | The card cost | Status | | Layer | What it is | The chip rows | The card cost | Status |
|---|---|---|---|---| |---|---|---|---|---|
| 5. The shadow placed per load, in its sound form (`mx8+shl4096x1`: one pass of a 256-instruction sub-block after every load, the same 4,096 shadow instructions per iteration as class v4) | the capex lever of the research file's section 16.2: the chip's core sits inside every read's dependency, so controller and core share one N5-class die or an interposer | the project about USD 60 M against 30 M, the break-even cap about USD 200 M against 100 M (modelled); `k` unchanged | measured on build-1 (10:4x UTC, this crate): 234 of 256 seeds accept within 32 attempts at 0.927 rejection per candidate (P(exhaust at 256) about 4e-9); the verifier 8.28 to 8.82 ms on core 40 with core 88 loaded against 8.33 to 8.63 for class v4's shape; the 5090 rows (ids 75ca9547da21b200 and bbfdfc1dcdda0b46) in the hash lane's v6 job at about 12:45 UK, the default at 16:30 the 16 x 27 v2 export's measured 13 to 14 W under the whole block as a labelled proxy; the Apple footprint of a 4,096-line block owed (the 1,024-line block cost the M5 Max 17 percent on 6 October) | a candidate class; the 16 x 27 iterated form stays dead (20.2a-close of the research file, corrected to the no-era figure) | | 5. The shadow placed per load, in its sound form (`mx8+shl4096x1`: one pass of a 256-instruction sub-block after every load, the same 4,096 shadow instructions per iteration as class v4) | the capex lever of the research file's section 16.2: the chip's core sits inside every read's dependency, so controller and core share one N5-class die or an interposer | the project about USD 60 M against 30 M, the break-even cap about USD 200 M against 100 M (modelled); `k` unchanged | measured on build-1 (10:4x UTC, this crate): 234 of 256 seeds accept within 32 attempts at 0.927 rejection per candidate (P(exhaust at 256) about 4e-9); the verifier 8.28 to 8.82 ms on core 40 with core 88 loaded against 8.33 to 8.63 for class v4's shape; **the 5090 rows MEASURED (the hash lane's v6 job, 11:14 to 11:20 UTC, every pack PASS at both states): mx8-genesis 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at 1,300 (0.599 MH/W); class v4's shape mx8_sh256x27 137.62 at 464.6 and 126.99 at 296.8 (0.428); the sound per-load form mx8_shl4096x1 135.99 at 428.6 and 126.08 at 271.8 (0.464); mx8_shl2304x3 135.85 at 483.6 and 125.88 at 303.4 (0.415). The rate within 1.3 percent of the control on both forms; the premium over class v3: shl4096x1 116.4 W unlocked and 59.2 W at the lock against the whole block's 152.4 and 84.2, so at class v4's instruction count the one-pass per-load form costs the 5090 24 percent LESS unlocked and 30 percent less at the knee (the 16-instruction block effect of 6 October, now with a sound construction); shl2304x3 (three passes of 2,304) costs 19 W MORE unlocked and 6.6 W more at the lock than the whole block, so the saving is the one-pass shape, not the placement**; the Apple footprint of a 4,096-line block owed (the 1,024-line block cost the M5 Max 17 percent on 6 October) | a candidate class; the 16 x 27 iterated form stays dead (20.2a-close of the research file, corrected to the no-era figure) |
| 6. The register-file width drawn per era in {8, 16, 32} | the link tax on layer 5: 4.5 to 9 TB/s of die-to-die traffic closes the interposer branch (modelled, the link figures approximate) | forces the single die | about 0 rate on every card by the occupancy arithmetic (unmeasured) | research | | 6. The register-file width drawn per era in {8, 16, 32} | the link tax on layer 5: 4.5 to 9 TB/s of die-to-die traffic closes the interposer branch (modelled, the link figures approximate) | forces the single die | about 0 rate on every card by the occupancy arithmetic (unmeasured) | research |
| 7. Warp-uniform data-dependent block selection (B drawn sub-blocks, one selected per iteration by a warp-folded register, no divergence) | moves the FPGA lane only (a per-program bitstream must carry every block) | nothing against the `f = 1` chip | B capped by the Apple compile footprint | research | | 7. Warp-uniform data-dependent block selection (B drawn sub-blocks, one selected per iteration by a warp-folded register, no divergence) | moves the FPGA lane only (a per-program bitstream must carry every block) | nothing against the `f = 1` chip | B capped by the Apple compile footprint | research |
| The reserve ordered by hardware orthogonality inside layer 3 | shfla first, the int8 tile last | section 4.2's order | | taken | | The reserve ordered by hardware orthogonality inside layer 3 | shfla first, the int8 tile last | section 4.2's order | | taken |