Class v6: the x16 quiet-core re-run (9.04 to 9.10 ms loaded on a quiet host, 10.1 to 14.8 under load): x16 stays admissible false at genesis, decided by the 2019-class core as ladder rung 3 is

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 0182c50dc (counter-asic-4) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 10:16:17 +00:00
parent f310ac4e86
commit b9dc2dac42

View file

@ -35,7 +35,7 @@ PENDING the hash lane's rows (16:00 UK): the x16 mixer verifier and build; the t
| Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number |
|---|---|---|---|---|---|---|
| Mixer applications per round `m` | **{4, 8}; 16 inadmissible on the loaded reference core** | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED 11:13 UK on build-1 (core 4, the frozen crate, today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16): x8 4.672 ms per warp cold alone and 5.413 with the sibling loaded; x16 8.688 and 10.112 ms; the multiplier is 1.86x on the verifier (the mixer IS the verifier's cost: 4.0 of the 8.7 ms), the sibling 16 percent on both; chip-model-v3's estimate (9 ms on a 2019-class core) stands within 4 percent on a current server core** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); x16 over the 10 ms loaded gate by 0.11 ms on the reference core, the ladder's rung-3 class of miss, so the band's top is x8 and x16 is `admissible: false` in the genesis list unless a quiet-core re-run admits it | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED 11:13 UK on build-1 (core 4, the frozen crate, today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16): x8 4.672 ms per warp cold alone and 5.413 with the sibling loaded; x16 8.688 and 10.112 ms; the multiplier is 1.86x on the verifier (the mixer IS the verifier's cost: 4.0 of the 8.7 ms), the sibling 16 percent on both; chip-model-v3's estimate (9 ms on a 2019-class core) stands within 4 percent on a current server core** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); x16 loaded: 10.11 ms on the first pass and 9.095, 9.043, 14.836 on the quiet re-run (12:15 UK, three reps of 60 warps, the box at load 26 of 96 threads), so the reference core reads 9.0 to 9.1 ms under the gate when the host is quiet and 10.1 to 14.8 when it carries other work: a host-load reading, not the mixer's; the margin on the chip model's 2019-class core (the 9 ms estimate) is about zero, so x16 stays `admissible: false` at genesis and the O-1.14 laptop run decides it, exactly as ladder rung 3 | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
| Op-mix weights (the ten non-load families) | each weight within B points of table 1.4.2, B = 2 today; the proposal: B = 4 with the shuffle and mulhi weights capped at their class v4 values | the microbench (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3 at the stock clock; a draw that raises shfl from 8 to 14 of 75 raises the premium per instruction by about 35 percent (modelled on those rows) | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix: pending the two packs | +0 ms | the two re-weighted packs' rows |
| Read width `W` | {1, 4} words | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15; w64 bandwidth-bound (71.9 MH/s on the 5090) | a 4-byte-granularity controller moves 4x the bytes at a w16 era (the Ren-Devadas lever reversed); the `f = 1` chip's energy per read rises toward the card's | 0 / pending / 0 (measured w16 on the M5 Max: within 1 percent) | 0 | none |
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |