Class v6 layer 4: lane D's 17:00 cut as 5.1b (24,000 drawn eras: the op-mix band's known-failed test reads 0 of 3,000 under the cap; the lossy curve's edge at +2 to +3; the (c''') per-width calibration; the bit-R bias law; the measured bound n)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 12:14:25 +00:00
parent 3f6ee5878d
commit dd692978a8

View file

@ -55,7 +55,7 @@ The band a card sees across the draw's range, from the hash lane's cost rows (12
| Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number |
|---|---|---|---|---|---|---|
| Mixer applications per round `m` | **{4, 8} at genesis; 16 in the list as `admissible: false` (the ladder's rule: a flag in the genesis list, flipped only by the 90 percent upgrade path once the 2019-class core measures it)** | **The card pays nothing for the draw, MEASURED (the hash lane's kit b on PC 1, 12:49 to 12:58 UK, generator 2 with the era, the multiplier the only difference, 250 batches per row, fingerprints PASS): x8 137.73 MH/s at 320.2 W unlocked and 127.44 at 212.0 W at the 1,300 lock; x16 137.72 at 320.1 and 127.46 at 211.7; within 0.1 MH/s and 0.5 W at either state, item derivation hidden under the read chain, so the verifier's 1.86x per doubling is the draw's whole cost.** x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); **x16 MEASURED (today's class v3 stream, the same id f5e904bc5d148926 for x8 and x16). Cold alone: build-1 core 4 x8 4.67 ms per warp, x16 8.69 (the hash lane, 10:13 UTC); build-3 core 4 (AX102, a faster core class) x8 3.48, x16 6.53 (the build-server lane, 11:01 UTC). With the SMT sibling loaded for the row's whole length (the method that counts: a sibling bench looped for the row, not run once, which ends in about a second and leaves the verify phase sibling-idle; the hash lane's earlier 5.41 / 10.11 / 9.04 to 9.10 ms rows were of that kind and are withdrawn): build-3 x8 6.24, x16 11.44 ms, about 1.8x on both classes. The multiplier is 1.86x on the verifier (the mixer IS the verifier's cost); chip-model-v3's estimate (9 ms on a 2019-class core) stands within the cold rows** | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured); the x16 build on the 5090 rides the 12:45 UK job | +1.9x per doubling measured (x8 to x16); **x16 at 11.4 ms loaded is over the 10 ms gate on the measured row, so `admissible: false` stands on the measurement, not only on the chip model's margin; the O-1.14 laptop run can only tighten it** | adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading; the band {4, 8} is the measured one |
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block (mx8+sh256x27, kit d, about 13:45 UK), and the microbench arithmetic stays the default until it lands** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner, which the capped band must read as 0 |
| Op-mix weights (the ten non-load families) | **B = 4 points on the injecting families only (add, sub, xor, mad, shfl, rotl, rotr); the lossy families (or, mul, mulhi) capped at their base, the ring-A rule `or + mul + mulhi` at most the table's 18 plus B, which keeps the per-candidate rejection r under 0.85 (r^256 under 1e-18)**; the shuffle weight capped at its class v4 value on the energy side | lane D's coverage run (4,900 drawn eras on three boxes, 11:3x UK, measured): at the lossy corner of B = 4 (or, mul and mulhi all at +4, 30 of 75 lossy against 18) r is 0.956 (the shipped 0.681) and 8 of 663 eras exhaust the 256-attempt cap (mean attempt 18.6, max 252), so 1.2 percent of that corner's epochs would take the last-resort program, which fails rule (a) in 9 percent of seeds (adv-accept-3); the random stratum with every weight drawn reads r 0.718, max attempt 107, 0 exhaustions in 1,444. The energy side (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3; a shuffle-heavy table raises the premium per instruction up to 2x (modelled), a multiply-heavy one 1.1x | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix within the capped band (the cost rows: at N = 100,000 the 5090's +146 W stock becomes up to +160 W multiply-heavy; shuffle-heavy excluded by the cap); the 5070 Ti scaled by 83/146, the M5 Max by 16/146; the two packs' measured rows replace this when they land. **First row MEASURED (kit b, PC 1, 12:5x UK): the shuffle-heavy table (shfl 13 of 48 against the stock 5) on class mx8 without a shadow block reads 137.65 MH/s at 312.9 W unlocked (7 W under the stock table's 320.2) and 127.33 at 212.4 at the lock (level): the multiply-heavy table (mulhi 11, mad 8, mul 5 of 48 against the stock 9, 4, 1; shfl 1 against 5; kit c, 13:03 to 13:06 UK) reads 137.77 at 313.5 W unlocked and 128.42 at 211.8 W at the lock, so the three tables sit within 7 W unlocked and 0.6 W at the lock: on the 48-op base program the weight move is a 2 percent term whichever way it goes; the row layer 1 needs is the same two tables inside the shadow block (mx8+sh256x27, kit d, about 13:45 UK), and the microbench arithmetic stays the default until it lands** | +0 ms | the known-failed test: the 8 of 663 exhaustions at the uncapped lossy corner (61 of 5,000 on the 17:00 cut), which the capped band must read as 0: **MEASURED 0 of 3,000 under the band (lane D's 17:00 cut, 5.1b; r 0.595, mean attempt 1.47, max 59)** |
| Read width `W` | **pinned at 4 words (16 bytes) at genesis, not drawn** (floor lane 3, 10.3: the width is the only wire lever on the SRAM die, 66x at 4 bytes against 44x at 16 at zero shadow; 8 words after the owed PC 1 and Mac rows; 16 never) | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15, the M5 Max within 1 percent; w64 bandwidth-bound (71.9 MH/s on the 5090) | the SRAM die's energy per read rises with the bits moved (0.25 to 0.38 nJ); the DRAM chip's toward the card's | 0 / 0 / 0 (measured) | 0 | the W = 8 rows (owed) |
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
| **The index fold (a ring-A design rule of layer 1, every era's draw passes through it): `load_index` folds a product's low bits before the stride rotation, so no era's R lands a biased product bit on an address bit** | every era; lane D's coverage (11:3x UK): the class is at HALF the family's epochs, not a corner: 52 to 58 percent of accepted programs in every stratum carry one site whose address bit R (or R+1, R+2) is biased over 6 sigma at 2^20, 33 to 40 percent over 100 sigma, the worst z 1,024, on programs (c''') passes; the F8 tail's bucket excess is the same mechanism at scale (+94 to +128 sigma at the biased bit) | lane D's family harness (build-1, 10:21 UTC, 16 drawn eras, every layer-1 parameter from the era's stream, the chain draw through the real rule, reads at the rule's own 2^20 sample): the index-bit bias fires HARD on 7 of 16 eras, abs z 130 to 511 at one site, every one at address bit R or R+1 (a product's bit 0 at P = 0.25 on era 15, R = 25, z -511; a product's bit 1 at 3/8 on era 13; an or-shaped source at 5/8 on era 6); the other 9 clean under abs z 3.8; the (c''') ratio sees none of it (0.9954 to 1.0000): adv-cache-2's era-stride class measured on class v5 accepted programs at the acceptance's own sample | a chip holding the favoured half of that site's window serves 75 percent of its reads instead of 50: about 1.6 percent of a hash's reads at f = 1/2 for one site, zero at f = 1 (the partial store already costs 1.26x the ops, chip-model 5.4): the f = 1 verdict does not move; the row is an auditor's flag on "uniform random reads", not a chip lever | nothing: the fold is one xor-rotate on the address path, measured as 0 on every card by the era-layout rows (the index form is the era draw's own) | 0 | the fold's form in `load_index` (fold the product's low bits before the rotation) and its vectors; a 6-sigma REFUSAL in layer 4 is not the lever: it would redraw about 40 percent of epochs (7 of 16 eras); the known-failed test is lane D's 7 of 16 eras at the rule's sample, which must read 0 of 16 with the fold |
@ -164,6 +164,21 @@ Every test the generator applies to a candidate program is today a function of t
The sampling bound: n drawn eras all passing bound the failing fraction at 3/n at 95 percent (190 eras for 1 in 64, 3,067 for 2^-10); the union bound over the tests' miss rates is bounded by their known-failed CASE counts (9 hot sets give the (c''') floor a miss rate under 0.33), so the proof of testing carries per test the count of fired cases, and tightening is by cases from the adversarial tails, not by eras. The record: one JSON per drawn era (the family id, the draw, the pinned binary sha, the box, the seeds, every verdict with its statistic, threshold, log sha256 and core-seconds) plus a family summary (eras per stratum, the worst era per test, the refuse rate per band, the bound's arithmetic, the corner cells not reached). The mixer m is a per-FAMILY axis (no per-era statistic sees it; the index tests run on the closed form), closed by adv-mixer-3's ladder at m = 4 (2 of 4 applications of margin against 6 of 8) and the x16 verifier row; the per-load placement has one value in the band; W = 16 is excluded. Main's word (11:4x UK): lane D's band is layer 1's op-mix band and the index fold with the bias test as its guard is layer 1's rule, not a next-class note.
### 5.1b The 17:00 cut, measured (lane D, `docs/analysis/class-v6/family-gate.md` section 6 on the mirror's master at 238100b0, 13:13 UK; 24,000 drawn eras of the family on build-1 and build-3, 54 core-hours, one era per seed with every layer-1 parameter from the era's own stream, the acceptance keyed on the family's shapes; census rows, logs, scripts and the harness diff under `docs/analysis/class-v6/logs/`)
What the cut settles, each row measured unless marked:
| Finding | The numbers | What it moves in this document |
|---|---|---|
| The op-mix band is settled by measurement | B = 4 with or, mul and mulhi free to rise exhausts the 256-attempt cap in 1.2 percent of eras (61 of 5,000 at the lossy corner); with or, mul and mulhi never raised above their base (section 2's band) r = 0.595, mean attempt 1.47, max 59, 0 exhausted in 3,000. The lossy-share curve (running, about 1,200 eras per point): r = 0.80, 0.88, 0.92, 0.96 at +1 to +4 points; exhaustion 0, 0, 0.2 percent, 1.4 percent; the edge between +2 and +3 points | the layer 1 op-mix row's known-failed test now reads 0 of 3,000 under the band (it owed a 0); the cap on the lossy families stays at the base, with +2 points the most the band could ever open to |
| What breaks at the lossy corner | class v5's last-resort scan passes at its first or second candidate on every exhausted era seen, so the corner costs liveness time, not an unchecked program; the per-era exhaustion is about 1,000x the independent-attempt estimate because one era's attempts share its weight table, which is the independence the scan's 1e-300 assumes | section 5.2's bound: the attempts within an era are not independent draws; the bound is per era from the census, not r^256 |
| The (c''') floor needs a per-width calibration | at width 4 (expectation divided by the width) it refuses 4.5 to 6.5 percent of candidates against 0.6 to 1.0 percent at width 1 | ring B's row: either a floor per width from the clean spread, or the statistic in sigma over its own expectation (the form the bucket statistic now takes); the full report (09:00 UK tomorrow) carries the per-width floor; W pinned at 4 at genesis (section 10.3) makes it one calibration, not a family of them |
| The shape axis moves the draw's cost, the mixer axis is invisible | 64 x 108: 1.8 attempts; 256 x 27: 3.8, through (a')'s fixpoint over the block; r 0.714 to 0.731 across m | m is closed per family by adv-mixer-3's ladder at m = 4 and the x16 verifier row, not by drawing eras (section 5.1a's reading stands, now measured); the block-shape row of layer 1 carries the attempt cost |
| The era-stride bias at the bit level | 48 to 58 percent of accepted programs in every stratum carry one site biased at over 6 sigma at 2^20; 73 percent of those at address bit R exactly (12 percent at R+1, 5 at R+2: the product law's bits 0, 1, 2 through `rotl(x*M, R)`); a third over 100 sigma, worst z 1,024; under R in 28..31 the over-100-sigma share falls from 37.7 to 7.5 percent; the bucket statistic sees the same mechanism (p99 +6.8 sigma on bit-clean eras, +44 on biased ones) | one value-level test covers both; the remedy is structural, the index fold of the product's low bits in `load_index` before the rotation (layer 1's rule, section 2), not a per-epoch refusal that would redraw half the epochs; the live-dataset price per site (ring C) is being read now (the F8 census at 2^24 on 64 seeds at two band points) |
| The bound arithmetic on the counts | 10,000 random eras and 3,000 band eras passing bound the failing fraction on the ring-B tests at 3.0e-4 and 1.0e-3 at 95 percent; the floors' miss rates on the live-dataset classes rest on 9 and 5 known-failed cases (under 0.33 and 0.60), tightened by cases from the adversarial tails, not by eras | section 5.1a's sampling bound now has its measured n; the gate-record JSON per era lands with the full report |
Owed from lane D by 09:00 UK tomorrow: the per-width floor, the bucket bound in sigma, the finished lossy curve, the two ring-C live rows, the gate-record JSON per era.
### 5.2 The cost of the redraw, and its bound
Lane D's measured coverage (11:3x UK; the full table in family-gate.md at 17:00): the (c''') refuse rate per stratum 2.56 percent of candidates in the random stratum, 3.41 at shape 64, 0.42 at the lossy corner; the rule-of-three coverage line a failing fraction under 3e-4 at 95 percent at 10,000 eras. The hash lane's draw costs: the base rules reject about two thirds of raw candidates per attempt (milliseconds), (c'') about 4 percent at 2.2 s per pass on a box core (2.3 s per epoch draw per node, 4 to 5 s on slower cores), (c''') 2.435 percent in the same pass, the bucket bound 1 to 3 percent (estimate) and the value-level test about the same cost again inside the same histogram; expected attempts per accepted program stay about 2.1, and a parameter set under which 32 attempts fail is redrawn.