Class v6 10.0o: the mixed FP32 candidate killed (determinism proved on CUDA; the determinism tax is the whole economics, +15 to +26 percent on every card for +7 percent of chip edge; the exponent-byte bias a new instrument reading)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
03029526da
commit
761ec73de0
1 changed files with 15 additions and 0 deletions
|
|
@ -527,6 +527,21 @@ The advantage, separated (the chip at 0.06): the board over a Blackwell owner 0.
|
|||
|
||||
**RESPONSE capability:** rotation costs a chip versatility, not life. The 18-family bank costs a chip firmware plus 43 percent of its core cells and 11 percent of its shadow energy, with zero obsolescence credit on any transition in the bank (10.0m); a passed boundary proves the rotation works (10.0d). The window: its form is not the lever (the gated flop file and the macro file agree within 5 percent; the residual over 32 registers 0.3 to 0.4 pJ); its cost to the card is under 1 percent, measured on a 5090 (+0.6 percent) and a 4090 (-0.9 percent) at stock on 8 October (the connected-state lane, replacing "about 0, unmeasured"); and the connected-state class is KILLED (the connected window moves the chip's edge 1.10x and 1.08x against the 1.25x gate on the k lane's re-optimised core; necessity costs a clock-gated file nothing, since it pays per write, not per live register; only the window's width reaches the chip, +0.14 of k).
|
||||
|
||||
#### 10.0o Amendment (16:3x UK): the mixed FP32 candidate, KILLED on the GPU budget and the full-board score (`docs/analysis/class-v6/mixed-fp32.md` on class-v6-mixedfp at 255be026; all measured unless marked; for class v7)
|
||||
|
||||
The candidate: the class v6 shape unchanged, four FP32 families (fadd, fmul, ffma, fcvt) drawn in the shadow block beside the ten integer families behind `IGNEUM_FG_FP32` (harness only), every result xor-injected, every operand a masked bitcast with the exponent field confined to 96..159 so no input or result is ever denormal, NaN or infinite; round to nearest even, no contraction, no fast-math, written in the emitter for CUDA, OpenCL and Metal; two weights, fp12 (18 percent of the shadow FP on the seed) and fp24 (30 percent).
|
||||
|
||||
| Reading | The numbers | Label |
|
||||
|---|---|---|
|
||||
| Determinism | bit-identical CPU reference against CUDA on Ada and Blackwell: the pack self-test PASS on three rented cards for every pack, the 2^24 fingerprint equal on the CPU and all three cards for ctrl (5203e444a20bc754), fp12 (d9ddef1fa7a7895a) and fp24 (8fdedbb54ad3614f); Metal and the AMD OpenCL row owed | measured (PROVED on CUDA) |
|
||||
| Census (sub-version 3, build-4, 256 seeds no era and 256 across eras 0 to 7) | both candidates 256 of 256 both ways, 0 exhausted, r 0.65 to 0.81 against the record's 0.67 to 0.83; the bias instruments fire 7x to 19x the record ((c'') 54 and 88 refusals against 8, (c''') 15 and 19 against 1, the era window-bit test 27 and 35 percent of candidates against 14.5); the F8-form read finds hot items (271 and 740 reads against the control's 29) on 1 of 16 and 2 of 16 seeds: an IEEE result's exponent byte carries 3 to 5 bits of entropy and the xor lands it on address bits 23 to 30 | measured |
|
||||
| The verifier | the quiet core +5.6 percent (fp12), +8.2 (fp24); loaded, fp24 sits on the 10 ms line | measured |
|
||||
| The card at stock (the class v5 kit worker, 250 x 2^24, nvidia-smi 1 Hz) | the 5090 3.553 microjoules per hash on ctrl, 4.078 on fp12 (+14.8 percent, 140.7 MH/s held, the card at its 575 W cap), 4.034 capped on fp24 (+13.5 with the clock down 108 MHz); the 4090 4.759, 5.664 (+19.0) and 6.022 (+26.5); v5-genesis +1 percent on both. 52 pJ per FP family op on the 5090, four fifths of it the determinism tax (the four integer ops per operand); the 10 percent budget allows about 11 percent of the shadow FP on the 5090, 9 on the 4090 | measured |
|
||||
| The chip side (modelled on the k lane's method; its synthesised rows owed) | an FP32 FMA lane on its own is the family a chip undercuts least, k 0.4 to 0.5 at the lock against mad's synthesised 0.20 (the hypothesis's grain of truth), but the drawn op is the FMA plus its masking, which is ARX work on both sides, so the blended k of an ffma family op is 0.19 and an fadd's 0.18: the integer families' own | modelled |
|
||||
| The full board, E_GPU over E_adversary at the 5090's lock | 3.0x on the record, 3.2x under fp12 (3.2x and 3.4x a node ahead): the candidate raises the chip's edge about 7 percent while costing every card 15 to 26 percent | modelled on measured card rows |
|
||||
|
||||
**KILL.** The meaning for v7: FP32 is deterministic across CUDA and the CPU under the stated rules; the determinism tax is the whole economics (any FP form a GPU runs bit-exactly on random registers needs the operand confined, and confinement is integer work at integer k); the exponent-byte bias is a new instrument reading (the F8-form max-item column, 271 and 740 against 29, which no rule reads today) worth a rule in layer 4. The FP32 candidate joins the long program, the select tree, the wide read and the scratchpad in the suite as a negative control (10.0g item 2).
|
||||
|
||||
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
|
||||
|
||||
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |
|
||||
|
|
|
|||
Loading…
Reference in a new issue