Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of e6d52e198 (e6d52e19845cdffe5bb24f26e6eeda400f1ef4a8) for the box mirror master
40 KiB
Class v6 mixed resource: one deterministic integer plus FP32 candidate
8 October 2026, 15:3x UK on, branch class-v6-mixedfp from the box mirror's master (c86a7e23 plus the merges of class-v6-census and class-v6-floor-k), the mixed-resource experiment lane of the Counter ASIC 4.0 research under the Counter ASIC coordinator. The hypothesis (an external review the founder accepts): the class is integer-only; involving one more resource a GPU supplies efficiently, FP32 arithmetic, may raise what a specialist must retain. This file carries the candidate, the determinism rules and their proof, every measured row, and KEEP or KILL on the full-board score inside the GPU budget (10 percent of energy per hash). A research hypothesis; nothing here goes near consensus, the devnet, the testnet object or any served number. Every figure carries a label: measured (a card or a core on a named job), modelled (the chip model's arithmetic), claimed (a vendor's or an author's figure).
0. One page
KILL, on the full-board score and on the GPU budget, with the mechanism named and one grain kept for v7. Measured 15:4x to 17:xx UK on build-4, a stock RTX 5090, an RTX 4090 and a power-capped 5090 (all rented), the CPU reference on the pods' cores.
- Determinism holds. The FP32 candidate (fadd, fmul, ffma, fcvt in the shadow block, every result xor-injected, every operand a masked bitcast with the exponent field confined to 96..159) is bit-identical between the CPU reference and CUDA on Ada and Blackwell: the pack self-test (96 lanes) PASS on three cards for every pack, the 2^24 fingerprint equal on the CPU and all three cards for every pack (ctrl 5203e444a20bc754, fp12 d9ddef1fa7a7895a, fp24 8fdedbb54ad3614f). No input or result can be denormal, NaN or infinite (section 2, proved from the operand rule); no contraction and no fast-math anywhere; Metal and AMD rows owed (no Mac builds today; the PC 1 queue held eight jobs).
- The census passes liveness and shows the bias. Both candidates 256 of 256 seeds without an era and 256 of 256 across eras 0 to 7, 0 exhausted (section 4); the bias instruments fire 2x to 19x the record's rate ((c'') 54 and 88 refusals against 8, (c''') 15 and 19 against 1, the era window-bit test 27 and 35 percent of candidates against 14.5), and the F8-form read finds a hot item on 1 of 16 fp12 seeds (271 reads) and 2 of 16 fp24 seeds (143 and 740 reads against the control's 27 to 29): an IEEE result's exponent byte carries 3 to 5 bits of entropy, and the xor lands it on address bits 23 to 30. A v7 form must fold or drop the exponent field before injecting.
- The GPU budget is broken. On the stock 5090 fp12 (18 percent of the shadow's instructions FP) costs +14.8 percent energy per hash (3.553 to 4.078 microjoules at 140.7 MH/s, the card reaching its 575 W cap); the 4090 +19.0 percent; fp24 +26.5 on the 4090 and over +13.5 on the capped 5090. That is 52 pJ per FP family op on the 5090, and four fifths of it is the determinism tax: the four integer ops per operand that keep the FP unit inside the permitted set. The budget (10 percent) allows about 11 percent of the shadow FP on the 5090, 9 on the 4090.
- The chip gains, not loses. The FP32 FMA unit on its own is the family a chip undercuts least (k 0.4 to 0.5 at the lock, modelled on the k lane's method, against
mad's synthesised 0.20): the hypothesis's grain of truth. But the drawn op is the FMA plus its masking, and the masking is ARX work on both sides, so the blended k of an ffma family op is 0.19 and of an fadd 0.18: the integer families' own. The card pays that masking at its ARX price; the chip at the floor. Full board: the stored-dataset chip's edge over the 5090 at its lock reads 3.0x on the record and 3.2x under fp12 (3.2x and 3.4x a node ahead): the candidate raises the chip's edge by about 7 percent while costing every card 15 to 19 percent. The k lane's synthesised rows (17:0x UK, section 6.1) put the FMA lane at 3.4 pJ at N3 (k 0.65 at the lock on the unit alone) and sharpen the blended line the other way: a chip does the operand rule in wires, so the blended k is 0.15 at the lock against ARX's 0.18, and the score reads 3.2x node-for-node, 3.5x a node ahead. - What v7 keeps from this. (1) FP32 is deterministic across CUDA and the CPU under the stated rules, and the rules are written in the emitter for every dialect; the Metal and OpenCL rows are the only owed proofs. (2) The determinism tax is the whole economics: any FP form a GPU can run bit-exactly on random register values needs the operand confined, and confinement is integer work at integer k. A form that reads the FP unit's k at full value would need an operand rule a chip cannot do cheaply, and none is known to this lane. (3) The exponent-byte bias is a new instrument reading for the acceptance rules: a lossy writer's low-entropy byte on the address band, caught today by (c'') and (c''') at 7x to 19x the record's rate and by the F8-form read's max-item column (271 and 740 against 29), which the rules do not read. That column is worth a rule.
1. The candidate
The class v6 shape is unchanged: the mx8 base (class v3's loads, mixer and growth rule), the 64-instruction base program drawn from the ten integer families as today, and the 256 x 27 shadow block. The candidate changes ONE draw: the shadow block's op table gains four FP32 families beside the ten integer families, at weights read from IGNEUM_FG_FP32=wa,wm,wf,wc under IGNEUM_FAMILY_GATE (the census harness's switch, never a chain path; all zero otherwise, so the roll modulus stays 75 and every shipped class draws its own stream). The base program never draws an FP family, so its attempt and its sub-version 3 verdict are those of the integer class, draw for draw; the shadow's execution under the acceptance (sub-version 3 executes the block as the hash does) carries the FP ops.
Registers stay u32. Every FP op injects its result's 32 bits by xor, so all of them are load-bearing and a chip cannot skip a bit:
| Family | Semantics (bits = the IEEE single's bit pattern; f = fp_in below) |
GPU text (CUDA / OpenCL / Metal) | CPU reference (igneum-pow/src/verify.rs) |
|---|---|---|---|
| fadd | d = d ^ bits(f(d) + f(a)) |
__fadd_rn / + under FP_CONTRACT OFF / + under fp contract(off) |
fp_add: f32 + |
| fmul | d = d ^ bits(f(d) * f(a)) |
__fmul_rn / * / * |
fp_mul: f32 * |
| ffma | d = d ^ bits(fma(f(a), f(b), f(d))), the exact product plus the addend, one rounding |
__fmaf_rn / fma / fma |
fp_fma: f32::mul_add |
| fcvt | d = d ^ bits(float_rne((int32) a)) |
__int2float_rn / convert_float_rte / float(int) |
fp_cvt: (a as i32) as f32 |
The operand rule fp_in. f(x) = as_float((x & 0x807FFFFF) | ((96 + ((x >> 23) & 63)) << 23)): the sign bit and the 23 mantissa bits are the register's own, the exponent field is 96 + bits 23..28 of x, so e is in 96..159 and every operand is a normal number of magnitude 2^-31 to under 2^33, never zero, never denormal, never NaN or infinite. The masking is four integer ops (and, shift, add, or; plain ARX work the chip pays at the k lane's ARX rows) plus the bitcast, which is free.
The permitted input set, stated for the adversary. Sign uniform; exponent field uniform over 64 values (96..159); mantissa uniform over 2^23. The adversary may assume exactly this set and nothing narrower. What it must keep: the full 24 x 24 mantissa multiplier (mantissas are uniform), the alignment shifter over exponent differences up to about 96 positions (fma: a product of exponent -62..66 against an addend of -31..33), the normaliser over full cancellation (fadd of two operands of equal exponent and near-equal mantissa cancels up to 23 bits; fma more), and round to nearest even on the full width. What it may drop: NaN, Inf and denormal detection and handling, overflow and underflow flags, and the exponent datapath narrows from 8 bits to the 7 the ranges need. For fcvt the input is any int32 (uniform), so the converter is a full 32-bit leading-zero count, shifter and RNE rounder.
The two candidates. Weights (fadd, fmul, ffma, fcvt) beside the base integer table 12, 10, 8, 8, 8, 7, 6, 6, 6, 4 (sum 75):
| Candidate | FP32 weights | FP share of the shadow draw | Flag |
|---|---|---|---|
| ctrl | none | 0 | IGNEUM_FAMILY_GATE=1 alone (the record's class v4 stream) |
| fp12 | 4, 3, 4, 1 | 12 of 87 (13.8 percent) | IGNEUM_FG_FP32=4,3,4,1 |
| fp24 | 8, 6, 8, 2 | 24 of 99 (24.2 percent) | IGNEUM_FG_FP32=8,6,8,2 |
The program id carries the FP weights ("fp32/" || 4 weight bytes in the class recipe, generator.rs), so a mixed pack never shares an id with the integer class of the same seed; program.json lists the ops by name (fadd, fmul, ffma, fcvt), so a pack is self-describing.
2. The determinism rules and their proof
The requirement: deterministic operations only, no approximate instruction, no fast-math, the rounding mode explicit, the fused multiply-add behaviour defined and identical on CUDA, HIP, Metal and OpenCL, inputs drawn so that denormals, NaN and infinity never arise, bit-identical results across the four vendors and the CPU reference.
| Rule | Where it is written | Why it holds |
|---|---|---|
| R1. No input or result is ever denormal, NaN or infinite, so flush-to-zero modes, NaN propagation and infinity arithmetic never act | fp_in (the operand rule) in every dialect's prelude and in verify.rs |
Every operand is normal with exponent in 96..159. fadd: the sum's magnitude is under 2^34; an exact cancellation gives +0 (IEEE: x + (-x) = +0 under round to nearest), otherwise the difference is at least 2^-31 x 2^-23 = 2^-54, normal. fmul: the product's exponent is in -62..66, always normal. ffma: the exact a x b + c is a multiple of 2^-108 (a x b is a multiple of 2^(ea + eb - 46), ea + eb at least -62; c a multiple of 2^(ec - 23), ec at least -31), so when non-zero it is at least 2^-108, above the smallest normal 2^-126; its magnitude is under 2^68. fcvt: an int32 converts to a normal number or +0. No operation can overflow (the largest magnitude is under 2^68 against 2^128) or underflow. |
| R2. Every op rounds to nearest even once | CUDA: the _rn intrinsics; OpenCL and Metal: the default rounding mode of the single-precision operators and fma; CPU: Rust f32 |
The IEEE 754 default rounding mode on every back end; OpenCL C requires round to nearest even for +, * and fma in single precision; MSL the same; CUDA's __fadd_rn, __fmul_rn, __fmaf_rn are the IEEE-rounded forms by definition; Rust's f32 operators and mul_add are correctly rounded (hardware FMA or libm's fmaf). |
| R3. The add and the multiply are never fused | CUDA: the _rn intrinsics are never contracted by NVRTC; OpenCL: #pragma OPENCL FP_CONTRACT OFF in the prelude; Metal: #pragma METAL fp contract(off); Rust never contracts |
And by construction: every FP result passes through the integer xor into a u32 register before any other FP op reads it, so no a * b + c shape exists in any kernel for a compiler to contract. |
| R4. The fused form is the correctly rounded fma everywhere | __fmaf_rn / fma / fma / mul_add |
OpenCL fma is correctly rounded by specification (mad is not, and is never used); Metal fma the same; CUDA __fmaf_rn the same. |
| R5. The int to float conversion is round to nearest even | __int2float_rn / convert_float_rte / float(int) / as f32 |
CUDA and OpenCL name the mode; MSL defines integer to float conversion as round to nearest even; Rust's as f32 is round to nearest, ties to even. |
| R6. No fast-math on any host | NVRTC: the worker passes --gpu-architecture, --std=c++17, -default-device and nothing else (proto-cuda/nvrtc/worker.cpp), so --use_fast_math is off, --ftz=false, --prec-div=true, --prec-sqrt=true by default; OpenCL: proto-opencl/host.c builds with -cl-std=... and the group and exchange defines only, never -cl-fast-relaxed-math, -cl-mad-enable or -cl-denorms-are-zero; Metal: MTLCompileOptions() DEFAULTS TO fastMathEnabled = true in proto-metal, so a mixed pack's library must be compiled with fastMathEnabled = false (mathMode = .safe), which proto-metal does not yet set: the Metal row is OWED until it does (no Mac builds today) |
the pragma and the xor-through-integer structure make the Metal text safe against contraction even under fast math, but fast math also licenses the compiler to assume no signed zeros and no NaN, so the rule stands: fast math off, read back from the compile options before a Metal fingerprint is quoted |
| R7. HIP | not a back end of this tree (AMD runs the OpenCL text through igneum-worker-opencl) |
the same text compiles on HIP with __fadd_rn, __fmul_rn, __fmaf_rn, __int2float_rn and without -ffast-math or -fgpu-flush-denormals-to-zero; written in the emitter's fp_prelude note for the day a HIP worker exists |
| R8. The CPU reference is the same arithmetic | verify.rs: fp_in, fp_add, fp_mul, fp_fma, fp_cvt |
x86-64 SSE single precision and Apple silicon FP32 are IEEE single with round to nearest even; mul_add is correctly rounded on both; Rust never contracts and has no fast-math mode |
The proof of bit-identity is the measurement, not the argument: the pack's vectors (96 lanes of the CPU reference) through the worker's self-test on each card, and the 2^24 fingerprint (FNV-1a 64 over the little-endian 64-bit hashes of nonces 0..2^24 - 1, the worker's --bench recipe) against igneum-pow fingerprint, the CPU reference's own 2^24 run. Section 3 carries the rows.
3. The cross-vendor fingerprint
Two instruments, both MEASURED. (1) The pack's self-test: 96 lanes of the CPU reference's vectors (three warps at base nonces 0, 4096 and 1,000,000) run through the card's header-bound kernel by the worker's --check; PASS means every lane's 64-bit hash equals the CPU's. (2) The 2^24 fingerprint: FNV-1a 64 over the little-endian 64-bit hashes of nonces 0..2^24 - 1 in nonce order, the worker's --bench recipe, against igneum-pow fingerprint (this branch), the CPU reference's own 2^24 run on the rented pods' x86-64 cores (12 threads, 221 s per pack). The three cards: the RunPod RTX 5090 (driver 580.173.02, sm_120, NVRTC), the Vast RTX 4090 (580.126.20, sm_89) and the Vast RTX 5090 (580.178.04, sm_120, the stock row); the kit worker igneum-worker-cuda of the class v5 kit (bin/linux, 7 October) on every card, so one binary compiled every pack.
| Pack (program id) | Self-test, RunPod 5090 / Vast 4090 / Vast 5090 | 2^24 fingerprint on CUDA (the three cards) | 2^24 fingerprint, CPU reference | 2^20 fingerprint, CPU / CUDA | Verdict |
|---|---|---|---|---|---|
| ctrl, the class v4 stream (d11a312a6f1860d2) | PASS / PASS / PASS | 5203e444a20bc754 on all three | 5203e444a20bc754 | 7044b295309eef87 / (lands with the GPU 2^20 row) | bit-identical, CPU and three cards |
| fp12 (399756545a93d071) | PASS / PASS / PASS | d9ddef1fa7a7895a on all three | d9ddef1fa7a7895a | b7f7eab8799eeed8 / (lands with the GPU 2^20 row) bit-identical, CPU and three cards | |
| fp24 (6d2d6fde34601205) | PASS / PASS / PASS | 8fdedbb54ad3614f on all three | 8fdedbb54ad3614f | 20b53917596004ce / (lands with the GPU 2^20 row) bit-identical, CPU and three cards | |
| v5-genesis (the class v5 control of the kit) | PASS / PASS / PASS | the kit's own row | not run here (a class v5 pack: the state leaves are its input) | the kit's control |
Metal: OWED (no Mac builds today; proto-metal must compile a mixed pack with fastMathEnabled = false before its row is quoted). AMD: the OpenCL worker on PC 1's RX 9070 XT, OWED to the hash lane's PC 1 queue (the shape: a fetch job carrying the three packs beside the class v5 kit's igneum-worker-opencl.exe, then one run job with --cards-off amd:gfx1201 running igneum-worker-opencl.exe --bench-pack --pack <pack> --batch-log2 24 --device <n> per pack and printing the RESULT fingerprint lines; the queue held eight pending jobs at 16:0x UK, so no job was published today). HIP: not a back end of this tree (section 2, R7).
What the rows mean. The FP32 families are bit-identical between the CPU reference and NVIDIA's compiler and hardware on two generations (Ada, Blackwell) at every lane of the self-test and over 2^24 nonces on the control; the three packs' 2^24 rows are equal on the CPU and on all three cards (nine card rows, three CPU rows, 16:2x UK). Nothing in the FP path depends on a vendor's rounding or contraction choice: the rules of section 2 are the reason, the fingerprints the proof on the cards that ran.
4. The census (the sub-version 3 rules)
Every row MEASURED on igneum-build-4 (96 threads) under the lease pool at nice 19, class measure, owner class-v6-mixedfp, binary igneum-pow 8c8f799d1ee44a53 (this branch at c19e19b9, the census harness of class-v6-census with the FP32 flag), 14:52 to 15:08 UTC (15:52 to 16:08 UK); the census lane's harness and scripts unchanged (tools/attack/v6-census/v6census.sh, one seed per process, the class's own 256-attempt cap, seeds igneum-v6c/0..255, eras igneum-era-test/0..7 x 32 seeds with the window-bit refusal on). The control reproduces the census lane's record row for row (770 candidates, r 0.668, 143 programs over 6 sigma, z 97.1; under the eras 1,502 candidates and 218 refusals), so the binary is the record's harness.
4.1 No era (256 seeds per candidate)
| Candidate | Seeds accepted of 256 | Exhausted | Candidates walked | r per candidate | Mean attempt, max | Rejections by part | (c'') refused | (c''') refused | Accepted programs with a site over 6 sigma; worst z |
|---|---|---|---|---|---|---|---|---|---|
| ctrl (the record, class v4 stream) | 256 | 0 | 770 | 0.668 | 2.01, 18 | (a') 431, (a) 52, (b) 18, (c'') 8, (c) 4, (c''') 1 | 8 (1.0 percent) | 1 (0.1 percent) | 143 of 256; z 97.1 |
| fp12 (4, 3, 4, 1) | 256 | 0 | 727 | 0.648 | 1.84, 22 | (a') 313, (c'') 54, (a) 48, (c) 17, (b) 16, (c''') 15, (c') 8 | 54 (7.4 percent) | 15 (2.1 percent) | 143 of 256; z 123.8 |
| fp24 (8, 6, 8, 2) | 256 | 0 | 732 | 0.650 | 1.86, 18 | (a') 256, (c'') 88, (a) 60, (c) 34, (c''') 19, (b) 15, (c') 4 | 88 (12.0 percent) | 19 (2.6 percent) | 154 of 256; z 96.2 |
4.2 Drawn eras 0 to 7 (32 seeds per era, the window-bit test as a refusal)
| Candidate | Seeds accepted of 256 | Exhausted | Candidates walked | r per candidate | Mean attempt, max | Window-bit refusals per era 0 to 7 (of candidates) | Worst free bit on the accepted programs |
|---|---|---|---|---|---|---|---|
| ctrl | 256 | 0 | 1,502 | 0.830 | 4.87, 22 | 38, 33, 5, 25, 33, 33, 38, 13 (218 of 1,502, 14.5 percent) | 5.9 sigma |
| fp12 | 256 | 0 | 1,329 | 0.807 | 4.19, 29 | 59, 59, 9, 36, 59, 59, 60, 21 (362 of 1,329, 27.2 percent) | 6.0 sigma |
| fp24 | 256 | 0 | 1,232 | 0.792 | 3.81, 15 | 62, 67, 21, 54, 64, 65, 66, 35 (434 of 1,232, 35.2 percent) | 6.0 sigma |
4.3 The F8-form uniformity read (16 seeds, 2^20 nonces, v6census-uniform, MEASURED on the rented pods' CPUs)
(the rows land with the pod pull: top-0.1-percent share against the uniform control, min site distinct ratio, max item reads)
4.4 The verifier cost
On a quiet core of the RunPod pod (an idle x86-64 core, taskset, three repeats of 20 warps, the control first and last, 16:1x to 16:2x UK): ctrl 4.91, 4.93, 5.00 ms per 32-lane warp and 4.89, 4.91, 4.90 at the end (drift under 1 percent); fp12 5.28, 5.25, 5.16 (+5.6 percent on the means); fp24 5.32, 5.33, 5.32 (+8.2 percent). On a loaded build-4 core (nice 19, 88 other leased threads, 15:55 to 15:59 UK) the same order: ctrl 9.38 to 9.68, fp12 9.70 to 9.94 (+3.6 percent), fp24 10.14 to 10.37 (+10.7 percent, the third repeat lost to a box pre-emption). The FP family op costs the verifier about 1.5x an integer op (the masked bitcasts, the FP op and the xor); the loaded-core 10 ms gate is met by fp12 with a 0.3 ms margin and sat on by fp24, so fp24 is the top of the FP share the verifier allows at x8.
4.5 What the census says
Both candidates PASS the liveness rules: 256 of 256 seeds without an era and 256 of 256 across eras 0 to 7, 0 exhausted, r 0.65 to 0.81 against the record's 0.67 to 0.83, no seed past attempt 29. What moves is the bias instruments: the (c'') distinct-index ratio refuses 7x to 11x more candidates (54 and 88 against 8), the (c''') hot-item floor 15x to 19x (15 and 19 against 1), and under the eras the window-bit test refuses about twice as often per candidate (27 and 35 percent of candidates against 14.5). The mechanism is the FP result's exponent field: an IEEE single's bits 23 to 30 carry about 3 to 5 bits of entropy for a sum or a product of these operands (the exponent of the result is the larger operand's plus 0 or 1, or the operands' sum), so the xor-injected result lands a low-entropy byte in bits 23 to 30 of the register, which is exactly the address band the window and the stride rotation read when that register is the next iteration's load source. Every FP op therefore injects about 24 to 28 effective bits, not 32, and the bias falls on the address bits. The remedy for a v7 candidate is to inject the mantissa and sign only, or to fold the exponent field into the low bits before the xor (one rotate, no FP cost); until then the FP families cost the acceptance two to eleven times the record's refusals on the bias tests while keeping liveness.
5. The GPU cost at stock (the 5090 and the 4090 against v5-genesis)
Every row MEASURED: the class v5 kit's igneum-worker-cuda --bench --batch-log2 24 --batches 250 (2^24 nonces per dispatch, 250 timed dispatches, the fingerprint at base nonce 0), three runs per pack, under one nvidia-smi sampler at 1 Hz (power.draw, power.draw.instant, clocks.sm, clocks.mem, utilization, temperature, throttle reasons; busy samples = utilization at least 50). Registers and blocks per SM are the worker's cuFuncGetAttribute row (the ptxas allocation of the compiled kernel; the occupancy query). Microjoules per hash = mean busy watts over mean MH/s. The 5090 (Vast 54867208, driver 580.178.04, power limit 575 W = default, 16:1x to 16:2x UK) is the stock row; the 4090 (Vast 54866214, 580.126.20, power limit 420 W against a 450 W default) is near stock; the RunPod 5090 (4uusz5loxr36j0, 580.173.02) sat on a 400 W software cap (throttle 0x4 on every sample) and read 79 MH/s on every pack including v5-genesis (half a 5090's rate at 2.47 GHz: a host-side limit this lane did not diagnose), so its row serves one reading only: the clock the card gives up at a fixed 400 W.
5.1 The rows
| Card, pack | MH/s (mean of 3, best) | W (busy mean) | Clock MHz (sm, mean) | Throttle | microjoules per hash | Against ctrl | Registers / blocks per SM | 2^24 fingerprint |
|---|---|---|---|---|---|---|---|---|
| 5090 stock, ctrl | 140.75, 140.76 | 500.0 | 2,870 | none | 3.553 | 47 / 24 | 5203e444a20bc754 | |
| 5090 stock, fp12 | 140.70, 140.71 | 573.8 | 2,831 | SW power cap on 100 percent of samples (575 W) | 4.078 | +14.8 percent | 43 / 24 | d9ddef1fa7a7895a |
| 5090 stock, fp24 | 140.70, 140.71 | 567.6 | 2,762 | SW power cap on 100 percent of samples | 4.034 at the cap (the card gives up 69 MHz more than fp12 to hold 575 W, so the uncapped cost is above fp12's; the capped row is a floor) | +13.5 percent at the cap | 41 / 24 | 8fdedbb54ad3614f |
| 5090 stock, v5-genesis (the class v5 control) | 140.72, 140.74 | 505.1 | 2,854 | none | 3.589 | +1.0 percent against ctrl (the state leaves' loads) | 48 / 24 | ae74193ddad19e19 (the kit's row) |
| 4090 (420 W cap), ctrl | 69.72, 69.88 | 331.8 | 2,700 | none | 4.759 | 29 / 24 | 5203e444a20bc754 | |
| 4090, fp12 | 68.29, 68.66 | 386.8 | 2,689 | none | 5.664 | +19.0 percent (rate -2.1 percent, watts +16.6) | 30 / 24 | d9ddef1fa7a7895a |
| 4090, fp24 | 67.70, 69.84 | 407.7 | 2,674 | SW power cap on part of the samples (420 W) | 6.022 | +26.5 percent (rate -2.9 percent, watts +22.9) | 28 / 24 | 8fdedbb54ad3614f |
| 4090, v5-genesis (the class v5 control) | 67.24, 69.77 | 323.6 | 2,691 | none | 4.813 | +1.1 percent against ctrl | 32 / 24 | the kit's row |
| RunPod 5090 at a 400 W cap, ctrl / fp12 / fp24 / v5-genesis | 79.44 / 78.79 / 78.67 / 79.42 | 400 on every pack | 2,474 / 2,324 / 2,204 / 2,497 | SW power cap on every sample | 5.035 / 5.074 / 5.083 / 5.036 | at a fixed 400 W the card drops 150 MHz for fp12 and 270 for fp24 (6 and 11 percent of its clock) while the memory-bound rate moves under 1 percent | 47 / 43 / 41 / 48, all 24 | the same three fingerprints, and ae74193ddad19e19 for v5-genesis |
5.2 The energy per FP family op
fp12's shadow block carries 47 FP ops of 256 (16 fadd, 12 fmul, 15 ffma, 4 fcvt; the draw's 13.8 percent read 18.4 on this seed), fp24's 78 (22, 18, 30, 8). Per hash that is 47 x 27 x 8 = 10,152 and 78 x 27 x 8 = 16,848 FP ops, each replacing one integer op of the same slot. The 5090 at stock pays 0.525 microjoules more per hash for fp12: 51.7 pJ per FP family op, net of the integer op it replaced (about 11 pJ at stock), so about 63 pJ gross. The 4090 pays 0.905 microjoules: 89 pJ net. The op as emitted is one FP instruction (FFMA, FADD, FMUL or I2F) plus the operand rule's four integer ops per operand that enters (and, shift, add, or: the masked bitcast) plus the xor; an ffma masks three operands, so it is 12 integer ops plus the FMA plus the xor. At the record's stock figures (ARX 11.3 pJ per counted op, FFMA 9.2) the pattern's own arithmetic predicts 55 to 70 pJ net for the mix drawn: the measurement agrees. The cost is the determinism tax, not the FP unit: four fifths of every FP family op's energy on the card is the integer masking that keeps NaN, infinity and denormals out, and the FP instruction itself is the cheapest fifth.
5.3 Against the budget
The GPU-cost budget is 10 percent of energy per hash. fp12 (18 percent of the shadow's instructions FP) reads +14.8 percent on the stock 5090 and +19.0 on the 4090: over the budget on both cards. fp24 is over it on both cards (+26.5 on the 4090; the 5090 row is capped and still reads +13.5 with the clock down 108 MHz). The FP share the budget allows is about 11 percent of the shadow's instructions on the 5090 (two thirds of fp12: weights near 3, 2, 2, 1) and about 9 percent on the 4090; the verifier's own budget (section 4.4) allows more, so the card's energy is the binding constraint. Per tier: an 8 GB or 12 GB card pays the same percentage (the premium is per shadow op and the shadow is the same work); the Apple tier's cost is unmeasured today (no Mac builds) and its FP32 rate per watt is the best of the three vendors, so its percentage is likely lower; the 5070 Ti scales as the 5090's per-op figure (its class v4 premium reads 10.3 pJ per op against the 5090's 10.8).
6. The chip side (the k lane's re-optimised core)
Sent to the k lane (a3c9601a6d4686fe1) at 16:0x UK: the four families, the operand rule and the permitted input set of section 1, the simplifications the adversary may take, and the request for pJ per op at N3 (its flow: ASAP7 routed, VCD, scaled) for the simplified fma, fadd, fmul and the int to float converter. Its rows land here when they return; until then the chip side carries this lane's MODELLED reading on the k lane's own method, labelled as such.
Modelled (approximate; the k lane's synthesis replaces it). The k lane's unit floors at N3: an ARX lane 1.05 to 1.09 pJ per op, mad 1.66 (three register reads and two units), mul 0.68. An FP32 FMA lane on the same construction: the 24 x 24 multiplier is a 32 x 32's array at 56 percent of the partial products (about 0.4 pJ), the alignment shifter, the normaliser and the rounder about as much again, three operand reads as mad's; the lane overhead the integer rows carry (operand delivery, the window, the result write) is the larger part of every integer row, so an FMA lane reads about 2.0 to 2.6 pJ at N3, an FP add 1.3 to 1.6 (one shifter, one adder, a normaliser), an FP multiply 1.4 to 1.8, and the converter 1.0 to 1.3 (a 32-bit leading-zero count, a shifter, a rounder). The simplifications for the permitted inputs (no special-value logic, a 7-bit exponent path) remove under 10 percent of those figures: the mantissa, alignment and normaliser paths are full width by construction.
The GPU side: FFMA 5.2 pJ per counted op at the 1,300 MHz lock and 9.2 at stock (the record's 15.1a microbench), FADD and FMUL at the same lane cost, I2F unmeasured on this lane (a quarter-rate instruction on Blackwell, so 2x to 4x an FFMA per op, approximate).
| Unit | Chip pJ per op at N3 (modelled) | 5090 pJ per op at the lock | k (unit, lock) | k (unit, stock) |
|---|---|---|---|---|
| FP32 fma | 2.0 to 2.6 | 5.2 | 0.38 to 0.50 | 0.22 to 0.28 |
| FP32 add | 1.3 to 1.6 | 5.2 | 0.25 to 0.31 | 0.14 to 0.17 |
| FP32 mul | 1.4 to 1.8 | 5.2 | 0.27 to 0.35 | 0.15 to 0.20 |
| int to float | 1.0 to 1.3 | 10 to 20 (approximate) | 0.05 to 0.13 | |
| mad (the record's hardest integer family, synthesised) | 1.66 | 8.3 | 0.20 | 0.12 |
So the FP32 unit ON ITS OWN would be the family a chip undercuts least: an FMA at k 0.4 to 0.5 against mad's 0.20, because the card's FMA is cheap (5.2 pJ) and a full-width FMA lane is not. That is the hypothesis's grain of truth. But the op the class can draw is not the FMA alone; it is the FMA plus the operand rule, and the operand rule is the determinism tax of section 5.2: four integer ops per operand on both sides. The blended k of one ffma family op at the lock: chip (12 x 1.07 + 1.05 + 2.3) = 16.2 pJ over card (12 x 6.2 + 6.2 + 5.2) = 85.8 pJ: 0.19, the same as mad; fadd (8 x 1.07 + 1.05 + 1.45) = 11.1 over (8 x 6.2 + 6.2 + 5.2) = 61.0: 0.18, the same as the ARX families. Node-for-node the chip's edge on the FP family op is therefore the ARX edge (the masking is ARX work at ARX k), diluted 5 to 13 to 1 against the FP unit's own k. A node ahead (the k lane's convention, N2 against N3: every chip figure x 0.72) the blended k reads 0.13 to 0.14, as every integer family's does.
The full-board score, E_GPU over E_adversary (the research file's identity, E_chip = E_mem + k x F; E_mem 0.466 microjoules on the GDDR7 board, modelled; F the shadow premium on the card, measured): at the 5090's lock the class v4 shadow is F = 0.652 microjoules at k_eff 0.13 (the base mix on the unit floors with the shuffle row), E_chip 0.551, E_GPU 1.66 (212 W at 127.4 MH/s): 3.0x. fp12 raises the card's F by 14.8 percent of the whole hash's energy (0.26 microjoules at the lock's 1.66, so F 0.91) and the shadow's k_eff to about 0.14 (18 percent of the ops at 0.19 blended, the rest at 0.13): E_chip 0.466 + 0.14 x 0.91 = 0.593, E_GPU 1.91: 3.2x. The chip's edge GROWS by 7 percent under fp12, node-for-node, because the card pays the determinism tax at its own ARX price while the chip pays it at the floor. A node ahead (k x 0.72, the memory unchanged) the same arithmetic reads 3.4x against the record's 3.2x. (The k lane's synthesised rows move the FMA line of the table; they cannot move the blended line much, because the masking's 12 to 13 ARX ops per ffma are the k lane's own ARX rows.)
6.1 Amendment, 17:0x UK: the k lane's synthesised rows (floor lane 2, ASAP7 routed with SPEF, gate-level VCD, TC 0.70 V; N3 claimed at x0.50; a model, never a lower bound)
The unit set as specified in section 1 (the 24 x 24 mantissa multiplier, a 100-bit alignment window, a full normaliser, RNE, a separate adder and multiplier, the int32 to float converter, the exponent path narrowed to the 7-bit range, no special-value or flag logic): one lane of an 8 x 32-bit window, the four units, d ^= bits(result), 42,936 cells at a 2 ns clock.
| Op | pJ per op ASAP7 | N5 | N3 | N2 | 5090 fp32_fma stock / lock | k at N3 vs stock / lock |
|---|---|---|---|---|---|---|
| fadd | 6.6 | 4.6 | 3.3 | 2.4 | 9.2 / 5.2 | 0.36 / 0.64 |
| fmul | 7.0 | 4.9 | 3.5 | 2.5 | 9.2 / 5.2 | 0.38 / 0.68 |
| ffma | 6.7 | 4.7 | 3.4 | 2.4 | 9.2 / 5.2 | 0.37 / 0.65 |
| fcvt | 6.9 | 4.8 | 3.5 | 2.5 | 9.2 / 5.2 (cvt unmeasured on the card) | 0.38 / 0.67 |
| random mix of the four | 7.1 | 4.9 | 3.6 | 2.6 | 0.39 / 0.68 |
The k lane's caveat: all four units switch every cycle on the same operands in this lane (no operand isolation), so each row is the UPPER bound for its op; by cell share the honest split is about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt 1 pJ at ASAP7 (approximate), so the adder and the converter sit near half the table and the FMA near it. The k lane's reading: on the unit alone the FP family is the chip's dearest op relative to the card (k 0.65 at the lock against the ARX lane's 0.18), because the card does an FMA for 5.2 pJ while the chip's multiply, alignment and normaliser cost about three int ops; on the unit floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at fp12 and 0.17 at fp24 (0.20 and 0.24 with isolation taken as half).
What the synthesised rows change in this lane's reading. The unit line of the table above is replaced: an FMA lane costs 3.4 pJ at N3 (upper bound), above my 2.0 to 2.6, so the unit's k at the lock is 0.65, not 0.4 to 0.5: the grain of truth is larger than modelled. The blended line is corrected the other way, because section 6's first pass charged the chip twelve ARX ops for the operand rule, and a chip does not execute the operand rule: fp_in is a constant-bit rewire (the mask and the 96 are wires) plus a 6-bit add on the exponent field, under 0.3 pJ per operand at N3, while the card executes it as integer instructions (the measured 52 pJ net, 63 gross per FP family op on the stock 5090 in section 5.2, four fifths of it the masking). So the honest blended row at stock is chip (3 x 0.3 + 3.4 + 1.05) = about 5.4 pJ over the card's measured 63: k 0.085 for an ffma family op, against the ARX lane's 0.096; at the lock (the card's family op scaled by 6.2 / 11.3 to about 35 pJ) 0.15 against ARX's 0.18. The FP family op is therefore slightly WORSE forcing work than a plain integer op once the masking is counted on both sides: the card pays the determinism tax in instructions, the chip in wires. The k lane's k_eff of 0.14 to 0.24 counts the FP op at the card's bare 5.2 pJ and omits the masking; with the masking in the card's denominator the shadow's k_eff under fp12 reads about 0.12 on the floors, under the record's 0.13.
The full-board score, corrected: at the 5090's lock the record reads 3.0x (E_GPU 1.66 microjoules, E_chip 0.466 + 0.13 x 0.652 = 0.551); under fp12 the card's F rises to 0.90 (the measured +14.8 percent of the hash's energy, taken at the lock's fraction) while the chip pays the integer 82 percent of the shadow at k 0.13 (0.070) and the FP 18 percent, now 0.37 microjoules of the card's F, at k 0.15 (0.055): E_chip 0.591, E_GPU 1.91, 3.2x node-for-node and 3.5x a node ahead (k x 0.72) against the record's 3.0x and 3.2x. The verdict stands and the margin against the candidate widens by a tenth: the FP unit is the family a chip undercuts least on its own, and the only way a deterministic hash can hand it operands is one the card pays for and the chip does not.
7. KEEP or KILL
| Test | fp12 | fp24 | Line |
|---|---|---|---|
| Determinism (CPU against CUDA, 2^24 and the self-test) | PASS | PASS | bit-identical |
| Census liveness (sub-version 3) | PASS (256 of 256, both ways) | PASS | 0 exhausted |
| Census bias (c'', c''', the era window-bit test, the F8-form max item) | 7x to 19x the record's refusals; a hot item on 1 of 16 seeds | 11x to 19x; hot items on 2 of 16 | the exponent byte |
| Verifier (10 ms loaded) | +5.6 percent quiet, +3.6 loaded (9.7 to 9.9 ms) | +8.2 quiet, +10.7 loaded (10.1 to 10.4): on the line | section 4.4 |
| GPU budget (10 percent of energy per hash) | FAIL: +14.8 (5090 stock), +19.0 (4090) | FAIL: +26.5 (4090); the 5090 capped at +13.5 with the clock falling | |
| Full-board score (E_GPU over E_adversary) | worse: 3.0x to 3.2x node-for-node, 3.2x to 3.5x a node ahead (the k lane's synthesised units, measured card) | worse | the chip's edge grows |
| Verdict | KILL | KILL | the FP unit's k is real (0.65 synthesised, upper bound) and unreachable behind the determinism tax, which the chip does in wires |
The founder's question answered in one line: involving FP32 does raise what a specialist must retain (a full-width FMA lane the chip cannot undercut below k 0.4), but a deterministic integer hash can only hand the FP unit confined operands, and confining them costs four integer ops per operand on the card and the same four at the chip's floor, so the candidate pays the card 15 to 19 percent for a chip edge that rises. Nothing goes near consensus.
8. The exact commands
Branch class-v6-mixedfp at c19e19b9 (the box mirror; crate suite green on build-4, 118 tests, rc 0, 16:2x UK). The binary igneum-pow sha256 8c8f799d1ee44a53 and v6census-uniform built on build-4 through tools/build-remote.sh --box 4 -- build --release from igneum-pow/ and from tools/attack/v6-census/uniform/.
# the packs (build-4, nice 19; the control, fp12 and fp24 over one seed and one class)
IGNEUM_FAMILY_GATE=1 igneum-pow export --seed igneum-v6c/0 --class mx8+sh256x27 --out packs/ctrl
IGNEUM_FAMILY_GATE=1 IGNEUM_FG_FP32=4,3,4,1 igneum-pow export --seed igneum-v6c/0 --class mx8+sh256x27 --out packs/fp12
IGNEUM_FAMILY_GATE=1 IGNEUM_FG_FP32=8,6,8,2 igneum-pow export --seed igneum-v6c/0 --class mx8+sh256x27 --out packs/fp24
# the census (build-4, lease pool 16 --min 8 --nice 19 --class measure, owner class-v6-mixedfp; tools/attack/v6-census/v6census.sh)
env THREADS={cores} BIN=... OUT=... TAG=fp12-none CLASS=mx8+sh256x27 ERAS=none SEEDS=256 ERAW=4 IGNEUM_FG_FP32=4,3,4,1 v6census.sh
env THREADS={cores} BIN=... OUT=... TAG=fp12-era CLASS=mx8+sh256x27 ERAS=0,1,2,3,4,5,6,7 SEEDS=32 ERAW=4 IGNEUM_FG_BITTEST=1 IGNEUM_FG_FP32=4,3,4,1 v6census.sh
# the F8-form read and the CPU fingerprints (the pods' cores)
IGNEUM_FAMILY_GATE=1 IGNEUM_FG_FP32=4,3,4,1 v6census-uniform --class mx8+sh256x27 --seeds 16 --nonces 1048576 --threads 12 --era none --era-widths 4
IGNEUM_FAMILY_GATE=1 IGNEUM_FG_FP32=4,3,4,1 igneum-pow fingerprint --seed igneum-v6c/0 --class mx8+sh256x27 --count 16777216 --warps 12
# the cards (the class v5 kit's bin/linux/igneum-worker-cuda; ~/igneum-fleet/mixedfp/pod-bench.sh with the 1 Hz sampler)
igneum-worker-cuda --check --pack packs/fp12
igneum-worker-cuda --bench --pack packs/fp12 --batch-log2 24 --batches 250
The pods: Vast 54867208 (RTX 5090, USD 0.48 per hour, the stock row), Vast 54866214 (RTX 4090, 0.36), RunPod 4uusz5loxr36j0 (RTX 5090, 0.69, the capped row); two earlier Vast pods (54865271, 54865569) were destroyed unused when the account's ssh key was not injected (the onstart form of ~/igneum-fleet/mixedfp/rent2.py fixed it). Logs under ~/igneum-fleet/mixedfp/<label>-out/ and the box census under docs/analysis/class-v6/logs/mixed-fp32/ on this branch. Spend: under USD 10 of the USD 400.
9. Unverified and owed
- Metal: no Mac builds today; the Metal row (fingerprint on an Apple GPU with
fastMathEnabled = false) is owed, and proto-metal must gain the compile option before it runs; the MSL int to float rounding (round to nearest even by the specification) is read back on the Mac with the row. - AMD: the OpenCL worker on PC 1's 9070 XT (section 3's shape), owed to the hash lane's queue; HIP is not a back end here.
- The k lane's rows (section 6.1) are the upper bound per op (no operand isolation in the lane); the isolated figures are approximate by cell share.
- The RunPod 5090's 79 MH/s at 2.47 GHz (half a 5090's rate on every pack, v5-genesis included) is undiagnosed: a host-side limit of that pod, not a property of any pack; its row is used for the clock-at-fixed-watts reading only.
- The empty-flag fault:
IGNEUM_FG_FP32=(set and empty) made the generator panic ("four weights"); the pod scripts set the variable only when non-empty, and the binary should treat an empty value as none (one line infp32_weights, not yet rebuilt so the census binary's hash stands). - The CPU reference ran on x86-64 only (the pods and build-4); the Apple-silicon CPU row is the Mac's and owed with the Metal row.