diff --git a/docs/analysis/class-v6/floor/shadow-k.md b/docs/analysis/class-v6/floor/shadow-k.md index 34f822a4d..6ea36c5ba 100644 --- a/docs/analysis/class-v6/floor/shadow-k.md +++ b/docs/analysis/class-v6/floor/shadow-k.md @@ -295,6 +295,32 @@ connected-state lane's measurement). The structure's other knobs do not reach th the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set. +### 4e. The mixed-resource lane's FP32 units (class-v6-mixedfp) on the adversary's lane + +The mixed-resource lane's candidate adds four FP32 families (fadd, fmul, ffma, fcvt) to the shadow's draw, every +result injected by xor, with inputs masked to a 7-bit exponent range (never zero, denormal, NaN or Inf). The +adversary's simplified units (`rtl/fp32_units.v`): an FMA with the 24 x 24 mantissa multiplier, a 100-bit +alignment window, a full normaliser and RNE; a separate adder and multiplier; the int32 to float converter; the +exponent path narrowed to the range; no NaN, Inf, denormal or flag logic. The lane: an 8 x 32-bit window, the +four units, `d ^= bits(result)`. Routed with SPEF, random-input VCD, 42,936 cells, a 2 ns clock. + +| Op (every unit evaluating each cycle: an UPPER bound per op, no operand isolation) | pJ per op ASAP7 | N5 | N3 | N2 | 5090 fp32_fma stock / lock | k at N3 vs stock / lock | +|---|---|---|---|---|---|---| +| fadd | 6.6 | 4.6 | 3.3 | 2.4 | 9.2 / 5.2 | 0.36 / 0.64 | +| fmul | 7.0 | 4.9 | 3.5 | 2.5 | 9.2 / 5.2 | 0.38 / 0.68 | +| ffma | 6.7 | 4.7 | 3.4 | 2.4 | 9.2 / 5.2 | 0.37 / 0.65 | +| fcvt | 6.9 | 4.8 | 3.5 | 2.5 | 9.2 / 5.2 (cvt unmeasured on the card) | 0.38 / 0.67 | +| random mix | 7.1 | 4.9 | 3.6 | 2.6 | | 0.39 / 0.68 | + +Reading: the four read alike because all four units switch every cycle on the same operands, so each row is the +upper bound for its op (a chip isolates the idle units; by cell share about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt +1 pJ at ASAP7, approximate). Even on the upper bound the FP family is the chip's dearest per op relative to the +card: k 0.65 at N3 at the lock against 0.18 for the integer ARX lane floor, because the card does an FMA for 5.2 +pJ (under its own int add at 6.2) while the chip's multiply, alignment and normaliser cost about three int ops. +On the units' floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at the fp12 mix and 0.17 at +fp24 (0.20 and 0.24 with isolation taken as half), the core's per-op overhead on top. The GPU-cost budget (10 +percent of energy per hash) is the binding side, and the vendor-rounding question is the class's, not the chip's. + ## 5. The chip edge at the measured k `E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash