shadow-k: the mixed-resource lane's FP32 units routed (6.6 to 7.1 pJ per op ASAP7 as an upper bound, k 0.63 to 0.68 at the lock at N3)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
08447ca97e
commit
089546fb12
1 changed files with 26 additions and 0 deletions
|
|
@ -295,6 +295,32 @@ connected-state lane's measurement). The structure's other knobs do not reach th
|
|||
the written register), and the chain's width sets lane count, which is free. On the GDDR7 board at the lock the
|
||||
window moves the chip's edge by about 1.1x (3.6x to 3.3x node-for-node), under the 1.25x gate that lane set.
|
||||
|
||||
### 4e. The mixed-resource lane's FP32 units (class-v6-mixedfp) on the adversary's lane
|
||||
|
||||
The mixed-resource lane's candidate adds four FP32 families (fadd, fmul, ffma, fcvt) to the shadow's draw, every
|
||||
result injected by xor, with inputs masked to a 7-bit exponent range (never zero, denormal, NaN or Inf). The
|
||||
adversary's simplified units (`rtl/fp32_units.v`): an FMA with the 24 x 24 mantissa multiplier, a 100-bit
|
||||
alignment window, a full normaliser and RNE; a separate adder and multiplier; the int32 to float converter; the
|
||||
exponent path narrowed to the range; no NaN, Inf, denormal or flag logic. The lane: an 8 x 32-bit window, the
|
||||
four units, `d ^= bits(result)`. Routed with SPEF, random-input VCD, 42,936 cells, a 2 ns clock.
|
||||
|
||||
| Op (every unit evaluating each cycle: an UPPER bound per op, no operand isolation) | pJ per op ASAP7 | N5 | N3 | N2 | 5090 fp32_fma stock / lock | k at N3 vs stock / lock |
|
||||
|---|---|---|---|---|---|---|
|
||||
| fadd | 6.6 | 4.6 | 3.3 | 2.4 | 9.2 / 5.2 | 0.36 / 0.64 |
|
||||
| fmul | 7.0 | 4.9 | 3.5 | 2.5 | 9.2 / 5.2 | 0.38 / 0.68 |
|
||||
| ffma | 6.7 | 4.7 | 3.4 | 2.4 | 9.2 / 5.2 | 0.37 / 0.65 |
|
||||
| fcvt | 6.9 | 4.8 | 3.5 | 2.5 | 9.2 / 5.2 (cvt unmeasured on the card) | 0.38 / 0.67 |
|
||||
| random mix | 7.1 | 4.9 | 3.6 | 2.6 | | 0.39 / 0.68 |
|
||||
|
||||
Reading: the four read alike because all four units switch every cycle on the same operands, so each row is the
|
||||
upper bound for its op (a chip isolates the idle units; by cell share about ffma 3.5 to 4, fmul 2.5, fadd 2, fcvt
|
||||
1 pJ at ASAP7, approximate). Even on the upper bound the FP family is the chip's dearest per op relative to the
|
||||
card: k 0.65 at N3 at the lock against 0.18 for the integer ARX lane floor, because the card does an FMA for 5.2
|
||||
pJ (under its own int add at 6.2) while the chip's multiply, alignment and normaliser cost about three int ops.
|
||||
On the units' floors the shadow's k_eff rises from 0.097 (class v4) to about 0.14 at the fp12 mix and 0.17 at
|
||||
fp24 (0.20 and 0.24 with isolation taken as half), the core's per-op overhead on top. The GPU-cost budget (10
|
||||
percent of energy per hash) is the binding side, and the vendor-rounding question is the class's, not the chip's.
|
||||
|
||||
## 5. The chip edge at the measured k
|
||||
|
||||
`E_chip = E_mem + N_ops x e_chip` (absolute: the chip's shadow cost is 102,100 x 3.5 pJ = 0.36 microjoules per hash
|
||||
|
|
|
|||
Loading…
Reference in a new issue