Counter ASIC 4.0 research: --bench honours --variant through the pinned race (the SM-sparse run of 00:22 to 00:41 UTC served base on all 48 rows); RESULT carries variant, sparse_blocks and block_warps, variant_not_installed is the known-failed line; 20.3a records the run and its knee rows (v4 309.9 W, v3 219.4 W, premium 90.5 W at 1,300 MHz)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 0b0151b55 (counter-asic-4) for the box mirror master; left on the branch: proto-cuda/nvrtc/worker.cpp
This commit is contained in:
igneum-labs 2026-10-08 00:44:35 +00:00
parent a3b0f01364
commit 19d64f2c86

View file

@ -377,6 +377,12 @@ The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, meas
PENDING. Per pack, unlocked and at the 1,300 knee when the helper answers: MH/s, watts, SM MHz, the fingerprint and the self-test verdict. The questions each row answers: `shl256x27` against `sh256x27`: equal watts and a rate inside the block-size effect (16-instruction blocks ran 2.5 to 3.5 percent faster than 256 on 6 October) means the per-load placement is free for the GPU and the capex row of 16.2 stands; `mm128`, `mm512`, `mm1430` against `mx8-genesis`: the premium per tile on the 5090 (the 4090 read 0.056 pJ per MAC at 4,096 tiles per hash), where the rate falls, and whether 11,440 tiles per hash carries the ALU shadow's premium inside the free band; a self-test FAIL on a tile pack is a layout finding, not a tuning.
### 20.3a The SM-sparse job's first run (PC 1, 00:22 to 00:41 UTC, 8 October 2026): no variant ran; the knee rows it did read
The hash lane's run `run-ca4-pc1-ca4sparse-5090-20261007` (the ca4sparse exe, the 5090 alone, class v3 and v4 packs, unlocked and at the 1,300 MHz lock, 48 rows, exit 0, every fingerprint equal to the Mac's, self-test PASS) served the BASE kernel on every row: the worker's `--bench` path turned the race off and built the pair without one, so `--variant sp43-w32` was parsed and never applied ("race 0 ms variant base" on all 48 rows, no NVRTC error because NVRTC never saw the variant). The numbers confirm it: every sparse row's rate equals base (v4 137.0 to 137.1 MH/s unlocked, 133.8 to 134.3 at the lock; v3 136.7 to 136.8 and 133.8 to 134.0) and the draw drifts with heat, not with the SM count (v4 unlocked 465.5 W at 69 C on the first row rising to 485.2 W at 76 C on the sixth; v3 331.6 falling to 319.2 W as the card cooled from 71 to 62 C). The SM-sparse question is unanswered by this run. The fix (worker.cpp, this branch, 8 October 2026, 00:xx UTC): a `--bench` with `--variant <name>` runs the pinned race (base and the named variant only, no timing, the variant installed whatever its speed) and the RESULT line carries `variant=`, `sparse_blocks=` and `block_warps=`; a run whose served variant is not the requested one prints a `variant_not_installed` line, which is the known-failed case of this fix. The rerun is the hash lane's queue slot.
What the run did read, and keeps (the 5090 at the 1,300 MHz lock, base kernel, measured): class v4 134.26 MH/s at 309.9 W (0.433 MH/W, SM 1,290 MHz), class v3 134.03 at 219.4 W (0.611 MH/W); the class v4 premium 90.5 W at the lock against 133.9 W unlocked on this run (145.3 on the efficiency pass; the unlocked v4 rows ran hot, 465 to 485 W). The premium per counted op at the lock: 90.5 W over 134.26 M x 102,100 = 6.6 pJ (6.5 on the efficiency pass).
### 20.4 Into the chip model (rows filled when 20.2 and 20.3 land)
| Row | Honest 5090 energy per hash | Premium over class v3 | Chip at k = 1 / 1.5 / 2 (GDDR7) | Break-even cap, years 1 to 2 | Status |