Counter ASIC 2.0 status 00:47: the M16 inline bench table, the repro run on PC 2

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 00:45:30 +00:00
parent a87db8d3e1
commit 724bd5aa00

View file

@ -745,3 +745,16 @@ floor-sweep-4 (00:34:19 to 00:37:49Z, 210 s, exit 0): the v3 server b37defef bes
| budget 12 GB (2^27) | fees-v1-shards2 | 4,717,439 | 14,786 | 17.4 | 12,915 MiB, 4.3 s |
Reading (the tier call is the prover-floor agent's write-up): the prover's own share beside the miner is 12,066 minus 3,833 = about 8.2 GB on the v1 shard at 2^26, the same as alone, so a 12 GB card fits mine-and-prove on paper (8.2 GB plus a miner's 1 GiB dataset and working set, about 10.0 GB of 12.3 before the display); 2^25 buys no memory (12,066 either way) and costs 1.8x the time. The cost is time, not memory: 24.4 s per shard beside the miner against 5.7 s alone (4.3x; the miner holds the card), 13.0 s on the empty shard against 3.3 s; that is inside proving v1's 600 DAA-second deadline by 25x, so a 12 GB card that mines and proves at once meets the deadline and simply takes fewer shards per hour (the per-prover throughput share, not a tier exclusion). The prover-floor agent's tier reading, on the 5090's allocation: a 12 GB card (12,288 MiB) mining and proving holds the server's 8.2 GB plus the miner's 1.7 GB plus its display (0.5 to 1 GB, not measured) = 10.4 to 10.9 GB, so it fits with 1.4 to 1.9 GB spare at 2^26 and 24 s a shard; the floor is now the Setup keys and the recursion stage (4.4 GB own after Setup plus about 3.8 GB at the recursion peak), so 2^25 is not a lever; a 16 GB card mines and proves at 2^27 (10.95 GB own plus 1.7) with 2.9 GB spare at 17 s. My earlier sentence here ("proves slower than the chain makes segments") was wrong against the deadline rule and is withdrawn. Not measured: any real 12 GB card (the on-order card runs the same two points). PC 2 released 00:37:49Z; "go PC 2" to m16-inline-pc2-1 at 00:39 (about 8 min), then the repro pair (fetch-repro-pc2-20261006, run-repro-pc2-20261006, about 14 min), then agg-cost-pc2-5 (about 12 min), then the ledger-fixes-0311 suites.
## 00:47 M16 inline bench on the 5090 (ledger-pc2 564acab); the repro run is on PC 2
m16-inline-pc2-1 (00:38:54 to 00:43:17Z, 263 s, exit 0; app 0.3.11; the miners resumed and the prover back on at 00:43:06Z). The recompute attacker on the same silicon, version-2 program, 15 s windows, power the median of 7 nvidia-smi samples:
| Setting (1 warp per block) | MH/s | of honest | W |
|---|---|---|---|
| honest, 1 GiB dataset | 132.20 | 1.00 | 326.6 (3,060 MHz) |
| inline, 256 MiB cache in VRAM | 11.26 | 0.085 | 415.7 |
| inline, 64 MiB (inside the L2: the SRAM-class emulation) | 33.88 | 0.256 | 431.0 (the power limit, 2,835 MHz) |
| inline, 32 MiB | 33.87 | 0.256 | 431.0 |
8 warps per block: 131.15 / 10.89 / 29.32 / 29.46. All three bit-exact checks passed on the card. Reading: a recompute attacker with the cache in SRAM-class memory reaches a quarter of the honest rate on the same silicon and is 5.1x worse per joule; the 1,024 dependent cache-line reads per hash bound it, not the integer budget (about 6 T op/s reached of the card's 50 T, approximate). This is the GPU-side measurement behind the chip model's recompute row (the on-die-cache chip in docs/analysis/chip-model-v3.md is the ASIC-side projection; both say the mixer's dependent reads, not the arithmetic, set the attacker's rate). Bench-log and ledger M16 and E17 updated on ledger-pc2 564acab. PC 2: the repro kit fetched 00:44:29Z (40,887,373 bytes, sha256 ok), run-repro-pc2-20261006 running (about 14 min), then agg-cost-pc2-5, then the ledger-fixes-0311 suites.