Counter ASIC 2.0: shard labels corrected (adopted = the v1 shard, prototype = block-338-shard1)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-06 00:18:50 +00:00
parent 7fdea00b20
commit c547a1ff4a
2 changed files with 4 additions and 4 deletions

View file

@ -71,7 +71,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash
HiveOS rigs (consequences C41, 23:33 UTC): untested on a GPU host tonight, and local mode never worked (the bundled node starts without the override file, so a local-mode rig has been refused by every devnet peer since the first height switch); the fix is the next cut's one HiveOS item (section 8a). The installer path with the hosted node is unaffected.
Prover memory tiers (the proving-methods agent, 23:42 UTC, from the prover-floor agent's patch v1 rows in docs/analysis/proving-methods.md on branch proving-methods, not pushed): the adopted shard falls from 20,516 to 12,708 MiB (idle included) with the 24 GB panic removed and the element threshold budgeted; the next floor is the server's Setup at 9,703 MiB (five recursion keys pre-built at 2^27 capacity). Per tier: the 16 GB tier opens now (prove-only at 2^27, 15.4 GB; mine-and-prove at 2^26 on paper, 14.4 GB, to be measured beside the miner); the 12 GB tier was not open on patch v1 (12,708 MiB against about 12,288 MiB reported by a 12 GB card). MEASURED on patch v3 (floor-sweep-3, 00:13 to 00:17Z 6 October, the v3 server b37defef on the RTX 5090, miners stopped, every proof verified by the unpatched host): with the element threshold at 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`) the v1 shard (4.7 M cycles) peaks at 10,291 MiB of device memory with the card's 2,089 MiB idle inside (about 8.2 GB the server's own) at 5.7 s against 4.3 s at 2^27; the empty shard 9,971 MiB; the server's floor after Setup 6,535 MiB (9,703 on v1, 11,623 stock). Under a plain 12 GB budget at 2^27 the v1 shard peaks at 12,915 MiB and the adopted prototype shard (60.4 M cycles) at 13,459 MiB in 16.8 s (28,295 MiB on the stock server). So the 12 GB tier profile is the 2^26 threshold on the v3 server, a third slower per shard; RISC Zero at po2 19 is no longer the route unless a real 12 GB card contradicts the 5090's reading. The 16 GB tier at 2^27: 16,115 MiB on the v1 shard, 4.0 s; upstream's own threshold (budget 32): 16,851 MiB. Not measured: any 12 GB card (a 4070 or 3060 is the on-order check), and mine-and-prove beside the miner (sweep 4, 4 points, queued on PC 2 after the aggregation-cost re-run). Recommended, for the project lead: provedefault.rs gains the 16 GB tier; the rig installer runs one server per card; a 4070 or 3060 enters the measurement loop, since no 12 GB card has run a prover here (every 12 GB number is a 32 GB card's reading with a budget).
Prover memory tiers (the proving-methods agent, 23:42 UTC, from the prover-floor agent's patch v1 rows in docs/analysis/proving-methods.md on branch proving-methods, not pushed): the adopted shard falls from 20,516 to 12,708 MiB (idle included) with the 24 GB panic removed and the element threshold budgeted; the next floor is the server's Setup at 9,703 MiB (five recursion keys pre-built at 2^27 capacity). Per tier: the 16 GB tier opens now (prove-only at 2^27, 15.4 GB; mine-and-prove at 2^26 on paper, 14.4 GB, to be measured beside the miner); the 12 GB tier was not open on patch v1 (12,708 MiB against about 12,288 MiB reported by a 12 GB card). MEASURED on patch v3 (floor-sweep-3, 00:13 to 00:17Z 6 October, the v3 server b37defef on the RTX 5090, miners stopped, every proof verified by the unpatched host): with the element threshold at 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`) the v1 shard (4.7 M cycles) peaks at 10,291 MiB of device memory with the card's 2,089 MiB idle inside (about 8.2 GB the server's own) at 5.7 s against 4.3 s at 2^27; the empty shard 9,971 MiB; the server's floor after Setup 6,535 MiB (9,703 on v1, 11,623 stock). Under a plain 12 GB budget at 2^27 the v1 shard (the adopted proving v1 shard) peaks at 12,915 MiB and the prototype shard (60.4 M cycles, not the adopted one) at 13,459 MiB in 16.8 s (28,295 MiB on the stock server). So the 12 GB tier profile is the 2^26 threshold on the v3 server, a third slower per shard; RISC Zero at po2 19 is no longer the 12 GB route (it stays the Apple route in proving-methods.md) unless a real 12 GB card contradicts the 5090's reading; the proving-methods agent's tier reading: 12 GB prove-only opens with about 2 GB spare and about 0.3 GB spare beside the miner (approximate), so mine-and-prove on 12 GB still points at a half shard; 16 GB holds the prototype shard alone. The 16 GB tier at 2^27: 16,115 MiB on the v1 shard, 4.0 s; upstream's own threshold (budget 32): 16,851 MiB. Not measured: any 12 GB card (a 4070 or 3060 is the on-order check), and mine-and-prove beside the miner (sweep 4, 4 points, queued on PC 2 after the aggregation-cost re-run). Recommended, for the project lead: provedefault.rs gains the 16 GB tier; the rig installer runs one server per card; a 4070 or 3060 enters the measurement loop, since no 12 GB card has run a prover here (every 12 GB number is a 32 GB card's reading with a budget).
## 7. Gates before any publish (all of them, no exceptions)

View file

@ -698,12 +698,12 @@ floor-sweep-3 (00:13:01 to 00:16:57Z, 237 s, exit 0; the v3 server b37defef, the
|---|---|---|---|---|
| budget 12 GB | block-83616 (empty) | 280,706 | 9,939 | 2.6 |
| budget 12 GB | block-56-transfers | 556,369 | 10,003 | 3.5 |
| budget 12 GB | fees-v1-shards2 (v1) | 4,717,439 | 12,915 | 4.3 |
| budget 12 GB | block-338-shard1 (adopted) | 60,415,376 | 13,459 | 16.8 |
| budget 12 GB | fees-v1-shards2 (v1, the adopted shard) | 4,717,439 | 12,915 | 4.3 |
| budget 12 GB | block-338-shard1 (prototype) | 60,415,376 | 13,459 | 16.8 |
| budget 12 GB, recursion alloc 2^26.6 | fees-v1-shards2 | 4,717,439 | 12,883 | 4.3 |
| budget 16 GB | fees-v1-shards2 | 4,717,439 | 16,115 | 4.0 |
| budget 32 GB | fees-v1-shards2 | 4,717,439 | 16,851 | 4.0 |
| element threshold 2^26 | block-83616 | 280,706 | 9,971 | 3.3 |
| element threshold 2^26 | fees-v1-shards2 | 4,717,439 | 10,291 | 5.7 |
The key-buffer growth fired as designed ("grow key buffer 36,700,160 -> 91,226,112 elements"), the panic class of sweep 2 is closed. Reading, pending the prover-floor agent's write-up: under a 12 GB budget the adopted shard peaks at 13,459 MiB and the v1 shard at 12,915 MiB with 2,089 MiB of idle inside, against about 12,288 MiB reported by a 12 GB card; the 2^26 element threshold brings the v1 shard to 10,291 MiB at 5.7 s against 4.3 (a third slower), which is the first row under a 12 GB card's size; the tier call is the agent's, in docs/analysis/proving-methods.md and the bench log. PC 2 released at 00:17Z; "go PC 2" to agg-cost-pc2-4 at 00:19 (about 20 min; the identities set back to 8 first), then M16, the repro job, the ledger-fixes-0311 suites.
The key-buffer growth fired as designed ("grow key buffer 36,700,160 -> 91,226,112 elements"), the panic class of sweep 2 is closed. Reading, pending the prover-floor agent's write-up: under a 12 GB budget the prototype shard peaks at 13,459 MiB and the adopted v1 shard at 12,915 MiB with 2,089 MiB of idle inside, against about 12,288 MiB reported by a 12 GB card; the 2^26 element threshold brings the v1 shard to 10,291 MiB at 5.7 s against 4.3 (a third slower), which is the first row under a 12 GB card's size; the tier call is the agent's, in docs/analysis/proving-methods.md and the bench log. PC 2 released at 00:17Z; "go PC 2" to agg-cost-pc2-4 at 00:19 (about 20 min; the identities set back to 8 first), then M16, the repro job, the ledger-fixes-0311 suites.