Counter ASIC 2.0 rollout 6b: mine-and-prove measured beside the miner (sweep 4)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
b33d33a0a9
commit
a87db8d3e1
1 changed files with 1 additions and 1 deletions
|
|
@ -71,7 +71,7 @@ AMD RDNA 4 (the RX 9070 XT) sits at about a seventh of an RTX 5090 on this hash
|
|||
|
||||
HiveOS rigs (consequences C41, 23:33 UTC): untested on a GPU host tonight, and local mode never worked (the bundled node starts without the override file, so a local-mode rig has been refused by every devnet peer since the first height switch); the fix is the next cut's one HiveOS item (section 8a). The installer path with the hosted node is unaffected.
|
||||
|
||||
Prover memory tiers (the proving-methods agent, 23:42 UTC, from the prover-floor agent's patch v1 rows in docs/analysis/proving-methods.md on branch proving-methods, not pushed): the adopted shard falls from 20,516 to 12,708 MiB (idle included) with the 24 GB panic removed and the element threshold budgeted; the next floor is the server's Setup at 9,703 MiB (five recursion keys pre-built at 2^27 capacity). Per tier: the 16 GB tier opens now (prove-only at 2^27, 15.4 GB; mine-and-prove at 2^26 on paper, 14.4 GB, to be measured beside the miner); the 12 GB tier was not open on patch v1 (12,708 MiB against about 12,288 MiB reported by a 12 GB card). MEASURED on patch v3 (floor-sweep-3, 00:13 to 00:17Z 6 October, the v3 server b37defef on the RTX 5090, miners stopped, every proof verified by the unpatched host): with the element threshold at 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`) the v1 shard (4.7 M cycles) peaks at 10,291 MiB of device memory with the card's 2,089 MiB idle inside (about 8.2 GB the server's own) at 5.7 s against 4.3 s at 2^27; the empty shard 9,971 MiB; the server's floor after Setup 6,535 MiB (9,703 on v1, 11,623 stock). Under a plain 12 GB budget at 2^27 the v1 shard (the adopted proving v1 shard) peaks at 12,915 MiB and the prototype shard (60.4 M cycles, not the adopted one) at 13,459 MiB in 16.8 s (28,295 MiB on the stock server). So the 12 GB tier profile is the 2^26 threshold on the v3 server, a third slower per shard; RISC Zero at po2 19 is no longer the 12 GB route (it stays the Apple route in proving-methods.md) unless a real 12 GB card contradicts the 5090's reading; the proving-methods agent's tier reading: 12 GB prove-only opens with about 2 GB spare and about 0.3 GB spare beside the miner (approximate), so mine-and-prove on 12 GB still points at a half shard; 16 GB holds the prototype shard alone. The 16 GB tier at 2^27: 16,115 MiB on the v1 shard, 4.0 s; upstream's own threshold (budget 32): 16,851 MiB. Not measured: any 12 GB card (a 4070 or 3060 is the on-order check), and mine-and-prove beside the miner (sweep 4, 4 points, queued on PC 2 after the aggregation-cost re-run). Recommended, for the project lead: provedefault.rs gains the 16 GB tier; the rig installer runs one server per card; a 4070 or 3060 enters the measurement loop, since no 12 GB card has run a prover here (every 12 GB number is a 32 GB card's reading with a budget).
|
||||
Prover memory tiers (the proving-methods agent, 23:42 UTC, from the prover-floor agent's patch v1 rows in docs/analysis/proving-methods.md on branch proving-methods, not pushed): the adopted shard falls from 20,516 to 12,708 MiB (idle included) with the 24 GB panic removed and the element threshold budgeted; the next floor is the server's Setup at 9,703 MiB (five recursion keys pre-built at 2^27 capacity). Per tier: the 16 GB tier opens now (prove-only at 2^27, 15.4 GB; mine-and-prove at 2^26 on paper, 14.4 GB, to be measured beside the miner); the 12 GB tier was not open on patch v1 (12,708 MiB against about 12,288 MiB reported by a 12 GB card). MEASURED on patch v3 (floor-sweep-3, 00:13 to 00:17Z 6 October, the v3 server b37defef on the RTX 5090, miners stopped, every proof verified by the unpatched host): with the element threshold at 2^26 (`SP1_GPU_ELEMENT_THRESHOLD=67108864`) the v1 shard (4.7 M cycles) peaks at 10,291 MiB of device memory with the card's 2,089 MiB idle inside (about 8.2 GB the server's own) at 5.7 s against 4.3 s at 2^27; the empty shard 9,971 MiB; the server's floor after Setup 6,535 MiB (9,703 on v1, 11,623 stock). Under a plain 12 GB budget at 2^27 the v1 shard (the adopted proving v1 shard) peaks at 12,915 MiB and the prototype shard (60.4 M cycles, not the adopted one) at 13,459 MiB in 16.8 s (28,295 MiB on the stock server). So the 12 GB tier profile is the 2^26 threshold on the v3 server, a third slower per shard; RISC Zero at po2 19 is no longer the 12 GB route (it stays the Apple route in proving-methods.md) unless a real 12 GB card contradicts the 5090's reading; the proving-methods agent's tier reading: 12 GB prove-only opens with about 2 GB spare and about 0.3 GB spare beside the miner (approximate), so mine-and-prove on 12 GB still points at a half shard; 16 GB holds the prototype shard alone. The 16 GB tier at 2^27: 16,115 MiB on the v1 shard, 4.0 s; upstream's own threshold (budget 32): 16,851 MiB. Mine-and-prove MEASURED beside the 5090 miner at full rate (floor-sweep-4, 00:34 to 00:38Z 6 October, the miner's 3,833 MiB resident inside every peak): at 2^26 the v1 shard peaks at 12,066 MiB (8,233 the server's own) in 24.4 s, the empty shard 11,586 MiB (7,753) in 13.0 s; at 2^25 the same 12,066 MiB in 43.5 s (no memory gained); at 2^27 14,786 MiB (10,953) in 17.4 s. The server's own working set does not move with the miner (8,233 beside it, 8,202 alone); the miner costs 4.3x in time, inside proving v1's 600 DAA-second deadline by 25x. Tier reading on the 5090's allocation: a 12 GB card mining and proving holds 8.2 + 1.7 (the miner) + 0.5 to 1 (its display, not measured) = 10.4 to 10.9 GB, 1.4 to 1.9 GB spare at 2^26 and 24 s a shard; a 16 GB card mines and proves at 2^27 with 2.9 GB spare at 17 s. Not measured: any real 12 GB card (the on-order 4070 or 3060 runs the same two points, its own display share replacing the 5090's). Recommended, for the project lead: provedefault.rs gains the 16 GB tier; the rig installer runs one server per card; a 4070 or 3060 enters the measurement loop, since no 12 GB card has run a prover here (every 12 GB number is a 32 GB card's reading with a budget).
|
||||
|
||||
## 7. Gates before any publish (all of them, no exceptions)
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue