From c6e03662a6a26bdf91db82ee511afd2e48e99c10 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:30:15 +0000 Subject: [PATCH] Counter ASIC 3.0 item 2: the RTX 5090 rows from PC 2 (bit-exact on CUDA, +1.1 s NVRTC per pack, the card-off key finding) Co-Authored-By: Claude Fable 5.1 --- docs/bench-log.md | 24 +++++++++++-- docs/plans/counter-asic-3-derivation.md | 46 ++++++++++++++++++++++--- 2 files changed, 64 insertions(+), 6 deletions(-) diff --git a/docs/bench-log.md b/docs/bench-log.md index 53a6b8e34..e0c2b94aa 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -2040,5 +2040,25 @@ docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare. -**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2 -lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED. +**RTX 5090 (PC 2, one job `run-ca3-derive-pc2-20261006`, `relay/playbooks/ca3-derive-pc2.ps1`, published 08:26:08Z +after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0):** +the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read +the key from settings.json (`nvidia:0:NVIDIA GeForce RTX 5090`, with the device index) where the 5 October jobs +posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a +loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid. + +| Pack | NVRTC | Cache | 1 GiB build | Self-test (64 samples, 96 lanes) | Fingerprint 2^24 | MH/s bw1 / bw8 (loaded) | +|---|---|---|---|---|---|---| +| v2-genesis-mh | 167 ms | 6 | 46 ms | PASS | 25f96e7dce90bd4e = Mac | 62.26 / 61.34 | +| mx8-genesis | 164 ms | 4 | 40 ms | PASS | 7c28cfb06c5c65a9 = Mac | 61.98 / 60.53 | +| dr736-genesis | 1,266 ms | 5 | 42 ms | PASS | 50e3eaa779da4f1e = Metal = Apple OpenCL | 61.08 / 58.51 | +| dr736-devnet-epoch0 | 1,266 ms | 6 | 32 ms | PASS | 9553f6d5c667205a = Metal | 62.15 / 61.44 | + +Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond +the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement +item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch, +approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a +day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that +fix and 1.1 s per variant per epoch before it; the Mac 0.75 s once a day. Filed: the card-off key form (post both +forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. **RX 9070 XT:** +OWED (PC 1). diff --git a/docs/plans/counter-asic-3-derivation.md b/docs/plans/counter-asic-3-derivation.md index 473f2539c..132fd494e 100644 --- a/docs/plans/counter-asic-3-derivation.md +++ b/docs/plans/counter-asic-3-derivation.md @@ -17,9 +17,10 @@ per item, the cache, the loads and the hash kernel are untouched. |---|---|---|---| | Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) | | The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes | -| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 | -| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is the project lead's desk today) | -| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change | +| Bit-exact: Metal, Apple OpenCL and CUDA (NVRTC, PC 2) against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples, 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, all three compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal and CUDA) | equal | passes on three compilers (5.2, 5.4a) | +| Daily 1 GiB build, M5 Max, Metal, measure lock; RTX 5090, PC 2 (the card still mining, 5.4a) | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%); 5090 42 / 32 ms against x8's 40 and v2's 46 on the loaded card | under 1 s on every discrete card | passes on both; the 9070 XT OWED (PC 1 is the project lead's desk today) | +| NVRTC compile per pack, RTX 5090 | 1,266 ms against x8's 164 (+1.1 s): the item function is inside every hash-kernel and race-variant compile | the 38 s compile-ahead budget at the 600-s epoch floor | the one-module-per-day fix is a requirement of the class (5.4a) | +| Hash rate, M5 Max, Metal, measure lock; RTX 5090 (loaded) | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16; 5090 61.1 / 62.1 against x8 62.0 and v2 62.3 | equal within noise | equal (0.3% Mac): the hash kernel does not change | | Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies | Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class @@ -351,6 +352,42 @@ NVIDIA card off in the app only under test (its key from settings.json), `--chec dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8. The rows land in the bench-log entry when the closing report is read. +### 5.4a RTX 5090, PC 2: the rows (job `run-ca3-derive-pc2-20261006`, published 08:26:08Z, ran 08:26:43 to 08:28:49Z, done exit 0, 126 s; lock taken 08:25:55Z, released 08:29:08Z) + +The installed 0.3.11 worker (sha256 2b3b8c92...c2674c, NVRTC 12.8, driver 13.3, sm_120, `--bench` present) on the +self-fetched zip (sha256 verified on the PC). THE CARD DID NOT COME OFF: the job read the card key from settings.json +as the brief asks (`nvidia:0:NVIDIA GeForce RTX 5090`), the mixer job of 5 October had used the state's key +(`nvidia:NVIDIA GeForce RTX 5090`, no index), and after the POST and 90 s one `igneum-worker-cuda` process was +still running (`worker processes left 1`; `gpu-before` 431 W, 3,896 MiB used, the prover's GPU server ON as the +clear file requires). So every row below was taken with the installed miner still on the card: the v2 control +reads 62.3 MH/s against its unloaded 136 to 137, and the same load sits under every pack, so the ratios between +the rows are the measurement and the absolute rates and build times are not (the key-format mismatch is filed in +section 8; the fix is to try both key forms and to confirm by the process list before measuring). + +| Pack | NVRTC compile | Cache | 1 GiB build | Self-test | Fingerprint 2^24 (the Mac's) | MH/s, block-warps 1 / 8 (loaded) | +|---|---|---|---|---|---|---| +| v2-genesis-mh (control) | 167 ms | 6 ms | 46 ms | PASS (cache FNV 48c4f5bf24166b2e, head, word [MASK], 64 samples, 96 lanes) | 25f96e7dce90bd4e (equal) | 62.26 / 61.34 | +| mx8-genesis (x8, class v3) | 164 ms | 4 ms | 40 ms | PASS | 7c28cfb06c5c65a9 (equal) | 61.98 / 60.53 | +| dr736-genesis | 1,266 ms | 5 ms | 42 ms | PASS (cache FNV 48c4f5bf24166b2e, 64 samples, 96 lanes) | 50e3eaa779da4f1e (equal to Metal and Apple OpenCL) | 61.08 / 58.51 | +| dr736-devnet-epoch0 | 1,266 ms | 6 ms | 32 ms | PASS (cache FNV 448274a57f508cbc, 64 samples, 96 lanes) | 9553f6d5c667205a (equal to Metal) | 62.15 / 61.44 | + +Reading. Bit-exactness on CUDA: PASS on both derivation packs, the 64 sampled words and the 96 vector lanes +against the Rust interpreter, and the 2^24 fingerprints equal to the Mac's, so the day program now agrees across +four compilers (Rust, Metal, Apple OpenCL, NVRTC). The build: 42 and 32 ms against x8's 40 and v2's 46 on the +loaded card (the unloaded x8 figure was 23 to 25 ms): the day program does not move the 5090's build beyond +noise, as on the Mac it moved it 7 ms. The hash rate: equal to v2 and x8 within the loaded noise (61 to 62 +against 62), as it must be. THE NUMBER THAT MOVED: the NVRTC compile, 1,266 ms against 164 for x8 (+1.1 s per +pack), because `kernel_bound.cu` includes memhard.h and the 6,624-statement item function is compiled into every +hash kernel and every race variant: at 17 variants that is about +19 s of compile-ahead per epoch on the 5090 +(approximate, 17 x 1.1 s; the race measured 232 to 300 ms for 17 variants today, epoch-length.md 6.1) unless the +item function is compiled once a day into its own module, which is the fix named in 5.5 and now a requirement +for the class, not an option: the 600-s epoch floor of the ladder leaves 38 s for the race today and this would +take half of it. + +Consequences per tier: a 5090 owner pays about 1.1 s of compile once a day (the item library) plus, until the +one-module fix lands, 1.1 s per epoch per variant; the Mac pays 0.75 s once a day; an AMD owner pays the OpenCL +build (OWED, PC 1); the integrated tier's compile is unmeasured and its build was already the open problem. + ### 5.5 The build per tier, with the integrated tier | Card | Build at x8 | Build with the day program | Source | @@ -419,7 +456,8 @@ chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x | Item | State | |---|---| -| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s | +| RTX 5090 unloaded: the job's card-off did not take (the settings.json key carries the device index, `nvidia:0:...`, the api/cards key of the earlier jobs did not; the fix is to post both forms and confirm by the process list) so the 5090's absolute build and rate rows are loaded figures with the ratios valid (5.4a); the fingerprints and the compile are not affected | a re-run after the key fix, on the next PC 2 slot | +| The item function as its own NVRTC module once a day (the +1.1 s per compile of 5.4a) | unimplemented; a requirement of the class before any activation | | RX 9070 XT (PC 1) | OWED: PC 1 is the project lead's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released | | The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it | | Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |