Counter ASIC 3.0 item 2: the RTX 5090 rows from PC 2 (bit-exact on CUDA, +1.1 s NVRTC per pack, the card-off key finding)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-06 09:30:15 +01:00
parent bdc07d3a1d
commit 54bc188a5d
2 changed files with 64 additions and 6 deletions

View file

@ -2040,5 +2040,25 @@ docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
**RTX 5090 (PC 2, one job `run-ca3-derive-pc2-20261006`, `relay/playbooks/ca3-derive-pc2.ps1`, published 08:26:08Z
after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0):**
the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read
the key from settings.json (`nvidia:0:NVIDIA GeForce RTX 5090`, with the device index) where the 5 October jobs
posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a
loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.
| Pack | NVRTC | Cache | 1 GiB build | Self-test (64 samples, 96 lanes) | Fingerprint 2^24 | MH/s bw1 / bw8 (loaded) |
|---|---|---|---|---|---|---|
| v2-genesis-mh | 167 ms | 6 | 46 ms | PASS | 25f96e7dce90bd4e = Mac | 62.26 / 61.34 |
| mx8-genesis | 164 ms | 4 | 40 ms | PASS | 7c28cfb06c5c65a9 = Mac | 61.98 / 60.53 |
| dr736-genesis | 1,266 ms | 5 | 42 ms | PASS | 50e3eaa779da4f1e = Metal = Apple OpenCL | 61.08 / 58.51 |
| dr736-devnet-epoch0 | 1,266 ms | 6 | 32 ms | PASS | 9553f6d5c667205a = Metal | 62.15 / 61.44 |
Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond
the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement
item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch,
approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a
day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that
fix and 1.1 s per variant per epoch before it; the Mac 0.75 s once a day. Filed: the card-off key form (post both
forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. **RX 9070 XT:**
OWED (PC 1).

View file

@ -17,9 +17,10 @@ per item, the cache, the loads and the hash kernel are untouched.
|---|---|---|---|
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is Josh's desk today) |
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
| Bit-exact: Metal, Apple OpenCL and CUDA (NVRTC, PC 2) against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples, 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, all three compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal and CUDA) | equal | passes on three compilers (5.2, 5.4a) |
| Daily 1 GiB build, M5 Max, Metal, measure lock; RTX 5090, PC 2 (the card still mining, 5.4a) | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%); 5090 42 / 32 ms against x8's 40 and v2's 46 on the loaded card | under 1 s on every discrete card | passes on both; the 9070 XT OWED (PC 1 is Josh's desk today) |
| NVRTC compile per pack, RTX 5090 | 1,266 ms against x8's 164 (+1.1 s): the item function is inside every hash-kernel and race-variant compile | the 38 s compile-ahead budget at the 600-s epoch floor | the one-module-per-day fix is a requirement of the class (5.4a) |
| Hash rate, M5 Max, Metal, measure lock; RTX 5090 (loaded) | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16; 5090 61.1 / 62.1 against x8 62.0 and v2 62.3 | equal within noise | equal (0.3% Mac): the hash kernel does not change |
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
@ -351,6 +352,42 @@ NVIDIA card off in the app only under test (its key from settings.json), `--chec
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
The rows land in the bench-log entry when the closing report is read.
### 5.4a RTX 5090, PC 2: the rows (job `run-ca3-derive-pc2-20261006`, published 08:26:08Z, ran 08:26:43 to 08:28:49Z, done exit 0, 126 s; lock taken 08:25:55Z, released 08:29:08Z)
The installed 0.3.11 worker (sha256 2b3b8c92...c2674c, NVRTC 12.8, driver 13.3, sm_120, `--bench` present) on the
self-fetched zip (sha256 verified on the PC). THE CARD DID NOT COME OFF: the job read the card key from settings.json
as the brief asks (`nvidia:0:NVIDIA GeForce RTX 5090`), the mixer job of 5 October had used the state's key
(`nvidia:NVIDIA GeForce RTX 5090`, no index), and after the POST and 90 s one `igneum-worker-cuda` process was
still running (`worker processes left 1`; `gpu-before` 431 W, 3,896 MiB used, the prover's GPU server ON as the
clear file requires). So every row below was taken with the installed miner still on the card: the v2 control
reads 62.3 MH/s against its unloaded 136 to 137, and the same load sits under every pack, so the ratios between
the rows are the measurement and the absolute rates and build times are not (the key-format mismatch is filed in
section 8; the fix is to try both key forms and to confirm by the process list before measuring).
| Pack | NVRTC compile | Cache | 1 GiB build | Self-test | Fingerprint 2^24 (the Mac's) | MH/s, block-warps 1 / 8 (loaded) |
|---|---|---|---|---|---|---|
| v2-genesis-mh (control) | 167 ms | 6 ms | 46 ms | PASS (cache FNV 48c4f5bf24166b2e, head, word [MASK], 64 samples, 96 lanes) | 25f96e7dce90bd4e (equal) | 62.26 / 61.34 |
| mx8-genesis (x8, class v3) | 164 ms | 4 ms | 40 ms | PASS | 7c28cfb06c5c65a9 (equal) | 61.98 / 60.53 |
| dr736-genesis | 1,266 ms | 5 ms | 42 ms | PASS (cache FNV 48c4f5bf24166b2e, 64 samples, 96 lanes) | 50e3eaa779da4f1e (equal to Metal and Apple OpenCL) | 61.08 / 58.51 |
| dr736-devnet-epoch0 | 1,266 ms | 6 ms | 32 ms | PASS (cache FNV 448274a57f508cbc, 64 samples, 96 lanes) | 9553f6d5c667205a (equal to Metal) | 62.15 / 61.44 |
Reading. Bit-exactness on CUDA: PASS on both derivation packs, the 64 sampled words and the 96 vector lanes
against the Rust interpreter, and the 2^24 fingerprints equal to the Mac's, so the day program now agrees across
four compilers (Rust, Metal, Apple OpenCL, NVRTC). The build: 42 and 32 ms against x8's 40 and v2's 46 on the
loaded card (the unloaded x8 figure was 23 to 25 ms): the day program does not move the 5090's build beyond
noise, as on the Mac it moved it 7 ms. The hash rate: equal to v2 and x8 within the loaded noise (61 to 62
against 62), as it must be. THE NUMBER THAT MOVED: the NVRTC compile, 1,266 ms against 164 for x8 (+1.1 s per
pack), because `kernel_bound.cu` includes memhard.h and the 6,624-statement item function is compiled into every
hash kernel and every race variant: at 17 variants that is about +19 s of compile-ahead per epoch on the 5090
(approximate, 17 x 1.1 s; the race measured 232 to 300 ms for 17 variants today, epoch-length.md 6.1) unless the
item function is compiled once a day into its own module, which is the fix named in 5.5 and now a requirement
for the class, not an option: the 600-s epoch floor of the ladder leaves 38 s for the race today and this would
take half of it.
Consequences per tier: a 5090 owner pays about 1.1 s of compile once a day (the item library) plus, until the
one-module fix lands, 1.1 s per epoch per variant; the Mac pays 0.75 s once a day; an AMD owner pays the OpenCL
build (OWED, PC 1); the integrated tier's compile is unmeasured and its build was already the open problem.
### 5.5 The build per tier, with the integrated tier
| Card | Build at x8 | Build with the day program | Source |
@ -419,7 +456,8 @@ chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x
| Item | State |
|---|---|
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
| RTX 5090 unloaded: the job's card-off did not take (the settings.json key carries the device index, `nvidia:0:...`, the api/cards key of the earlier jobs did not; the fix is to post both forms and confirm by the process list) so the 5090's absolute build and rate rows are loaded figures with the ratios valid (5.4a); the fingerprints and the compile are not affected | a re-run after the key fix, on the next PC 2 slot |
| The item function as its own NVRTC module once a day (the +1.1 s per compile of 5.4a) | unimplemented; a requirement of the class before any activation |
| RX 9070 XT (PC 1) | OWED: PC 1 is Josh's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |