Counter ASIC 3.0 item 2: the RTX 5090 rows from PC 2 (bit-exact on CUDA, +1.1 s NVRTC per pack, the card-off key finding)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
bdc07d3a1d
commit
54bc188a5d
2 changed files with 64 additions and 6 deletions
|
|
@ -2040,5 +2040,25 @@ docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core
|
|||
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
|
||||
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
|
||||
|
||||
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
|
||||
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
|
||||
**RTX 5090 (PC 2, one job `run-ca3-derive-pc2-20261006`, `relay/playbooks/ca3-derive-pc2.ps1`, published 08:26:08Z
|
||||
after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0):**
|
||||
the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read
|
||||
the key from settings.json (`nvidia:0:NVIDIA GeForce RTX 5090`, with the device index) where the 5 October jobs
|
||||
posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a
|
||||
loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.
|
||||
|
||||
| Pack | NVRTC | Cache | 1 GiB build | Self-test (64 samples, 96 lanes) | Fingerprint 2^24 | MH/s bw1 / bw8 (loaded) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2-genesis-mh | 167 ms | 6 | 46 ms | PASS | 25f96e7dce90bd4e = Mac | 62.26 / 61.34 |
|
||||
| mx8-genesis | 164 ms | 4 | 40 ms | PASS | 7c28cfb06c5c65a9 = Mac | 61.98 / 60.53 |
|
||||
| dr736-genesis | 1,266 ms | 5 | 42 ms | PASS | 50e3eaa779da4f1e = Metal = Apple OpenCL | 61.08 / 58.51 |
|
||||
| dr736-devnet-epoch0 | 1,266 ms | 6 | 32 ms | PASS | 9553f6d5c667205a = Metal | 62.15 / 61.44 |
|
||||
|
||||
Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond
|
||||
the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement
|
||||
item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch,
|
||||
approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a
|
||||
day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that
|
||||
fix and 1.1 s per variant per epoch before it; the Mac 0.75 s once a day. Filed: the card-off key form (post both
|
||||
forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. **RX 9070 XT:**
|
||||
OWED (PC 1).
|
||||
|
|
|
|||
|
|
@ -17,9 +17,10 @@ per item, the cache, the loads and the hash kernel are untouched.
|
|||
|---|---|---|---|
|
||||
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
|
||||
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
|
||||
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is Josh's desk today) |
|
||||
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
|
||||
| Bit-exact: Metal, Apple OpenCL and CUDA (NVRTC, PC 2) against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples, 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, all three compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal and CUDA) | equal | passes on three compilers (5.2, 5.4a) |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock; RTX 5090, PC 2 (the card still mining, 5.4a) | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%); 5090 42 / 32 ms against x8's 40 and v2's 46 on the loaded card | under 1 s on every discrete card | passes on both; the 9070 XT OWED (PC 1 is Josh's desk today) |
|
||||
| NVRTC compile per pack, RTX 5090 | 1,266 ms against x8's 164 (+1.1 s): the item function is inside every hash-kernel and race-variant compile | the 38 s compile-ahead budget at the 600-s epoch floor | the one-module-per-day fix is a requirement of the class (5.4a) |
|
||||
| Hash rate, M5 Max, Metal, measure lock; RTX 5090 (loaded) | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16; 5090 61.1 / 62.1 against x8 62.0 and v2 62.3 | equal within noise | equal (0.3% Mac): the hash kernel does not change |
|
||||
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
|
||||
|
||||
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
|
||||
|
|
@ -351,6 +352,42 @@ NVIDIA card off in the app only under test (its key from settings.json), `--chec
|
|||
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
|
||||
The rows land in the bench-log entry when the closing report is read.
|
||||
|
||||
### 5.4a RTX 5090, PC 2: the rows (job `run-ca3-derive-pc2-20261006`, published 08:26:08Z, ran 08:26:43 to 08:28:49Z, done exit 0, 126 s; lock taken 08:25:55Z, released 08:29:08Z)
|
||||
|
||||
The installed 0.3.11 worker (sha256 2b3b8c92...c2674c, NVRTC 12.8, driver 13.3, sm_120, `--bench` present) on the
|
||||
self-fetched zip (sha256 verified on the PC). THE CARD DID NOT COME OFF: the job read the card key from settings.json
|
||||
as the brief asks (`nvidia:0:NVIDIA GeForce RTX 5090`), the mixer job of 5 October had used the state's key
|
||||
(`nvidia:NVIDIA GeForce RTX 5090`, no index), and after the POST and 90 s one `igneum-worker-cuda` process was
|
||||
still running (`worker processes left 1`; `gpu-before` 431 W, 3,896 MiB used, the prover's GPU server ON as the
|
||||
clear file requires). So every row below was taken with the installed miner still on the card: the v2 control
|
||||
reads 62.3 MH/s against its unloaded 136 to 137, and the same load sits under every pack, so the ratios between
|
||||
the rows are the measurement and the absolute rates and build times are not (the key-format mismatch is filed in
|
||||
section 8; the fix is to try both key forms and to confirm by the process list before measuring).
|
||||
|
||||
| Pack | NVRTC compile | Cache | 1 GiB build | Self-test | Fingerprint 2^24 (the Mac's) | MH/s, block-warps 1 / 8 (loaded) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2-genesis-mh (control) | 167 ms | 6 ms | 46 ms | PASS (cache FNV 48c4f5bf24166b2e, head, word [MASK], 64 samples, 96 lanes) | 25f96e7dce90bd4e (equal) | 62.26 / 61.34 |
|
||||
| mx8-genesis (x8, class v3) | 164 ms | 4 ms | 40 ms | PASS | 7c28cfb06c5c65a9 (equal) | 61.98 / 60.53 |
|
||||
| dr736-genesis | 1,266 ms | 5 ms | 42 ms | PASS (cache FNV 48c4f5bf24166b2e, 64 samples, 96 lanes) | 50e3eaa779da4f1e (equal to Metal and Apple OpenCL) | 61.08 / 58.51 |
|
||||
| dr736-devnet-epoch0 | 1,266 ms | 6 ms | 32 ms | PASS (cache FNV 448274a57f508cbc, 64 samples, 96 lanes) | 9553f6d5c667205a (equal to Metal) | 62.15 / 61.44 |
|
||||
|
||||
Reading. Bit-exactness on CUDA: PASS on both derivation packs, the 64 sampled words and the 96 vector lanes
|
||||
against the Rust interpreter, and the 2^24 fingerprints equal to the Mac's, so the day program now agrees across
|
||||
four compilers (Rust, Metal, Apple OpenCL, NVRTC). The build: 42 and 32 ms against x8's 40 and v2's 46 on the
|
||||
loaded card (the unloaded x8 figure was 23 to 25 ms): the day program does not move the 5090's build beyond
|
||||
noise, as on the Mac it moved it 7 ms. The hash rate: equal to v2 and x8 within the loaded noise (61 to 62
|
||||
against 62), as it must be. THE NUMBER THAT MOVED: the NVRTC compile, 1,266 ms against 164 for x8 (+1.1 s per
|
||||
pack), because `kernel_bound.cu` includes memhard.h and the 6,624-statement item function is compiled into every
|
||||
hash kernel and every race variant: at 17 variants that is about +19 s of compile-ahead per epoch on the 5090
|
||||
(approximate, 17 x 1.1 s; the race measured 232 to 300 ms for 17 variants today, epoch-length.md 6.1) unless the
|
||||
item function is compiled once a day into its own module, which is the fix named in 5.5 and now a requirement
|
||||
for the class, not an option: the 600-s epoch floor of the ladder leaves 38 s for the race today and this would
|
||||
take half of it.
|
||||
|
||||
Consequences per tier: a 5090 owner pays about 1.1 s of compile once a day (the item library) plus, until the
|
||||
one-module fix lands, 1.1 s per epoch per variant; the Mac pays 0.75 s once a day; an AMD owner pays the OpenCL
|
||||
build (OWED, PC 1); the integrated tier's compile is unmeasured and its build was already the open problem.
|
||||
|
||||
### 5.5 The build per tier, with the integrated tier
|
||||
|
||||
| Card | Build at x8 | Build with the day program | Source |
|
||||
|
|
@ -419,7 +456,8 @@ chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x
|
|||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
|
||||
| RTX 5090 unloaded: the job's card-off did not take (the settings.json key carries the device index, `nvidia:0:...`, the api/cards key of the earlier jobs did not; the fix is to post both forms and confirm by the process list) so the 5090's absolute build and rate rows are loaded figures with the ratios valid (5.4a); the fingerprints and the compile are not affected | a re-run after the key fix, on the next PC 2 slot |
|
||||
| The item function as its own NVRTC module once a day (the +1.1 s per compile of 5.4a) | unimplemented; a requirement of the class before any activation |
|
||||
| RX 9070 XT (PC 1) | OWED: PC 1 is Josh's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
|
||||
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
|
||||
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
|
||||
|
|
|
|||
Loading…
Reference in a new issue