Merge branch 'ca3-derive' into ca3-coord
# Conflicts: # docs/bench-log.md
This commit is contained in:
commit
bc148a85ec
3 changed files with 69 additions and 7 deletions
|
|
@ -2040,8 +2040,29 @@ docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core
|
|||
the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate
|
||||
laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.
|
||||
|
||||
**RTX 5090 (PC 2, one job `relay/playbooks/ca3-derive-pc2.ps1`):** PENDING the proving agent's clear and the PC 2
|
||||
lock; the rows are appended below when the closing report is read. **RX 9070 XT:** OWED.
|
||||
**RTX 5090 (PC 2, one job `run-ca3-derive-pc2-20261006`, `relay/playbooks/ca3-derive-pc2.ps1`, published 08:26:08Z
|
||||
after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0):**
|
||||
the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read
|
||||
the key from settings.json (`nvidia:0:NVIDIA GeForce RTX 5090`, with the device index) where the 5 October jobs
|
||||
posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a
|
||||
loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.
|
||||
|
||||
| Pack | NVRTC | Cache | 1 GiB build | Self-test (64 samples, 96 lanes) | Fingerprint 2^24 | MH/s bw1 / bw8 (loaded) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2-genesis-mh | 167 ms | 6 | 46 ms | PASS | 25f96e7dce90bd4e = Mac | 62.26 / 61.34 |
|
||||
| mx8-genesis | 164 ms | 4 | 40 ms | PASS | 7c28cfb06c5c65a9 = Mac | 61.98 / 60.53 |
|
||||
| dr736-genesis | 1,266 ms | 5 | 42 ms | PASS | 50e3eaa779da4f1e = Metal = Apple OpenCL | 61.08 / 58.51 |
|
||||
| dr736-devnet-epoch0 | 1,266 ms | 6 | 32 ms | PASS | 9553f6d5c667205a = Metal | 62.15 / 61.44 |
|
||||
|
||||
Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond
|
||||
the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement
|
||||
item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch,
|
||||
approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a
|
||||
day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that
|
||||
fix and 1.1 s per variant per epoch before it; the Mac 0.75 s once a day. Filed: the card-off key form (post both
|
||||
forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. **RX 9070 XT:**
|
||||
OWED (PC 1).
|
||||
|
||||
## 6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow
|
||||
|
||||
Branch `ca3-shadow`, worker "shadow" (`docs/analysis/latency-shadow-2026-10-06.md` carries the design, the chip side and the consequences; this entry carries the measurements). The knob: `LoadClass::shadow`, class name `<class>+sh<S>x<R>`, a block of `S` ALU instructions drawn from the program stream after the 64 base instructions and run `R` times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, `cargo test` in igneum-pow 54 + 4 + 19 + 7 green). Packs `proto-cuda/packs-ca3-shadow/*` over `mx8` for seed igneum-genesis; the control is the pinned class v3 pack `packs-ca2-mixer/mx8-genesis`. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads).
|
||||
|
|
|
|||
|
|
@ -17,9 +17,10 @@ per item, the cache, the loads and the hash kernel are untouched.
|
|||
|---|---|---|---|
|
||||
| Verifier per 32-lane unit, one M5 Max core, `with-lock.sh measure`, load average 4.9 / 4.5 / 5.3 | 4.875 / 4.944 ms (two rounds of 50), worst cold unit 5.241; the devnet seeds 4.872 | 10 ms | passes, 5.1 ms of margin (x8 reads 2.061 / 2.063 in the same session: 2.37x) |
|
||||
| The same on a 2019-class laptop core (2.5x, approximate, the design document's ratio; O-1.14 unmeasured) | about 12.2 ms steady, 13.1 worst cold | 10 ms | FAILS on the approximate row; the half-length class `dr368` (the x4-equivalent op count) reads 2.692 ms here, about 6.7 ms on that row, and passes |
|
||||
| Bit-exact: Metal and Apple OpenCL against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples (OpenCL), 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, both compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal) | equal | passes on two compilers; CUDA (PC 2) section 5.4 |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%) | under 1 s on every discrete card | passes on the Mac; the 5090 section 5.4; the 9070 XT OWED (PC 1 is Josh's desk today) |
|
||||
| Hash rate, M5 Max, Metal, measure lock | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16 | equal within noise | equal (0.3%): the hash kernel does not change |
|
||||
| Bit-exact: Metal, Apple OpenCL and CUDA (NVRTC, PC 2) against the Rust CPU interpreter, two packs | cache FNV, dataset head and word [MASK], 64 samples, 96 vector lanes, 2^24 fingerprint 50e3eaa779da4f1e (dr736-genesis, all three compilers) and 9553f6d5c667205a (dr736-devnet-epoch0, Metal and CUDA) | equal | passes on three compilers (5.2, 5.4a) |
|
||||
| Daily 1 GiB build, M5 Max, Metal, measure lock; RTX 5090, PC 2 (the card still mining, 5.4a) | 29.0 / 29.1 / 28.9 ms GPU against mx8's 22.1 / 22.1 ms in the same session (+32%); 5090 42 / 32 ms against x8's 40 and v2's 46 on the loaded card | under 1 s on every discrete card | passes on both; the 9070 XT OWED (PC 1 is Josh's desk today) |
|
||||
| NVRTC compile per pack, RTX 5090 | 1,266 ms against x8's 164 (+1.1 s): the item function is inside every hash-kernel and race-variant compile | the 38 s compile-ahead budget at the 600-s epoch floor | the one-module-per-day fix is a requirement of the class (5.4a) |
|
||||
| Hash rate, M5 Max, Metal, measure lock; RTX 5090 (loaded) | dr736-genesis 27.06 to 27.13 MH/s GPU, mx8-genesis 27.08 to 27.16; 5090 61.1 / 62.1 against x8 62.0 and v2 62.3 | equal within noise | equal (0.3% Mac): the hash kernel does not change |
|
||||
| Chip model (section 7) | 1,278,976 chip ops per hash, 39.1 MH/s at 50 T op/s, 0.29x bare; 0.34x at a 1.2x allowance, 0.43x at 1.5x, 0.86x at the old 3x | under 1x | the allowance is the result: the 3x of the fixed shape no longer applies |
|
||||
|
||||
Go / no-go: GO as reserve entry R0 (section 6), NO-GO for genesis-live at 736 instructions until the 2019-class
|
||||
|
|
@ -351,6 +352,42 @@ NVIDIA card off in the app only under test (its key from settings.json), `--chec
|
|||
dataset .. ms` line and the self-test, `--bench` at 2^24 for the fingerprint and the rate, block-warps 1 and 8.
|
||||
The rows land in the bench-log entry when the closing report is read.
|
||||
|
||||
### 5.4a RTX 5090, PC 2: the rows (job `run-ca3-derive-pc2-20261006`, published 08:26:08Z, ran 08:26:43 to 08:28:49Z, done exit 0, 126 s; lock taken 08:25:55Z, released 08:29:08Z)
|
||||
|
||||
The installed 0.3.11 worker (sha256 2b3b8c92...c2674c, NVRTC 12.8, driver 13.3, sm_120, `--bench` present) on the
|
||||
self-fetched zip (sha256 verified on the PC). THE CARD DID NOT COME OFF: the job read the card key from settings.json
|
||||
as the brief asks (`nvidia:0:NVIDIA GeForce RTX 5090`), the mixer job of 5 October had used the state's key
|
||||
(`nvidia:NVIDIA GeForce RTX 5090`, no index), and after the POST and 90 s one `igneum-worker-cuda` process was
|
||||
still running (`worker processes left 1`; `gpu-before` 431 W, 3,896 MiB used, the prover's GPU server ON as the
|
||||
clear file requires). So every row below was taken with the installed miner still on the card: the v2 control
|
||||
reads 62.3 MH/s against its unloaded 136 to 137, and the same load sits under every pack, so the ratios between
|
||||
the rows are the measurement and the absolute rates and build times are not (the key-format mismatch is filed in
|
||||
section 8; the fix is to try both key forms and to confirm by the process list before measuring).
|
||||
|
||||
| Pack | NVRTC compile | Cache | 1 GiB build | Self-test | Fingerprint 2^24 (the Mac's) | MH/s, block-warps 1 / 8 (loaded) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| v2-genesis-mh (control) | 167 ms | 6 ms | 46 ms | PASS (cache FNV 48c4f5bf24166b2e, head, word [MASK], 64 samples, 96 lanes) | 25f96e7dce90bd4e (equal) | 62.26 / 61.34 |
|
||||
| mx8-genesis (x8, class v3) | 164 ms | 4 ms | 40 ms | PASS | 7c28cfb06c5c65a9 (equal) | 61.98 / 60.53 |
|
||||
| dr736-genesis | 1,266 ms | 5 ms | 42 ms | PASS (cache FNV 48c4f5bf24166b2e, 64 samples, 96 lanes) | 50e3eaa779da4f1e (equal to Metal and Apple OpenCL) | 61.08 / 58.51 |
|
||||
| dr736-devnet-epoch0 | 1,266 ms | 6 ms | 32 ms | PASS (cache FNV 448274a57f508cbc, 64 samples, 96 lanes) | 9553f6d5c667205a (equal to Metal) | 62.15 / 61.44 |
|
||||
|
||||
Reading. Bit-exactness on CUDA: PASS on both derivation packs, the 64 sampled words and the 96 vector lanes
|
||||
against the Rust interpreter, and the 2^24 fingerprints equal to the Mac's, so the day program now agrees across
|
||||
four compilers (Rust, Metal, Apple OpenCL, NVRTC). The build: 42 and 32 ms against x8's 40 and v2's 46 on the
|
||||
loaded card (the unloaded x8 figure was 23 to 25 ms): the day program does not move the 5090's build beyond
|
||||
noise, as on the Mac it moved it 7 ms. The hash rate: equal to v2 and x8 within the loaded noise (61 to 62
|
||||
against 62), as it must be. THE NUMBER THAT MOVED: the NVRTC compile, 1,266 ms against 164 for x8 (+1.1 s per
|
||||
pack), because `kernel_bound.cu` includes memhard.h and the 6,624-statement item function is compiled into every
|
||||
hash kernel and every race variant: at 17 variants that is about +19 s of compile-ahead per epoch on the 5090
|
||||
(approximate, 17 x 1.1 s; the race measured 232 to 300 ms for 17 variants today, epoch-length.md 6.1) unless the
|
||||
item function is compiled once a day into its own module, which is the fix named in 5.5 and now a requirement
|
||||
for the class, not an option: the 600-s epoch floor of the ladder leaves 38 s for the race today and this would
|
||||
take half of it.
|
||||
|
||||
Consequences per tier: a 5090 owner pays about 1.1 s of compile once a day (the item library) plus, until the
|
||||
one-module fix lands, 1.1 s per epoch per variant; the Mac pays 0.75 s once a day; an AMD owner pays the OpenCL
|
||||
build (OWED, PC 1); the integrated tier's compile is unmeasured and its build was already the open problem.
|
||||
|
||||
### 5.5 The build per tier, with the integrated tier
|
||||
|
||||
| Card | Build at x8 | Build with the day program | Source |
|
||||
|
|
@ -419,7 +456,8 @@ chip ops per hash, 78.2 MH/s, 0.57x bare, 0.69x at 1.2x, 0.86x at 1.5x: under 1x
|
|||
|
||||
| Item | State |
|
||||
|---|---|
|
||||
| RTX 5090: the daily build, NVRTC compile per pack, the self-test and 2^24 fingerprints, the rate | PENDING the PC 2 job (section 5.4); the playbook and the zip are ready; the clear file is polled every 60 s |
|
||||
| RTX 5090 unloaded: the job's card-off did not take (the settings.json key carries the device index, `nvidia:0:...`, the api/cards key of the earlier jobs did not; the fix is to post both forms and confirm by the process list) so the 5090's absolute build and rate rows are loaded figures with the ratios valid (5.4a); the fingerprints and the compile are not affected | a re-run after the key fix, on the next PC 2 slot |
|
||||
| The item function as its own NVRTC module once a day (the +1.1 s per compile of 5.4a) | unimplemented; a requirement of the class before any activation |
|
||||
| RX 9070 XT (PC 1) | OWED: PC 1 is Josh's desk today; the same job shape runs there with `igneum-worker-opencl.exe --bench-pack` when released |
|
||||
| The 2019-class laptop core (O-1.14) | unmeasured; the approximate row decides against 736 at genesis and for 368, and a measurement replaces it |
|
||||
| Cryptanalysis of random ARX programs | none; item 3's brief should name the day program as a target beside `M_r` |
|
||||
|
|
|
|||
|
|
@ -17,7 +17,7 @@ The test every result is judged against (Josh, 6 October): a chip maker must hav
|
|||
| # | Item | Worker branch | State | Close |
|
||||
|---|---|---|---|---|
|
||||
| 1 | Partial-store chip and the time-memory curve | ca3-analysis | CLOSED, merged (71df794) | verdict OVER 2x: the f = 1 chip (the dataset stored in DRAM, nothing recomputed) is 5.1x per joule on GDDR7 and 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent, $2.8 per MH/s against the 5090's $14.7; the curve is monotone toward f = 1, so the partial-store chip is never built; the mixer and item 2 do not touch it; chip-model-v3.md section 5 |
|
||||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | INTERIM merged (acb96ee, dd5041b); 5090 job queued on PC 2 | GO as reserve R0; NO-GO for genesis-live at dr736 (verifier 4.9 ms per unit on one M5 Max core, about 12 ms on a 2019-class core by the 2.5x rule, over the gate; dr368 2.69 ms passes both); urgency LOW after item 1 (the stored-dataset chip derives no item) |
|
||||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | CLOSED, merged (acb96ee, dd5041b, bdc07d3, 54bc188); 5090 job run-ca3-derive-pc2-20261006 exit 0 in 126 s | GO as reserve R0; NO-GO for genesis-live at dr736 (verifier 4.9 ms per unit on one M5 Max core, about 12 ms on a 2019-class core by the 2.5x rule, over the gate; dr368 2.69 ms passes both); urgency LOW after item 1 (the stored-dataset chip derives no item) |
|
||||
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | CLOSED, merged (c0642af, merge ed06814) | detector.mjs + 7 of 7 tests + one observer hook, dry run quiet on the devnet (max correlation 0.53 against the 0.8 edge; excess spread 0 to 5.2 percent); funding.md rule 5: bounty escrowed and benchmark live before daily issuance crosses USD 20,000 a day; epoch-length.md sections 11 and 12: signal trigger M = 6 windows; FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the bank-bound ceiling (unmeasured); layer 9 ranks above layer 7. The live observer is NOT restarted yet: the write path is untested; one restart after item 7's hook merges, then the first live detector row is recorded here |
|
||||
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
|
||||
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | merged to 6422b6f (merge fcce185); the 5090 family job queued on PC 2 | proposed reserve order R1 byte permute, R2 popcount and clz, R3 indexed shuffle, R4 bit-field extract, R5 variable shifts, R6 select, R7 andn, R8 mm8 (`docs/plans/counter-asic-3-reserve.md`, a decision for Josh, not in the spec); vendor-share.mjs with tests and one observer hook; repro.md carried from repro-bench (dda9fa3) with section 8: today's devnet NVIDIA 0.986 of blue blocks, Intel 0.014, coverage 1.00 at 135.8 MH/s with PC 1 off |
|
||||
|
|
@ -50,6 +50,8 @@ What moves the f = 1 rows (chip-model-v3.md section 5.7): not the dataset size (
|
|||
|
||||
Verifier headroom under the 10 ms gate (bdc07d3, the budget item 8's shadow ops may spend; steady / worst cold): on this M5 Max core x8 7.9 / 7.8 ms, dr368 7.3 / 7.1, dr736 5.1 / 4.8; on the approximate 2019-class core x8 4.8 / 4.6, dr368 3.3 / 2.8, dr736 none (over by 2.2 / 3.1). `cargo test -p igneum-pow` on acb96ee under the build lock: 95 passed, 0 failed; the pinned v2 and v3 packs regenerate byte for byte.
|
||||
|
||||
The 5090 rows (job run-ca3-derive-pc2-20261006, 08:26:43 to 08:28:49Z, exit 0; the card was LOADED: the miner stayed up because the job posted the settings key without the device index, see corrections; ratios valid, absolutes not): bit-exact on CUDA for both dr736 packs (64 samples, 96 lanes, fingerprints 50e3eaa779da4f1e and 9553f6d5c667205a equal to the Mac's); 1 GiB build 42 / 32 ms against x8 40 and v2 46; rates 61 to 62 MH/s on every pack (the v2 control 62.3 against 136 unloaded); NVRTC compile 1,266 ms against x8's 164 ms, +1.1 s per pack because memhard.h's item function sits inside every hash-kernel and race-variant compile, so a once-a-day derivation module is a requirement of the class, not an option. Prover left ON, app untouched.
|
||||
|
||||
Consequence: the derivation costs no hash rate on any card and 7 ms a day of build on the Mac (29 against 22 ms; the 5090 and the 9070 XT rows owed, the loaded-iGPU tier is the one to watch); it costs the verifier, and the verifier budget is shared with item 8's shadow ops under the one 10 ms gate, so the class v4 candidate is the pairing that fits, not either lever alone (both workers told).
|
||||
|
||||
### Item 8, the Mac rows (interim, ca3-shadow a050a54; knob d4b7300; `docs/analysis/latency-shadow-2026-10-06.md`; Metal packbench, IOReport GPU + DRAM watts without root; the 5090 rows queued on PC 2; the 9070 XT OWED)
|
||||
|
|
@ -96,6 +98,7 @@ The live observer runs the shared checkout, which autosync fast-forwards from or
|
|||
| Found by | What was wrong | Fixed |
|
||||
|---|---|---|
|
||||
| item 3 (43c3ead) | spec 01 section 1.13.1's era table and `docs/plans/mixer-x4.md` section 2 still said `mixer_mult = 4`; the code (`LoadClass::MX8`, `V3_CLASS`) and spec 1.8.5 say 8 | both lines corrected on ca3-coord, 6 October 2026 |
|
||||
| item 2's PC 2 job | the settings.json card key is `nvidia:0:NVIDIA GeForce RTX 5090` (with the device index); the 5 October job scripts posted the state's key without the index, so `POST api/cards` switched nothing and the miner stayed up through the 90 s wait: the job's 5090 rows are loaded-card figures | items 6 and 8 told to post both key forms and confirm by the process list; the class fix (one key form everywhere, a check that fails a job script posting a card key without the index) is owed to the job tooling |
|
||||
| the coordinator's merge of ca3-shadow | `git add -A docs` staged `docs/bench-log.md` with its conflict markers inside (a conflicted path is marked resolved by `git add`), so 45f3019 carried markers into HEAD; found at the ca3-reserve merge as nested markers | the three 6 October entries kept in order with every marker removed (fcce185); the class guard already exists, `tools/ci/no-conflict-markers.sh` in CI, and it would have failed the push; it passes on the merged tree |
|
||||
| the merge of ca3-derive and ca3-shadow | both added a field to `LoadClass` (`derive_len`, `shadow`); resolved as the union, every other literal spreads `..`; `cargo test -p igneum-pow` on the merged tree (cargo 1.99 at `~/.cargo/bin`; the Homebrew 1.69 on PATH cannot read the lock file): 59 + 7 + 4 + 19 + 7 = 96 passed, 0 failed, the pinned v2 and v3 packs byte for byte | merged 09:10 UTC |
|
||||
| item 3 | the mixer's op count: 144 integer ops per application as written in `memhard.rs` (128 with RC and rk hoisted) against the 130 the chip model prices (`chip-model-v3.md` section 1, from spec 1.8.4) | items 1 and 2 asked to state which figure their rows use and why; the status close carries the answer |
|
||||
|
|
|
|||
Loading…
Reference in a new issue