diff --git a/docs/design/class-v5-stored-state.md b/docs/design/class-v5-stored-state.md index c8096d539..bbc391aa0 100644 --- a/docs/design/class-v5-stored-state.md +++ b/docs/design/class-v5-stored-state.md @@ -44,6 +44,7 @@ Clock rule for this page: every time is UK time (UK time tonight, UTC+1). The bu | 22:10 | Fix commit 8ca66afa on both mirrors (AP-F4-1 agreed form, verified last resort; suite 74 + packs 20 green, gate GREEN). The class-walk harness case: FAIL on the unfixed fork (e0:v3 e1:v3), PASS on the node lane's pair 4 c8f9b383 (e0:v4 e1:v4, v5 from epoch 8) | | 01:5x | Tonight's record landed on both mirrors: main's sentence (nine of nine hot sets refused), the 61-row era reading, the class v5 attempts census, AP-F8-3 and AP-F8-6, spec 1.4.7 and 1.8.6 with the constants and id tables, the AMD fingerprint, the master merges (the program-id recipe form, the audit lane's spec rewrite and checks). Proofs on the tree: gate GREEN 71 checks; igneum-pow suite on box 2 74 unit, derivation 2, derive 7, mixer 4, packs 20 (the pinned ids and 82b19cbde8557ea5 under master's recipe form), recheck 2, scratch 7, spec_readback 2. Two corrections found by the proofs: the pinned packs' derivation text (master's export writes the sub-version suffix for generator 4; class v5's rung-0 text follows its plain id, the missing arm added) and the public-export scrub (the founder's name on the page, the zone name in three files) | | 05:2x | F1 PASS on the freeze (100,000 class v5 programs, max saved share 4.688 percent, 0 over 5, 0 mismatches, 0 panics): the attack-pass families on class v5 read F8 PASS with the known residue, F4 PASS (AP-F4-1 fixed and passed), F9 PASS, F1 PASS | +| 09:4x | The Arc B580 read (09:36 UK): no fingerprint, the kit worker's self-test fails on both classes' vector lanes with the dataset right: an Intel OpenCL per-nonce evaluation fault, not class v5's; the Intel kit HELD out of 0.3.24 with the crossing time 09:36 UK; the kit stands on five of six platforms | ## 1. The claim, in one paragraph @@ -282,7 +283,7 @@ Every item of this page that has no code, no test or no measurement yet, with th | The kit for every platform: Metal | done (branch `v5-kits`, 7 October 2026, 19:3x UK): `proto-metal/main.swift` `servePackDataset` reads `leaves.bin`, checks the pack's count and FNV-1a 64, binds buffers 2 and 3 as `packbench.swift` does, self-tests the dataset head and last word against `vectors.json` and drops the leaf buffer after the build; a v5 prepare and job served on the M5 Max at 18:37:32Z | kits lane | 0 | | The kit: CUDA | done (`v5-kits`): `packfile.h` reads `IGNEUM_STATE_LEAVES`, the FNV, file, root and block (generator 5 = class v5; a v5 pack without the count or another class with one is refused), `pf_load_leaves` reads and checks `leaves.bin`; `worker.cpp` uploads it (`cuMemcpyHtoD`) for `igneum_build(ds, cache, leaves, nLeaves, nItems)` and frees it after; the CPU emulation runs `--check` on v5-dn3-epoch0: PASS on igneum-build-1 at 18:40:52Z (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 lanes); the loader test carries the class v5 cases known-failed first (46 ok) | kits lane | 0 | | The kit: OpenCL (AMD), Intel | host done (`v5-kits`): `proto-opencl/host.c` sets the build arguments for both kernel shapes in one place and uploads the leaves in `--bench-pack`, `--serve` and the compiled-in bench (the host reference derives words from the same leaves); the Linux and Windows workers are in the kit zip; the 9070 XT run placed on PC 1 as one lock slot after the 9070 XT G1 and ladder job (the hash lane publishes `tools/class-v5/pc1-amd-v5-bench.ps1`, expected 82b19cbde8557ea5); Intel: not measured tonight, no Arc B580 on PC 1 or PC 2 and the only Arc path needs a driver click, which no PC job may raise (the coordinator's ruling, 19:5x UK), so the card's holder and a click-free driver path are owed | coordinator (PC 1 queue), the B580's holder | 1 | -| Fingerprints equal across platforms, G1 on the 5090 and the Mac, AMD and Intel after | v5-dn3-epoch0 reads 82b19cbde8557ea5 on Metal (packbench, 18:34:47Z, 14.163 MH/s GPU time) and on Apple OpenCL (the kit host `--bench-pack`, 18:34:59Z, 15.265 MH/s wall), both under the Mac's measure lock, cache FNV and vectors PASS; CUDA: 82b19cbde8557ea5 on a one-shot RunPod RTX 4090 (driver 580.159.04, NVRTC 12.8, sm_89, 450 W) from the kit zip through `tools/class-v5/fleet-cuda-v5-bench.sh`: the kit's Linux NVRTC worker `--bench` at 18:57:45Z (check PASS, 62.759 MH/s over 5 dispatches of 2^24) and `bench.cu` at 18:57:50Z (the fleet lane's own run 62.775 MH/s); no 5090 was quiet (p1-5090 is a live voter), so the 5090 row stays the fleet's when one frees; AMD and Intel as the row above AMD READ (the hash lane's PC 1 queue job run-ca3-pc1-v5-amd-bench-20261007, 01:15 to 01:16 UK, exit 0, the card loaded beside the miners): RX 9070 XT gfx1201 (AMD-APP 3683.0, 32 CUs) on the kit zip packs-ca3-v5-20261007T221001Z (worker sha256 27faa253..., kernel_bound.cl 6b40f2dd..., leaves.bin 704 bytes 50687c8d...): self-test PASS (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 vector lanes), fingerprint of the 2^24 outputs at base 0 82b19cbde8557ea5 = the expected, match True; the v4-genesis control 892b6d55a7ddcfcb PASS. The rates (6.8 and 7.7 MH/s) are loaded-card figures, not the row. Intel: NOT MEASURED tonight (the coordinator's word, 03:4x UK): the Arc B580 job on the second PC (run-ca3-pc2-v5-intel-bench-20261007, arc-v5-bench.ps1's content, the kit on that PC since 01:12 UK) stays queued until a Windows entry exists; the Intel kit rides 0.3.25 with the crossing time stated when it reads. | coordinator (AMD, Intel), fleet (5090) | 1 | +| Fingerprints equal across platforms, G1 on the 5090 and the Mac, AMD and Intel after | v5-dn3-epoch0 reads 82b19cbde8557ea5 on Metal (packbench, 18:34:47Z, 14.163 MH/s GPU time) and on Apple OpenCL (the kit host `--bench-pack`, 18:34:59Z, 15.265 MH/s wall), both under the Mac's measure lock, cache FNV and vectors PASS; CUDA: 82b19cbde8557ea5 on a one-shot RunPod RTX 4090 (driver 580.159.04, NVRTC 12.8, sm_89, 450 W) from the kit zip through `tools/class-v5/fleet-cuda-v5-bench.sh`: the kit's Linux NVRTC worker `--bench` at 18:57:45Z (check PASS, 62.759 MH/s over 5 dispatches of 2^24) and `bench.cu` at 18:57:50Z (the fleet lane's own run 62.775 MH/s); no 5090 was quiet (p1-5090 is a live voter), so the 5090 row stays the fleet's when one frees; AMD and Intel as the row above AMD READ (the hash lane's PC 1 queue job run-ca3-pc1-v5-amd-bench-20261007, 01:15 to 01:16 UK, exit 0, the card loaded beside the miners): RX 9070 XT gfx1201 (AMD-APP 3683.0, 32 CUs) on the kit zip packs-ca3-v5-20261007T221001Z (worker sha256 27faa253..., kernel_bound.cl 6b40f2dd..., leaves.bin 704 bytes 50687c8d...): self-test PASS (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 vector lanes), fingerprint of the 2^24 outputs at base 0 82b19cbde8557ea5 = the expected, match True; the v4-genesis control 892b6d55a7ddcfcb PASS. The rates (6.8 and 7.7 MH/s) are loaded-card figures, not the row. Intel: HELD out of 0.3.24, crossing time 09:36 UK (the Arc B580 job on the second PC, run-ca3-pc2-v5-intel-bench-20261007, read at 09:36 UK): no fingerprint; the kit worker (igneum-worker-opencl.exe sha256 27faa253...) fails its self-test on the Arc before any batch, on the v5 pack AND on the class v4 control alike, with every cache and dataset check passing (cache head, last line and FNV 7334fa46e5d972eb and 48c4f5bf24166b2e; dataset head, last word and 64 samples) and 96 of 96 vector lanes wrong (v5 lane 0 at base 0: device 729ebd46376e2851, expected e552166a03298f7f; v4 control lane 0: device 11bdacb6ee4108c2, expected dfbc8db1c06dacd8). So the fault is the bound kernel's per-nonce evaluation on Intel OpenCL (Intel(R) OpenCL Graphics, OpenCL 3.0 platform, driver 32.0.101.6733, the kernel compiled as OpenCL C 1.2, 160 CUs, sub-group shuffle present), on both classes, not class v5's leaves; the RX 9070 XT matched on the same kit. The kit stands on five of six platforms. Open: the shipper's read of the installed worker's self-test on that card; the prime suspect is the program's cross-lane shfl in a sub-group the Intel compiler sized under 32 while the size query answered 32 (the fix: intel_reqd_sub_group_size(32) with the compiled size verified, else the local-memory exchange), which waits on the job's exchange and kernel lines from the hash lane. The Intel kit rides 0.3.25. | coordinator (AMD, Intel), fleet (5090) | 1 | | The kit zip | `igneum-build-1:/srv/artefacts/packs/packs-ca3-v5-20261007T183921Z.zip`, sha256 e6c088bb34fecdc3ff297dbb06438a14ade7d8c55273357726d28f7a1334a25e, 919,129 bytes, 56 files (v4-genesis, v5-genesis, v5-dn3-epoch0 with `leaves.bin`; `bin/linux` and `bin/windows` NVRTC and OpenCL workers; sources; the two bench scripts; SHA256SUMS), built by `tools/class-v5/kits-remote.sh` on build-1 at 18:41:15Z after the loader test and the emulation check | kits lane | 0 | | The crossing on Devnet 3 by height after its gate | NOT started: the floor in the Devnet 3 override, the Devnet 2 style gate (zero rejected across the flip, exec roots agreeing, every node holding the epoch's stream); the harness's flip case is the rehearsal (PASS twice) | node lane (a283f5f0d364ceef0) owns the fork, the shipper (ae892a8b0f78fe31c) the cut; this lane the gate's v5 checks | 3 | | Miner cost rows per card (5090, M5 Max, 9070 XT, 4070) | the 4090 row only (rate equal within 0.01 percent, build +0.24 ms); the four named cards owed, each labelled measured with the date | fleet (5090, 4070), this lane (M5 Max under the lock), PC 1 (9070 XT) | 3 |