class v5 design page: the Intel row moves to its fingerprint (the Arc B580 reads 82b19cbde8557ea5 on the rebuilt kit; the rotate fold was the whole fault; six platforms equal)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 13:49:15 +00:00
parent 8146215d8b
commit 676ef54ee3

View file

@ -52,6 +52,7 @@ Clock rule for this page: every time is UK time (UK time tonight, UTC+1). The bu
| 14:1x | The pin's FAIL seed re-read (ec0a8cf9...8212, its era 234e082d...6062, day 1244071, the run's own epoch-9 stream from /tmp/igneum-fast-time-v5x-cross-0325-d5b68fae): igneum-pow at exactly 1c420786 draws attempt 3, program id 1f1cf82877f46ee6 (8f481459 the same). Neither id of the FAIL was the freeze's: the cross harness's "cli v5 ebf64b32e2d84c5b" came from the release worktree's pre-freeze igneum-pow copy and the pin miner's 65b57e3b847d362e from "igneum-pow-v5/src"; rule 19's fingerprint pairing catches both. The miner's side of a past seed cannot be re-read (igneum-miner reads ids from a node's template only), so the pin goes on the fingerprint equality and the record says the attempt-3 seed's miner id was not re-read; the seeded attempt-3 case (a miner `program-id` read on a given seed, era, day and stream, asked of the node lane by 15:00 UK, default a kaspa-pow test binary on class-v5-node) is the next harness item |
| 14:1x | The seeded attempt-3 case closed EQUAL: kaspa-pow `program-id` (class-v5-node f912e614, consensus/pow/src/bin: the template prepare's own path on a given seed, era, day, genesis day and stream file; a stream-less v5 read refuses with the engine's line, exit 3) reads the pin's FAIL seed ec0a8cf9...8212 (era 234e082d...6062, day 1244071, genesis day 1243740, the run's epoch-9 stream) as attempt 3, id 1f1cf82877f46ee6, equal to igneum-pow at exactly 1c420786 (build-1, 14:07 UK); the class v4 control reads attempt 0, 17d03dc8ddfe865c. The harness now reads that seed on every run (`class-v5-harness/attempt-3-seed/{seed.json,state.igsd1}`, check `attempt_3_seed_reads_equal`: the engine path beside the run's CLI, attempt and id against the record's expectation; an unreadable side is a FAILED CHECK, never a skip); the node lane lifts the read into `igneum-miner program-id` on its next line |
| 14:2x | The harness's attempt-3 check read live: on the rebuilt pair 6ccaf9e9 with the CLI at exactly 1c420786 (build-1, a 3-epoch no-flip case, 14:16 UK, `class-v5-harness/attempt3-check.{log,json}`): "ATTEMPT-3 SEED ec0a8cf9e1c6874f: cli attempt 3 id 1f1cf82877f46ee6 | engine attempt 3 id 1f1cf82877f46ee6 | expected attempt 3 id 1f1cf82877f46ee6 | EQUAL", `attempt_3_seed_reads_equal` true; the standalone path (`--a3-only`, seconds, no nodes) reads the same, and its known-failed flag (`--a3-known-failed`, the record's id with one digit flipped) reads DIFFER. A first short run died at 14:10 UK from this lane's own fault: the runner script was edited while that run's bash was reading it (the mid-run edit class build-remote's comment names); the second run is the record |
| 14:4x | The Arc B580 reads EQUAL on the rebuilt kit 65b47211 (14:46 UK): 82b19cbde8557ea5, self-test PASS 96 of 96, the v4 control PASS; the rotate fold was the whole Intel fault. The kit is one fingerprint on six platforms. The 0.3.26 upgrade line (8f481459's line with the Intel kit in) goes to the shipper and the coordinator |
## 1. The claim, in one paragraph
@ -290,7 +291,7 @@ Every item of this page that has no code, no test or no measurement yet, with th
| The kit for every platform: Metal | done (branch `v5-kits`, 7 October 2026, 19:3x UK): `proto-metal/main.swift` `servePackDataset` reads `leaves.bin`, checks the pack's count and FNV-1a 64, binds buffers 2 and 3 as `packbench.swift` does, self-tests the dataset head and last word against `vectors.json` and drops the leaf buffer after the build; a v5 prepare and job served on the M5 Max at 18:37:32Z | kits lane | 0 |
| The kit: CUDA | done (`v5-kits`): `packfile.h` reads `IGNEUM_STATE_LEAVES`, the FNV, file, root and block (generator 5 = class v5; a v5 pack without the count or another class with one is refused), `pf_load_leaves` reads and checks `leaves.bin`; `worker.cpp` uploads it (`cuMemcpyHtoD`) for `igneum_build(ds, cache, leaves, nLeaves, nItems)` and frees it after; the CPU emulation runs `--check` on v5-dn3-epoch0: PASS on igneum-build-1 at 18:40:52Z (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 lanes); the loader test carries the class v5 cases known-failed first (46 ok) | kits lane | 0 |
| The kit: OpenCL (AMD), Intel | host done (`v5-kits`): `proto-opencl/host.c` sets the build arguments for both kernel shapes in one place and uploads the leaves in `--bench-pack`, `--serve` and the compiled-in bench (the host reference derives words from the same leaves); the Linux and Windows workers are in the kit zip; the 9070 XT run placed on PC 1 as one lock slot after the 9070 XT G1 and ladder job (the hash lane publishes `tools/class-v5/pc1-amd-v5-bench.ps1`, expected 82b19cbde8557ea5); Intel: not measured tonight, no Arc B580 on PC 1 or PC 2 and the only Arc path needs a driver click, which no PC job may raise (the coordinator's ruling, 19:5x UK), so the card's holder and a click-free driver path are owed | coordinator (PC 1 queue), the B580's holder | 1 |
| Fingerprints equal across platforms, G1 on the 5090 and the Mac, AMD and Intel after | v5-dn3-epoch0 reads 82b19cbde8557ea5 on Metal (packbench, 18:34:47Z, 14.163 MH/s GPU time) and on Apple OpenCL (the kit host `--bench-pack`, 18:34:59Z, 15.265 MH/s wall), both under the Mac's measure lock, cache FNV and vectors PASS; CUDA: 82b19cbde8557ea5 on a one-shot RunPod RTX 4090 (driver 580.159.04, NVRTC 12.8, sm_89, 450 W) from the kit zip through `tools/class-v5/fleet-cuda-v5-bench.sh`: the kit's Linux NVRTC worker `--bench` at 18:57:45Z (check PASS, 62.759 MH/s over 5 dispatches of 2^24) and `bench.cu` at 18:57:50Z (the fleet lane's own run 62.775 MH/s); no 5090 was quiet (p1-5090 is a live voter), so the 5090 row stays the fleet's when one frees; AMD and Intel as the row above AMD READ (the hash lane's PC 1 queue job run-ca3-pc1-v5-amd-bench-20261007, 01:15 to 01:16 UK, exit 0, the card loaded beside the miners): RX 9070 XT gfx1201 (AMD-APP 3683.0, 32 CUs) on the kit zip packs-ca3-v5-20261007T221001Z (worker sha256 27faa253..., kernel_bound.cl 6b40f2dd..., leaves.bin 704 bytes 50687c8d...): self-test PASS (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 vector lanes), fingerprint of the 2^24 outputs at base 0 82b19cbde8557ea5 = the expected, match True; the v4-genesis control 892b6d55a7ddcfcb PASS. The rates (6.8 and 7.7 MH/s) are loaded-card figures, not the row. Intel: HELD out of 0.3.24, crossing time 09:36 UK (the Arc B580 job on the second PC, run-ca3-pc2-v5-intel-bench-20261007, read at 09:36 UK): no fingerprint; the kit worker (igneum-worker-opencl.exe sha256 27faa253...) fails its self-test on the Arc before any batch, on the v5 pack AND on the class v4 control alike, with every cache and dataset check passing (cache head, last line and FNV 7334fa46e5d972eb and 48c4f5bf24166b2e; dataset head, last word and 64 samples) and 96 of 96 vector lanes wrong (v5 lane 0 at base 0: device 729ebd46376e2851, expected e552166a03298f7f; v4 control lane 0: device 11bdacb6ee4108c2, expected dfbc8db1c06dacd8). So the fault is the bound kernel's per-nonce evaluation on Intel OpenCL (Intel(R) OpenCL Graphics, OpenCL 3.0 platform, driver 32.0.101.6733, the kernel compiled as OpenCL C 1.2, 160 CUs, sub-group shuffle present), on both classes, not class v5's leaves; the RX 9070 XT matched on the same kit. The kit stands on five of six platforms. CAUSE (the Intel lane, 10:4x UK): the kit worker (27faa253, built from class-v5 1095eaa8) lacks proto-opencl/intel_rotr.h, the Intel rotate-fold rewrite of 26e135a3 (Intel's compiler turns rotr_var's `rotate(x, (0u - n) & 31u)` into a left rotate, so every variable right-rotate of every program is wrong on Intel; host.c calls igneum_intel_rotr_patch before the build on an Intel platform); master at 3a4ba893 lacks it too, release-0.3.23 and 0.3.24 carry it, and the Intel lane lands it on the mirror's master (intel-rotr-master, about 10:30 UK). The installed 0.3.21 worker passed 96 of 96 on the same Arc on 7 October and mined 54 re-checked blocks. Not the sub-group size (that patch, drafted, is held). The fix for the kit: rebuild the worker from a tree with intel_rotr.h (master after the landing), the Arc job again through the shipper's second-PC queue; a no-op on NVIDIA and AMD, so the five matched fingerprints stand. The rebuilt kit (packs-ca3-v5-20261008T085619Z.zip, sha256 65b47211..., from 8f481459 with the fix) is with the hash lane for the Arc re-read; the second PC has been dark since 10:46 UK (a whole-PC fault, a hand at the box needed), so the read cannot land before 0.3.25's pin: 0.3.25 pairs with the freeze 1c420786 and ships that zip with the Intel worker held out; the Intel kit and 8f481459 ride the first cut after an equal Arc read, 0.3.26 at the earliest. | coordinator (AMD, Intel), fleet (5090) | 1 |
| Fingerprints equal across platforms, G1 on the 5090 and the Mac, AMD and Intel after | v5-dn3-epoch0 reads 82b19cbde8557ea5 on Metal (packbench, 18:34:47Z, 14.163 MH/s GPU time) and on Apple OpenCL (the kit host `--bench-pack`, 18:34:59Z, 15.265 MH/s wall), both under the Mac's measure lock, cache FNV and vectors PASS; CUDA: 82b19cbde8557ea5 on a one-shot RunPod RTX 4090 (driver 580.159.04, NVRTC 12.8, sm_89, 450 W) from the kit zip through `tools/class-v5/fleet-cuda-v5-bench.sh`: the kit's Linux NVRTC worker `--bench` at 18:57:45Z (check PASS, 62.759 MH/s over 5 dispatches of 2^24) and `bench.cu` at 18:57:50Z (the fleet lane's own run 62.775 MH/s); no 5090 was quiet (p1-5090 is a live voter), so the 5090 row stays the fleet's when one frees; AMD and Intel as the row above AMD READ (the hash lane's PC 1 queue job run-ca3-pc1-v5-amd-bench-20261007, 01:15 to 01:16 UK, exit 0, the card loaded beside the miners): RX 9070 XT gfx1201 (AMD-APP 3683.0, 32 CUs) on the kit zip packs-ca3-v5-20261007T221001Z (worker sha256 27faa253..., kernel_bound.cl 6b40f2dd..., leaves.bin 704 bytes 50687c8d...): self-test PASS (cache FNV 7334fa46e5d972eb, dataset head, word [268435455], 64 samples, 96 of 96 vector lanes), fingerprint of the 2^24 outputs at base 0 82b19cbde8557ea5 = the expected, match True; the v4-genesis control 892b6d55a7ddcfcb PASS. The rates (6.8 and 7.7 MH/s) are loaded-card figures, not the row. Intel: READ EQUAL on the rebuilt kit (65b47211, the Intel rotate-fold fix in): the Arc B580 on the second PC through the hash lane's queue (run-ca3-pc2-v5-intel-bench-20261008, 14:45 to 14:46 UK, the job keeping the host's whole stdout): v5-dn3-epoch0 fingerprint 82b19cbde8557ea5 = expected, match True, self-test PASS 96 of 96 lanes, 10.794 MH/s quiet; the v4-genesis control 892b6d55a7ddcfcb PASS, 10.718 MH/s; the host's lines: "intel: rotr_var rewritten to the shift form before the build", exchange sub_group_shuffle_xor (cl_khr_subgroup_shuffle) with sub-group size 32 for a 32-item work-group (queried through clGetKernelSubGroupInfoKHR), the kernel's compiled sub-group size 32, build 924.5 ms (v5) and 518.7 ms (v4); device Intel(R) OpenCL Graphics (OpenCL 3.0), driver 32.0.101.6733, 160 compute units. So the rotate fold was the whole fault (the sub-group reading was not needed; that patch stays unapplied) and the kit reads one fingerprint on six platforms: CUDA (RTX 4090), Metal and Apple OpenCL (M5 Max), AMD OpenCL (RX 9070 XT), Intel OpenCL (Arc B580), the CPU verifier. The 09:36 UK read before the fix, for the record: Intel: HELD out of 0.3.24, crossing time 09:36 UK (the Arc B580 job on the second PC, run-ca3-pc2-v5-intel-bench-20261007, read at 09:36 UK): no fingerprint; the kit worker (igneum-worker-opencl.exe sha256 27faa253...) fails its self-test on the Arc before any batch, on the v5 pack AND on the class v4 control alike, with every cache and dataset check passing (cache head, last line and FNV 7334fa46e5d972eb and 48c4f5bf24166b2e; dataset head, last word and 64 samples) and 96 of 96 vector lanes wrong (v5 lane 0 at base 0: device 729ebd46376e2851, expected e552166a03298f7f; v4 control lane 0: device 11bdacb6ee4108c2, expected dfbc8db1c06dacd8). So the fault is the bound kernel's per-nonce evaluation on Intel OpenCL (Intel(R) OpenCL Graphics, OpenCL 3.0 platform, driver 32.0.101.6733, the kernel compiled as OpenCL C 1.2, 160 CUs, sub-group shuffle present), on both classes, not class v5's leaves; the RX 9070 XT matched on the same kit. The kit stands on five of six platforms. CAUSE (the Intel lane, 10:4x UK): the kit worker (27faa253, built from class-v5 1095eaa8) lacks proto-opencl/intel_rotr.h, the Intel rotate-fold rewrite of 26e135a3 (Intel's compiler turns rotr_var's `rotate(x, (0u - n) & 31u)` into a left rotate, so every variable right-rotate of every program is wrong on Intel; host.c calls igneum_intel_rotr_patch before the build on an Intel platform); master at 3a4ba893 lacks it too, release-0.3.23 and 0.3.24 carry it, and the Intel lane lands it on the mirror's master (intel-rotr-master, about 10:30 UK). The installed 0.3.21 worker passed 96 of 96 on the same Arc on 7 October and mined 54 re-checked blocks. Not the sub-group size (that patch, drafted, is held). The fix for the kit: rebuild the worker from a tree with intel_rotr.h (master after the landing), the Arc job again through the shipper's second-PC queue; a no-op on NVIDIA and AMD, so the five matched fingerprints stand. The rebuilt kit (packs-ca3-v5-20261008T085619Z.zip, sha256 65b47211..., from 8f481459 with the fix) is with the hash lane for the Arc re-read; the second PC has been dark since 10:46 UK (a whole-PC fault, a hand at the box needed), so the read cannot land before 0.3.25's pin: 0.3.25 pairs with the freeze 1c420786 and ships that zip with the Intel worker held out; the Intel kit and 8f481459 ride the first cut after an equal Arc read, 0.3.26 at the earliest. | coordinator (AMD, Intel), fleet (5090) | 1 |
| The kit zip | `igneum-build-1:/srv/artefacts/packs/packs-ca3-v5-20261007T183921Z.zip`, sha256 e6c088bb34fecdc3ff297dbb06438a14ade7d8c55273357726d28f7a1334a25e, 919,129 bytes, 56 files (v4-genesis, v5-genesis, v5-dn3-epoch0 with `leaves.bin`; `bin/linux` and `bin/windows` NVRTC and OpenCL workers; sources; the two bench scripts; SHA256SUMS), built by `tools/class-v5/kits-remote.sh` on build-1 at 18:41:15Z after the loader test and the emulation check | kits lane | 0 |
| The crossing on Devnet 3 by height after its gate | NOT started: the floor in the Devnet 3 override, the Devnet 2 style gate (zero rejected across the flip, exec roots agreeing, every node holding the epoch's stream); the harness's flip case is the rehearsal (PASS twice) | node lane (a283f5f0d364ceef0) owns the fork, the shipper (ae892a8b0f78fe31c) the cut; this lane the gate's v5 checks | 3 |
| Miner cost rows per card (5090, M5 Max, 9070 XT, 4070) | the 4090 row only (rate equal within 0.01 percent, build +0.24 ms); the four named cards owed, each labelled measured with the date | fleet (5090, 4070), this lane (M5 Max under the lock), PC 1 (9070 XT) | 3 |