diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 90af33b10..5745224e7 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -77,6 +77,8 @@ jobs: run: bash tools/ci/second-engine-check.sh - name: no playbook quits, pauses or resumes the installed app (self-test first, then the tree) run: bash tools/ci/playbook-quit-check.sh --self-test && bash tools/ci/playbook-quit-check.sh + - name: no script writes into another worktree or walks Projects (tools/ci/no-foreign-tree-writes.sh) + run: bash tools/ci/no-foreign-tree-writes.sh --self-test && bash tools/ci/no-foreign-tree-writes.sh - name: the signer is never piped into head run: bash tools/ci/signer-pipe-check.sh - name: bash bodies in PowerShell job scripts pass bash -n, the lost-quote class (self-test first, then the tree) diff --git a/site/bench.html b/site/bench.html index 61d393be1..7a3d050cc 100644 --- a/site/bench.html +++ b/site/bench.html @@ -172,12 +172,12 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
-
83 entries, newest at the bottom
+
88 entries, newest at the bottom

Engineering log

Every measurement the project has made, newest at the bottom, written by the people and agents who ran it, with the commands and hardware. Prototype numbers are not mining numbers and say so.

- +

Igneum bench log

Append-only. Every number here was measured on the machine named, on the date given.

2026-10-03 proto-metal / igneum-bench, first run

@@ -577,6 +577,9 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:

Dry run 3, 6 October 2026, 14:56 to 15:02Z (job ember-dryrun-pc1-3, unelevated, no prompt, measure only; engine from ember-tune 07d5a72, kit sha256 36b522c9...): the first measurement engine on PC 1 that mined. Both cards, one 60 s row each at the installed app's 80% cap, clocks unlocked, rate = the worker's STATUS wall rate, draw = nvidia-smi every 5 s:

CardMH/sWMH/WcorememoryGPU Climit
RTX 5090127.31316.50.4022,850 MHz13,801 MHz68460 W of 575
RTX 407028.68102.70.2792,805 MHz10,251 MHz46160 W of 200

Nothing set; the installed app's miners back after 350 s. Why every earlier run (5 and 6 October, runs 1 to 4 and dry runs 1 and 2) read its copied settings as defaults, measured on PC 1 (collect ember-acl-2): the engine's own start locks its app folder with icacls /inheritance:r /grant:r <user>:F; cutting the folder's inheritance propagates down, the non-inheritable grant gives the children nothing, so a file COPIED in before the start (settings.json, machine-id, wallet.json) is left with no access entry and its owner cannot read it (ReadAllText: access denied), while the engine's own files written after the lock inherit fine, which hid it for a day. A first fix with (OI)(CI)F /T left the file empty too: /T re-applies /inheritance:r to each file after the propagation and an (OI)(CI) entry on a file is inherit-only. The right form is the inheritable grant without /T (07d5a72). Consequence for every tier on Windows: nothing changes for the installed app (its files were always its own); any tool that drops files into the app folder before the app starts (an installer's seed, a migration, a support script) was unreadable to the app until now and is readable from 0.3.13 on.

+

Run 5, 6 October 2026, 15:28 to 15:39Z (job ember-tune-pc1-5, elevated on the maintainers' click, PC 1 on 0.3.13, the tune engine = kit ember-kit-5 from 07d5a72, mode=direct): the first run that set limits. The 5090's power ladder, 75 s a step, the clock unlocked (2,850 MHz core, 13,801 MHz memory), the rate = the worker's STATUS wall rate, the draw = nvidia-smi every 5 s:

+
CapLimitMH/sWMH/WGPU C
100%575 W127.38309.90.41164
90%518 W99.32313.60.31765
80%460 W123.11312.20.39465
70%403 W127.38310.90.41065
60% (floor 400 W)400 W127.38311.30.40965
+

Reading: the cap does not bind on this hash (310 to 314 W under every limit, as the 4 October stability line said), so the power knob is flat at 0.41 MH/W on the 5090 and the saving must come from the clocks; the 90% and 80% rows' rate dips at the same draw are stalls inside those holds (a worker restart or a template wait), not the cap. The clock ladder's first step (2,781 MHz at 575 W) was requested at 15:37:47Z and never measured: the playbook's own watchdog killed the live engine at 15:39:03Z (all three cards mining at 177 MH/s) because its idle clause sampled one log line and read "idle" from a missing match; the 4070's and the 9070 XT's plans never ran. Consequences: the 5090 is probably left with its core clock locked at 2,781 MHz (an -lgc lock persists until -rgc or a reboot) under the 575 W cap, which costs little rate; freeing it needs administrator rights; and the Power Helper was NOT registered by this run (the registration lived only in the installed app's cap path, which a --sweep engine skips). Fixed the same hour: the watchdog's idle rule (three consecutive status lines reading 0.00 MH/s and 300 s, never a missing match), the after snapshot and the engine-log dump on every exit, and an elevated tune engine registering the task itself before its first step. The decided way out: the maintainers switches Power control ON in the 0.3.13 app (its one prompt registers the task from the install folder), a job frees the clock through the task (rgc), the tune runs unelevated through the task.

Consequence for the tiers: an AMD card is tuned on its power limit alone until its stock core clock is read (a 9070 XT at -30% is the floor the driver allows, 4 steps, 5 minutes); every NVIDIA card's two-knob plan waits on the user's one click on Power control; the re-run on PC 1 is held until the quit's source is named (the event-log collect) and follows the 0.3.11 rollout (the update clears the jobs folder, so the engine and the helper are fetched again), with the scheduler's slot.

5 October 2026 (night), read width of the lottery hash: 4, 16 and 64-byte loads, a per-load mix, a written scratch; three cards (gate 1 experiment, cryptographer)

Branch readwidth (commits 019b014, b970dda, 4badcee, a9e002c, d0018cf and the entry commit); plan and recommendation in docs/plans/read-width.md. Nothing here changes consensus: every class sits behind igneum-pow --class and the default class is generator version 2 byte for byte (igneum-pow/tests/packs.rs passes on the four pinned packs after every commit). Question (the maintainers, after "the 9070 XT on the eGPU" above): would wider reads keep the latency-bound random-access property while closing the vendor gap. Additions from the coordinator: a per-load width drawn from an era-fixed mix, and a written per-warp scratch (measurement only, no soundness claim).

@@ -694,6 +697,53 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:
ProgramCompile ms (library + pipeline)Mhash/s (GPU)Verify
epoch018.8 (9.3 + 9.5)27.2PASS
epoch117.8 (8.8 + 9.1)27.9PASS
epoch217.6 (8.6 + 9.1)27.8PASS
epoch316.1 (8.0 + 8.2)27.7PASS
epoch418.6 (9.0 + 9.6)27.5PASS
epoch515.9 (7.7 + 8.2)28.4PASS
epoch617.6 (8.6 + 9.1)27.8PASS
epoch717.6 (8.7 + 9.0)30.1PASS
epoch820.4 (9.8 + 10.6)28.4PASS
epoch918.2 (8.7 + 9.5)29.3PASS
min / median / max15.9 / 17.7 / 20.410 of 10

Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256 (the pack's two libraries, memhard.metal and program.metal): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.

Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Apple M5 Max's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: host.c times clBuildProgram only in the prepare path and no prepared line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Apple M5 Max) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.

+

6 October 2026, Counter ASIC 3.0 item 2: the per-day derivation

+

Branch ca3-derive (worker "derive", from ca3-coord 50df751; commits acb96ee and after), design, spec text and the chip row in docs/plans/counter-asic-3-derivation.md and docs/analysis/chip-model-v3.md section 6. Question (the plan's item 2): replace the fixed-shape mixer (the chip model's 3x fixed-function allowance, 0.31x to 0.92x) with a random item-derivation program drawn per day from the day key stream (RandomX's SuperscalarHash idea, superscalar.cpp read at upstream 7607fb2), keep the 8 dependent cache reads per item exactly, keep the op count per item at or above x8's, and measure the verifier against the 10 ms gate, bit-exactness, the daily build and the hash rate. Machine: Apple M5 Max, 64 GiB, Darwin 25.6.0; every timing row names its lock and load average.

+

The construction (class dr736, igneum-pow/src/derive.rs): nine straight-line programs of 736 instructions per item (one before each cache read, one after the last), four draws per instruction from the mixer's own SplitMix64 stream after its 40 draws, twelve two-register forms (add, sub, xor, mul-lo by c|1, rotate-add, xor-rotate, add-constant, xor-constant, the M_r form (d ^ c) * odd, d * odd + c, d ^= c & b, d += c | b), every instruction reading the register the previous one wrote (the chain, s[0] first) and writing another, every form a bijection on the state; the acceptance test rejects a register never written, fewer than 8 distinct rotations, or a draw under the x8 mixer's counts from the code (72 x 128 = 9,216 chip ops, 72 x 144 = 10,368 as written, 1,152 multiplies; the coordinator's correction of the 130-per-application figure). The genesis day draws 6,624 instructions, 10,659 GPU ops, 9,992 chip ops, 1,461 multiplies per item; the verifier runs it with a word-major interpreter over the 32 items of a load, dispatching on instruction pairs, no JIT.

+

Verifier per 32-lane unit, one core (with-lock.sh measure, one session 07:42:20 to 07:42:33 UTC, load average 4.91 / 4.53 / 5.34 at the start, 4.46 / 4.44 / 5.30 at the end; igneum-pow bench --seed igneum-genesis --day 2026-10-03 --class <c> --warps 50, two rounds, then the devnet seeds once):

+
Classms per unit, avg of 50 (round 1 / 2)Worst cold of threeAgainst x8Ops per item (GPU / chip / mul)
v20.598 / 0.5940.6971,296 / 1,152 / 144
x8 (mx8, class v3)2.061 / 2.0632.179110,368 / 9,216 / 1,152
dr7364.875 / 4.9445.2412.37x10,659 / 9,992 / 1,461
x8, devnet seeds2.0782.155
dr736, devnet seeds4.8725.2412.34x10,701 / 10,083 / 1,362
dr368 (half length, the x4-equivalent fallback)2.6922.8981.31x5,350 / 5,004 / 752
+

The v2 row reads the quiet nights' 0.60 (readwidth 0.604 to 0.626; 6.4a 0.607 to 0.611), so these are quiet-core figures. The interpreter's cost split (examples/derive_perf.rs, a functional run under the run lock, load 3.9 to 4.9): 10.7 ns per instruction per 32-item batch cold, 7.18 with one dispatch per instruction, 4.98 with pair dispatch; a uniform program (predictable dispatch) 3.4 to 4.3 ns, so about 1.5 ns is dispatch and 3.5 ns the vector body (NEON, 1,180 .4s instructions in the binary).

+

Bit-exactness (with-lock.sh run): dr736-genesis on Metal (packbench --batches 1 --batch-log2 24) cache FNV 48c4f5bf24166b2e PASS, dataset head and word [MASK] PASS, vectors 3/3 standalone and 3/3 in batch, fingerprint 2^24 50e3eaa779da4f1e, compile 784 ms cold; on Apple OpenCL (igneum-bench-cl-dr736-genesis --bench-pack) the self-test PASS with the 64 samples and 96 of 96 lanes, fingerprint 50e3eaa779da4f1e (equal); dr736-devnet-epoch0 on Metal PASS, fingerprint 9553f6d5c667205a. Two compilers agree with the Rust interpreter on the derived dataset and on 2^24 outputs.

+

Daily 1 GiB build and hash rate, Metal (with-lock.sh measure, the same session, packbench --batches 2 --batch-log2 22 --group 256, three rounds):

+
PackCompile (1 / 2 / 3)Build, GPU ms (1 / 2 / 3)MH/s GPU (1 / 2 / 3)
mx8-genesis (x8, the control)80 / 1 / 1 ms31.3 / 22.1 / 22.127.155 / 27.076 / 27.123
dr736-genesis751 / 1 / 1 ms28.9 / 29.0 / 29.127.125 / 27.129 / 27.063
+

Chip model (chip-model-v3.md section 6): 1,278,976 chip ops per hash on the genesis day, 39.1 MH/s at 50 T op/s, 0.29x bare (0.31x at the floor, x8's figure); the fixed-function allowance of the wired mixer (3x) no longer applies to a chip that must run the day's program: at ProgPoW's claimed 1.2x the row reads 0.34x, at a cautious 1.5x 0.43x, at the old 3x 0.86x; equal silicon 0.29x / 0.36x. dr368: 0.57x bare, 0.69x / 0.86x.

+

Consequences per tier. The verifier: no miner tier runs it; a node on any 2026 core verifies a block in 5 ms (x8: 2.1), a pool core serves 205 shares per second (x8: 485; a 22,000-member pool at one share per 10 s needs 11 cores against 4.5), IBD over 108,000 headers is 8.8 min on one core (x8: 3.7); on a 2019-class laptop core (2.5x, approximate, O-1.14 unmeasured) 736 reads about 12 ms, over the gate, and 368 about 6.7 ms, under it. The build: the Apple M5 Max pays 7 ms more per day (29 against 22 ms), nothing to any tier; the 5090 is the RTX 5090 machine 2 job below; the 9070 XT is OWED (PC 1 is the maintainers' desk today; its x8 build was 72 to 77 ms); the integrated gfx1036 tier already misses the per-prepare rule at x8 (epoch-length.md 6.1: 6.9 / 9.4 / 11.7 s prepares at x1, about 55 to 94 s at x8, approximate) and the day program leaves that need (per-day dataset reuse in the workers, 0.3.12) the same in kind. The compile: the Metal item library is 0.75 to 0.8 s cold once a day and 1 ms from the shader cache; the CUDA worker compiles memhard.h into every per-epoch kernel and every race variant, so the 5090's nvrtc line is the number to read. The hash rate: unchanged within 0.3% on the Apple M5 Max, as the hash kernel only loads. Packs grow by about 550 KB (memhard.h 196 KB, program.json 156 KB): nothing to any tier.

+

Go / no-go: GO as reserve entry R0 (the PROPOSED text in the derivation document's section 6, not in docs/spec); NO-GO for genesis-live at 736 instructions until the 2019-class core measurement lands under 10 ms; the number that decides it is 4.88 ms per unit on one M5 Max core (pass) against about 12 ms on the approximate laptop row (fail); dr368 passes both rows at 2.69 ms with the chip at 0.57x bare.

+

RTX 5090 (PC 2, one job run-ca3-derive-pc2-20261006, relay/playbooks/ca3-derive-pc2.ps1, published 08:26:08Z after the proving agent's clear at 08:24:27Z, lock 08:25:55 to 08:29:08Z; ran 08:26:43 to 08:28:49Z, done exit 0): the installed 0.3.11 worker through NVRTC 12.8 on the self-fetched zip. The card did NOT come off: the job read the key from settings.json (nvidia:0:NVIDIA GeForce RTX 5090, with the device index) where the 5 October jobs posted the state's key without it, and one worker process stayed up through the 90 s wait, so every row is a loaded-card figure (the v2 control 62.3 MH/s against its unloaded 136 to 137) with the ratios valid.

+
PackNVRTCCache1 GiB buildSelf-test (64 samples, 96 lanes)Fingerprint 2^24MH/s bw1 / bw8 (loaded)
v2-genesis-mh167 ms646 msPASS25f96e7dce90bd4e = Mac62.26 / 61.34
mx8-genesis164 ms440 msPASS7c28cfb06c5c65a9 = Mac61.98 / 60.53
dr736-genesis1,266 ms542 msPASS50e3eaa779da4f1e = Metal = Apple OpenCL61.08 / 58.51
dr736-devnet-epoch01,266 ms632 msPASS9553f6d5c667205a = Metal62.15 / 61.44
+

Reading: bit-exact on CUDA (four compilers now agree on both packs); the build and the rate do not move beyond the loaded noise; the number that moved is the NVRTC compile, +1.1 s per pack, because memhard.h's 6,624-statement item function is inside every hash-kernel and race-variant compile (17 variants: about +19 s per epoch, approximate, against a 38 s compile-ahead budget at the 600-s epoch floor), so compiling the item function once a day into its own module is a requirement of the class. Consequences: a 5090 owner pays 1.1 s once a day after that fix and 1.1 s per variant per epoch before it; the Apple M5 Max 0.75 s once a day. Filed: the card-off key form (post both forms, confirm by the process list) before the next PC 2 round; the unloaded 5090 rows re-run then. RX 9070 XT: OWED (PC 1).

+

6 October 2026, Counter ASIC 3.0 item 8: program work in the latency shadow

+

Branch ca3-shadow, worker "shadow" (docs/analysis/latency-shadow-2026-10-06.md carries the design, the chip side and the consequences; this entry carries the measurements). The knob: LoadClass::shadow, class name <class>+sh<S>x<R>, a block of S ALU instructions drawn from the program stream after the 64 base instructions and run R times at the end of every iteration (no load; the base program, its attempt and the acceptance verdict are the class's without the shadow; v2 and v3 byte-identical, cargo test in igneum-pow 54 + 4 + 19 + 7 green). Packs proto-cuda/packs-ca3-shadow/* over mx8 for seed igneum-genesis; the control is the pinned class v3 pack packs-ca2-mixer/mx8-genesis. Ops per hash = shadow instructions x 1.83 (counted from the emitted statements: add 5, rotr 2, shfl 2, the rest 1, weighted over the non-load weights) + 930 (the base program's 384 ALU instructions and 128 loads).

+

Apple M5 Max, Metal (proto-metal/packbench --pack <dir> --batches 60 --batch-log2 24 --group 256 under with-lock.sh measure, two sessions 07:51 to 08:03 UTC, GPU time; power = the IOReport "Energy Model" GPU and DRAM channels at 2 Hz through a dlopen of libIOReport, no root, the mean over each run after its first 6 s; idle GPU 0.44 W, DRAM 0.64 W; the package is about 17 W more, Ember Tune's 38 W row, approximate; load averages 3.3 to 6.8, the GPU idle: the Apple M5 Max mines nothing):

+
PackShadow instrs per hashOps per hashMH/sAgainst the controlGPU WDRAM WMicrojoules per hash (GPU + DRAM)Verifier ms per warp, one core, avg of 20 (worst cold)Bit-exact, fingerprint 2^24Load at start
mx8-genesis (control; runs 1, 2, 3)093027.07, 27.07, 27.1011.2, 9.2, 12.310.2, 10.0, 10.20.782.062 (2.188)yes, 7c28cfb06c5c65a94.6, 6.8, 4.2
sh256x24,0968,40026.85-0.8%15.110.40.952.080 (2.249)yes, 33e8bbe4c35b2e544.6
sh256x714,33627,20026.74-1.3%18.010.61.072.112 (2.224)yes, 6cfb70911007520a4.2
sh256x1326,62449,70026.75-1.2%20.610.51.162.212 (2.237)yes, 59ac286fe2a5a9ef3.7
sh64x52 (64-instruction block; runs 1, 2)26,62449,70027.72, 27.79+2.5%19.8, 20.310.2, 10.31.092.137 (2.237)yes, 9dd010f79d8ca9f44.4, 4.1
sh1024x3 (1,024-instruction block)24,57645,90022.53-16.8%20.29.71.332.150 (2.274)yes, a05399c819b79aad6.0
sh256x27 (runs 1, 2)55,296102,10026.48, 26.86-1.5%26.9, 26.910.4, 10.21.402.229 (2.276)yes, 3d2e8245cc084d073.3, 5.6
sh256x4081,920150,80025.10-7.3%26.710.11.462.329 (2.396)yes, 0844b706302f1c9c5.0
sh256x53 (runs 1, 2)108,544199,60024.40, 24.09-10.4%28.8, 28.59.5, 9.71.582.427 (2.461)yes, 4f824b15cf2b124a3.9, 4.6
sh256x88180,224330,70021.39-21.0%31.78.41.882.619 (2.771)yes, 0572522e39a94d8a3.8
+

Reading: the M5 Max stays latency-bound to about 100,000 ops per hash and its 5 percent point is about 130,000 (between the 102,100 and 150,800 rungs), 2.2x under the chip model's 290,000 (from memory); the block size matters on Apple (64 instructions +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N: the instruction footprint); the GPU rises from 11 to 27 W at 100,000 ops (marginal 6.9 pJ per counted op; 2.9 pJ at the compute-bound end) and the energy per hash from 0.78 to 1.40 microjoules, GPU plus DRAM. The verifier's law on this core: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp (0.1 ns per lane-instruction), so 330,700 ops cost 0.56 ms here and about 1.4 ms on a 2019-class core by the 2.5x rule: inside every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none, M5 Max / 2019-class, item 2's figures), so the node never binds the shadow before the cards do. The control is 2.5 percent under the 5 October figure for this pack (27.7, mixer-x4.md 6.2) on three runs; every row is read against today's 27.08.

+

RTX 5090 (PC 2, 1ccfe586), CUDA through NVRTC (job run-ca3-shadow-pc2-20261006, 08:30:26Z to 08:38:57Z, 511 s, exit 0, --stop-miners, the prover off for the run and back on, the card EMPTY before the ladder by nvidia-smi's compute-apps list; the installed worker 0.3.11 --bench --batches 250 --batch-log2 24 --block-warps 1, wall time; power = nvidia-smi -l 1 means over each bench's window; the card under the app's 431 W limit, 73.8 W idle, 56 to 71 C):

+
PackShadow instrs per hashOps per hashMH/sAgainst the controlWatts, meanSM MHzMicrojoules per hashBit-exact, fingerprint 2^24 = the Apple M5 Max'sNVRTC ms
mx8-genesis (control, first and last)0930131.94, 132.47342.4, 357.43,052, 3,0372.65yes, 7c28cfb06c5c65a9158, 160
sh256x24,0968,400132.34+0.1%354.43,0372.68yes248
sh256x714,33627,200132.32+0.1%385.33,0372.91yes250
sh256x1326,62449,700132.28+0.1%424.63,0343.21yes240
sh64x52 (64-instruction block)26,62449,700136.77+3.5%431.5 (the cap)3,0243.15yes181
sh1024x3 (1,024-instruction block)24,57645,900132.02-0.1%428.33,0303.24yes510
sh256x2755,296102,100131.95-0.2%431.0 (the cap)2,8243.27yes242
sh256x4081,920150,800131.75-0.3%431.02,4273.27yes242
sh256x53108,544199,600128.67-2.7%431.01,7533.35yes241
sh256x88180,224330,70086.39-34.7%431.01,8344.99yes241
+

Reading: the 5090 holds to 150,800 ops (-0.3 percent) and loses 2.7 percent at 199,600, under a 431 W cap that the control never reaches (342 to 357 W) and that binds from 102,100 ops up: the SM clock falls from 3,037 to 1,834 MHz and at 330,700 ops the card is compute-bound at the capped clock (28.6 T counted op/s, the 45.2 T budget scaled by the clock). The 5 percent point at 431 W is about 210,000 ops. The marginal ALU energy at the shipping clock, read on the three rungs under the cap: 10.2 to 13.2 pJ per counted op, twice the 5.5 pJ the chip model assumed. The 64-instruction block runs 3.5 percent above the control here too. Clock rows (-lgc): OWED, nvidia-smi refused the lock without administrator rights and the job did not ask for them. Power-cap rows: the 5 October sweep's (floor 400 W, so -pl 200 and 250 cannot be set; the cap never binds at the control).

+

RX 9070 XT (PC 1, ae432dc7): OWED (PC 1 is the maintainers' desk and not released today); the OpenCL kernels are in every pack.

+

Consequences per tier (the file's section 8 in short): at the recommended N = 100,000 ops per hash (sh256x27) the Apple card loses 1.5 percent of its rate and pays 16 W more (income per watt 0.56x, per pound unchanged), the 5090 holds its rate at its 431 W cap (350 W at the control: per watt 0.81x, measured), the 9070 XT holds by its budget (owed), a rig pays about 30 percent more electricity for the same hash, a pool user sees nothing, and the f = 1 chip's edge per joule falls from 1.6x to 0.9x against the M5 Max and from 5.6x to 2.1x against the 5090 at k = 1, where k is the chip core's energy per op over the 5090's measured 11 pJ: the number that decides the item. Verdict: GO as a class v4 candidate at N = 100,000 (mx8+sh256x27), subject to the 9070 XT row and the gates; NO-GO above 130,000 or with a block over 256 instructions. No card we own may lose more than 5 percent (the 2.0 rule): the M5 Max caps N at 130,000.

+

6 October 2026, Counter ASIC 3.0 item 6: the reserve families' step costs

+

Branch ca3-reserve, worker "reserve" (docs/plans/counter-asic-3-reserve.md carries the proposed order and spec text; this entry carries the measurements). Method: the dot4 probe's dependent chain (docs/analysis/int8-matrix-family.md section 4), one op of the family per step per lane, 1,048,576 lanes x 4,096 steps, best of 3 dispatches per run, three runs, bit-exact against a CPU reference on two whole 32-lane warps (the shuffle rows need the whole warp). Every chain has the same glue (acc = OP(acc, x, y); x = x * K + acc; y = rotl(y, 7) ^ (acc + s)), so the "step cost" is the family's one op plus four glue ops against the add-xor-rotate chain of the 9070 XT bench-log entry (alu: x = x * K + rotl(y, 7); y = (y ^ x) + s, 5 ops per step counted, no acc). Reference rows are live families (alu; rotr = the live rotr_var text; shflx = the live shfl, lane XOR 8). Candidate rows are the seven families of spec 1.13.2 (shl, shr, bfe with the vendor's extract function and bfec the C form (y >> 7) & 0x1fff, andn, perm = bytes (b1, b3, b0, b2), popc and clz folded by add, sel on bit 5, shfla = lane + 3 mod 32). Comparison rows: dot4u and dot4s (Apple, emulated), dot4i (__dp4a) and mm8 (one mma.sync.m8n8k16 u8 per step per warp, inline PTX) on CUDA. Sources: proto-metal/family-probe.swift, proto-cuda/family-probe.cu, the RTX 5090 machine 2 job tools/ca3-reserve/pc2-family-probe.ps1 (made by make-pc2-playbook.sh). G steps/s is the whole card's dependent-step throughput; ops per step counted = the family's op plus 4 glue (alu 5, shfla and shflx 6: shuffle plus xor, mm8 1 mma plus 4).

+

Apple M5 Max, Metal (swiftc -O -o family-probe family-probe.swift -framework Metal under with-lock.sh build; three runs of with-lock.sh measure ./family-probe --reps 3, 07:29:05 to 07:29:08 UTC, load average 7.64 / 7.59 / 7.14 before and after every run (the Apple M5 Max was loaded by other agents' builds the whole morning; the measure lock held, the GPU idle: the Apple M5 Max mines nothing), GPU start-to-end time):

+
kernelbest ms, runs 1 / 2 / 3best of the three, msG steps/s (best)ns per step (best)ops per step countedstep cost (ratio to alu, best)bit-exact, 3 runs
alu5.039 / 4.918 / 4.8734.8738811,19051.00yes
rotr (live)5.610 / 5.533 / 5.4925.4927821,34151.13yes
shflx (live)4.196 / 4.259 / 4.2234.1961,0241,02460.86yes
shl4.120 / 4.202 / 4.1194.1191,0431,00650.85yes
shr4.296 / 4.194 / 4.3194.1941,0241,02450.86yes
bfe (extract_bits)3.731 / 3.805 / 3.8163.7311,15191150.77yes
bfec (C form)3.762 / 3.818 / 3.6463.6461,17889050.75yes
andn3.680 / 3.732 / 3.6763.6761,16889850.75yes
perm5.500 / 5.548 / 5.5245.5007811,34351.13yes
popc4.261 / 4.223 / 4.2624.2231,0171,03150.87yes
clz4.918 / 4.916 / 4.9144.9148741,20051.01yes
sel3.718 / 3.831 / 3.7183.7181,15590850.76yes
shfla (lane + 3)9.337 / 9.323 / 9.2879.2874622,26761.91yes
dot4u (emulated)7.818 / 8.044 / 7.9257.8185491,90951.60yes
dot4s (emulated)23.264 / 23.254 / 23.04523.0451865,62654.73yes
+

Reading of the Apple M5 Max rows. The run-to-run spread is under 4% on every row. The dot4 rows reproduce the 5 October figures (1.6x unsigned, 4.7x signed), which is the check on the method. A step cost under 1.00 means the family's op plus the glue is cheaper than the five-op reference chain: the reference's two registers are a tighter dependency than the three-register candidate chains, and Apple's shifts, extract, andn and select each cost about what an xor costs. Three rows cost more than the reference: perm (1.13: no byte-permute function in MSL; the uchar4 swizzle compiles to shifts and masks, so a byte permute is emulated on Apple at about the price of the live rotr), clz (1.01) and shfla (1.91: a shuffle by a computed lane index costs 2.2x the live xor shuffle on Apple, simd_shuffle against simd_shuffle_xor; the second shuffle form is the one candidate Apple pays for). mm8 as a chain on Apple is owed (Metal 4 matmul2d; this toolchain is Swift 5.8 without the tensor API).

+

RTX 5090 (PC 2, 1ccfe586), CUDA (job run-ca3-family-pc2-20261006, a signed run job with --stop-miners, published 08:41:45Z after /tmp/igneum-devnet/pc2-ca3.clear (08:24:27Z) under the mkdir lock (taken 08:41:26Z, released 08:43:24Z the moment the closing report was read); ran 08:42:27Z to 08:43:03Z, done, exit 0, 36 s; node tools/jobs.mjs run-ca3-family-pc2-20261006 --all). The card to itself: the app had stopped the miner before the script started (workers_before: no CUDA compute app, the card at 847 MHz SM and 72 W), the script posted the card off through api/cards in both key forms (settings.json carries two NVIDIA keys, nvidia:0:NVIDIA GeForce RTX 5090 with 8 identities and the older nvidia:NVIDIA GeForce RTX 5090 with 2) and read the card quiet by nvidia-smi's compute-apps list and the process list after 30 s; prover off at 08:42:27Z and back on at 08:43:02Z ({"ok":true}, in the finally block); the cards restored to their settings. nvcc 12.8 in WSL2, -arch=sm_120, the source sha256 1a3d07b8...0cf90d equal on the Apple M5 Max, the Windows side and inside WSL. Three runs of ./family-probe --reps 3, CUDA event time; the SM clock ramped from 862 MHz at run 1 to 2,572 MHz at run 3 (gpu_before per run; 129 W at the end), so the best of the three runs is the card's warm figure and the table carries it; the run-to-run spread of the best values is under 3% on every row except shl (12%: run 2 caught the ramp). Driver 610.47:

+
kernelbest ms, runs 1 / 2 / 3best of the three, msG steps/s (best)ns per step (best)ops per step countedstep cost (ratio to alu, best)bit-exact, 3 runs
alu0.553 / 0.541 / 0.5530.5417,94113251.00yes
rotr (live)0.729 / 0.719 / 0.7150.7156,00517551.32yes
shflx (live)0.808 / 0.817 / 0.8170.8085,31519761.49yes
shl0.687 / 0.770 / 0.6960.6876,25316851.27yes
shr0.698 / 0.697 / 0.6910.6916,21216951.28yes
bfe (bfe.u32, inline PTX)0.837 / 0.851 / 0.8350.8355,14220451.54yes
bfec (C form)0.836 / 0.837 / 0.8350.8355,14420451.54yes
andn0.680 / 0.700 / 0.6930.6806,31316651.26yes
perm (__byte_perm)0.706 / 0.715 / 0.7180.7066,08417251.30yes
popc0.819 / 0.813 / 0.8350.8135,28319851.50yes
clz0.897 / 0.898 / 0.8840.8844,85621651.63yes
sel0.723 / 0.713 / 0.7290.7136,02417451.32yes
shfla (lane + 3)0.845 / 0.848 / 0.8290.8295,17920261.53yes
dot4i (__dp4a)0.677 / 0.656 / 0.6250.6256,87015351.16yes
mm8 (mma.sync.m8n8k16.u8, one per warp per step)1.313 / 1.314 / 1.3501.3133,2723211 mma + 42.43yes
+

Reading of the 5090 rows. Every row is bit-exact, mm8 included, so the m8n8k16 fragment layout of the CPU reference (PTX ISA 9.4 section 9.7.16.5.3) is the layout the hardware uses. The alu chain reads 7,941 G steps/s here against 8,754 through OpenCL event time on 5 October: a CUDA event pair around a 0.54 ms kernel carries about 0.05 ms of launch, which also compresses every ratio toward 1 (approximate: the ratios are the card's at the 2% level, not better). On this card every candidate costs more than the reference chain, unlike Apple: the 5090 runs the two-register add-xor-rotate chain at one IMAD and one funnel shift per step, and the three-register candidate chains pay their glue. Against the live rotr (1.32), the candidates read: andn 0.95x, shl 0.96x, shr 0.97x, perm 0.98x, sel 1.00x, popc 1.14x, shfla 1.16x (the same as the live xor shuffle, 1.49: the indexed shuffle costs NVIDIA nothing extra), bfe 1.17x, clz 1.23x. bfe.u32 and the C form cost the same to the nanosecond (0.835 ms), so the compiler emits the same code for both and no single-instruction bit-field extract is in play on this architecture (not checked by cuobjdump; the equal times are the evidence). dp4a reads 1.16x (1.17x on 5 October). mm8 is the most expensive row on NVIDIA too (2.43x the reference: one tensor-core mma per warp per dependent step, latency-bound), which supports its place at the end of the reserve on the honest-card side as well as on the chip side.

+

RX 9070 XT (PC 1, ae432dc7), OpenCL: OWED. PC 1 is the maintainers' desk and not released today (the brief's rule); the OpenCL twin of the probe (__builtin_amdgcn_* paths for v_bfe_u32, v_perm_b32, v_bcnt_u32_b32, v_cndmask_b32, ds_bpermute_b32) is the next job on that card.

+

Consequences per tier, Mac rows (the hash is latency-bound by 128 dependent DRAM reads; a family at W_new = 4 points is about 4% of the 64 instructions, so these per-op costs bound a family's hash-rate cost and are not hash rates; the 5% rule of 1.13.2 is argued from them, not measured, until a family is live):

+
TierWhat the rows meanWhat is being done
Apple user (M-series laptop or desktop, the app's Metal worker)six of the seven candidates (shl, shr, bfe, andn, popc, sel) cost at most the live rotr step; clz the same as the reference; perm 1.13x (emulated); shfla 1.91x, the only candidate over the live shfl's cost by more than 2x on this card. At 4 points of 64 a 1.91x op costs under 1% of the program's ALU time, itself a small share of a latency-bound hash (approximate: argued, measured when live)the proposed order puts shfla after the plain datapath families (R6), so Apple pays it last; mm8 stays last
NVIDIA user (8 to 32 GB card)every candidate is native and costs 0.95x to 1.23x the live rotr step (andn, the shifts, perm, sel under 1.0x; popc 1.14x; shfla 1.16x; bfe 1.17x; clz 1.23x); nothing on this card is emulated above a compiler sequence; at 4 points of 64 no family moves the ALU time by over 1% (argued) on a hash bound by DRAM readsthe family-live measurement at each unlock rehearsal; nothing to change in the order for NVIDIA
AMD user (RX 9070 XT, 16 GB)owed: no row todaythe RTX 5090 machine 1 job when the desk is free
A rig or a pool userthe same per-card figures; no family changes the dependent-read boundnothing until a family is live
A chipevery candidate but mm8 is a 32-bit datapath structure (barrel shifter, byte crossbar, popcount tree, 32-lane shuffle crossbar: docs/plans/counter-asic-3-reserve.md section 3 names them with approximate areas); none is licensable as a block the way an int8 matrix unit isthe reserve order of that document
+

6 October 2026, Counter ASIC 3.0 gate run, the hash side

+

Branch ca3-v4-hash, worker "v4-hash", on ca3-coord 3213ee9. The candidate class v4 mx8+sh256x27 composed with the era as the chain composes it (mx8-era<hex>+sh256x27, generator 3), on the devnet epoch-0 and day seeds at the genesis era (v4-devnet-epoch0) and at the six test eras of 2.0 (v4-era-0 to v4-era-5), with the class v3 control mx8-devnet-epoch0 re-exported beside them (proto-cuda/packs-ca3-v4/). Gates G1, G2 and G3 of docs/plans/counter-asic-2-rollout.md section 7 and the pinned verifier benchmark; the evidence tables and one JSON per run in docs/plans/counter-asic-3-gate/ (hash-gates.md). G4, G5 and G6 are the node's and the release's.

+

G1, bit-exact (with-lock.sh run, 15:49:29 to 15:50:02Z, load 6.8 to 6.7). packbench --pack <dir> --batches 1 --batch-log2 24 --group 256 (Metal) and igneum-bench-cl --bench-pack --pack <dir> --batches 1 --batch-log2 24 (Apple OpenCL), both built from this branch: on all eight packs the cache FNV 448274a57f508cbc PASS, the dataset head and last PASS, vectors 3 of 3 standalone and 3 of 3 in batch (Metal), 96 of 96 lanes (Apple OpenCL), and one fingerprint per pack across both harnesses. The control reproduces 2.0's 90f794dd556f7a3b (the harness fired on a known case).

+
PackClassFingerprint 2^24 at base 0 (Metal = Apple OpenCL)RTX 5090 (CUDA NVRTC, PC 2)RX 9070 XT
mx8-devnet-epoch0 (control)mx8-erad810f22d90f794dd556f7a3b90f794dd556f7a3bPC 1 job 1
v4-devnet-epoch0mx8-erad810f22d+sh256x27f410c731b6bc2d31f410c731b6bc2d31PC 1 job 1
v4-era-0mx8-erab2ed8a89+sh256x27b115c410e08be6cab115c410e08be6caPC 1 job 1
v4-era-1mx8-era676a17fc+sh256x27edc2b18fc67e9d1cedc2b18fc67e9d1cPC 1 job 1
v4-era-2mx8-era843155d7+sh256x27604ed87109570559604ed87109570559PC 1 job 1
v4-era-3mx8-erad6367bfe+sh256x279541e2a41dde2ee69541e2a41dde2ee6PC 1 job 1
v4-era-4mx8-era4488f3ed+sh256x27a9ffa2b67bd2e366a9ffa2b67bd2e366PC 1 job 1
v4-era-5mx8-eraf897c84e+sh256x271f34e9c4659452491f34e9c465945249PC 1 job 1
+

The PC 2 job (run-ca3-v4-gates-pc2-20261006, tools/ca3-v4/pc2-v4-gates.ps1, 15:52:23 to 15:52:49Z, exit 0, 26 s, the installed 0.3.11 igneum-worker-cuda.exe through NVRTC on each pack's own text, beside the app's miner, the prover untouched, no card switched off: bit-exactness is not load sensitive) ran G1 (--bench --batches 5 --batch-log2 24) and G2 on all eight packs in one job. The intake's report keeps the last 200 KB of job.log and the 8 x 1,024 G2 found lines filled it; the G1 lines came home through a collect whose command filters job.log (collect-ca3-v4-g1only-20261006, 16:12Z): self-test PASS on every pack, every fingerprint equal to the Apple M5 Max's (the column above), NVRTC 180 to 274 ms per pack, the 1 GiB build 38 to 45 ms. G1 is GREEN on NVIDIA and Apple; AMD is PC 1 job 1.

+

G2, the verifier exact on 1,024 hashes per card (with-lock.sh run, 15:53:54 to 15:54:22Z, load 6.8 to 8.9). One serve-mode job per card and pack, job g2 00..01 ffffffffffffffff 0 1024 <epoch> <day> class=v3 era=<hex> (Metal after prepare <epoch> <day> <pack> class=v3 era=<hex>), every nonce a found line, re-hashed with igneum-pow hash-bound --prehash 00..01 --nonce 0 --count 1024 on the same pack (--count ported from ca2-era). Metal 1,024 of 1,024 on all eight packs; Apple OpenCL 1,024 of 1,024 on all eight; RTX 5090 (igneum-worker-cuda.exe --serve --pack <dir> --race off, the same job) 1,024 of 1,024 on all eight packs, every found line equal to the Apple M5 Max's verifier; RX 9070 XT: PC 1 job 1. G2 is GREEN on NVIDIA and Apple.

+

G3, the soundness suite on the class. cargo test -j4 --release in igneum-pow (with-lock.sh build, nice 19, cargo 1.99.0, 15:49:29 to 15:50:01Z, load 6.8): 59 + 7 + 4 + 19 + 7 = 96 of 96. The mixer harness now takes the class (IGNEUM_MIXER_CLASS) in the fuzz, the stats and the determinism test, an era to compose (IGNEUM_MIXER_ERA=igneum-era-test/<n>), and a shadow contract (S instructions, no load, every field in range, the base program equal to the class's without the shadow draw for draw); the edge test is the dataset's alone and takes no class. On mx8+sh256x27 (15:54:25 to 15:55:39Z, load 9.1): fuzz 200 programs, 800 units, 200 in the top 256 nonces; stats avalanche 49.96 and 49.98 percent, worst bit z 2.65 and 3.56, 0 duplicates (v2 49.87 and 49.98, z 2.25 and 2.30); edge pass; determinism two builds equal and equal to the pinned packs-ca3-shadow/sh256x27. With era test/0 composed (50 programs, 200 units, 15:56:35 to 15:57:03Z, load 12.6): 4 of 4, avalanche 50.03 and 50.11, z 2.65 and 3.76. The GPU fuzz on the written packs (packbench --batches 1 --batch-log2 9 --batch-base 4294967040, every tenth on Apple OpenCL --batch-log2 10, with-lock.sh run, 15:55:42 to 15:58:21Z, load 11.3 to 12.3): Metal 200 of 200 and 50 of 50, Apple OpenCL 20 of 20 and 5 of 5. Known-failed case: the first era-composed run FAILED (1 failed, 15:55:42Z) on assert_ne!(p.program_id(), base.program_id()), which is the finding below.

+

The verifier (with-lock.sh measure, one session, 15:54:27 to 15:54:31Z, load 9.1 to 9.4; one core on a loaded box, within 3 percent of the quiet figures). igneum-pow bench --seed igneum-genesis --day 2026-10-03 --warps 20 --class <c>, two rounds:

+
Classms per 32-lane warp, avg of 20 (round 1 / round 2)Worst cold single runAgainst x82019-class core by the 2.5x rule (approximate)Gate
v20.655 / 0.6480.674about 1.6 ms10 ms
x8 (mx8)2.149 / 2.1202.5261about 5.4 ms10 ms
mx8+sh256x272.326 / 2.3072.559+0.18 ms, +8.5 percentabout 5.8 ms (worst cold about 6.4)10 ms, about 4 ms spare
+

Per tier: 73 microseconds per hash on an M5 Max core, so a pool checks about 13,700 shares per second per such core and about 5,500 per 2019-class core (approximate); a node on any tier verifies a warp inside the gate; no tier is slower than under x8 by more than 8.5 percent.

+

Findings. (1) Under the chain's path a generator-3 program's id is program_id(3, seed, attempt), class-independent: the seven exported packs carry 73bcbfe8ccf988f1 with and without the shadow, and 50 of 50 fuzz seeds agree. A node and a miner could agree on the id while running different classes, and 2.0's G4 check program_ids_differ_across_the_switch would not fire across a v4 activation: the v4 seam (G4, G6) must stamp its own generator version or put the class in the id. The hash needs nothing for it; the cut must not go without it. (2) The report cap: a G2 playbook that prints 8,192 found lines loses its own G1 lines; the tooling fix is a found-lines file plus a count and digest on stdout, with a collect. Owed: AMD (PC 1 job 1), the 2019-class core (O-1.14).

6 October 2026, 00:4xZ, the empty /api/state reply (proving v1 branch)

Reported by the aggregation-cost agent: PC 2's /api/state answered {} (2 bytes) at 22:22Z, 22:41Z and 00:18Z. Not measured on PC 2 (no job); derived from the app source and node 1's RPC, read-only on the Apple M5 Max:

FigureValueSource
Paid shards, devnet, all provers663curl -s 127.0.0.1:26790 -d '{"jsonrpc":"2.0","id":1,"method":"igneum_getProvingStatus","params":[]}' at tip DAA 0x22caf
Paid wei, all provers0x2c2961a69990745400 = 814.64 IGNsame call
Average per paid shard1.23 IGN (approximate: the mean over 663)814.64 / 663
u64::MAX in IGN18.452^64 - 1 over 1e18
Paid shards per app start before the reply empties15 (approximate: at the mean payout)18.45 / 1.23
@@ -737,7 +787,11 @@ th{font-family:var(--f-mono);font-size:12px;letter-spacing:.12em;text-transform:

A fresh node on the Apple M5 Max against the seed only (the 0.3.12 binary 83089544, 300 s, kaspa_p2p_flows=debug): 11,700 Finality relay: certificate lines, every one new=false, 15 distinct indices, 240 per second at the peak (2,530 per 10 s), 203 votes; no route error on the Apple M5 Max (it drains 240/s with a 256-deep route) and one connection, where the fleet's slower boxes filled the route and lost the peer.

The fix (four changes, 5a339733): ingest_certificate ignores an index below keep_from (counted, debug: the echo stops at its source once the seed runs it); the router's overflow policy for IgneumFinality is Drop with a counted warn once per 10 s per peer, never a disconnect; the finality route is subscribed with 4,096 (a checkpoint's worst case is MAX_VOTES_PER_BLOCK 48 votes on each of 30 blocks plus the certificates); the relay flow skips votes while IBD runs (counted, said once per 30 s; certificates still go in and land pending). No consensus change, no digest change. Tests: the overflow-policy table (p2p 33 of 33), the flows crate (19 of 19), a certificate below the window submitted twice (ignored, no gossip, counter 2; an index inside goes the normal way) with the finality tests (12 of 12).

After, on the fixed binary against the still-unfixed seed (203ae727, same run, 13:15:31 to 13:20:31Z): 63,628 certificates received (the seed's echo had grown to 3,032 per 10 s at the peak as more fleet nodes joined), 0 route errors, 0 drops, 4 connections kept (the seed and three peers learned from it), 168 votes skipped during IBD. The receiver side of the fix holds under a storm five times the morning's; the source side (the guard) cannot show on the seed until 0.3.13 runs there, and the fresh node's own guard never fires during IBD (its window starts at genesis), which is correct. Harness s7 on the fixed binary (--quick --live-only): PASS, 192 blocks accepted in 60 s under a 50 blocks/s flood from one peer, honest template p50/p95/max 0.4/0.6/1.4 ms, rss 306 to 321 MB.

-

Per tier: a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of finality.rs (taken out of the commit).

+

Per tier: a home miner joining today sees the warning and the peers=0 flicker every checkpoint until the seed runs 0.3.13; a rig the same once; a pool user nothing; a fleet operator gets a node that keeps its only peer, and a seed that stops amplifying old certificates to every peer (13,354 lines of work it did not need in seven minutes). Owed: the fleet agent's synced-node reading; a receiver-side limit on certificates per index per minute as a second belt once the seed is fixed; the formatter's reflow of finality.rs (taken out of the commit).

+

6 October 2026, 16:01Z: Ember run 6 on PC 1 (ember-tune-pc1-6, 0.3.13 + kit-6 = 564bdea, elevated, one click)

+

The helper registered inside the run but on the scratch copy (fixed: task_exe, the reregister verb; see the plan's run 6 notes). RTX 5090: chosen 1854 MHz at 100% = 127.71 MH/s at 226.8 W, 0.563 MH/W, against 127.9 at 311.0 W (0.411) untuned: 84 W saved for 0.15% of rate. Every cap step 60 to 100% read 311 to 313 W (the cap never binds). The clock ladder: 2781 MHz 298.8 W 0.428; 2472 MHz 262.0 W 0.488; 2163 MHz 239.9 W 0.533; 1854 MHz 226.8 W 0.563 (the floor, not the optimum: the next cut's ladder goes to 45%). RTX 4070: caps 100 to 60% all 106.0 W 28.71 MH/s (0.271); 50% 99.1 W 28.70 (0.290); clock 2794 MHz at 50% 99.1 W (0.290); 2484 MHz 81.1 W 28.73 (0.354); further rows and the 9070 XT ladder below once the run closes.

+

Run 6 closed 16:39:37Z, exit 0, 2317 s, 19 rows. RTX 4070 chosen 1863 MHz at 50% = 28.78 MH/s at 75.6 W (0.381) against 28.72 at 106.0 W (0.271): 30 W saved for no rate lost; its clock ladder at 50%: 2794 MHz 99.1 W 0.290; 2484 MHz 81.1 W 0.354; 2173 MHz 77.7 W 0.370; 1863 MHz 75.6 W 0.381 (the floor). RX 9070 XT: aborted at step 1, "card reports 0 W, acknowledged true" = the applied rule demanded watts from a card that reports offsets (fixed bd7fcf4); the draw itself was read on every tick (amd_watts_source=engine_telemetry, 363 samples). The helper registered on the scratch copy (fixed 200362a: task_exe, the reregister verb). PC 1 mined through the installed app again by 16:44Z: 170.6 MH/s over the three cards, 0 faults.

+

Re-point and proof, 16:48 to 16:53Z: the Power Helper task re-pointed by the helper itself ("1 reregister ok: the task now runs ...Programs\Igneum Miner\igneum-app.exe --power-helper"), then the installed app's caps through the task with no prompt ("-pl 460: set to 460.00 W from 575.00 W", "-pl 160: set to 160.00 W from 100.00 W"). Task Running, Highest, user Admin. Cards: 5090 221 W at 1845 MHz, 4070 75.8 W at 1860 MHz, 9070 XT 202 W, all mining through the installed app. Ember closed 16:53Z.

Generated from the repository at build time. Times are UTC. Machine names are model names.

diff --git a/site/build.mjs b/site/build.mjs index 709eae186..102e557db 100644 --- a/site/build.mjs +++ b/site/build.mjs @@ -275,6 +275,11 @@ function stampProduct(html, file) { const TESTNET_OPEN = false; const DL_HOST = 'https://dl.igneum.network'; const DL_SNAPSHOT = join(here, 'downloads.json'); +// The snapshot is a tracked file: the build rewrites it only on an explicit refresh (SITE_DOWNLOADS_REFRESH=1 or +// --refresh-downloads, the ship pipeline's step, committed with the release), or in CI (a throwaway checkout). A plain +// local build, and the pre-push hook, read live but never write: on 6 October 2026 the hook's build stamped 0.3.14's rows +// into five worktrees that had nothing to do with the release, as uncommitted edits to tracked files. +const DL_REFRESH = process.env.SITE_DOWNLOADS_REFRESH === '1' || process.argv.includes('--refresh-downloads') || !!process.env.CI; async function loadDownloads() { const snapshot = existsSync(DL_SNAPSHOT) ? JSON.parse(readFileSync(DL_SNAPSHOT, 'utf8')) : { files: {} }; if (process.env.SITE_DOWNLOADS_OFFLINE) return { ...snapshot, source: 'snapshot (offline)' }; @@ -286,7 +291,8 @@ async function loadDownloads() { const live = await r.json(); if (!live || typeof live.files !== 'object') throw new Error('no files object'); for (const [k, f] of Object.entries(live.files)) if (!/^\/public\/[a-z0-9.-]+$/.test(f.alias) || !/^\d+\.\d+\.\d+$/.test(String(f.version)) || !(f.size > 0)) throw new Error(`bad entry ${k}`); - writeFileSync(DL_SNAPSHOT, JSON.stringify(live, null, 2) + '\n'); + if (DL_REFRESH) writeFileSync(DL_SNAPSHOT, JSON.stringify(live, null, 2) + '\n'); + else if (JSON.stringify(live) !== JSON.stringify(snapshot)) console.warn('downloads: the live index differs from site/downloads.json; the release step refreshes it (SITE_DOWNLOADS_REFRESH=1 node site/build.mjs)'); return { ...live, source: 'live' }; } catch (e) { console.warn(`downloads: using the snapshot (${e.message})`); diff --git a/site/index.html b/site/index.html index 344848da9..d5235a86e 100644 --- a/site/index.html +++ b/site/index.html @@ -676,7 +676,7 @@ pre{margin:0;font-family:var(--f-mono);font-size:13px;line-height:1.6;color:var( - +