Merge remote-tracking branch 'build/master' into class-v5

This commit is contained in:
igneum-labs 2026-10-08 08:38:09 +00:00
commit 80caad07b6
15 changed files with 160 additions and 30 deletions

View file

@ -291,7 +291,7 @@ attack-pass) by the coordinator's exception rule; nothing touches GitHub.
| F8 hot-set gate (pre-freeze reading) | 64 seeds at 2^24, chain path, v5 with the dn3 state, box 2: hand-started at 18:47 UTC (11 seeds, killed on the coordinator's rule: every hand-started run off the boxes, loads 601 and 401), re-queued through `lease pool 64` at 19:23 UTC (17 more seeds), released at 19:4x UTC on the Counter ASIC lane's yield so the class v5 (c''') census, the 0.3.24 board's critical path, could take the pool | every seed read equals sub-version 3's seed for seed (0.9915x to 1.144x, p10 1.50x), as the v5 lane predicted: the leaves change the words, not the read addresses | READING, not the gate line |
| F8 hot-set gate, the frozen tip (THE GATE LINE) | igneum-pow class-v5 1c420786 (the 0.995 per-site floor (c''') on sub-version 3's rules; binary sha256 0f5c98dc41a1b3aa..., run from a copy); pairing: the library draws the dn3 epoch-0 program as e5a4ac5978462156, the harness validates bit for bit against the library on the dn3 state (0 mismatches on 66 validation lines); 64 seeds p2 to p65 at 2^24 nonces, chain path, the v5 dataset from v5-dn3-epoch0's state.igsd1, window-model control, build-2 under `lease pool` class v5 as two halves of 32 (ended 21:58:43Z and 22:03:21Z) | 61 of 64 under 1.2x of the window model (0.9919x to 1.144x, p75 1.0024x); 3 over, all inside the named four-seed residue and none new: p10 1.5047x (0x4018f5, 346 reads of 2^31, no predicted source), p8 1.3787x (0x839d33, 419), p4 1.2166x (0x400197, 363); p34 0.9997x under the (c''') floor; p23 1.0000x, p19 0.9997x, p15 0.9998x, p18 1.0001x, p56 1.0000x. Seed for seed the ratios equal sub-version 3's within 0.001 except where the floor moved a draw: the state leaves change the words, not the read addresses. Nothing to a chip | PASS (the known residue p4, p8, p10 unattributed and chased) |
| F9 exhaustion count, the frozen tip (A GATE LINE for the 0.3.24 move) | 10^5 chain-shaped seeds on the v5 chain path (era-composed class, `--chain`) with the dn3 state at igneum-pow class-v5 1c420786, pairing e5a4ac5978462156, build-1, ten chunks of 10,000 under `lease pool 4` (chunks 1, 3, 5 to 9 under class v5; chunks 0, 2, 4 re-leased under class release on the coordinator's order; 14,477 to 14,480 s per chunk, about 1.45 s per seed); binary copied into the run dir; interim line sent at 00:55 UTC (seeds drawn, 0 exhausted, 0 panics, max 30, F1 0 failures), which cleared the move; last chunk written 02:34:54 UTC | 100,000 of 100,000 seeds drawn, 0 exhausted, 0 panics, 0 past attempt index 31, max attempt index 30; histogram by attempt index (0 = accepted on the first draw) 0: 31,454; 1: 21,460; 2: 14,660; 3: 10,263; 4: 7,047; 5: 4,701; 6: 3,297; 7: 2,256; 8: 1,532; 9: 1,027; 10: 702; 11: 509; 12: 365; 13: 216; 14: 153; 15: 103; 16: 80; 17: 56; 18: 39; 19: 20; 20: 24; 21: 9; 22: 6; 23: 9; 24: 3; 25: 4; 26: 2; 27: 1; 29: 1; 30: 1; first-draw acceptance 0.3145, mean attempt index 2.185 (3.185 draws per seed), 4,862 seeds (4.86 percent) at index 8 or above, 255 (0.255 percent) at 16 or above; the 256-attempt cap and the deterministic last resort never reached. Meaning per tier: no epoch seed in 10^5 fails to draw a program, so the liveness halt of AP-F8-2 has no observed case on the frozen tip at this count (the bound it supports is under 3e-5 per seed at 95 percent, about one epoch in 33,000 at worst; a halt a node operator would see as a stuck epoch, a miner as a dead epoch, a holder as a paused chain), and the draw cost stays at about 3.2 candidates per epoch for every node | PASS (0 of 10^5; the record `f9-grind.md`, section (d)) |
| F1 shadow redundancy, the frozen tip (A GATE LINE for the 0.3.24 move) | 10^5 class v5 programs through the string-seed path (`generate_from_seed_bytes_program_class`, class V5, every candidate draw through the (c''') floor over 2^20) at igneum-pow class-v5 1c420786, build-1, `lease pool 16` re-leased under class release at 22:34:05 UTC (cores 8 to 23), binary copied into `frozen-1c420786-f1/bin` | Running at the 03:06 UTC reading (8 October 2026): the census process (pid 3151604) at 15 busy cores over a 45 s sample, 2 d 19 h of CPU banked over 4 h 30 min of wall, 0 on its live panic path (`census.log`), no end marker; projected end about 09:00 UTC from F9's measured draw cost (about 1.8 core-s per candidate under the night's load, 3.2 draws per program). The interim zeros stand as the partial. Harness gap, named: the census collects every Report in memory and writes `out/census.csv` only at its end, with no progress line, so a count before the end is not an observable on this harness (the v4 10^5 reference of 07 Oct 09:33 UTC, 0.13 core-s per program, ran before sub-version 3's acceptance rule, which is the fifteenfold). Owed on this branch before the harness's next 10^5 run on any class (coordinator's ruling, 04:1x UK): a progress line every 1,000 programs (count, elapsed, running failure count), a partial `census.csv` flushed at the same cadence, and a known-failed test of the flush (a kill after 2,000 programs leaves 2,000 rows) | running; the record line in a second merge when `census.csv` writes |
| F1 shadow redundancy, the frozen tip (A GATE LINE for the 0.3.24 move) | 10^5 class v5 programs through the string-seed path (`generate_from_seed_bytes_program_class`, class V5, every candidate draw through the (c''') floor over 2^20) at igneum-pow class-v5 1c420786, pairing e5a4ac5978462156, build-1, `lease pool 16` class release (cores 8 to 23, re-leased 22:34:05 UTC), binary sha256 bb70bbf69a4b3223... copied into `frozen-1c420786-f1/bin`; 20,774 s of census (5 h 46 min; about 5 core-s per program, the (c''') draw cost), census.csv (sha256 4e34b669f2f5c680..., 100,000 rows) written 04:20 UTC on 8 October 2026 | 100,000 of 100,000 programs; instructions saved min 0.000, mean 0.623, max 4.688 percent (worst `attack-f1/95060` at attempt 0, 6,912 to 6,588 per iteration; `81748`, `66933`, `3006` at the same 4.688; next 4.311); chip-view ops saved mean 0.520, max 4.783; programs over 5 percent 0, over 10 percent 0; soundness: differential mismatches 0 of 100,000 (8 random states each), verifier mismatches 0 of 100,000; 0 panics; histogram of saved, 0.5 percent bins from 0: 55,241; 20,597; 11,762; 9,851; 1,484; 656; 259; 133; 13; 4; 0; 0; draw attempts per program: index 0 31,630, max 30 (the same shape as F9's). Against the v4 10^5 (max 5.078, the AP-F1-1 letter miss): the v5 tip's worst sits 0.39 points under the letter, the two top bins are empty, and the mean is unchanged (0.617 to 0.623), so the v5 leaves add no redundancy and remove the one letter miss. Meaning per tier: the shadow block of every drawn program stays within 5 percent of its naive count under the harness's rules and the honest compiler finds the same shortcuts, so no chip gets a shadow-side discount (a miner on a card pays the full block, a hypothetical ASIC gains nothing here) and the chain's verifier agrees with the harness on every program (0 mismatches), so no node disagrees with another on any drawn block. Harness gap and its fix: section 12 of the record (the running census was the old binary; the flush is in 18a9c04a for every census after it) | PASS (0 of 10^5 over the letter; AP-F1-1 FIXED-AND-PASSED on v5 at this count; the record `f1-shadow.md`, section 13) |
| F4 weak-day census, the post-freeze commit | 2^24 chain days at class-v5 8ca66afa (AP-F4-1 in the agreed form: cost A at most 205, k >= 1 or all-ROT-equal rejected, the forty redrawn), the harness on the agreed w32 convention (digits 0 to 31) and median 226; known-failed day 29,337 redrawn under the rule (cost 203 to 228), day 20,729 at 219 unchanged; build-1 `lease pool 12` class adv, 379 s, ended 22:3x UTC | 0 of 2^24 days over 1.1x on M1 (median 226; minimum cost 206 at day 27,016, 1.097x, one adder above the reject line) and 0 on M2; mean 225.79, sd 6.07 (pre-rule: 5.69e-4 over, min 203). The redraw rule removes the LUT tail by construction and the measurement agrees | PASS; AP-F4-1 FIXED-AND-PASSED against class v5 at 8ca66afa |
| F9 exhaustion count | 10^5 chain-shaped seeds on the v5 chain path with the dn3 state, box 1, ten parallel chunks (the chain draw costs about 2.2 s per candidate through (c''), so 10^6 is about fifty hours); killed before its first chunk closed, re-queued through `lease pool` | pending | pending the re-queue |
| F1 shadow redundancy | 10^5 class v5 programs through the string-seed path, box 1; the known firings fire under v5 (planted 50 of 256: 19.53 percent; the real block 0.000; the must-not-fire 1.157); the census killed before its end, re-queued through `lease pool` | pending | pending the re-queue |

View file

@ -276,3 +276,69 @@ exclude), which is the workaround until the build-server lane's check lands.
| The CPU verifier | runs the block as written (`verify.rs` interprets every instruction), so on a 4.7 percent program it does 4.7 percent of the shadow work a compiled miner skips: 0.03 ms of the 0.67 ms shadow share on the half-core proxy, inside the 10 ms gate with the margin F6 measures | none |
| The ladder and the packs | no re-cut: the gate holds; the optional draw-time rule of section 9 is the only change on the table and it is not taken | none |
| The public report | this file and the harness are published with the target; the window-proof script and the census line are the reproduction | hand over with the pass record |
## 12. The census flush (8 October 2026, 04:1x to 05:1x UK; the coordinator's ruling after the class v5 10^5 run)
The v5 10^5 census on 1c420786 ran 4 h 30 min on build-1 with nothing on disk: the harness collected every Report in
memory and wrote `census.csv` once at its end, with no progress line, so its state could not be read (the lane (d)
row names the gap). Ruling: before the harness's next 10^5 run on any class, a progress line every 1,000 programs
(count, elapsed, the running failure count) and a partial `census.csv` flushed at the same cadence, with a
known-failed test of the flush. Done on attack-v5-frozen at 18a9c04a (`tools/attack/f1-shadow/src/main.rs`,
`census`): the crossing thread rewrites `out/census.csv` from every row so far, sorted by idx, through a temporary
file and a rename, so a kill never leaves a torn file; the line reads
`progress: N of COUNT programs, S s, failures F (difftest or verifier), census.csv N rows`.
Known-failed test (box 2, `flush-test/`, binary sha256 14180ef40d55eb74..., `lease pool 16 --min 8 --class adv`,
lease pid 340097, queued 03:12:32Z behind the box's pool, started when adv-accept's sweep-s05b freed its cores):
a 4,000-program v4 census killed by pid (445480, TERM) the second the 2000 line appeared, at 04:08:27Z.
| Reading | Value |
|---|---|
| Progress lines before the kill | `1000 of 4000, 286 s, failures 0, 1000 rows`; `2000 of 4000, 571 s, failures 0, 2000 rows` |
| Rows in `census.csv` after the kill (header excluded) | 2,000 |
| `census.csv.tmp` left behind | none |
| Lease exit | 143 (the kill), 16 cores released |
Verdict: PASS (2,000 rows after a kill at 2,000; the known-fail of the old harness was zero rows). A side reading:
1,000 programs per 286 s on 16 cores is about 4.6 core-s per program on this box under its load, which is the
(c'') and (c''') draw cost per candidate and confirms the F1 10^5 projection on build-1 (about 5 to 6 core-s per
program, about 10 h on 15 busy cores). Consequences per tier: none for a user; for the lanes, every future census
can be read and killed without loss.
## 13. The frozen class v5 tip, 10^5 programs (1c420786; the 0.3.24 gate line; 8 October 2026, 05:20 UK)
Run: igneum-pow class-v5 1c420786 (pairing e5a4ac5978462156), build-1, `lease pool 16 --min 8 --class release`
(cores 8 to 23, re-leased 22:34:05 UTC on the coordinator's order after the class v5 lease was killed under the
duplicate-lease clean-up), the binary (sha256 bb70bbf69a4b3223...) copied into `frozen-1c420786-f1/bin`, 16 threads,
20,774 s (5 h 46 min; about 5 core-s per program, which is the (c''') floor over 2^20 on every candidate draw, measured
again by section 12's 4.6 core-s on box 2), `out/census.csv` (sha256 4e34b669f2f5c680..., 100,000 rows) and
`out/summary.txt` written 04:20 UTC. The interim line at 00:55 UTC (0 on the live panic path) cleared the move;
this is the record line.
| Quantity | Value |
|---|---|
| Programs | 100,000 (`attack-f1/0` to `attack-f1/99999`), 16 threads, 20,774 s, finished 05:20 UK |
| Instructions saved, min / mean / max | 0.000 / 0.623 / 4.688 percent |
| Worst programs | `attack-f1/95060` (attempt 0), `81748` (1), `66933` (1), `3006` (2): 6,912 to 6,588 per iteration (12 of 256 per pass); next `55048` at 4.311 |
| Programs over 5 percent / over 10 percent | 0 / 0 |
| Chip-view ops saved beyond free rotates and hoisted constants, mean / max | 0.520 / 4.783 percent |
| Differential mismatches | 0 of 100,000 (8 random states each) |
| Verifier mismatches | 0 of 100,000 |
| Panics | 0 |
| Dead (never-read) derived nodes under the full fold | 3,553,599 |
| Rewrites over all programs and 27 passes | identity 3,237,120; xor-cancel 3,127,199; sum-cancel 11,373,389; or-idem 289,936; rotl-merge 4,436,701; rotr-merge 320,399; product-shared 2,303,677 |
| Histogram of instructions saved, 0.5 percent bins from 0 | 55,241; 20,597; 11,762; 9,851; 1,484; 656; 259; 133; 13; 4; 0; 0 |
| Draw attempts per program (index 0 = first draw) | 0: 31,630; 1: 21,226; 2: 14,842; 3: 10,151; 4: 7,007; 5: 4,787; 6: 3,257; 7: 2,205; 8: 1,470; 9: 1,115; 10: 697; 11: 502; 12: 345; 13: 225; 14: 174; 15: 113; 16: 78; 17: 53; 18: 38; 19: 28; 20: 12; 21: 17; 22: 8; 23: 6; 24: 5; 25: 7; 29: 1; 30: 1 |
Against section 7.2 (class v4 at 10^5, max 5.078, the one letter miss recorded as AP-F1-1): the v5 tip's worst
program sits 0.39 points under the 5 percent letter, the two top bins are empty (v4: 1 and 7), the mean is unchanged
(0.617 to 0.623) and the shape of the shortcut is the one of section 7.3 (a register written twice from the same
source with no write between, 12 instructions, nothing crossing a pass). The attempt histogram has F9's shape
(first-draw acceptance 0.316 against F9's 0.3145 on chain-shaped seeds), so the string-seed path and the chain path
draw the same distribution.
Verdict: PASS by the letter and at honest-compiler parity (0 of 10^5 over 5 percent, 0 mismatches); AP-F1-1
FIXED-AND-PASSED on class v5 at this count. Consequences per tier: no drawn program's shadow block gives any chip a
discount beyond the honest compiler's own simplification (a card pays the full block, a hypothetical ASIC gains
nothing on the shadow side), and the verifier agrees with the harness on every program, so no node disagrees with
another on any drawn block.

View file

@ -363,6 +363,8 @@ Per tier: a miner on class v4 pays the premium and gets the 2.1x to 3.9x chip ce
The k column. k is the chip core's energy per op over the GPU's at the same operating point, and the model's 3.4x row takes k about 0.33 for an ALU-shaped core. On the public figures the int8 tensor tile is the one GPU block whose energy per op a chip at the same node cannot undercut with certainty: the 4090 measures 0.056 pJ per MAC; NVIDIA's 5 nm INT4 test chip reads 0.021 pJ per MAC at 0.46 V and about 0.1 at nominal (JSSC 2023, via Dally's NASEM slides; claimed), so INT8 at 2x to 4x that gives k 0.7 to 3 with the centre near 1; every ALU-shaped block reads k 0.3 to 0.8 on the same sources. A shadow built of tensor tiles at the ALU shadow's premium (about 11,400 u8 tiles per hash) therefore gives 2.1x at k = 1 and 1.6x at k = 1.5 and removes the k 0.3 column from the table; it needs a SIMD byte-dot verifier (the scalar one at 12.4 ms fails the 10 ms gate). This is a design candidate, not the shipped stream: the shipped shadow is ALU-shaped and its row stays 2.1x at k = 1 and 3.4x at k about 0.33.
Correction, 8 October 2026 (the research lane's microbench on the 5090, counter-asic-4-research.md 15.1a and the corrected 20.3 and 20.4 at 71fd465b): the per-MAC figures above are wrong by a factor of 32. A `mma.m8n8k16` tile is 1,024 multiply-adds per warp, 32 per lane, so a hash does 32 MACs per tile, not 1,024; the 4090's "0.056 pJ per MAC" is 1.8 pJ, and the 5090 at the ALU shadow's premium reads 2.9 pJ per MAC unlocked and 1.5 pJ at the 1,300 MHz lock (the packs job, 366,080 MACs per hash; the microbench's dependent u8 tile 4.1 and 2.2, the wide s8 m16n8k32 tile 1.36 and 0.83). Against the same 5 nm MAC array figures (0.04 to 0.4 pJ per INT8-class MAC, claimed) a chip's k on tile work is therefore 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, not near 1: at the same premium a tensor-shaped shadow leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. The tensor-tile column (2.1x at k = 1, 1.6x at k = 1.5) is withdrawn as a candidate; its premise, that a chip's MAC is no cheaper than the GPU's, is false by 4x to 30x on the public figures. The shipped row is unchanged: 2.1x at k = 1 and 3.4x at k about 0.33, the ALU shadow at the operating point's knee, measured four times at 82 to 90 W.
The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job.
## 6. The per-day derivation (item 2)

View file

@ -189,7 +189,7 @@ Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and e
What the rows say:
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised) [corrected 8 October 2026: these per-MAC figures count 1,024 multiply-adds per tile per lane where a tile is 1,024 per warp and 32 per lane, so they are low by 32x: 22 pJ at R = 8 falling to 1.8 pJ at R = 512; the watts and microjoules per hash stand; counter-asic-4-research.md 15.1a at 71fd465b], 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
@ -273,7 +273,7 @@ Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `
| Scheme | Verdict | Why, in one line |
|---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add as first counted, 1.8 to 22 pJ with the per-lane count corrected on 8 October 2026, 15 W for 4,096 tiles per hash) is the reason, and the correction strengthens it: a chip's MAC at 0.04 to 0.4 pJ (claimed) against the GPU's 1.8 pJ gives a chip k of 0.03 to 0.3 on tile work, below the ALU shadow's. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.

View file

@ -2758,3 +2758,26 @@ RTX 5080 (the same shape; 8 October 2026, 00:45 to 01:54 UTC):
Reading (5080): the rate holds within 0.3 percent of unlocked down to 1,000 MHz on both classes and falls 5.2 percent at 900 MHz on class v4, so the knee is between 1,000 and 900 MHz, a third of the 2,963 MHz boost and lower than the 5090's (the 5080's 84 SMs have more compute headroom per unit of its memory bandwidth, so the memory wait hides the shadow down to a lower clock). Best MH per watt within the 1 percent rate tolerance: class v4 at 1,100 MHz (71.20 MH/s at 146.6 W, 0.486 MH/W; 106.5 W recovered for 0.29 percent rate), class v3 at 1,000 MHz (71.11 at 103.7 W, 0.686; 66.0 W for 0.25 percent). The class v4 premium is 83.4 W unlocked and 41 W at the best points (146.6 W against 105.6 W at 1,100). The draw floors from 1,500 MHz down (about 147 W on v4, 104 W on v3) with the power governor as the only throttle reason say the clock lever is spent by 1,500 MHz on this card. Consequence per tier: a 5080 owner on class v4 locked near 1,100 MHz pays 147 W instead of 253 for 0.3 percent less rate (MH per watt up 72 percent).
What the lever is and is not: nvidia-smi exposes no voltage offset (that is NVAPI's); the clock lock walks the driver's V/F curve, which is where the watts come from; the Mac has no lever (no clock cap on Apple silicon); AMD has the ADLX tune line through igneum-gpu-telemetry or nothing. The knob goes into Ember Tune for 0.3.24 (src/ember.rs: the clock ladder continues below 45 percent of the maximum in 100 MHz steps to a 20 percent floor, the search stops at the first row more than the tolerance under the cap point's rate or on a faulted row, the best MH per watt within tolerance is the point, the fingerprint checked on every step, the result stored per card as the lock_* fields). The 5080 stock rows against the rented 5080 of 7 October (71.16 MH/s at 145 W on driver 580): the rate agrees to 0.4 percent, the watts do not (253 W here); PC 1's three power fields agree, so the difference sits with the rented card's sampler or its cap, the fleet lane's re-measure owed.
## 8 October 2026, the hot-table packs on the RTX 5090: ld.global.cs against the plain load, unlocked and at the 1,300 MHz lock (branch ca3-v4-amend, the hash lane, job run-ca4-pc1-hot-ldcs-5090-20261008b)
PC 1, RTX 5090 (170 SMs, driver 13.4, NVRTC 12.8), igneum-worker-cuda 1.0 (4 October 2026). The eight 5 October hot-table packs (proto-cuda/packs-ca2-hot) and the mx8-genesis control (proto-cuda/packs-ca2-mixer), each run twice per state: variant base (plain loads) and variant ldcs (ld.global.cs on the hot-table loads). 16,777,216 nonces per run, nvidia-smi 1 Hz sampler (power.draw; instant and average agreed within 0.5 W on every row), 19 to 30 samples per row. Every one of the 36 rows PASS on its pinned fingerprint. Unlocked 06:34 to 06:45 UTC at 2,865 MHz; lock1300 06:45 to 06:56 UTC at 1,290 MHz through the helper; clocks reset (rgc) at the end. The 5090 was switched off in the app by the runner for the job and back on after.
| pack | unlocked base MH/s | unlocked base W | unlocked ldcs MH/s | lock1300 base MH/s | lock1300 base W | lock1300 base MH/W | lock1300 ldcs MH/s | lock1300 ldcs MH/W |
|---|---|---|---|---|---|---|---|---|
| mx8-genesis (control) | 137.7 | 312.0 | 137.7 | 127.3 | 211.4 | 0.602 | 127.3 | 0.603 |
| hot32k4 | 147.4 | 318.9 | 147.4 | 144.4 | 215.3 | 0.671 | 144.4 | 0.672 |
| hot32k4a | 118.9 | 320.1 | 118.9 | 116.6 | 214.9 | 0.542 | 116.6 | 0.542 |
| hot64k2 | 137.7 | 318.1 | 137.7 | 135.0 | 214.3 | 0.630 | 135.0 | 0.630 |
| hot64k4 | 141.0 | 318.8 | 141.0 | 138.2 | 214.4 | 0.645 | 138.2 | 0.644 |
| hot64k4a | 115.7 | 319.6 | 115.8 | 113.4 | 214.8 | 0.528 | 113.4 | 0.528 |
| hot64k8 | 164.2 | 326.4 | 164.2 | 160.9 | 219.2 | 0.734 | 160.8 | 0.733 |
| hot96k4 | 139.0 | 318.3 | 139.0 | 136.2 | 214.5 | 0.635 | 136.3 | 0.635 |
| hot96k4a | 114.6 | 318.8 | 114.5 | 112.3 | 214.8 | 0.523 | 112.3 | 0.523 |
What the rows say.
1. The ldcs variant changes nothing: every pack reads the same rate and the same watts as its base run within 0.1 MH/s and 1 W, both states. The streaming hint on the hot-table loads is dead as a lever; the table's cache behaviour is already what the hardware gives. No further ldcs rows are owed.
2. The hot packs hold their rate under the lock far better than mx8: the 1,300 MHz lock costs mx8 7.5 percent of rate (137.7 to 127.3) and costs the hot packs 2 percent (hot32k4 147.4 to 144.4, hot64k8 164.2 to 160.9). The hot packs are bound by the table's latency, not by the core; the lock takes a third of the watts off every pack (319 to 215 W) and the hot packs pay almost no rate for it.
3. Per watt at the lock the hot family runs 0.52 to 0.73 MH/W against the control's 0.60: hot64k8 is 22 percent cheaper per hash than mx8 on this card, hot32k4 11 percent cheaper, the "a" packs 10 to 13 percent dearer. Whether a cheaper hash on the GPU is a gain or a loss for resistance is the research lane's call: it is a gain only if the saving comes from the memory path an ASIC would have to buy too.
4. The 5090's locked class v4 reading from the efficiency pass (1,300 MHz: 134.6 MH/s at 223 W, 0.60 MH/W) sits level with the mx8 control here (0.602), so the two passes agree on the control and the hot rows are comparable to the v4 grid.

File diff suppressed because one or more lines are too long

View file

@ -190,3 +190,11 @@ The 0.3.23 node's heights move Devnet 3's digest to ba75bf6f, and the one-box-at
Main's word on route (A) or (B) did not come (asked 01:41, 01:50, 01:53, 01:56 BST), so nothing applied, nothing placed, no minute: build-1's /fleet/move.json still serves the 22:30 BST 2720d8d2 move (m2720-1); every Devnet 3 node is on the 0.3.23 pin 2720d8d2, digest ba75bf6f (dn3-j1 behind its dead proxy unverified since 22:30); the 0.3.24 Mac entry stands staged (DMG 1aa301cc, both token folders) and unpublished; the live manifest is 0.3.23 (Mac and HiveOS), the 0.3.23 and 0.3.24 Windows entries and the Discord card unpublished; the shipper's stand-down at 02:15 held by default (its line was not sent at 02:15, the shipper's miss; the state was unchanged). **Finding:** build-1's Devnet 3 seed (the --go process on 26631 with JSON RPC 27632, the node lane's DAA reader) is DOWN at 02:56 BST (no process with --appdir=/home/build/dn3seed; node1-dn3 26671 and the observer 26651 run on 2720d8d2; the node lane's 0.3.24 reader on 28690 runs); the last DAA read held is 25,169 at 00:52:38 BST, and at 1.0 DAA/s the chain passed 32,400 at about 02:53 BST, so the 39,600 floor is lost and the fourth re-cut is from a morning minute main names (before 12:50 BST, or the three heights move with the floor in the same commit); the fleet's nodes dial 26631, so the seed's restart on the pin is part of the morning's move. **The morning's shape on main's word:** the fleet lane and the build-server lane respawned (the puller and the gate script r0324/move/pair-gate-aa8c2978.py ready; 33 of 35 boxes reachable at 01:55); build-1's seed restarted on the pin; the pin re-cut from the minute; the pairs and the hive with the v5 kit e6c088bb; the one-minute move with the gate line first; the Mac entry at the minute; the Windows chain (the 0.3.24 host on PC 1, the installer and smoke on PC 2 in the simplest shape, 0.3.23's take 3 first) and the card after; the attack-pass record lines (F9's full 10^5, F1's census) as they land. The night's clean lines: the interim zeros at 01:55 (71,292 seeds, 0 exhausted, 0 panics, max attempt 30); the kit's fingerprint equal on four platforms; PC 1's lock-free queue (the 9070 XT G1 14 of 14 and its v5 fingerprint, the family rows, the 5080 grid through the cleared helper, the CA4 rows); dfbd1e10 fully gated with the fast-time PASS. Three rules from the night for release-rules: a digest-moving release's minute is named from the pairs on the boxes, never from a clock the pairs have not met (the three lost floors); a lane that owns a gate in the critical path answers within ten minutes or its work is reassigned by main, not waited on (the two dark lanes); the move file carries the pair's miner sha into the pack gate (the puller rule above).
**build-1's seed, the reading and the repair (03:0x BST):** /home/build/dn3seed.log ends at 02:09:05 BST at DAA 29,732 mid-stream with no stop, shutdown or panic line (an igneumd stop writes "igneumd has stopped"), so it was killed abruptly; it ran under nohup from a shell, so no journal names the killer, and the OOM record needs sudo (the morning's build-server lane reads it); the datadir intact (13 GB). The node lane's DAA reads 25,169 (00:52:38) and 28,906 (01:55:09 BST) came from it while it lived; the chain read 32,659 at 02:57:50 BST from node1-dn3's JSON on 28670 (the observer 32,660), past 32,400 at about 02:53 as computed. Restarted on the shipper's word at 03:00:09 BST on the kept datadir with the 2720d8d2 hands pair (igneumd f2cf6a87; the same flags; setsid nohup with stdout appended behind a dated banner, the stop line in its own log; the restart script at /tmp/dn3seed-restart.sh for the morning's move): digest ba75bf6f, object version 7, the N15 lines (resumed from the exec snapshot at tip 11,901, the records continuous), DAA 32,905 in step with node1-dn3, 22 peers at 03:02 (20 inbound, the fleet's dials returning; 38 before the kill). A repair of the hub the fleet dials, not a release change; the DAA reads come from 27632 again. dn3-floor-cut.sh (the fourth re-cut from a named minute: the DAA from 28670, the floor = the publish DAA plus 7,200 to the next 3,600, the commit, both mirrors, the gate set, about 20 minutes plus the fast-time pair's 14; the latest minute before the heights move is about 12:50 BST) stands ready for main's minute.
**The attack-pass record lines on the frozen class v5 1c420786 (pairing e5a4ac5978462156), both PASS on the full 10^5:** F9 exhaustion (the last chunk 03:34:54 BST): 100,000 of 100,000 seeds through the chain path, 0 exhausted, 0 panics, 0 past attempt 31, max attempt 30, first-draw acceptance 0.3145, mean attempt index 2.185 (3.185 draws per seed); F1 redundancy (census.csv 05:20 BST): 100,000 of 100,000 programs, saved max 4.688 percent (0 over 5, 0 over 10), differential mismatches 0, verifier mismatches 0, 0 panics. Nothing from the attack lane holds the 0.3.24 move; the interim at 01:55 BST was the gate as main set it. **PC 1 overnight (the hash lane, 03:4x to 05:1x BST):** the 5090 floor2 pass ended by its 45 min cap with the rows to 400 MHz (every fingerprint matched; the card left at the 400 lock until the restore job at 03:46 BST brought it back to 2,880 MHz); the CA4 sparse runs read base rows (the research lane's second exe ran base on every variant, its third failed NVRTC on the card, "identifier d is undefined"; the finding with that lane); the microbench exe exited at once (rerun as -b after the packs); the packs job fell on a read-only $pid (fixed, rerun as -b from 05:12 BST); the 5080 Ember tune, the 9070 XT tune pass and the hot table follow; PC 2 held all night for main's word ("PC 2 clear" never given, take 3 never ran).
## 21. The core-clock knob in 0.3.24 and the measured Ember line (07:1x BST, 8 October)
**The knob:** the hash lane's 74585c91 (main's order of 7 October: the Ember clock ladder continues below 45 percent in 100 MHz steps to a 20 percent floor; the search stops at the knee, the first row more than the tolerance under the cap point's rate, or on a faulted row, the fingerprint check; the best MH per watt within tolerance is the point; lock_result and the card's lock_* fields for the UI; tests known-failed first on a fake helper) sat on the mirror's master, not on the release line; cherry-picked onto release-0.3.24 as e181f497 with ember.rs resolved as the union (the knob's floor and fine step beside the efficient-point ceiling of 7 October: EFFICIENT_W, DEFAULT_CAP_PCT, power_ceiling), the plan-count test updated to the knob's ladder on the 5090 (1 + 6 + 7; b6e2845f); app gate GREEN on build-1 (294 + 35 + 8), pre-push 60; the DMG re-cut under the lock on b6e2845f; the UI lane's drawing of the lock fields asked onto that tip. The rule: the knob never sets a lock below the knee without the user's own choice.
**The measured Ember line (the Counter lane, read on PC 1 overnight):** the installed app's stock power-limit climb on the RTX 5080 lands at 60.3 MH/s at 123 W (0.489 MH/W, clock_cap 2,936; run-ca3-pc1-ember-5080-20261007 at 07:05 BST, the app's own tune complete before the script's cast fault), while the clock-lock grid on the same card gives 71.1 MH/s at 103.7 W at the 1,000 MHz lock (0.686 MH/W) and 71.2 at 146.6 W at 1,100 on class v4 (0.486) (run-ca3-pc1-v4-eff-5080-20261007-d at 02:54 BST), so the core-clock lock is worth about 40 percent more per watt and 18 percent more rate than the climb alone on the 5080; on the 5090 the knob's reference rows are class v4 at 1,200 MHz (133.8 MH/s at 305 W, 0.439) and class v3 at 1,300 (134.6 at 223 W, 0.603), the knee at 1,300 on both, measured four times (docs/bench-log.md, the 7 to 8 October entry). The 9070 XT tune row follows. Nothing else changes in the cut: the pin dfbd1e10 and the kit e6c088bb stand; the move on main's morning minute.

View file

@ -206,23 +206,25 @@
"card": "AMD Radeon RX 9070 XT (16 GB)",
"generator": "v2",
"mh_s": 18.92,
"watts": 199,
"mh_per_w": 0.095,
"watts": 202,
"mh_per_w": 0.093,
"miner": "igneum-worker-opencl bench (installed worker), class v3 control",
"date": "2026-10-06",
"source": "Counter ASIC 3.0 status: item 8, the RX 9070 XT rows (job run-ca3-pc1-amd-g1-shadow-20261006); watts from the app's telemetry read of 5 October 2026 (the job's ADLX sample parsed 0 rows)",
"date": "2026-10-08",
"source": "Counter ASIC 3.0 status: the 9070 XT Ember tune (run-ca3-pc1-ember-9070-20261007, 06:11Z, app 0.3.20) for the watts; item 8's RX 9070 XT rows (the G1 ladder, 6 October 2026) for the rate",
"by": "measured by the team",
"note": "the card sits at 87 to 95 percent of its dependent random-read ceiling (2.42 to 2.68 G loads/s), in a Thunderbolt enclosure; AMD OpenCL 3683.0",
"note": "the watts are the app's own power reading at the stock point, no external sampler (8 October 2026, the status row): the Ember tune on this card measured the stock point and stopped, 0.3.20 having no AMD knob (power_pct 0, clock_cap 0, limit 0.0 W)",
"v4_cost": "+2 percent of rate (19.29 against 18.92 MH/s) at 102,100 ops per hash, watts owed, measured 6 October 2026",
"tuned": "stock, bench only (the AMD tune pass runs 7 October: set only if the ADLX tune line reads, else measure-only)",
"tuned": "no lever: the app has no AMD knob today (an AMD core-clock knob through rocm-smi or ADL is the path, a morning item)",
"driver_os": "Adrenalin 26.9.2, Windows 11",
"hive": {
"core_mhz": null,
"mem_mhz": null,
"pl_w": null,
"label": "stock (no measured tune point; the 5080 and 9070 XT passes and the class v4 efficiency pass land theirs when read)"
"label": "stock (no lever on AMD in 0.3.20; no measured tune point)"
},
"group": "buy"
"group": "buy",
"mh_s_display": "18.9 (18.8 to 19.2 on the G1 ladder)",
"watts_display": "202 stock"
},
{
"card": "Apple M5 Max (40 GPU cores, Metal)",

File diff suppressed because one or more lines are too long

View file

@ -17,8 +17,8 @@ function Summary([string] $status, [hashtable] $extra) { $o = [ordered]@{ job =
$jobDir = $env:IGNEUM_JOB_DIR; if (-not $jobDir) { $jobDir = Join-Path $env:TEMP 'igneum-ca4-mb' }
New-Item -ItemType Directory -Force -Path $jobDir | Out-Null
$jobs = Split-Path $jobDir
$packs = Join-Path (Join-Path $jobs 'fetch-ca4-hot-packs-20261008') 'packs'
if (-not (Test-Path $packs)) { "RESULT error packs missing at $packs (the fetch job fetch-ca4-hot-packs-20261008 runs first)"; Summary 'failed' @{ error = 'packs missing' }; exit 2 }
$packs = Join-Path (Join-Path $jobs 'fetch-ca4-hot-packs-20261008b') 'packs'
if (-not (Test-Path $packs)) { "RESULT error packs missing at $packs (the fetch job fetch-ca4-hot-packs-20261008b runs first)"; Summary 'failed' @{ error = 'packs missing' }; exit 2 }
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
if (-not $inst) { 'RESULT error no installed igneum-worker-cuda.exe'; Summary 'failed' @{ error = 'no worker' }; exit 2 }
$exe = Join-Path $inst 'igneum-worker-cuda.exe'
@ -82,7 +82,7 @@ function RunState([string] $state) {
"RESULT packs-state $state start $(Stamp) $((& $smi -i $idx --query-gpu=clocks.sm,clocks.applications.graphics,power.draw --format=csv,noheader 2>&1 | Out-String).Trim())"
foreach ($pk in $packOrder) { foreach ($variant in $variants) {
$d = Join-Path $packs $pk
$cls = ''; $pid = ''; try { $pj = Get-Content -LiteralPath (Join-Path $d 'program.json') -Raw | ConvertFrom-Json; $cls = [string] $pj.load_class; $pid = [string] $pj.program_id } catch { }
$cls = ''; $progId = ''; try { $pj = Get-Content -LiteralPath (Join-Path $d 'program.json') -Raw | ConvertFrom-Json; $cls = [string] $pj.load_class; $progId = [string] $pj.program_id } catch { }
$t0 = Get-Date
$extra = @(); if ($variant -ne 'base') { $extra = @('--variant', $variant) }
$lines = @(& $exe --bench --pack $d --batches 250 --batch-log2 24 --block-warps 1 @extra --device 0 2>&1 | ForEach-Object { "$_" })
@ -98,7 +98,7 @@ function RunState([string] $state) {
$wi = Mean ($w | Where-Object { $_.Count -ge 8 -and $_[6] -match '^\s*[\d.]+\s*$' } | ForEach-Object { $_[6] }); $wa = Mean ($w | Where-Object { $_.Count -ge 8 -and $_[7] -match '^\s*[\d.]+\s*$' } | ForEach-Object { $_[7] })
$eff = if ($watts -and $mhs) { [math]::Round([double]$mhs / $watts, 4) } else { '' }
$err = ''; if (-not $res -or $rc -ne 0) { $e = ($lines | Where-Object { $_ -match 'FAIL|error|MISMATCH|compile' } | Select-Object -First 1); $err = if ($e) { ($e -replace '\s+', ' ') } else { "exit $rc, no RESULT line" } }
"RESULT ca4-research pack=$pk dir=$d program_id=$pid class=$cls variant=$variant state=$state mhs=$mhs watts=$watts watts_instant=$wi watts_average=$wa mh_per_w=$eff sm_mhz=$sm mem_mhz=$mem util=$util fingerprint=$fp check=$check exit=$rc wall_s=$([int]($t1 - $t0).TotalSeconds) samples=$($w.Count) error=$err race=[$race]"
"RESULT ca4-research pack=$pk dir=$d program_id=$progId class=$cls variant=$variant state=$state mhs=$mhs watts=$watts watts_instant=$wi watts_average=$wa mh_per_w=$eff sm_mhz=$sm mem_mhz=$mem util=$util fingerprint=$fp check=$check exit=$rc wall_s=$([int]($t1 - $t0).TotalSeconds) samples=$($w.Count) error=$err race=[$race]"
Start-Sleep -Seconds 3
} }
"RESULT packs-state $state end $(Stamp)"

View file

@ -19,6 +19,14 @@ $jobs = Split-Path $jobDir
$exe = Join-Path (Join-Path $jobs 'fetch-ca4-microbench-exe-20261007') 'igneum-worker-cuda-ca4mb.exe'
if (-not (Test-Path $exe)) { "RESULT error exe missing at $exe (the fetch job fetch-ca4-microbench-exe-20261007 runs first)"; Summary 'failed' @{ error = 'exe missing' }; exit 2 }
"RESULT start $(Stamp) job=$jobName exe=$exe sha256=$((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) label=ca4-research"
# the loader (the research lane, 8 October 2026, 04:1x UTC): an exit 2 after one line is a loader line; NVRTC's two DLLs must sit beside the
# exe, and this exe was fetched into its own job folder; the installed worker's copies are copied beside it (the install folder is read, never
# written) and the exe runs from its own folder
$inst = @("$env:LOCALAPPDATA\Programs\Igneum Miner", "$env:ProgramFiles\Igneum Miner") | Where-Object { Test-Path (Join-Path $_ 'igneum-worker-cuda.exe') } | Select-Object -First 1
$exeDir = Split-Path $exe
if ($inst) { foreach ($dll in (Get-ChildItem -Path $inst -Filter 'nvrtc*.dll' -ErrorAction SilentlyContinue)) { if (-not (Test-Path (Join-Path $exeDir $dll.Name))) { Copy-Item -LiteralPath $dll.FullName -Destination $exeDir; "RESULT dll copied beside the exe: $($dll.Name) sha256 $((Get-FileHash -Algorithm SHA256 $dll.FullName).Hash.ToLower())" } } }
"RESULT dlls beside the exe: $((Get-ChildItem -Path $exeDir -Filter '*.dll' -ErrorAction SilentlyContinue | ForEach-Object { $_.Name }) -join ', ')"
Set-Location $exeDir
$smi = (Get-Command nvidia-smi -ErrorAction SilentlyContinue).Source
if (-not $smi) { foreach ($p in @((Join-Path $env:ProgramFiles 'NVIDIA Corporation\NVSMI\nvidia-smi.exe'), (Join-Path $env:SystemRoot 'System32\nvidia-smi.exe'))) { if (Test-Path $p) { $smi = $p; break } } }
$row = @(& $smi --query-gpu=index,name,uuid,pci.bus_id --format=csv,noheader 2>&1 | ForEach-Object { "$_" }) | Where-Object { $_ -match 'RTX 5090' } | Select-Object -First 1
@ -73,9 +81,17 @@ function ParseUtc([string] $s) { try { return [datetime]::Parse($s, [cultureinfo
function RunState([string] $state) {
$env:CUDA_VISIBLE_DEVICES = $uuid
"RESULT microbench-state $state start $(Stamp) $((& $smi -i $idx --query-gpu=clocks.sm,clocks.applications.graphics,power.draw --format=csv,noheader 2>&1 | Out-String).Trim())"
$lines = @(& $exe --microbench --mb-seconds 60 --device 0 2>&1 | ForEach-Object { "$_" })
# the research lane's exact command (8 October 2026, 03:54Z: with `--device 0` added the exe exited 2 after one line at both states and the
# script kept only probe-matching lines, so the refusal text was lost); CUDA_VISIBLE_DEVICES selects the card; every line is echoed when the
# exe exits non-zero or prints fewer than five lines
# the argument parser's pack check runs before the microbench branch on this build (04:25Z: "--pack <dir> is required"); the probes read
# no pack, so the sub-version 3 kit's class v4 pack satisfies the check and changes nothing (the research lane's word, 04:4x UTC)
$packArg = Join-Path (Join-Path (Join-Path $jobs 'fetch-ca3-v4-sub3-pc1-20261007') 'packs') 'v4-devnet-epoch0'
if (-not (Test-Path $packArg)) { "RESULT error the pack for the argument check is missing at $packArg"; Summary 'failed' @{ error = 'pack argument missing' }; exit 2 }
$lines = @(& $exe --microbench --mb-seconds 60 --pack $packArg 2>&1 | ForEach-Object { "$_" })
$rc = $LASTEXITCODE
foreach ($l in $lines) { if ($l -match '^RESULT microbench') { "$l state=$state" } elseif ($l -match 'error|fail|compile|skipped') { "RESULT microbench-note state=$state $l" } }
if ($rc -ne 0 -or $lines.Count -lt 5) { foreach ($l in $lines) { "RESULT microbench-raw state=$state exit=$rc $l" } }
foreach ($l in $lines) { if ($l -match '^RESULT microbench') { "$l state=$state" } elseif ($l -match 'error|fail|compile|skipped|usage') { "RESULT microbench-note state=$state $l" } }
foreach ($l in ($lines | Where-Object { $_ -match '^RESULT microbench probe=' })) {
$name = ''; $g = ''; $ck = ''; $st = ''; $su = ''; $eu = ''
if ($l -match 'probe=(\S+)') { $name = $Matches[1] }; if ($l -match 'G_ops_s=(\S+)') { $g = $Matches[1] }; if ($l -match 'checksum=(\S+)') { $ck = $Matches[1] }; if ($l -match 'status=(\S+)') { $st = $Matches[1] }

View file

@ -86,7 +86,7 @@ function RunState([string] $state) {
"RESULT packs-state $state start $(Stamp) $((& $smi -i $idx --query-gpu=clocks.sm,clocks.applications.graphics,power.draw --format=csv,noheader 2>&1 | Out-String).Trim())"
foreach ($pk in $packOrder) {
$d = Join-Path $packs $pk
$cls = ''; $pid = ''; try { $pj = Get-Content -LiteralPath (Join-Path $d 'program.json') -Raw | ConvertFrom-Json; $cls = [string] $pj.load_class; $pid = [string] $pj.program_id } catch { }
$cls = ''; $progId = ''; try { $pj = Get-Content -LiteralPath (Join-Path $d 'program.json') -Raw | ConvertFrom-Json; $cls = [string] $pj.load_class; $progId = [string] $pj.program_id } catch { }
$t0 = Get-Date
$lines = @(& $exe --bench --pack $d --batches 250 --batch-log2 24 --block-warps 1 --device 0 2>&1 | ForEach-Object { "$_" })
$rc = $LASTEXITCODE
@ -101,7 +101,7 @@ function RunState([string] $state) {
$eff = if ($watts -and $mhs) { [math]::Round([double]$mhs / $watts, 4) } else { '' }
$err = ''; if (-not $res -or $rc -ne 0) { $e = ($lines | Where-Object { $_ -match 'FAIL|error|MISMATCH|compile' } | Select-Object -First 1); $err = if ($e) { ($e -replace '\s+', ' ') } else { "exit $rc, no RESULT line" } }
$tag = if ($pk -like 'mx8_shl256x27*') { ' construction=unsound-energy-reading-only' } else { '' }
"RESULT ca4-research pack=$pk dir=$d program_id=$pid class=$cls$tag state=$state mhs=$mhs watts=$watts watts_instant=$wi watts_average=$wa mh_per_w=$eff sm_mhz=$sm mem_mhz=$mem util=$util fingerprint=$fp check=$check exit=$rc wall_s=$([int]($t1 - $t0).TotalSeconds) samples=$($w.Count) error=$err"
"RESULT ca4-research pack=$pk dir=$d program_id=$progId class=$cls$tag state=$state mhs=$mhs watts=$watts watts_instant=$wi watts_average=$wa mh_per_w=$eff sm_mhz=$sm mem_mhz=$mem util=$util fingerprint=$fp check=$check exit=$rc wall_s=$([int]($t1 - $t0).TotalSeconds) samples=$($w.Count) error=$err"
Start-Sleep -Seconds 3
}
"RESULT packs-state $state end $(Stamp)"

View file

@ -30,6 +30,8 @@ Write-Output ('RESULT TUNE job_id=' + $jobId + ' card_match=' + $want)
$cards = @($s.mining.cards | Where-Object { $_.enabled -and ($_.vendor -eq 'nvidia' -or $_.vendor -eq 'amd') -and $_.name -notmatch 'Radeon\(TM\) Graphics' -and $_.name -match $want })
if ($cards.Count -eq 0) { Write-Output ('RESULT TUNE error=no enabled card matching ' + $want + ' in api/state (names: ' + (($s.mining.cards | ForEach-Object { $_.name }) -join ', ') + ')'); exit 1 }
Write-Output ('RESULT TUNE cards ' + (($cards | ForEach-Object { $_.key + ' (' + $_.name + ', ' + $_.vendor + ', sweep_state ' + $_.sweep_state + ', line "' + $_.tune_line + '")' }) -join ' | '))
# every card's stored curve from the last tune (the 06:0x UTC 5080 run's rows were lost to a cast fault in this script's curve line; the app keeps them)
foreach ($c0 in $s.mining.cards) { if ($c0.tune_curve -and $c0.tune_curve.Count -gt 0) { Write-Output ('RESULT TUNE stored-curve key=' + $c0.key + ' source=' + $c0.tune_source + ' line="' + $c0.tune_line + '" rows=' + (($c0.tune_curve | ForEach-Object { "$($_.clock_mhz)MHz/$($_.mem_mhz)mem/$($_.power_pct)pct/$($_.limit_w)W/$($_.watts)W/$($_.mhs)MHs/$($_.eff)MHW/$($_.mark)" }) -join ' | ')) } }
foreach ($c in $cards) { Write-Output ('RESULT TUNE before key=' + $c.key + ' power_w=' + $c.power_w + ' power_pct=' + $c.power_pct + ' clock_cap=' + $c.clock_cap_mhz + ' mhs=' + $c.mhs + ' line=' + $c.tune_line) }
# the climb on (NVIDIA cards; an AMD card has no clock knob and runs its ladder regardless)
if (-not (Post 'api/tune/goal' '{"goal":"balanced","climb":true}')) { Write-Output 'RESULT TUNE error=tune_goal_refused'; exit 1 }
@ -62,10 +64,12 @@ foreach ($c in $order) {
}
}
$st = State
if ($st) { $cc = $st.mining.cards | Where-Object { $_.key -eq $c.key } | Select-Object -First 1; if ($cc) { Write-Output ('RESULT TUNE after key=' + $c.key + ' power_w=' + $cc.power_w + ' power_pct=' + $cc.power_pct + ' clock_cap=' + $cc.clock_cap_mhz + ' line=' + $cc.tune_line + ' source=' + $cc.tune_source + ' note=' + $cc.sweep_note); if ($cc.tune_curve) { Write-Output ('RESULT TUNE curve key=' + $c.key + ' ' + (($cc.tune_curve | ForEach-Object { ($_.clock_mhz) + 'MHz/' + ($_.mem_mhz) + 'mem/' + ($_.power_pct) + '%:' + $_.mhs + 'MH/s ' + $_.watts + 'W ' + $_.eff + ' ' + $_.mark }) -join ' | ')) } } }
if ($st) { $cc = $st.mining.cards | Where-Object { $_.key -eq $c.key } | Select-Object -First 1; if ($cc) { Write-Output ('RESULT TUNE after key=' + $c.key + ' power_w=' + $cc.power_w + ' power_pct=' + $cc.power_pct + ' clock_cap=' + $cc.clock_cap_mhz + ' line=' + $cc.tune_line + ' source=' + $cc.tune_source + ' note=' + $cc.sweep_note); if ($cc.tune_curve) { Write-Output ('RESULT TUNE curve key=' + $c.key + ' ' + (($cc.tune_curve | ForEach-Object { "$($_.clock_mhz)MHz/$($_.mem_mhz)mem/$($_.power_pct)pct/$($_.limit_w)W/$($_.watts)W/$($_.mhs)MHs/$($_.eff)MHW/$($_.mark)" }) -join ' | ')) } } }
}
# the climb switch back to what it was
Post 'api/tune/goal' ('{"climb":' + $(if ($climbBefore) { 'true' } else { 'false' }) + '}') | Out-Null
# a card the app cannot steer (AMD in 0.3.20: no power limit, no clock cap) logs no progress lines and stores one baseline row; the stored curve is the count, not the log
if ($rows -eq 0) { $st2 = State; if ($st2) { foreach ($c2 in $cards) { $cc2 = $st2.mining.cards | Where-Object { $_.key -eq $c2.key } | Select-Object -First 1; if ($cc2 -and $cc2.tune_curve) { $rows += $cc2.tune_curve.Count } } } }
Write-Output ('RESULT TUNE climb_restored=' + $climbBefore + ' rows=' + $rows)
$smi = Join-Path $env:ProgramFiles 'NVIDIA Corporation\NVSMI\nvidia-smi.exe'
if (-not (Test-Path $smi)) { $smi = Join-Path $env:SystemRoot 'System32\nvidia-smi.exe' }

View file

@ -26,20 +26,20 @@ case "$STEP" in
family-kit) $P add --kind fetch --target ae432dc7 --id fetch-ca3-pc1-amd-family-20261007 --file "$KITS/igneum-ca3-pc1-amd-family-kit-20261007.zip" --dir jobs --extract --title "CA3 PC 1: the family probe kit (7 Oct)" --expires-hours 36 $DEPLOY ;;
eff-5090) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v4-eff-5090-20261007-b --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 75 --title "CA3 PC 1: class v4 efficiency pass on the RTX 5090 (clock locks, v4 against v3, the card alone)" --expires-hours 36 $DEPLOY ;;
eff-5090-floor) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v4-eff-5090-floor-20261007 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 60 --title "CA3 PC 1: class v4 efficiency pass 2 on the RTX 5090 (1,400 MHz down to the driver floor, v4 against v3, the card alone)" --expires-hours 36 $DEPLOY ;;
ca4-exe) $P add --kind fetch --target ae432dc7 --id fetch-ca4-sparse3-exe-20261008 --file "$KITS/ca4/igneum-worker-cuda-ca4sparse3.exe" --dir jobs --title "CA4 research: the SM-sparse worker exe with the race fix (scratch build, counter-asic-4, 00:46 UTC)" --expires-hours 36 $DEPLOY ;;
ca4-sparse) $P add --kind run --target ae432dc7 --id run-ca4-pc1-ca4sparse-5090-20261008 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 70 --title "CA4 research: SM-sparse worker variants on the RTX 5090, unlocked and at the 1,300 MHz knee, v4 and v3 (the card alone)" --expires-hours 36 $DEPLOY ;;
ca4-exe) $P add --kind fetch --target ae432dc7 --id fetch-ca4-sparse5-exe-20261008 --file "$KITS/ca4/igneum-worker-cuda-ca4sparse5.exe" --dir jobs --title "CA4 research: the SM-sparse worker exe with the capture fix (scratch build, counter-asic-4, 03:55 UTC)" --expires-hours 36 $DEPLOY ;;
ca4-sparse) $P add --kind run --target ae432dc7 --id run-ca4-pc1-ca4sparse-5090-20261008-c --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 70 --title "CA4 research: SM-sparse worker variants on the RTX 5090, unlocked and at the 1,300 MHz knee, v4 and v3 (the card alone)" --expires-hours 36 $DEPLOY ;;
eff-5090-floor2) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v4-eff-5090-floor2-20261007 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 45 --title "CA3 PC 1: class v4 efficiency pass 3 on the RTX 5090 (1,100 MHz down to the driver floor, v4 against v3, the card alone)" --expires-hours 36 $DEPLOY ;;
eff-5080) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v4-eff-5080-20261007-d --cards-off "nvidia:NVIDIA GeForce RTX 5080" --script tools/ca3-v4-amend/pc1-v4-efficiency.ps1 --timeout-minutes 75 --title "CA3 PC 1: class v4 efficiency pass on the RTX 5080 (clock locks, v4 against v3, the card alone)" --expires-hours 36 $DEPLOY ;;
helper-diag2) $P add --kind run --target ae432dc7 --id run-ca3-pc1-helper-diag2-20261007 --script tools/ca3-v4-amend/pc1-helper-diag2.ps1 --timeout-minutes 5 --title "PC 1: Power Helper diagnostic 2 (the task exe against the running app, scheduler history, crash log, a 20 s unelevated helper probe)" --expires-hours 12 $DEPLOY ;;
ca4-mb-exe) $P add --kind fetch --target ae432dc7 --id fetch-ca4-microbench-exe-20261007 --file "$KITS/ca4/igneum-worker-cuda-ca4mb.exe" --dir jobs --title "CA4 research: the microbench worker exe (scratch build)" --expires-hours 36 $DEPLOY ;;
ca4-mb) $P add --kind run --target ae432dc7 --id run-ca4-pc1-microbench-5090-20261007 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-microbench.ps1 --timeout-minutes 60 --title "CA4 research: per-block micro-benchmark on the RTX 5090, unlocked and at the 1,300 MHz knee (the card alone)" --expires-hours 36 $DEPLOY ;;
ca4-mb) $P add --kind run --target ae432dc7 --id run-ca4-pc1-microbench-5090-20261008-c --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-microbench.ps1 --timeout-minutes 60 --title "CA4 research: per-block micro-benchmark on the RTX 5090, unlocked and at the 1,300 MHz knee (the card alone)" --expires-hours 36 $DEPLOY ;;
ca4-packs-kit) $P add --kind fetch --target ae432dc7 --id fetch-ca4-packs-20261007 --file "$KITS/igneum-ca4-packs-20261007.zip" --dir jobs --extract --title "CA4 research: the two prototype classes, both per-load exports, plus the mx8 control" --expires-hours 36 $DEPLOY ;;
ca4-packs) $P add --kind run --target ae432dc7 --id run-ca4-pc1-packs-5090-20261007 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-packs.ps1 --timeout-minutes 60 --title "CA4 research: the per-load shadow and the int8 tile prototypes against mx8 and sh256x27 on the RTX 5090 (the card alone)" --expires-hours 36 $DEPLOY ;;
ca4-packs) $P add --kind run --target ae432dc7 --id run-ca4-pc1-packs-5090-20261008-b --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-packs.ps1 --timeout-minutes 60 --title "CA4 research: the per-load shadow and the int8 tile prototypes against mx8 and sh256x27 on the RTX 5090 (the card alone)" --expires-hours 36 $DEPLOY ;;
v5-kit) $P add --kind fetch --target ae432dc7,1ccfe586 --id fetch-ca3-v5-kit-20261007 --file "$KITS/v5kit/packs-ca3-v5-20261007T221001Z.zip" --dir jobs --extract --title "Class v5 kit (the v5 lane, 7f58af97): packs, workers, SHA256SUMS" --expires-hours 36 $DEPLOY ;;
v5-amd) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v5-amd-bench-20261007 --script tools/ca3-v4-amend/pc1-v5-amd-bench-20261007.ps1 --timeout-minutes 12 --title "Class v5 kit fingerprint on the RX 9070 XT (beside the miners; the v5 lane's script)" --expires-hours 36 $DEPLOY ;;
v5-intel) $P add --kind run --target 1ccfe586 --id run-ca3-pc2-v5-intel-bench-20261007 --script tools/ca3-v4-amend/pc1-v5-intel-bench-20261007.ps1 --timeout-minutes 12 --title "Class v5 kit fingerprint on PC 2's Arc B580 eGPU (the v5 lane's script, class-v5 a4b08245; beside the miners; only on the shipper's PC 2 clear)" --expires-hours 36 $DEPLOY ;;
hot-kit) $P add --kind fetch --target ae432dc7 --id fetch-ca4-hot-packs-20261008 --file "$KITS/igneum-ca4-hot-packs-20261008.zip" --dir jobs --extract --title "CA4 research: the eight 5 October hot-table packs plus the mx8 control" --expires-hours 36 $DEPLOY ;;
hot-ldcs) $P add --kind run --target ae432dc7 --id run-ca4-pc1-hot-ldcs-5090-20261008 --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-hot-ldcs.ps1 --timeout-minutes 75 --title "CA4 research: the hot table with the streaming hint (base against ldcs) on the RTX 5090, unlocked and at the knee (the card alone)" --expires-hours 36 $DEPLOY ;;
hot-kit) $P add --kind fetch --target ae432dc7 --id fetch-ca4-hot-packs-20261008b --file "$KITS/igneum-ca4-hot-packs-20261008b.zip" --dir jobs --extract --title "CA4 research: the eight 5 October hot-table packs plus the mx8 control" --expires-hours 36 $DEPLOY ;;
hot-ldcs) $P add --kind run --target ae432dc7 --id run-ca4-pc1-hot-ldcs-5090-20261008b --cards-off "nvidia:NVIDIA GeForce RTX 5090" --script tools/ca3-v4-amend/pc1-ca4-hot-ldcs.ps1 --timeout-minutes 75 --title "CA4 research: the hot table with the streaming hint (base against ldcs) on the RTX 5090, unlocked and at the knee (the card alone)" --expires-hours 36 $DEPLOY ;;
amd-g1) $P add --kind run --target ae432dc7 --id run-ca3-pc1-v4-sub3-amd-g1-20261007 --cards-off amd:gfx1201 --script tools/ca3-v4-amend/pc1-v4-sub3-amd-g1.ps1 --timeout-minutes 45 --title "CA3 PC 1: G1 on the sub-version 3 kit and the AMD ladder on the 9070 XT (the card alone)" --expires-hours 36 $DEPLOY ;;
family) $P add --kind run --target ae432dc7 --id run-ca3-pc1-amd-family-20261007-e --cards-off amd:gfx1201 --script tools/ca3-v4-amend/pc1-amd-family-20261007.ps1 --timeout-minutes 20 --title "CA3 PC 1: item 6 family step costs on the 9070 XT (run e, the card alone)" --expires-hours 36 $DEPLOY ;;
tune-5080) $P add --kind run --target ae432dc7 --id run-ca3-pc1-ember-5080-20261007 --script tools/ca3-v4-amend/pc1-ember-card.ps1 --timeout-minutes 35 --title "CA3 PC 1: Ember Tune on the RTX 5080 (the installed app tunes the one card)" --expires-hours 36 $DEPLOY ;;

View file

@ -70,10 +70,19 @@ $exe = Join-Path $inst 'igneum-worker-cuda.exe'
if ($ca4) {
# the research lane's exe with the race fix (counter-asic-4 after 2d0013d1, 00:46 UTC; the first exe parsed --variant and never applied it:
# every row of the 00:22Z run read "variant base"); the rows carry its variant= and sparse_blocks= fields and any variant_not_installed line
$exe = Join-Path (Join-Path $jobs 'fetch-ca4-sparse3-exe-20261008') 'igneum-worker-cuda-ca4sparse3.exe'
if (-not (Test-Path $exe)) { "RESULT error the research worker is missing at $exe (the fetch job fetch-ca4-sparse3-exe-20261008 runs first)"; Summary 'failed' @{ error = 'research worker missing' }; exit 2 }
$exe = Join-Path (Join-Path $jobs 'fetch-ca4-sparse5-exe-20261008') 'igneum-worker-cuda-ca4sparse5.exe'
if (-not (Test-Path $exe)) { "RESULT error the research worker is missing at $exe (the fetch job fetch-ca4-sparse5-exe-20261008 runs first)"; Summary 'failed' @{ error = 'research worker missing' }; exit 2 }
}
"RESULT worker $exe sha256 $((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower()) bytes $((Get-Item $exe).Length)"
if ($ca4) {
# the research lane's card-free check of the race order (ca4sparse4, 03:04 UTC): the run would build "variants=2 names=base,sp43-w32"
$d0 = Join-Path $packs 'v4-devnet-epoch0'
$lr = @(& $exe --bench --pack $d0 --batches 250 --batch-log2 24 --block-warps 32 --variant sp43-w32 --list-race 2>&1 | ForEach-Object { "$_" })
foreach ($l in $lr) { if ($l -match 'list-race|error|compile') { "RESULT list-race $l" } }
if (-not ($lr | Where-Object { $_ -match 'variants=2' })) { 'RESULT ca4 STOP: the card-free race check does not list the sparse variant (variants=2 absent); the sparse rows would be base again, nothing run'; Summary 'failed' @{ error = 'sparse variant not in the race order' }; exit 4 }
# ca4sparse5 (03:55 UTC): --list-race also compiles the rewritten kernel through the PC's nvrtc; a compile error or compiled=0 stops the job with nothing run
if (($lr | Where-Object { $_ -match 'compiled=0|NVRTC_ERROR|error: ' })) { 'RESULT ca4 STOP: the card-free check does not compile the rewritten kernel on this PC (the list-race lines above carry the error); nothing run'; Summary 'failed' @{ error = 'rewritten kernel does not compile' }; exit 4 }
}
$appExe = Join-Path $inst 'igneum-app.exe'
if (Test-Path $appExe) { try { "RESULT app-version $((& $appExe --version 2>&1 | Out-String).Trim())" } catch { } }