72 KiB
Counter ASIC 3.0: status
Started 6 October 2026, 07:15 UTC, on the project lead's order "run it all"; closed 09:5x UTC with every item measured on the Mac and the 5090 and every AMD row owed (PC 1 not released). Final tree: ca3-coord, 45 commits over the base, no-conflict-markers.sh clean, observer tests 10 of 10, igneum-pow 96 of 96 with the pinned v2 and v3 packs byte for byte. Coordinator worktree igneum-wt-ca3-coord, branch ca3-coord from 50df751 (ca2-coord 9206aea merged with master ddfcdac). The plan: docs/plans/counter-asic-3.md. The audit behind it: docs/analysis/asic-resistance-history.md. The 2.0 record: docs/plans/counter-asic-2-status.md, docs/bench-log.md "Counter ASIC 2.0, the numbers", docs/analysis/chip-model-v3.md. Class v3 is live on the devnet since DAA 154,800 (crossing 03:51:42Z, verdict PASS 05:21:48Z). Nothing in this run publishes to the devnet; what passes is a class v4 candidate behind program_class_v4_activation_daa under the six gates and the two-publish rollout of 2.0.
The test every result is judged against (the project lead, 6 October): a chip maker must have to build a better GPU than NVIDIA to launch an ASIC, the way Bitmain's Antminer X5 reached CPU parity per joule only at many times the price. The binding rule from 2.0: the hash stays bound by dependent random memory reads.
1. Machines and rules in force today
| Machine | State | Rule |
|---|---|---|
| Mac M5 Max | mining paused; Metal worker free | measurements under with-lock.sh measure only |
| PC 2 (1ccfe586, RTX 5090) | HELD for the 0.3.14 shipper's Rust suite from 17:5x UK (main): nothing of 3.0 goes to PC 2 until the shipper reports the suite done (the P2 G6 run waits on it). Earlier: RELEASED at 08:43:24Z after the three jobs (derive 08:26 to 08:29Z, shadow 08:29 to 08:41Z, family 08:41 to 08:43Z, all exit 0; prover ON, app untouched); then the prover-floor agent, then the 0.3.12 release engineer's build. Before that: CLEAR at 08:24:27Z (the clear file carries the proving agent's constraints: prover left ON, no quit or restart, /opt/igneum, /opt/igneum-segal and settings.json untouched); the three 3.0 jobs run one at a time through the mkdir lock (items 8, 2, 6); "PC 2 released" to the proving agent after the last; the prover-floor agent next. Before that: the proving agent's jobs segments-pc2-pv1 runs b and c (run b claimed nothing on a PowerShell key bug; run c from about 07:53Z); "PC 2 clear" expected about 08:35Z; then three 3.0 jobs one at a time (items 2, 6, 8); the prover-floor agent queues after "PC 2 released" |
nothing published until "PC 2 clear"; one job at a time (/tmp/igneum-devnet/pc2-ca3.lock); released by message when done |
| PC 1 (ae432dc7, RTX 5090 + RTX 4070 + RX 9070 XT) | FREE at 18:01:05Z after the queue (reset, family runs a to e, G2, derive, the 4070 ladder, the watts job that failed and left the 9070 XT off, the restore and the flag job); handed to main for the project lead's 0.3.14 click and the Ember window. Earlier: the project lead's desk; the AMD queue opens on "Ember closed" (reset, family, G2, derive, the 4070, watts, about 25 min), then the Ember agent's 6-minute window, then the 0.3.14 update on the project lead's click (after the last job, not between jobs: an update wipes the jobs folder and the kits) | not used today; every AMD row OWED |
Standing rule from main (17:5x UTC; CLAUDE.md on master 630b537, docs/plans/build-server.md): from the next build, every Linux and Windows cargo build and every Linux test suite from the 3.0 lanes runs on igneum-build-1 through tools/build-remote.sh and tools/cross-remote.sh from the worktree's crate directory (artefacts in target-remote/, IGNEUM_AGENT=ca3-<lane>, one slot, a 2 h cap; the clean node build 87 s there against 9 to 18 min on the Mac). PC 2 keeps only GPU and Windows-runtime jobs (G6's engine test on the card stays a PC job; the crate suites move). Passed to the node lane, the hash lane and the PC 1 worker.
2. The items
| # | Item | Worker branch | State | Close |
|---|---|---|---|---|
| 1 | Partial-store chip and the time-memory curve | ca3-analysis | CLOSED, merged (71df794) | verdict OVER 2x: the f = 1 chip (the dataset stored in DRAM, nothing recomputed) is 5.1x per joule on GDDR7 and 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent, $2.8 per MH/s against the 5090's $14.7; the curve is monotone toward f = 1, so the partial-store chip is never built; the mixer and item 2 do not touch it; chip-model-v3.md section 5 |
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | CLOSED, merged (acb96ee, dd5041b, bdc07d3, 54bc188); 5090 job run-ca3-derive-pc2-20261006 exit 0 in 126 s | GO as reserve R0; NO-GO for genesis-live at dr736 (verifier 4.9 ms per unit on one M5 Max core, about 12 ms on a 2019-class core by the 2.5x rule, over the gate; dr368 2.69 ms passes both); urgency LOW after item 1 (the stored-dataset chip derives no item) |
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | CLOSED, merged (c0642af, merge ed06814) | detector.mjs + 7 of 7 tests + one observer hook, dry run quiet on the devnet (max correlation 0.53 against the 0.8 edge; excess spread 0 to 5.2 percent); funding.md rule 5: bounty escrowed and benchmark live before daily issuance crosses USD 20,000 a day; epoch-length.md sections 11 and 12: signal trigger M = 6 windows; FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the bank-bound ceiling (unmeasured); layer 9 ranks above layer 7. The live observer is NOT restarted yet: the write path is untested; one restart after item 7's hook merges, then the first live detector row is recorded here |
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | CLOSED, merged (b85f0f1, merge 820f4d4); 5090 job run-ca3-family-pc2-20261006 exit 0 in 36 s, the card to itself | proposed reserve order R1 byte permute, R2 popcount and clz, R3 indexed shuffle, R4 bit-field extract, R5 variable shifts, R6 select, R7 andn, R8 mm8 (docs/plans/counter-asic-3-reserve.md, a decision for the project lead, not in the spec); vendor-share.mjs with tests and one observer hook; repro.md carried from repro-bench (dda9fa3) with section 8: today's devnet NVIDIA 0.986 of blue blocks, Intel 0.014, coverage 1.00 at 135.8 MH/s with PC 1 off |
| 8 (added by item 1's finding, coordinator 08:xx UTC) | Program work in the latency shadow: the hash rate, watts and verifier cost at N = 50,000, 100,000, 200,000 ops per hash on the M5 Max and the 5090; the 5090 power-cap rows | ca3-shadow | CLOSED, merged (70eaf06, merge 12d8a93); 5090 job run-ca3-shadow-pc2-20261006 (lock 08:29:22 to 08:41:03Z), the card empty | GO as a class v4 candidate at mx8+sh256x27 (100,000 ops per hash): the chip's per-joule edge over the 5090 falls from 5.6x to 2.1x at k = 1 for 0.2 percent of the 5090's rate and 1.5 percent of the Mac's; NO-GO above 130,000 ops or a block over 256 instructions; the chip question is decided by k, the chip core's energy per op against the 5090's measured 11 pJ: over 2x only at k under 0.5 |
3. Measured numbers
Item 1 (analysis, no new measurement; every chip figure is arithmetic on cited memory figures, approximate where marked)
| Row | Rate against the 5090's 136.1 MH/s | Per joule against 2.40 microjoules per hash | $ per MH/s | Binding |
|---|---|---|---|---|
| f = 0 on-die recompute chip (sections 1 to 3; 9,360 ops per item hoisted, 10,512 unhoisted) | 0.31x bare, 0.92x with the 3x factor (0.27x / 0.82x unhoisted) | 1.86x (1.75x unhoisted; 1.3x to 2.4x over the on-die read energy) | $16.8 | compute |
| f = 1 on GDDR7 (the 5090's own memory system, 21.3 G reads/s activate ceiling) | 1.22x | 5.1x | $2.8 | memory |
| f = 1 on one HBM3 stack | 0.61x | 7.5x | $6.6 | memory |
| f = 1 on eight HBM3 stacks | 4.9x | 9.2x | $4.0 | memory |
| f = 0.25 to 0.75, both memories | between the ends and worse than both on $ per MH/s |
The whole case for the f = 1 chip: the 5090's memory system draws about 55 W of its 326 at the hash (17 percent, approximate); the rest is the GPU spinning on loads at 0.15 percent of its integer budget. The op count per mixer application, counted from memhard.rs: 144 as written, 128 with the RC and rk adds hoisted; 72 x 128 + 144 = 9,360 per item, which is the spec's "about 130" x 72 exactly; the chip model's rows stand at 9,360 and are given at 10,512 beside them.
What moves the f = 1 rows (chip-model-v3.md section 5.7): not the dataset size (one HBM3 stack holds 24 GB, the schedule reaches 4 GiB at year 4), not the read width (the decision to stay at 4 B stands), not the chain length; only (a) the honest card's watts at the hash (a 5090 holding 136 MH/s at a 250 W cap reads 3.9x on GDDR7, at 200 W 3.2x) and (b) program work in the latency shadow (the card hides 512 ops per hash behind 128 reads and could hide about 330,000 before compute binds; at N = 330,000 the model reads 1.85x at a chip core as efficient as the GPU's ALU, 2.5x at 1.5x worse), which is the design item this run adds (item 8 below) and the one that answers the project lead's test directly: a chip must carry the memory system AND the ALU budget.
Item 2, the Mac rows (interim, ca3-derive dd5041b; measure lock, load average 4.5 to 5.3; the 5090 rows queued on PC 2; the 9070 XT OWED)
| Class | Ops per item (chip, RC + rk hoisted / GPU as written / multiplies) | Verifier ms per unit, one M5 Max core (worst cold) | 2019-class core, 2.5x rule, approximate | Metal 1 GiB build | Hash rate, M5 Max | Bit-exact | Chip row (bare / at a 1.2x / 1.5x allowance; the 3x no longer applies) |
|---|---|---|---|---|---|---|---|
| v2 (x1) | 1,152 / 1,296 / 144 | 0.598 | 1.5 | 2.45x | |||
| v3 (x8), live | 9,216 / 10,368 / 1,152 (the acceptance floors) | 2.061 to 2.063 | 5.2 | 22.1 ms | 27.1 MH/s | yes | 0.31x / 0.92x with the 3x factor |
| dr368 (9 x 368 instructions) | about half of dr736 | 2.69 | 6.7, passes | ||||
| dr736 (9 x 736 instructions, the genesis-day draw: 9,992 / 10,659 / 1,461) | 4.875 to 4.944 (5.241) | 12, over the gate | 29.0 ms | 27.1 MH/s (equal) | yes (Metal and Apple OpenCL, fingerprint 50e3eaa779da4f1e) | 0.29x / 0.34x / 0.43x |
Verifier headroom under the 10 ms gate (bdc07d3, the budget item 8's shadow ops may spend; steady / worst cold): on this M5 Max core x8 7.9 / 7.8 ms, dr368 7.3 / 7.1, dr736 5.1 / 4.8; on the approximate 2019-class core x8 4.8 / 4.6, dr368 3.3 / 2.8, dr736 none (over by 2.2 / 3.1). cargo test -p igneum-pow on acb96ee under the build lock: 95 passed, 0 failed; the pinned v2 and v3 packs regenerate byte for byte.
The 5090 rows (job run-ca3-derive-pc2-20261006, 08:26:43 to 08:28:49Z, exit 0; the card was LOADED: the miner stayed up because the job posted the settings key without the device index, see corrections; ratios valid, absolutes not): bit-exact on CUDA for both dr736 packs (64 samples, 96 lanes, fingerprints 50e3eaa779da4f1e and 9553f6d5c667205a equal to the Mac's); 1 GiB build 42 / 32 ms against x8 40 and v2 46; rates 61 to 62 MH/s on every pack (the v2 control 62.3 against 136 unloaded); NVRTC compile 1,266 ms against x8's 164 ms, +1.1 s per pack because memhard.h's item function sits inside every hash-kernel and race-variant compile, so a once-a-day derivation module is a requirement of the class, not an option. Prover left ON, app untouched.
The RX 9070 XT rows (PC 1 job run-ca3-pc1-amd-derive-20261006, 17:23:56 to 17:29:13Z, exit 0, beside the miners so the absolutes are loaded-card figures and the ratios stand): bit-exact on AMD OpenCL for both dr736 packs (fingerprints 9553f6d5c667205a and 50e3eaa779da4f1e equal to the Mac's and the 5090's, two passes each, self-tests PASS); the OpenCL compile 2,070 to 2,092 ms per dr736 pack against 41 ms for mx8 (+2.0 s per pack, the AMD twin of NVRTC's +1.1 s: the once-a-day derivation module is a requirement on every vendor, and on a one-click AMD miner the shadow-free x8 pack compiles in 0.04 s where the derivation pack takes 2.1); the 1 GiB daily build 168 to 189 ms against 205 to 273 for mx8 in the same loaded session (a build the same size or smaller, inside the noise); the rate ratio to mx8 0.982 and 1.006 (no hash-rate cost). With these the derivation's bit-exactness is on all three vendors and its hash-rate cost is zero on all three.
Consequence: the derivation costs no hash rate on any card and 7 ms a day of build on the Mac (29 against 22 ms; the 5090 and the 9070 XT rows owed, the loaded-iGPU tier is the one to watch); it costs the verifier, and the verifier budget is shared with item 8's shadow ops under the one 10 ms gate, so the class v4 candidate is the pairing that fits, not either lever alone (both workers told).
Item 8, the Mac rows (interim, ca3-shadow a050a54; knob d4b7300; docs/analysis/latency-shadow-2026-10-06.md; Metal packbench, IOReport GPU + DRAM watts without root; the 5090 rows queued on PC 2; the 9070 XT OWED)
| N, shadow ops per hash | M5 Max MH/s (against the control 27.07 to 27.10) | Watts (GPU + DRAM) | Microjoules per hash | Verifier per warp, one M5 Max core | Reading |
|---|---|---|---|---|---|
| 512 (class v3 today) | 27.1 | 21 | 0.78 | 2.06 ms | the honest best per joule we own: 3x better than the 5090's 2.34 at 290 W |
| about 100,000 | -1.5 percent | 37 | 1.40 | 2.06 + 0.32 | still latency-bound |
| about 130,000 | -5 percent | the 2.0 rule's edge on this card | |||
| 151,000 | -7.3 percent | ||||
| 200,000 | -10 percent | compute binds | |||
| 331,000 | -21 percent | 40 | 1.88 | 2.06 + 0.56 (about 1.4 ms more on a 2019-class core) |
The verifier is not the constraint: 3.2 microseconds per 1,000 shadow instructions per warp, so no pairing of item 2's class with any N binds the gate before the cards do; the cards bind. Block size on Apple: a 64-instruction block +2.5 percent, 256 holds, 1,024 costs 17 percent at the same N. Marginal ALU energy on the Mac 6.9 pJ per counted op at N = 100,000. Chip side at N = 100,000 and a chip core equal to the GPU's ALU (k = 1): 1.4x over the M5 Max per joule, 2.7x over the 5090 on the model watts (5.0x today at 290 W with the shadow empty); k decides it. Power caps: the 5 October sweep re-used (floor 400 W, the cap never binds), no -pl steps; the PC 2 job carries the N ladder and -lgc clock rows at the four Ember caps (OWED if refused).
Consequence per tier (interim): the M5 Max at 21 W is the per-joule best honest miner we own (0.78 against the 5090's 2.34 microjoules), so against Apple silicon the stored-dataset chip of item 1 reads 1.7x per joule on GDDR7 (0.47 against 0.78) and 2.4x on one HBM3 stack, inside the "1 to 2x" band on GDDR7; per pound the Mac stays the worst tier (list price against 27 MH/s). Filling the shadow to about 100,000 ops costs an Apple miner 16 W more for 1.5 percent of rate (0.78 to 1.40 microjoules) and buys 1.4x against a chip at k = 1; the N the project can pick without any card we own losing over 5 percent is set by the Mac at about 130,000 until the 5090 and 9070 XT rows land. An NVIDIA owner's number is the PC 2 job; an AMD owner's is owed.
Item 8, the 5090 rows and the chip side (ca3-shadow 70eaf06; job run-ca3-shadow-pc2-20261006, the card confirmed empty by nvidia-smi compute-apps and the process list, the installed worker's --bench, nvidia-smi at 1 Hz, the app's 431 W cap; every pack bit-exact on Metal, CUDA, clang emulation and Apple OpenCL)
| N ops per hash | M5 Max MH/s (delta) | M5 Max W, microjoules | RTX 5090 MH/s (delta) | 5090 W, SM MHz, microjoules | Verifier ms per warp, one M5 Max core |
|---|---|---|---|---|---|
| 930 (class v3 today: 512 instructions, 384 ALU, about 1.83 counted ops each) | 27.08 | 21.0, 0.78 | 132.2 | 350, 3,040, 2.65 | 2.06 |
| 49,700 | 26.75 (-1.2%; a 64-instruction block +2.5%) | 31.1, 1.16 | 132.3 (+0.1%; 64-block +3.5%) | 425, 3,034, 3.21 | 2.21 |
102,100 (sh256x27, the candidate) |
26.67 (-1.5%) | 37.2, 1.40 | 131.95 (-0.2%) | 431 cap, 2,824, 3.27 | 2.23 |
| 150,800 | 25.10 (-7.3%) | 36.7, 1.46 | 131.75 (-0.3%) | 431, 2,427, 3.27 | 2.33 |
| 199,600 | 24.25 (-10.4%) | 38.2, 1.58 | 128.67 (-2.7%) | 431, 1,753, 3.35 | 2.43 |
| 330,700 | 21.39 (-21%) | 40.1, 1.88 | 86.39 (-35%, compute-bound at the capped clock, 28.6 T op/s) | 431, 1,834, 4.99 | 2.62 (worst cold 2.77) |
Where each card leaves the latency bound (the 5 percent rule): M5 Max about 130,000 ops; RTX 5090 about 210,000 at its 431 W cap (the cap binds from 102,100 ops and the governor lowers the clock); RX 9070 XT OWED. The verifier's law: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp, so 330,700 ops add 0.56 ms (about 1.4 ms on a 2019-class core): it fits every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none), and the node never binds before the cards. Marginal ALU energy per counted op: the 5090 10 to 13 pJ at its shipping clock (twice the 5.5 pJ the item 1 model assumed), the M5 Max 6.9 pJ. Power caps: -pl below 400 W cannot be set, the 5 October sweep rows re-used (302 to 316 W at every cap); the clock rows are OWED (nvidia-smi refused -lgc without rights; the job did not ask; an elevated job is a decision for the project lead, section 6). Item 1's denominator: the 5090 at the hash is 290 W in the app and 350 W in the bench, not 326; the f = 1 rows move 2 to 11 percent.
Chip side (approximate: the f = 1 stored-dataset chip of item 1 plus an ALU core at k times the 5090's measured 11 pJ per op): at N = 100,000 its per-joule edge over the 5090 falls from 5.6x (shadow empty, 350 W) to 2.1x on GDDR7 and 2.3x on one HBM3 stack at k = 1, 1.5x at k = 1.5, 3.2x at k = 0.5, 4.1x at k = 0.3; over the M5 Max from 1.6x to 0.9x at k = 1. The chip needs a 14,000-lane ALU array (about 30 mm^2 at N5, 150 W at k = 1) at 100,000 ops and 22,600 lanes (60 to 100 mm^2, 500 W) at 331,000; at 28 nm the array is reticle-class. That is the project lead's test as a number: with the shadow filled the chip must carry the memory system and a GPU-class datapath, and its edge is the ratio k of its datapath's energy per op to the GPU's.
Consequences per tier at N = 100,000: an Apple miner loses 1.5 percent of rate and pays 16 W more (0.56x per watt, per pound unchanged); a 5090 loses 0.2 percent and goes from 350 to 431 W (0.81x per watt), so a 5090 rig pays about 23 percent more electricity for the same blocks; a pool user sees nothing; a small NVIDIA card (4060 class) binds near 100,000 by its ALU budget (model, owed); AMD holds by its budget (owed); every verifier tier is untouched (+0.17 ms per warp).
Item 8, the RTX 4070 rows (PC 1 job run-ca3-pc1-4070-shadow-20261006, 17:36:05 to 17:48:10Z, exit 0; the 4070 alone through api/cards, the installed CUDA worker's --bench, nvidia-smi at 1 Hz with 12 of 12 idle samples carrying power; the card at its Ember tune point: a 1,860 MHz core lock and the 160 W cap, memory 10,251 MHz; the 5090 and 9070 XT mining beside it; every pack's fingerprint equal to the Mac's)
| N ops per hash | Pack | 4070 MH/s (delta) | Watts (min to max) | Microjoules per hash | SM MHz | Reading |
|---|---|---|---|---|---|---|
| 930 (class v3) | mx8-devnet-epoch0, twice | 30.95 | 79.3 to 79.8 | 2.56 to 2.58 | 1,860 | the control: 2.5x the M5 Max's energy per hash, a third more than the 5090's at the hash |
| 49,700 | sh256x13 | 31.08 (+0.4%) | 93.5 | 3.01 | 1,860 | |
| 49,700, 64-instruction block | sh64x52 | 32.09 (+3.7%) | 92.6 | 2.89 | 1,860 | the 64-block gain seen on the Mac and the 5090 |
| 102,100 (the candidate) | sh256x27 | 31.08 (+0.4%) | 109.0 | 3.51 | 1,860 | latency-bound; +30 W for no rate |
| 199,600 | sh256x53 | 31.07 (+0.4%) | 138.2 | 4.45 | 1,860 | still latency-bound |
| 330,700 | sh256x88 | 27.14 (-12.3%) | 159.9 (the cap) | 5.89 | 1,846 | the 160 W cap binds; compute-bound under it |
Reading: the model's "a small NVIDIA card binds near 100,000" is replaced by the measurement: the 4070 holds its rate to about 200,000 ops per hash at its tune point and binds only at 330,700 when its power cap does. Consequence per tier (main's reading, 17:5x UTC): class v4's 100,000 ops per hash costs the 12 GB NVIDIA tier nothing in rate, and the shadow has 2x headroom on that card; the 4070 owner pays 30 W more (79 to 109, 0.73x per watt); the N no card we own loses 5 percent at stays set by the M5 Max (about 130,000), not by the small card. One line to check: the restore printed "no entry was enabled before this job" and waited for no worker, while the card had been mining at 28.8 MH/s before it; the card's state after the job is read from the intake below.
Item 8, the RX 9070 XT rows (PC 1 job run-ca3-pc1-amd-g1-shadow-20261006, 15:46:27 to 15:56:54Z, exit 0; the 9070 XT alone through api/cards in every key form (amd:3:gfx1201 is the installed key today), confirmed by the process list, restored in finally and mining again 15 s later; the 5090 and 4070 mining beside it, the 5090 probably clock-locked at about 2,781 MHz from the aborted Ember step; AMD OpenCL 3683.0, device-event time, 2^24 per dispatch)
| N ops per hash | Pack | 9070 XT MH/s (window s) | OpenCL build ms / dataset ms | Watts | Fingerprint = the Mac's |
|---|---|---|---|---|---|
| 930 (class v3) | mx8-devnet-epoch0 | 18.92 (82), 18.92 again at the end | 41 to 51 / 74 | OWED (the ADLX helper ran, 0 samples parsed) | yes |
| 49,700 | sh256x13 | 19.31 (80) | 429 / 74 | OWED | yes |
| 49,700, 64-instruction block | sh64x52 | 19.07 (81) | 342 / 74 | OWED | yes |
| 102,100 (the candidate) | sh256x27 | 19.29 (80) | 428 / 74 | OWED | yes |
| 199,600 | sh256x53 | 19.11 (71) | 424 / 74 | OWED | yes |
| 330,700 | sh256x88 | 19.60 (53) | 418 / 70 | OWED | yes |
Reading: the AMD card never leaves the latency bound on this ladder (its ALU budget is about 650,000 ops per hash); the +2 to 4 percent with the shadow is inside its clock and noise band. Consequence per tier: an AMD RDNA 4 owner loses nothing at the candidate and nothing at 330,700; the N the project can pick stays set by the M5 Max (about 130,000) and the capped 5090 (about 210,000); the one-click AMD miner pays 0.42 s of OpenCL compile per pack at the boundary against 0.05 on class v3, under half a second; the daily build is unchanged at 74 ms. The AMD per-joule row (item 1) and MH per W stay OWED until the sampler reads watts.
Item 6, the 5090 rows (ca3-reserve b85f0f1; job run-ca3-family-pc2-20261006, 08:42:27 to 08:43:03Z, exit 0, the card quiet, prover off and back on; nvcc 12.8 sm_120; best of 3 runs; every row bit-exact)
| Family | M5 Max, Metal (ratio to the alu chain, 881 G steps/s) | RTX 5090, CUDA (ratio to the alu chain, 7,941 G steps/s) | RX 9070 XT | Native or emulated |
|---|---|---|---|---|
| rotr (live) | 1.13 | 1.32 | OWED | native everywhere |
| shflx (live shuffle) | 0.86 | 1.49 | OWED | native |
| variable shifts shl / shr | 0.85 / 0.86 | 1.27 / 1.28 | OWED | native |
| bfe (bit-field extract) | 0.77 | 1.54 (a two-instruction sequence on NVIDIA) | OWED | native on Apple and AMD, sequence on NVIDIA |
| andn | 0.75 | 1.26 | OWED | native |
| perm (byte permute) | 1.13, EMULATED | 1.30 (prmt) |
OWED | emulated on Apple |
| popc / clz | 0.87 / 1.01 | 1.50 / 1.63 | OWED | native |
| sel | 0.76 | 1.32 | OWED | native |
| shfla (lane + delta) | 1.91 (2.2x the xor shuffle) | 1.53 | OWED | native; the 32-lane crossbar is the chip's cost (about 6x the xor butterfly, approximate) |
| dot4 (comparison) | 1.60 unsigned / 4.73 signed, emulated | 1.16 | 1.06 (5 October) | |
| mm8 (comparison, R8) | OWED (Metal 4 matmul2d) | 2.43 (mma.m8n8k16.u8, bit-exact) |
OWED | the licensable block |
The RX 9070 XT column (PC 1 job run-ca3-pc1-amd-family-20261006-e, 17:31:57 to 17:34:47Z, exit 0; the card ALONE: every gfx1201 entry posted off, its worker pid 19788 gone, three runs at 17:32:11 / 15 / 19Z with card_state=alone, every entry restored with its own flag and identities and the card mining again 20 s later under pid 13040; the gfx1036 and old-platform columns run after the restore beside the miners; AMD OpenCL 3683.0, gfx1201, 32 CUs; the range below is over the three runs' best-of-3): alu 1,074 to 1,186 G steps/s (ratio 1.00); rotr 0.99 to 1.13; shflx (xor shuffle, ds_bpermute) 0.82 to 1.21; shl 1.00 to 1.02; shr 0.91 to 1.02; bfe 1.00 to 1.09 NATIVE (amd_bfe); andn 0.91 to 1.01; perm 1.73 to 1.93 EMULATED (no byte-permute path in AMD's OpenCL C); popc 0.91 to 1.01; clz 1.19 to 1.30; sel 0.90 to 1.01; shfla (lane + delta, ds_bpermute) 0.75 to 0.84 NATIVE; dot4 1.00 to 1.11 NATIVE (sudot4); mm8 1.68 to 1.83 NATIVE (the WMMA iu8 builtin reaches gfx12; exact UNVERIFIED: the fragment layout is not in any source at hand, so the CPU reference was not attempted rather than guessed). The AMD column of item 6 is CLOSED. Reading: on RDNA 4 every 32-bit datapath family sits within 0.75x to 1.30x of the alu chain, the shuffles cheaper than it (an LDS operation overlapping the dependent chain), so the ds_bpermute number that could have moved R3 leaves the order as proposed; the two dear rows on AMD are the same two as on Apple and NVIDIA, the byte permute and mm8.
Reading: against the live rotr every candidate is 0.95x to 1.23x on NVIDIA; on Apple only perm (emulated) and shfla cost more than rotr; the 8x emulation bound of 1.13.2 holds everywhere by 4x or more; the 5 percent hash-rate bound at W_new = 4 is argued from the step cost (under 1 percent of ALU time on a read-bound hash), not measured, since no reserve family is live. Proposed order R1 perm, R2 popc and clz, R3 shfla, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8, W_new = 4 each, family n at era n (mm8's era-4 unlock kept as a named exception or moved to era 8: the project lead's call); the full proposed 1.13.2 text with edge vectors per family is docs/plans/counter-asic-3-reserve.md section 6. Consequence per tier: an Apple miner pays the emulated perm at 1.13x a step and shfla at 1.91x, under 1 percent of its hash rate at W_new = 4 (argued); an NVIDIA miner pays nothing measurable; an AMD miner's row is owed and its ds_bpermute_b32 cost is the one number that could move R3; a chip pays a barrel shifter, a byte crossbar, a popcount tree and a 32-lane crossbar per lane, which is the point.
Item 6, the Mac rows (interim, ca3-reserve 192a683; the 5090 job waits on PC 2)
Step cost per family on the M5 Max as a ratio to the add-xor-rotate chain (881 G steps/s; best of 3; load average 7.64; all bit-exact): shl 0.85, shr 0.86, bfe 0.77, andn 0.75, byte permute 1.13 (emulated on Apple), popcount 0.87, clz 1.01, select 0.76, indexed shuffle 1.91 (2.2x the xor shuffle), live rotr 1.13, dot4 unsigned 1.60, dot4 signed 4.73. Every 32-bit datapath family costs an Apple lane under 1.2x a step, inside the 8x emulation bound of 1.13.2 with room; the matrix family is the only one past 1.6x. The rows feed item 8's ALU pricing. The full table with consequences lands at item 6's close.
Items 4 and 7 on the live tables (dry mode, read-only, 09:2x UTC; the live observer still runs master's code)
| Module | Reading | Consequence |
|---|---|---|
| Detector (window epochs 41 to 46, tip DAA 171,165, 20 ids) | network 125 to 152 MH/s per epoch, settled; bands from the fleet log 5090 99.4 / 115.6 / 123 MH/s (n 849), M5 Max 26.5 / 27.4 / 28.5 (n 586), Intel UHD 1.7 / 1.8 / 2.2 (n 758); correlation pairs 6, max r 0.94, edges 1 (the two 5090 ids, PC 1 back since 07:2x), groups none; alert inactive, held 0 of 6; events none | quiet on the devnet; two findings sent to the detector worker: two honest same-model cards cross the 0.8 edge (the group rule of 3 or more ids is what keeps the alert off; the edge may be a network-estimate artefact), and each 5090 id reads about 50 MH/s on chain against the 115 MH/s band (a factor of two to explain before an unknown miner is banded). ANSWERED (ca3-detector 9c7838e, merged): the r 0.94 edge was an artefact of the epoch common factor (the mean over every present id let the paused-and-resumed Mac push the other residuals together); the factor is now the median over the steady core ids, and the same window reads max r 0.10, no edge, tests 7 of 7, the fabricated design still alerts. The factor of two is identities=2 on PC 2's worker (2 x 50.4 = 100.8 MH/s on chain against the 115.6 band, 0.87x, the ratio the whole network shows: reds, pending blocks, template latency); the key count joins the coinbase tag with the card model (owed, 0.3.12). The devnet's max pairwise r over the windows run is 0.10 to 0.53 against the 0.8 edge; the alert stands as a trigger with the 3-id group rule |
| Vendor share (10-minute window, network 253 MH/s) | fleet-reported: NVIDIA 241.7 MH/s (2 workers, 0.920), AMD 18.8 (1, 0.072), Intel 2.2 (1, 0.008), Apple 0 (the Mac paused); chain-attributed: NVIDIA 0.914 (529 of 579 blue blocks, 10 ids), AMD 0.078 (45 blocks, 8 ids), Intel 0.009, unknown 0 (every id maps to a fleet key, 30 mapped) | the two readings agree within 1 percent; the devnet is a one-vendor fleet at 91 percent NVIDIA, which is what the metric is for: an AMD owner's share of income is 0.078 for 1 of 4 cards, an Apple owner's 0 while paused; the public testnet's number is the one that matters |
The live observer runs the shared checkout, which autosync fast-forwards from origin/master; new code reaches it only through a push to master, which is the project lead's call (CLAUDE.md: push only when the project lead asks). So the one restart main asked for waits on that push; the dry rows above are the evidence until then.
Consequences per tier (item 1)
| Tier | Meaning | Being done |
|---|---|---|
| Home miner, 8 or 12 GB card | Nothing today (no chip exists; a 28 nm controller project is $5M to $30M and about 32 months by the Ethash precedent); when one lands it runs 0.3 to 0.5 microjoules per hash against 10 to 20 for this tier (approximate), the first tier out | the power-cap rows and the item 8 measurement; the detector (item 4) is what tells this miner a chip has arrived |
| 16 GB AMD (9070 XT) | 7x worse per joule than the 5090 and 30x worse than the chip (approximate); a chip ends AMD home mining first | vendor-share metric (item 7); nothing in the hash fixes AMD's 2.4 G dependent reads/s |
| 24 or 32 GB (5090, M5 Max) | the honest best at 2.40 microjoules; 5x to 9x behind the chip in the model, 2x to 5x by the precedent | a 200 W cap on the 5090, if the rate holds, halves the gap |
| Rig | per joule it is its cards; at $2.8 against $14.7 per MH/s a chip fleet is cheaper per dollar too | the issuance trigger: bounty and benchmark live before daily issuance crosses about $50K (item 4b) |
| Pool user | a chip fleet is a few operators; Monero's was found at 85 percent by its share pattern | the detector, before the public testnet |
| The public claim "under 2x" | held for the recompute chip at the op budget; per joule and against the stored-dataset chip the model reads over 2x on both memories and the precedent reads 2.1x to 4.8x; NOT SAFE TO PUBLISH as worded | decision for the project lead (section 6); nothing on the site or the devnet changes from this run |
3a. Corrections found by the run
| Found by | What was wrong | Fixed |
|---|---|---|
| item 3 (43c3ead) | spec 01 section 1.13.1's era table and docs/plans/mixer-x4.md section 2 still said mixer_mult = 4; the code (LoadClass::MX8, V3_CLASS) and spec 1.8.5 say 8 |
both lines corrected on ca3-coord, 6 October 2026 |
| PC 1 job 2 (family, run-ca3-pc1-amd-family-20261006, 16:56 to 16:59Z) | the probe ran on device 0, the integrated gfx1036, not the 9070 XT (device 1): every row dev=0 name= empty, alu 40.59 G steps/s (a one-CU figure); the rows are RDNA 2 iGPU ratios (shifts 1.15, bfe 1.15 native, andn 1.16, popc 1.55, clz 1.75, sel 1.81, shfla and shflx 1.89 through ds_bpermute, perm 2.43 and dot4 2.57 emulated, mm8 none: AMD's OpenCL C compiles no byte-permute, dot4 or WMMA builtin) and the 9070 XT column stays owed |
the probe picks the device by name (gfx1201 on the newest AMD platform) and prints it in every row; re-run in the next PC 1 slot |
| the watts job (run-ca3-pc1-amd-watts-20261006, 17:49:46 to 17:52:54Z, FAILED exit 1) | the sampler proved itself (12 of 12 9070 XT watts lines), the 90 s app-state window ran (the card mining at 18.87 MH/s, 203 W, identities 8 before it), every gfx1201 entry was posted off, the card went quiet and the sh256x27 bench started; then the script exited with no APPROW, no LADDER, no watts error= and NO RESTORE line in any upload (no enabled=True post, no card_workers_after): a finally that did not run or did not print, so the 9070 XT may have been left OFF |
CONFIRMED and restored: api/state at 17:55:10Z read the card enabled false, state off, 23 W idle; the identities job run d posted it back and at 17:57:10Z it mined under pid 3436 at 18.9 MH/s, 195 W, identities 8; the entry amd:1:gfx1201 was set back to enabled at 17:59:05Z (job run-ca3-pc1-amd-flag-amd1-20261006) with the card mining through it (18.86 MH/s, 200 W); PC 1 free at 18:01:05Z; the AMD watts row stays OWED; the cause and a fixed script (an unconditional finally with its own first line, no exit inside the try, the device dumps out of the report) with the script worker before any re-run |
| the 4070 ladder's restore line | "no entry was enabled before this job" and no wait for the worker, while the card had mined at 28.8 MH/s before it | the intake shows the 4070's worker restarted 15 s after the restore and racing at 30.9 MH/s: the card came back; the script's settings read, not the card, was wrong |
PC 1 family run d (run-ca3-pc1-amd-family-20261006-d, 17:17 to 17:20Z, exit 0; run b had failed because the script's probe parameter was named $args, PowerShell's automatic variable, so the splat was empty: renamed, with a rule added to tools/ci/ps-drive-ref-check.sh that fires on the old signature and stays quiet on the new) |
the probe chose the 9070 XT by name (gfx1201, 32 CUs); every variant built on it: mm8 NATIVE through the WMMA builtin on gfx12 (exact unverified), dot4 native (sudot4), bfe native, the shuffles native (ds_bpermute, ds_swizzle), the byte permute EMULATED (no amd_perm path in AMD's OpenCL C). The step costs are unusable: the card was mining beside the probe, the alu chain read 226 then 195 G steps/s and the ratios swung from 5.9 to 11.8x (run 1) to 0.17 to 1.3x (run 2), the loaded card's scheduler | the 9070 XT column is taken with the card alone (run e, job 1's switch path), after the G2 job |
| the identities job, first run (run-ca3-pc1-amd-identities-20261006, 17:02:10Z, failed in 0 s) | my own script: "... under $appDir: nothing posted" is a PowerShell 5.1 parse error ($appDir: reads as a drive-qualified variable), so the script never started; the script worker checks this shape by hand, CI did not |
${appDir}:; the class guard tools/ci/ps-drive-ref-check.sh added to ci.yml (67 .ps1 files clean; a backtick-escaped $ in a bash-generating here-string is ignored); the job republished as -b |
| the identities job, run c (run-ca3-pc1-amd-identities-20261006-c, 17:05:53 to 17:07:53Z, done) | the 9070 XT's live entry amd:gfx1201 read identities 2 at 18.91 MH/s and 203 W before; the one POST (the app's cards-array shape) and a 120 s settle; after: identities 8, mining under a new worker pid, 18.9 MH/s (avg 18.77), 199 W. The identities count moves no rate (18.92 at 2 was the number to beat). Run b of the same job had failed on the flat body shape (400) | closed: PC 1's 9070 XT runs as it did before job 1 |
| the reset job (16:55Z) | the 9070 XT's ADLX state read factory 0 after Ember run 6 (offsets zero, the flag cleared by the engine's set of 0); the app runs the card under amd:gfx1201 with identities 2 where it ran 8 before job 1's restore | the reset to factory 1 applied and read back; the one POST api/cards back to 8 identities running as its own job, the rate before and after recorded |
| item 2's PC 2 job | the settings.json card key is nvidia:0:NVIDIA GeForce RTX 5090 (with the device index); the 5 October job scripts posted the state's key without the index, so POST api/cards switched nothing and the miner stayed up through the 90 s wait: the job's 5090 rows are loaded-card figures |
items 6 and 8 told to post both key forms and confirm by the process list; the class fix (one key form everywhere, a check that fails a job script posting a card key without the index) is owed to the job tooling |
| the coordinator's merge of ca3-shadow | git add -A docs staged docs/bench-log.md with its conflict markers inside (a conflicted path is marked resolved by git add), so 45f3019 carried markers into HEAD; found at the ca3-reserve merge as nested markers |
the three 6 October entries kept in order with every marker removed (fcce185); the class guard already exists, tools/ci/no-conflict-markers.sh in CI, and it would have failed the push; it passes on the merged tree |
| the merge of ca3-derive and ca3-shadow | both added a field to LoadClass (derive_len, shadow); resolved as the union, every other literal spreads ..; cargo test -p igneum-pow on the merged tree (cargo 1.99 at ~/.cargo/bin; the Homebrew 1.69 on PATH cannot read the lock file): 59 + 7 + 4 + 19 + 7 = 96 passed, 0 failed, the pinned v2 and v3 packs byte for byte |
merged 09:10 UTC |
| item 3 | the mixer's op count: 144 integer ops per application as written in memhard.rs (128 with RC and rk hoisted) against the 130 the chip model prices (chip-model-v3.md section 1, from spec 1.8.4) |
items 1 and 2 asked to state which figure their rows use and why; the status close carries the answer |
4. The chip model, before and after
Every chip figure is arithmetic on cited memory and logic figures and is approximate; every GPU figure is measured and names its entry. "Per chip" is rate per chip against the 5090's rate; "per joule" is energy per hash, the Ethash chips' metric.
| Chip | Before 3.0 (the 2.0 record, 5 October) | After 3.0 (6 October) | Source |
|---|---|---|---|
| On-die 256 MiB cache recompute chip (f = 0), class v3 | 0.31x bare, 0.92x with the 3x fixed-function factor per chip; per joule not priced; the public claim "under 2x" rested on this row | per chip unchanged; per joule 1.86x (1.3x to 2.4x over the on-die read energy); with the per-day derivation (item 2, dr736) the 3x factor goes: 0.29x bare, 0.34x at a 1.2x allowance, 0.43x at 1.5x | chip-model-v3.md sections 2, 5.4, 6 |
| Stored-dataset memory-controller chip (f = 1), the Ethash class | not priced (O-1.6 open, the curve never drawn) | per chip 1.22x on GDDR7, 0.61x on one HBM3 stack, 4.9x on eight; per joule 5.1x (GDDR7) to 9.2x (eight HBM3 stacks) at the 326 W denominator, 5.6x at the measured 350 W bench control; $2.8 per MH/s against the 5090's $14.7; the Ethash precedent for this class 2.1x to 4.8x per joule; the curve is monotone toward f = 1 so the partial-store chip is never built; the mixer, item 2 and x16 do not touch it | chip-model-v3.md section 5; history rows 3 and 4 |
The same chip with the latency shadow filled (item 8, N = 100,000 ops per hash, class v4 candidate mx8+sh256x27) |
not a lever anyone had priced | per joule over the 5090 2.1x (GDDR7) and 2.3x (HBM3) at k = 1, 1.5x at k = 1.5, 3.2x at k = 0.5; over the M5 Max 0.9x at k = 1; the chip needs a 14,000-lane ALU array, about 30 mm^2 at N5 and 150 W at k = 1, reticle-class at 28 nm; k, the chip core's energy per op against the 5090's measured 11 pJ, decides it, and 2x is crossed only at k under 0.5 | latency-shadow-2026-10-06.md sections 6 and 8 |
| The honest denominators | the 5090 at 136.1 MH/s and 326 W (a peak with the prover on) | the 5090 at 290 W in the app and 350 W in the bench (2.34 to 2.65 microjoules per hash), cap floor 400 W so no power cap binds; the M5 Max at 21 W GPU plus DRAM (0.78 microjoules), three times better per joule than the 5090 and the honest best we own | item 8; miner-eff's 4 October log |
What 3.0 did to the model in one line: the chip that matters is not the one 2.0 priced; it is the Ethash-class memory-controller chip, over 2x per joule today; the lever that answers it is program work in the latency shadow, which makes the chip carry a GPU-class datapath, and the measured candidate takes its edge from 5.6x to 2.1x at a chip core no better than the GPU's, for 0.2 percent of the 5090's rate.
5. What passes as class v4
Judged on measurements against the six gates of docs/plans/counter-asic-2-rollout.md section 7. Nothing is published; nothing is cut.
| Candidate | Measured | The six gates | Verdict |
|---|---|---|---|
mx8+sh256x27: class v3 plus a 256-instruction ALU block run 27 times per iteration, about 100,000 ops per hash (item 8) |
hash rate -0.2 percent on the 5090, -1.5 percent on the M5 Max, under the 5 percent rule; verifier +0.17 ms per warp (2.23 ms, gate 10 ms, 2019-class core about 5.6 ms); bit-exact on Metal, CUDA, clang emulation and Apple OpenCL; the 5090 at 431 W (its cap) from 350; the 9070 XT owed | G1 bit-exact: NVIDIA and Apple GREEN, the AMD vendor NOT RUN (gfx1036 or the 9070 XT); G2 (1,000 random hashes per card re-hashed on the CPU): NOT RUN; G3 (crate suite green: 96 of 96; the Metal fuzz, edge and stats runs on the class): NOT RUN; G4 (the fast-time 3-node network across a v4 activation): NOT RUN; G5, G6: NOT RUN | READY FOR THE GATE RUN as THE class v4 candidate, on the project lead's word; not ready for a cut. Two design decisions ride with it: the acceptance rule does not yet interpret the block (cost stated in the analysis), and the block stays at 64 to 256 instructions |
dr368 or dr736: the per-day derivation (item 2) |
dr736 verifier 4.9 ms per unit on the M5 Max core, about 12 ms on a 2019-class core (over the gate); dr368 2.69 ms, passes both; bit-exact on Metal, Apple OpenCL and CUDA; build +7 ms a day on the Mac, +2 ms on the 5090; NVRTC +1.1 s per pack (a once-a-day module is a requirement) | not a v4 candidate: a reserve entry | GO as reserve R0 (proposed text, counter-asic-3-derivation.md section 6); NO-GO genesis-live at dr736 until O-1.14; urgency LOW (the f = 1 chip derives no item) |
| The reserve order R1 to R8 (item 6) | step costs on two cards, every family under 1.91x a step, bit-exact | a spec ordering, no activation | GO for the order as proposed; a decision for the project lead |
| Layer 9 (epoch length) ranked above layer 7 (mm8) (item 5) | FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the unmeasured bank-bound ceiling | reserve ranking | GO for the ranking; the rented FPGA hour owed |
So: one class v4 candidate, mx8+sh256x27, defined and measured on the hash's own numbers, with gates G1 (AMD), G2, G3 (Metal runs), G4, G5 and G6 still to run before any publish, by the two-publish rollout of 2.0 and a program_class_v4_activation_daa at tip + 14,400.
5a. The gate run (the project lead, 16:50 UTC: "3.0 run it now"; opened 16:5x UTC)
No publish, no manifest, nothing on the live devnet; PC 2 one job at a time with the installed app untouched; the Mac's miner stays paused. Two lanes: the hash side (branch ca3-v4-hash: G1 Metal, Apple OpenCL and CUDA, G2, G3 and the verifier benchmark; one PC 2 job first) and the node side (branch ca3-v4-node and the fork worktree igneum-node-ca3v4 from release-0.3.13-node: the v4 switch through igneum-pow, the fork, the workers, the miner and the app, then G4, G4b, G6 with --features igneum-pow, G5; its PC 2 jobs after the hash lane's). G1 AMD rides in PC 1 job 1 on "go PC 1 AMD". Evidence lands in docs/plans/counter-asic-3-gate/.
| Gate | What | State | Evidence (full lines in docs/plans/counter-asic-3-gate/hash-gates.md and node-gates.md, JSON per run beside them) |
|---|---|---|---|
| G1 | bit-exact v4 on every vendor against the Mac reference (2^24 fingerprint, self-test) | GREEN on all three vendors | Apple (Metal and Apple OpenCL, 15:49 to 15:50Z) and NVIDIA (the 5090, PC 2 job, 15:52Z): eight packs, three harnesses, one fingerprint per pack; AMD (PC 1 job run-ca3-pc1-amd-g1-shadow-20261006, 15:46 to 15:56Z): 7 of 7 packs equal to the Mac's on gfx1201 (sh256x27 3d2e8245cc084d07) |
| G2 | the CPU verifier exact on 1,024 random hashes per card | GREEN on all three vendors | AMD (PC 1 job run-ca3-pc1-amd-g2-20261006, 17:22Z, the kit worker's serve mode on gfx1201 through the class=v3 and class=v4 tokens, beside the miners): 1,024 of 1,024 found and distinct on the control mx8-devnet-epoch0 (digest 2a1824a2...) and on the generator-4 candidate v4-devnet-epoch0 (435b976a...), both re-hashed on the Mac through tools/ca3-v4/g2-recheck.sh --digest: MATCH 1,024 of 1,024 each. The first kit's sh256x27 is a string-seed pack the worker's job protocol refuses, so G2 ran the chain-seed packs the 5090 and the Mac ran |
| G3 | the soundness suite green on the class | GREEN | the crate suite 96 of 96 (97 on the merged tree at 17:17Z, cargo 1.99); the Metal fuzz 200 of 200 and 50 of 50; Apple OpenCL 20 of 20 and 5 of 5; 4 of 4 CPU tests on the class and on the era-composed class |
| Verifier benchmark | ms per warp on one M5 Max core, v2 / x8 / v4 in one session, gate 10 ms | GREEN | 2.33 ms on the candidate (x8 2.06), about 5.8 ms on a 2019-class core by the 2.5x rule (approximate); the 2019-class core itself still unmeasured (O-1.14) |
| G4 | the fast-time 3-node network across a v4 activation, plus the known-failed case | GREEN (and GREEN again on the id fix with the id assertion and its failed case, 491131a) | run 1 (15:54 to 15:59Z, class-v4.mjs, activation 150 rounded to epoch 3 at DAA 180, the v3 switch at 60): 3 of 3 switch lines, templates class 2 / 3 / 4 by epoch, 181 blocks before and 128 after DAA 180, program ids agree on all three miners and no v2 or v3 id reappears under v4, rejected 0/0/0, one sink at 308/308/308, one digest; the first v4 epoch's cache ready in 2 ms (the v3 day cache reused across the boundary); run 2 the known-failed case (the switch at never: no v4 epoch reported) |
| G4b | a real Metal miner across a v4 boundary through --prepare-packs | GREEN | run 3 (16:02 to 16:07Z, igneum-bench from the committed tree): 5 PREPARE lines 5 to 6 DAA before each boundary, three class=v4 era=<hex>, the worker's prepared lines (430 ms the first with the v4 day built on the GPU, 139 and 175 ms with the day resident; a v3 prepare 62 ms), every swap with no pause, 301 blocks accepted on the Metal miner (123 after the switch) all re-checked on the CPU mismatched 0, need 0, no mismatch or out-of-date line, no exit 42 or 44 |
| G5 | the PC-built Windows workers and the Mac workers from the same commit | GREEN, one gap named (re-done from the fix 7c22d0d: cuda.exe 3bc8ad8f..., opencl.exe 16ef0154..., igneum-bench f9ca4b07...) | the Windows workers cross-built on the Mac from a522d04 as 0.3.11's G5 did (the build job has no worker unit: a tooling gap): igneum-worker-cuda.exe e563126e... (1,536,512 bytes), igneum-worker-opencl.exe 7fce1249... (478,208), both with the resource block, mingw not bit-reproducible (a link-time stamp); the Mac igneum-bench 30c70754... from the same tree, the binary G4b mined with; the node and app from PC 2 job build-20261006-155958, every sha256 verified |
| G6 | the node suites on PC 2 with --features igneum-pow, from a fork branch off the current release tip | GREEN | fork ca3-v4-node 5f7e0543 on release-0.3.13-node; PC 2 job build-20261006-155958 (15:59 to 16:08Z under the lock, the app untouched, the prover on): every stage ok in 392 s; kaspa-consensus 98 (2.0's flake did not recur), consensus-core 109 + 7 (override_params_carry_the_program_class_v4_activation), kaspa-pow 15 with the v4 engine test under the feature, p2p-flows 33, igneum-app 112 + 26 + 8 |
BLOCKER before any cut (found by the hash lane, 17:0x UTC): program_id(3, seed, attempt) is class-independent, so all seven v4 packs carry the same program id as the v3 control of their seed (73bcbfe8ccf988f1), and a stale worker across the activation would see no id mismatch (2.0's G4 id check cannot fire on it; G4's per-epoch ids differ only because each epoch has its own seed). The fix is on the v4 seam, assigned to the node lane: a generator-4 stamp in the id so a v3 and a v4 program of one seed differ, v2 and v3 ids byte-identical, the packs re-exported (fingerprints unchanged, ids moved), G4's id assertion and the fuzz re-run. FIX MERGED (ca3-v4-node 7c22d0d, 17:3x UTC): the trap was the CLI's --era path (stamp_era stamped generator 3 on any class), not the chain seam (the fork already stamped generator 4 and its kaspa-pow test asserts the same-seed v3 and v4 ids differ); now generate_era and the CLI stamp the generator from the class, the shadow block marks class v4 in packcheck.rs, packfile.h and main.swift (a generator-3 pack with a shadow block is refused as "a v4 program stamped v3", a generator-4 pack without one refused; the ladder packs stay loadable); v2 and v3 ids byte-identical; the crate suite 97 of 97 on the merged tree; the seven gate packs re-exported with generator 4, class "v4" and program id c120d7963abdcd96 (the v3 control keeps 73bcbfe8ccf988f1), only the generator, id, class and comment lines changed, the kernels and vectors byte-identical. The hash lane's re-run on the fix is GREEN (ca3-v4-hash ca61dec, merged; hash-gates.md "Follow-up 1"): the crate suite 97 of 97; the seven packs re-exported here equal the tree's (0 differing files); the fingerprints unchanged on the rebuilt Metal and Apple OpenCL harnesses (all eight); Mac G2 through the class=v4 token 1,024 of 1,024 on all eight packs, both harnesses; the fuzz 200 programs and 800 units, stats, edge and determinism 4 of 4 on the class and 4 of 4 era-composed, the same-seed v3 and v4 ids now differ (assert_ne). The node lane's re-run on the fix is GREEN (ca3-v4-node 491131a, merged): G4 re-run 16:32 to 16:38Z on igneumd and igneum-miner rebuilt on the fixed crate, SUMMARY PASS, 181 / 125 blocks across DAA 180, rejected 0/0/0 and 0/0/0, one sink at 305/305/305, and the id assertion: every v4 epoch's id on all three miners equals the CLI's class v4 id for that seed and era and differs from the same-seed v3 id (e3 30544487d1289d6e against 1ae1c9f0eda0cef5, e4 53e36801cc6fafcf against af81f6e844959460, e5 876e155e3fb59983 against b4f25c4678496a78); the assertion's own failed case (--id-check-against v4) reports exactly FAILED CHECK v4_ids_differ_from_the_same_seed_v3_id, exit 1. The seven re-exported packs' Metal fingerprints by packbench equal the table (only the ids moved). G5 re-done from 7c22d0d because packfile.h and main.swift moved: igneum-worker-cuda.exe 3bc8ad8f..., igneum-worker-opencl.exe 16ef0154... (resource block verified), igneum-bench f9ca4b07...; kaspa-pow with the feature 15 of 15 on the rebuilt fork. THE BLOCKER IS CLOSED: the gate table is green on the hash and on the cut, with AMD G2 and the AMD watts the owed rows (queued on PC 1). Two conditions ride with the fix: every kit sent to a PC carries the re-exported packs (the AMD G2 kit is being rebuilt from the merged tree), and the mixer.rs harness change rides with any merge of 7c22d0d (it does, on ca3-coord). The PC 1 AMD rows taken on the old-id packs stand: the id is not an input to the hash.
The per-tier cost line of the candidate (item 8, measured): the 5090 -0.2 percent of rate at 350 to 431 W (a rig about 23 percent more electricity), the M5 Max -1.5 percent at 21 to 37 W, pool users nothing, the verifier +0.17 ms per warp; the 9070 XT and the 4070 rows land with the PC 1 jobs. Owed before the cut, besides the blocker: the AMD watts (the sampler fix is in; the re-run queued), the 2019-class core (O-1.14; the US laptop's CPU could answer it with a Windows igneum-pow build, a proposal), the G2 found-lines file and digest and the 200 KB report cap (tooling, in hand on ca3-v4-hash).
Two preconditions on the cut, from main (17:0x UTC, both from today's incident), assigned to the node lane:
| # | Precondition | What passes it | State |
|---|---|---|---|
| P1 | The cut rehearses first on the rented fleet as a staging network (the fleet agent owns the boxes; the node lane supplies the v4 override object and the rehearsal plan docs/plans/counter-asic-3-rehearsal.md) |
a fleet-only chain crosses the v4 activation with every box switched: zero rejected blocks, blocks on both sides, the G4 id assertion live on every box, one sink, one digest; the stale-box case refused at the handshake with no fork | PLAN WRITTEN (ca3-v4-node f39a8eb, merged): fleet-only chain igneum-devnet-400 from its own genesis state, 12 or more mining boxes, a seed box, one stale 0.3.13 box; the signal flip at DAA 7,200 (the first full 3,600-DAA window) with the floor at 14,400 not reached; ten steps, the report fields and pass rules; the id assertion from the boxes' miner lines. Two objects as files in docs/plans/counter-asic-3-gate/: the publish object (the live 13 fields plus program_class_v4_activation_daa = N6, the floor, tip + 14,400 rounded up, 219,600 at the 16:57Z DAA of 202,919, and program_class_v4_signal_window_daa = 86,400; digest ac8e60ce...) and the rehearsal object (every earlier switch at 0, window 3,600, floor 14,400; digest bc2142b1...); the no-file devnet digest on this binary 7f2e49be... (3c505021 superseded: the window field joined the digest). THE RUN STARTED 18:22Z (the fleet agent): igneum-devnet-400 from the devnet genesis with the re-cut 1-block/s object override-v4-rehearsal-1bps.json (600-DAA epochs, lead 100, window 600, the flip at epoch 2 = DAA 1,200, the floor at 2,400; node digest d23394e7...; the file written byte for byte on every box) and the signalling fork's Linux binaries from PC 2 job build-20261006-174823 (igneumd 847ddfd1..., igneum-miner 37178aeb..., verified on the Mac and every box); the seed on the RunPod pod dn2-seed (public p2p 213.173.110.229:11927, no miner); 15 mining boxes beside their live-devnet nodes on separate ports with fresh appdirs and 1-thread CPU miners (dn2-1, dn2-2, dn2-3, hub-1, 3080, 4070, 4090, 5090, A5000, 3090-2, 3090-3, 4070-1, 4090-1b, 4090-3, the 8x 4090 rig); the stale box p2-3090-4 on 0.3.13's igneumd d6350586 with the same file and no miner, started last; the 38 wave pods join about 18:50Z. The clock from the seed's genesis at about 18:23Z: the window full 18:41Z, the flip about 18:43Z, the floor about 19:03Z (run to the floor so its line is exercised). A sampler on every box every 5 minutes; the id lines go to the node lane for the CLI check; the report in section 4's table plus rehearsal- |
| P2 | No fixed-height activation on the live devnet again: miner signalling for class changes as the v4 seam's activation rule (the block carries the miner's object version; the class flips at the first epoch boundary after 95 percent of mining weight over a window signals the new object, with a floor height after which it flips regardless), PROPOSED spec text in counter-asic-3-node.md, implemented behind the override, with the fast-time harness test (67 percent does not flip; 100 percent flips at the next boundary; nobody signals and the floor flips it; each with its failed case) and G6 again on PC 2 |
the three harness runs PASS with their failed cases; G6 green on the fork change | GREEN (ca3-v4-node 0635c04 and 3892035, fork 0562a7f2 on 5f7e0543; PROPOSED text in counter-asic-3-node.md section 6: the header version's high byte carries the node's object version, the low byte stays the block version so later objects pass the version check; weight = blue blocks by the finality walk over the window ending at each epoch's seed block; 95 percent = 9,500 bps; window 86,400 DAA in the file and the digest, 0 = off; monotone, memoised per seed block; the fixed height stays as the floor; IGNEUM_CLASS_SIGNAL lowers a node's byte on devnet and simnet only; RPC fields 20 to 23). The fast-time gate infra/fast-time/class-v4-signal.mjs (three CPU miners, window 120, v3 from 60): two of three signalling PASS with no v4 epoch over epochs 0 to 7 (share 6,166 to 7,583 bps, 0 rejected, one sink 424/424/424); all three PASS with the flip at epoch 3, the first full window, 10,000 bps, 3 of 3 switch lines, 181 / 124 blocks, 0 rejected, one sink, the id assertion on epochs 3 to 5 (e3 e48e6be7c6699824 against the v3 6847355c88e8215f); nobody with the floor at 300 PASS, 0 bps, the flip at epoch 5 by the floor and not before; the failed case (two of three told to expect a flip) FAIL on eight checks. G6 on the signalling fork: three PC 2 jobs under the lock, prover on, app untouched (build-20261006-173017 hit 2.0's flake, 97 of 98; build-20261006-174027 kaspa-consensus alone exit 0; build-20261006-174823 the five crates exit 0 in 35 s with consensus-core 110 + 7, igneum-exec 18, kaspa-pow 15 with the v4 engine and the signal-rule tests, igneum-miner 18, p2p-flows 33, and the app 112 + 26 + 8 exit 0; its closing line "upload incomplete" is the relay blob fault after both test stages closed exit 0). Owed and named: a class-signal witness in the pruning-proof format (a proof-synced node falls back to the floor rule for epochs whose window reaches below its pruning point and logs it), the same class as the era witness |
Facts for the cut from the node lane (docs/plans/counter-asic-3-node.md): the devnet digest moves (c562d70e... to 3c505021...), so the cut is a one-sweep binary rollout and a 0.3.13 node is refused at the handshake afterwards (intended, fleet-wide); a 0.3.13 miner reads the v4 height as never (an optional proto field), so miners and nodes move together; infra/fast-time/override-60x.json as committed carried the proving-v1 block twice and lacked four 0.3.12 and 0.3.13 fields (the node refused the file; fixed, with a new CI check override-json-check.sh); the 48 GiB target clone vendor/igneum-node-ca3v4/target-ca3v4 can go after the cut.
THE ONE LINE FOR [user]: class v4 (mx8+sh256x27, 100,000 ops per hash in the latency shadow) has passed every gate on the fixed tree (the program-id blocker closed with its own failed case) and is ready for a whole-fleet one-sweep cut (the digest moves to ac8e60ce... at the floor N6) once P1 passes: P2, miner-signalled activation (95 percent of blue-block weight over a one-day window, the floor height after which it flips regardless), is designed, implemented and GREEN on the fast-time gate with its failed case and on G6; P1, the rehearsal on the rented fleet as a staging network, has its plan and objects written and waits for the fleet agent's run; the clean-day wait removed by the project lead ("can we run the v4 class now?", 18:2x UTC): the P1 rehearsal runs NOW on the fleet (15 prover boxes plus the four Devnet 2 boxes, the wave joining about 18:50Z) and the 0.3.15 cut (class v4 plus the fourteenth field, one digest move, miners first, hands last) goes tonight on the project lead's go the minute P1 passes; its cost is the 5090 at 431 W instead of 350 for 0.2 percent less rate (a rig pays about 23 percent more electricity), the M5 Max at 37 W instead of 21 for 1.5 percent, the 9070 XT no rate at all (its watts pending), the verifier +0.27 ms per warp; what it buys is the stored-dataset chip's per-joule edge over the 5090 falling from 5.6x to 2.1x at a chip core equal to the GPU's; proposed for a day when no other cut is in flight, not tonight (0.3.14 and the fleet night come first).
6. Decisions for the project lead
- The public claim. "Under 2x" is true of the recompute chip per chip at the op budget and false of the stored-dataset chip per joule (item 1). Nothing on the site or in the litepaper changes from this run; the two drafts below are for the project lead's decision, with the rows that bound them.
| Row | Per chip (rate) | Per joule | Source |
|---|---|---|---|
| On-die recompute chip, class v3 (f = 0) | 0.92x with the 3x factor (0.31x bare) | 1.86x (1.3x to 2.4x over the on-die read energy) | chip-model-v3.md sections 2 and 5.4, model |
| Stored-dataset memory-controller chip, GDDR7 (f = 1) | 1.22x | 5.1x | chip-model-v3.md 5.4, model |
| Stored-dataset chip, HBM3 one stack / eight stacks | 0.61x / 4.9x | 7.5x / 9.2x | chip-model-v3.md 5.4, model |
| Ethash precedent, the same chip class: Linzhi Phoenix 2020, Antminer E9 2022, Jasminer X4 2021 | 2.1x / 2.9x / 4.8x per joule | asic-resistance-history.md rows 3 and 4 | |
| With the latency shadow filled (item 8, model until measured) | 1.22x | 1.85x at N = 330,000 and a chip core equal to the GPU's ALU; 2.5x at 1.5x worse | chip-model-v3.md 5.7 |
Draft (a), scoped: "The strongest recompute chip we can price, holding the whole 256 MiB cache on-die, reaches under 1x per chip against an RTX 5090. A memory-controller chip that stores the whole dataset reaches 1.2x per chip and, in our model, 5x to 9x per joule; the Ethash chips of this class reached 2.1x to 4.8x. The lever against it, program work in the latency shadow, is being measured (Counter ASIC 3.0 item 8)."
Draft (b), the measured fact only: "An RTX 5090 mines this hash at 136 MH/s and 17.5 billion dependent 4-byte reads a second, 82 percent of its memory's random-read ceiling, with its integer units at 0.15 percent of their budget. A chip beats it only by reading per watt what a 512-bit GDDR7 board reads, and the gap is the card's own idle logic."
Either replaces "under 2x" on the site and in the litepaper once the project lead chooses; until then the claim stays as it is and this file records that it is not safe as worded.
- The class v4 candidate: run the six gates on
mx8+sh256x27(100,000 ops per hash) and, if green, cut it by the 2.0 rollout shape. Cost to the tiers: the 5090 goes from 350 to 431 W for the same blocks (a rig pays about 23 percent more electricity), the M5 Max from 21 to 37 W for 1.5 percent of rate; the verifier +0.17 ms per warp. Gain: the stored-dataset chip's per-joule edge over the 5090 falls from 5.6x to 2.1x at a chip core equal to the GPU's. Alternative: wait for the 9070 XT and the 4060-class rows (both owed) before the gates, since a small card binds near 100,000 by its ALU budget (model). - The reserve order R1 perm, R2 popc and clz, R3 shfla, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8 (W_new = 4 each; mm8's era-4 unlock kept as the named exception or moved to era 8), and R0 the per-day derivation as a reserve family (dr368 the safe draw; dr736 after O-1.14).
- Commission the mixer cryptanalysis: USD 80,000 to 160,000, the brief and the reviewer shortlist in funding.md (item 3); it should name random ARX programs (item 2's class) beside M_r. No outreach has been made.
- The bounty trigger: escrowed and the benchmark live before daily issuance crosses USD 20,000 a day (funding.md rule 5); the detector is the clock (quiet on the devnet, max pairwise r 0.10 to 0.53 against the 0.8 edge).
- The live observer: the detector and the vendor-share hooks reach it only through a push to master (autosync); the project lead's word on the push, then one restart and the first live rows into this file.
- The 5090 clock rows (
-lgcat 2,781 / 2,472 / 2,163 / 1,854 MHz): an elevated job on PC 2 would ask for administrator rights at the keyboard (the 5 October prompt class); not run without the project lead's word. - PC 1: the 9070 XT rows for items 2, 6 and 8, the detector's band and the FPGA ranking's AMD line are all owed on PC 1's release.
- The job tooling: one card-key form everywhere (the settings.json key carries the device index) and a CI check that fails a job script posting a key without it (the class fix for item 2's loaded-card run).
Main's decisions on section 6 (18:0x UTC, 6 October 2026)
| # | Decision | Record |
|---|---|---|
| 1 | The public claim: the sentence deployed on the hero at 16:21Z ("In our public model the strongest chip reaches 5x to 9x per joule against an RTX 5090 today, about 2x once the lever now in its gates ships") plus the litepaper's long form; "under 2x" nowhere | docs/evidence.md row 17 (the claim, its sources, draft (a) of this file chosen); the review finding behind it: docs/review/round-4-reddit-2026-10-06.md item 3 (Serious), flipped by the measured fact |
| 2 | The cut: yes, after P1's rehearsal passes and 0.3.14 has run a clean day; the fleet agent owns P1's run | section 5a |
| 3 | R0 (the per-day derivation, dr368 the safe draw) and the reserve order R1 perm to R8 mm8: GO as proposed, with the once-a-day item module recorded as a requirement on every vendor (NVRTC +1.1 s, AMD OpenCL +2.0 s per pack) | counter-asic-3-derivation.md, counter-asic-3-reserve.md |
| 4 | The cryptanalysis spend, USD 80,000 to 160,000: AWAITING [user] (his money decision; put to him with the brief on a quieter day, not tonight) | funding.md |
| 5 | Push: ca3-coord merged into master through CI (nothing activates: v4 sits behind the override; the detector and vendor-share hooks go live on the observer's next restart) | the merge commit and the CI run, below |
| 6 | The AMD watts: on the runner's --cards-off mechanism, next cut |
section 6a |
| 7 | The 2019-class core (O-1.14): a Windows igneum-pow job on the US laptop when it next appears on the relay | owed list |
6a. A rule from the run (main, 17:5x UTC, from the watts job's failure)
A script never switches the installed app's cards: a POST of enabled false to /api/cards from a script is the same class as /api/pause from a script (a script that dies leaves the box degraded, unattended). A job that needs a card alone asks the runner for it: the runner's --stop-miners grows a --cards-off <keys> that posts the exact settings entries off before the script and restores them with their own flags and identities on ANY exit, as it restarts miners; tools/ci/playbook-quit-check gains the api/cards enabled-false pattern so the shape fails CI. Both land tonight (the PC 1 worker, branch ca3-pc1-amd); the four PC 1 scripts that carry the shape move to the flag; the AMD watts row re-runs only on the runner mechanism, after the app that carries it is installed (owed to the next cut).
7. Unverified and owed
| Item | Owed | Why |
|---|---|---|
| 1 | the 5090's watts at the hash alone: LANDED from Ember Tune run 5 (job ember-tune-pc1-5, PC 1, 15:30 to 15:38Z, elevated, nvidia-smi -pl set directly, the clock unlocked at 2,850 MHz core and 13,801 MHz memory, 75 s a step, MH/s wall from the worker's STATUS lines): 575 W cap 127.38 MH/s at 309.9 W (0.411 MH per W); 518 W 99.32 at 313.6 (a stall inside the hold); 460 W 123.11 at 312.2; 403 W 127.38 at 310.9; 400 W (the floor) 127.38 at 311.3; 64 to 65 C. The cap never binds on this hash (310 to 314 W under every limit), the power knob is flat at 0.41 MH per W, 2.44 microjoules per hash in the app on PC 1 beside two other cards; the clock caps were not measured (the watchdog ended the run at the first clock step, fixed on ember-tune 5e4ef43) and the card was left clock-locked at about 2,781 MHz until a reset (not a card under test in the PC 1 AMD jobs). The earlier state: the run of 07:21 to 07:56Z produced no 5090 step (its second engine never mined; re-run pending the project lead); the 4 October PC 1 log (miner-eff record, run win-ae432dc7-20261004-164723) gives p95 290 W at 124 MH/s in the app under a 460 W limit, and the card's cap floor is 400 W, so a power cap CANNOT bind on this kernel: item 1's 326 W denominator is a peak with the prover on, the hash alone is about 290 W (2.34 microjoules at 124 MH/s), and the lever that can move the watts is the core clock, now in item 8's PC 2 job (-lgc steps at 2,781 / 2,472 / 2,163 / 1,854 MHz); every DRAM energy figure is a streaming figure applied to random reads; the 9070 XT rows |
PC 1 not released; no chip measured |
| 4 | the 9070 XT rate band for the detector; the coinbase card-model tag (0.3.12) so the chain-attributed band needs no fleet log; the public testnet's first-week baseline (history check 2) | PC 1 not released; a miner change |
| 5 | a rented FPGA hour (the soft-overlay reads-in-flight number is a model on cited HBM figures) | no FPGA in the fleet |
| 6 | the 9070 XT step costs | PC 1 not released |
| 4 + 7 | the live observer restart with both hooks, and the first live detector and vendor-share rows | needs a push to master (the project lead's word) |
| 2 | the 2019-class core (O-1.14), which decides dr736 against dr368; the once-a-day NVRTC module for the item function (required before any activation); cryptanalysis of random ARX programs; the loaded-iGPU tier's build with the day program; the 5090 absolutes re-run with the card quiet (ratios stand) | unmeasured; unimplemented |
| 8 | the 9070 XT and 4060-class rows (where a small card binds); the 5090 clock rows (an elevated job); the 5090 at a 575 W cap (model only); the Mac package watts (IOReport gives GPU + DRAM, Ember's 38 W approximate); the chip side's k, lane area and 28 nm scaling; the Metal fuzz, edge and stats runs and gates G2, G4 to G6 on the class; the acceptance rule's reading of the block | PC 1 not released; no elevated job; the gates are the next step on the project lead's word |
| 6 | mm8 as a chain on Apple (Metal 4 matmul2d; the Mac's Swift toolchain has no tensor API); mm8's exactness on AMD (the gfx12 WMMA fragment layout; an empirical layout probe would settle it in one job); the 5 percent rule per family with the family live (argued only); the RDNA ISA guides unread (mnemonics from LLVM's tables). The 9070 XT step costs themselves are CLOSED (run e) | |
| 1 | the DRAM energy figures are streaming figures applied to random 32-byte reads; the GDDR7 burst and HBM3 tFAW are behind the JEDEC paywall; no chip has been built or torn down | |
| all | every AMD RDNA 4 number in this run | PC 1 is the project lead's desk today. 16:0x UTC: the RX 9070 XT is back on PC 1 (amd:gfx1201, 16,304 MiB) beside the 5090 and an RTX 4070; main has asked for the PC 1 jobs to be PREPARED, not published: (1) G1 AMD plus item 8's ladder on the 9070 XT (about 15 min, the card alone), (2) item 6's family step costs on AMD (about 3 min), (3) item 2's dr736 build and compile on AMD (about 3 min), (4) optional: the 4070 ladder (about 10 min); branch ca3-pc1-amd (1f33cc6, e054ed7, merged): the four scripts pass every CI check, the kit zip sha256 a70fce5b... verified, the AMD watts readback is igneum-gpu-telemetry.exe (ADLX board watts); the commands and the row map in tools/ca3-pc1-amd/README.md; published one at a time on "go PC 1 AMD" after the 0.3.13 update and the Ember table run, about 33 min of PC 1 in all |
7a. The two fast-forwards to master (decision 5)
| Push | Commit | ci | windows-ci |
|---|---|---|---|
| 1 | 88f5026 (ca3-coord merged with master: Counter ASIC 3.0 complete, the two CI check fixes, the close) | GREEN (37508680115) | red on "payload inputs (payload-inputs.zip from the downloads host ...)", red since 630b537 at 17:51Z, the shipper's 0.3.14 payload on the downloads host; the last green windows-ci 2cf8851 at 15:53Z; untouched by this branch |
| 2 | 0396580 (35776fe: the runner's --cards-off with the restore on any exit, the quit-check's rule 2 with the dated allow list, the stripped PC 1 scripts; plus the executable bit on the quit-check) |
GREEN (37509530709) | cancelled (37509530516): superseded by the build-server fix 3e6a488 pushed to master minutes later; the payload-inputs step is the shipper's to turn green with 0.3.15's payload |
8. Close
Closed 6 October 2026, 18:1x UTC. The lanes: item 1 (ca3-analysis), item 2 (ca3-derive), items 4 and 5 (ca3-detector), item 3 (ca3-crypto-brief), items 6 and 7 (ca3-reserve), item 8 (ca3-shadow), the PC 1 AMD jobs (ca3-pc1-amd), the hash gates (ca3-v4-hash) and the node gates with both preconditions (ca3-v4-node and the fork): thank you, every one of you, for the numbers and for the faults you found in your own work and in mine before they reached a cut.