From e8d1e8f8a48150263e2dd2214ba080348adb0a73 Mon Sep 17 00:00:00 2001 From: igneum-josh <337424239+igneum-josh@users.noreply.github.com> Date: Wed, 7 Oct 2026 23:33:15 +0100 Subject: [PATCH] Counter ASIC 4.0 research: 20.2a the per-load fault fixed with clock readings, the drawn-era census (0 pairs across 16 eras), the Metal fingerprints (the tile reference bit-exact; the Apple emulation costs 35 and 78 percent of rate at 1,024 and 4,096 tiles per hash); 20.2b the biased-product-bits requirement from adv-cache-2 Co-Authored-By: Claude Fable 5.1 --- docs/analysis/counter-asic-4-research.md | 20 +++++++++++++ igneum-pow/tests/ca4_trace.rs | 36 ++++++++++++++++++++++++ 2 files changed, 56 insertions(+) diff --git a/docs/analysis/counter-asic-4-research.md b/docs/analysis/counter-asic-4-research.md index 2ad32d23c..8f1c6746c 100644 --- a/docs/analysis/counter-asic-4-research.md +++ b/docs/analysis/counter-asic-4-research.md @@ -349,6 +349,26 @@ The tile cost law from these rows: AVX2 (5.82 - 5.33) ms over (1,430 - 128) x 8 **A finding on the per-load placement (not a tuning).** With the same base program and the same loads, the per-load class derives 3,595 to 3,633 distinct items per warp where the class v4 shape derives all 4,096: 11 to 12 percent of the warp's 4,096 reads land on an item another lane of the same warp has already read. The cause (read from the construction, not yet from a trace): a sub-block of 16 drawn ALU instructions ends right before the next load, and when its last writer of the load's source register is a `shfl` (29 of 256 shadow instructions are shuffles), two lanes carry a neighbour's value into the same address; in the class v4 shape the base program's own instructions re-randomise every register between a shuffle and a load, and the sub-version 3 source rule (a load's source last written by an injecting op or a rotate, judged over base then shadow) does not see the per-load order. Consequences: a within-warp duplicate read is served from L1 or L2 on the GPU and from a lane buffer on a chip, so it costs neither side DRAM energy but it is 11 percent fewer dependent reads per hash, which is exactly the uniformity bound the attack pass gates (F8's 1.2x on the hot set); the per-load class as drawn FAILS that spirit and is not a candidate as exported. The fix is one rule: a sub-block's instructions that would be the last writer of the next load's source are drawn from the injecting families only (or the sub-block ends one instruction before the load with a rotate), the sub-version 3 freshness fixpoint run over the real execution order; 2 to 3 agent hours plus a re-export and the F8 census at 2^24 on 64 seeds. The card rows of the exported pack still answer the energy question (the same instruction count per hash, 11 percent fewer DRAM reads, so the pack reads slightly FASTER than the control and that part of the delta is the duplicate reads, not the placement). +### 20.2a The per-load fault, fixed (22:16 to 22:33 UTC) + +| Reading | Clock (UTC) | Value | +|---|---|---| +| The known-failed record (the first export, id 854050a4293f0615), traced | 22:16 | 10,728 distinct items of 12,288 over three units (the class v4 shape 12,286); 1,482 same-iteration duplicate lanes at sites 8, 10 and 15 only; the sub-block last writers of those loads' sources `mul`, `mul`, `mul`; shuffles and rotates clean | +| The mechanism, from the 64-seed census of the first rule | 22:16 | 29 of 64 seeds failing, up to 620 duplicate lanes a seed, load sources collapsed to 1 to 17 distinct values in 32 lanes; the last base writers named `mulhi`, `mul`, `or`, `rotl`, `load`: a lossy BASE writer followed by 27 passes of a 16-instruction map collapses the register before the next load, so a last-writer rule alone does not cover it | +| The fix, committed (2f718001) | 22:24 | two layers: (1) generator: a per-load sub-block instruction writing the NEXT load's source is redrawn from the injecting families when its op is `mul`, `mulhi` or `or`; (2) `accept.rs`: the dynamic test steps the per-load sub-blocks inside `run_unit` in the order the class executes, and a new rejection `DuplicateLanes` (this class only) refuses a candidate whose load reads one address in two lanes of a unit; a rejected candidate redraws the attempt | +| After, the trace | 22:23 | 12,287 distinct of 12,288, 0 duplicate lanes on the genesis seed (accepted at attempt 3) | +| After, the 64-seed census (2 units each on a second dataset, 16,384 load rows) | 22:23 | 1 duplicate pair in all (seed ca4-census/49, site 3): the chance floor of a 2^24 index space (32 x 31 / 2 / 2^24 per row, about 0.5 pairs expected; the class v4 shape's own trace carries 2 of 12,288 from the same floor) | +| After, the drawn-era split (16 eras `igneum-era-test/0..15` over the per-load class, 2 units each) | 22:32 | R under 28: 12 eras, 3,072 rows, 0 duplicate pairs; R 28 and up: 4 eras, 1,024 rows, 0 pairs; every era accepted at attempt 3 | +| The suite | 22:23 | 64 + 2 + 7 + 4 + 19 + 2 + 7 passed, 0 failed (igneum-build-2) | +| The fixed pack `mx8_shl256x27_v2` | 22:29 | attempt 3, id bd64b207a30413fb, OVERALL PASS; on build-1 and in the hash lane's kit beside the first export (both run on PC 1, the row names its directory and id) | +| Metal fingerprints (M5 Max, `packbench`, 3 batches of 2^24, under the measure lock) | 22:30 | `mx8_sh256x27` 3d2e8245cc084d07 (the 6 October fingerprint, 27.01 MH/s); `mx8_shl256x27_v2` ee5d7c71180e5ea7, vectors 3 of 3 standalone and in batch, 26.88 MH/s (-0.5 percent against the control: the per-load placement costs the Apple GPU nothing); `mx8_mm128` 270e4ae36b37e9a1, 3 of 3, **17.67 MH/s (-35 percent)**; `mx8_mm512` a1c1ff3148d775d1, 3 of 3, **5.98 MH/s (-78 percent)**, compile 15.0 s | + +Two readings from the Metal row that move section 15.2's tile verdict. First, the reference tile (`mm8_ref`: 12 index shuffles and 32 byte products per lane) is bit-exact against the Rust verifier on all three vector warps of both tile packs, so the tile semantics are pinned on two implementations before any card runs the PTX form. Second, the Apple cost is far above the "10 ALU steps per tile" estimate: 1,024 tiles per hash cost the M5 Max 35 percent of its rate and 4,096 tiles 78 percent, so a tile shadow at the ALU shadow's premium (11,440 tiles per hash) would take the Apple tier out entirely unless Metal gains an integer matrix path reachable from the toolchain (`mpp::tensor_ops::matmul2d` with `uchar` operands, unverified). The tile shadow therefore stands as a class only with an Apple exemption nobody has designed, which moves it from rank 3 to beside rank 5 until that path is measured; the k question it answers is unchanged. + +### 20.2b A named requirement for every CA4 prototype: no biased product bits in an address (the crypto lane's adv-cache-2 reading, 7 October 2026, 22:1x UTC) + +The crypto lane's finding: a product's low bits are biased (P(bit 0) = 1/4, measured exactly), the bias survives the odd stride multiplier, and the stride rotation places the biased bits at address bits R and up, inside the 28-bit item index unless R is 28 or more. The devnet era draws R = 29, which cuts them off, so 31 of 32 devnet-era programs read clean while 6 of 16 drawn-era programs (R from 3 to 22) show a site over 1.04x (2 over 1.2x, the worst 1.51x); under the 2 GiB genesis dataset (D = 29) R = 29 would show it too. The devnet's cleanliness is an accident of its era draw; the chain prevalence is the drawn-era figure; the price to a partial-store chip stays under 0.1 percent of a hash's reads per site, so no chip number moves. The requirement for any class this file proposes (the per-load shadow, the tile block, a re-weighted shadow): (1) a load whose source register's last writer is a product (`mul`, `mulhi`, `mad`) carries biased low bits into the address, and the acceptance must judge it at the VALUE level (the bit bias of the index at the site over the units), not by the index-distinctness ratio alone, which the duplicate-lane test above is; (2) the census of any candidate reads across the drawn eras split by R (3 to 22 against 28 to 31), as 20.2a now does for the per-load class, never the devnet era alone. The per-load class's dynamic rule covers distinctness, not bias; the value-level test is owed and is the same item for class v5's acceptance. Nothing in class v4 or v5 moves on this without main's word. + ### 20.3 The card rows (PC 1, the 5090; the hash lane's job) PENDING. Per pack, unlocked and at the 1,300 knee when the helper answers: MH/s, watts, SM MHz, the fingerprint and the self-test verdict. The questions each row answers: `shl256x27` against `sh256x27`: equal watts and a rate inside the block-size effect (16-instruction blocks ran 2.5 to 3.5 percent faster than 256 on 6 October) means the per-load placement is free for the GPU and the capex row of 16.2 stands; `mm128`, `mm512`, `mm1430` against `mx8-genesis`: the premium per tile on the 5090 (the 4090 read 0.056 pJ per MAC at 4,096 tiles per hash), where the rate falls, and whether 11,440 tiles per hash carries the ALU shadow's premium inside the free band; a self-test FAIL on a tile pack is a layout finding, not a tuning. diff --git a/igneum-pow/tests/ca4_trace.rs b/igneum-pow/tests/ca4_trace.rs index 120ba2170..e59483247 100644 --- a/igneum-pow/tests/ca4_trace.rs +++ b/igneum-pow/tests/ca4_trace.rs @@ -113,3 +113,39 @@ fn per_load_shadow_census_64_seeds_has_no_duplicate_reads() { let total: usize = failures.len(); assert!(worst.0 <= 4 && total <= 3, "seeds with duplicate reads beyond the chance floor: {failures:?} worst {worst:?}"); } + +#[test] +fn per_load_shadow_census_over_16_drawn_eras_splits_by_stride_rotation() { + // the Counter lane's ask (adv-cache-2, 7 October 2026, 23:1x UK): read the per-load class's duplicate reads across + // drawn eras, split by the stride rotation R (R 28 and up cuts a product's biased low bits out of the index; + // R 3 to 22 keeps them in), not the devnet era alone + use igneum_pow::generator::{generate_era, EraParams, V3_ALLOWED}; + let ds = DatasetSource::new("2026-10-03", DatasetMode::ClosedForm, 24); + let class = LoadClass::parse("mx8+shl256x27").unwrap(); + let mut low = (0usize, 0usize, 0usize); // eras, rows, duplicate pairs with R under 28 + let mut high = (0usize, 0usize, 0usize); + let mut lines = Vec::new(); + for n in 0..16u32 { + let label = format!("igneum-era-test/{n}"); + let eb = EraParams::test_era_bytes(&label); + let p = generate_era("igneum-genesis", b"igneum-genesis", class, &eb, &V3_ALLOWED); + let era = p.class.era.expect("an era class"); + let seed = igneum_pow::seed::seed_words_from_bytes(b"igneum-genesis"); + let mut dups = 0usize; + let mut rows = 0usize; + for base in [0u32, 1 << 20] { + for row in trace_load_indices(&p, &seed, base, &ds) { + let s: HashSet = row.iter().copied().collect(); + dups += 32 - s.len(); + rows += 1; + } + } + let bucket = if era.stride_rot >= 28 { &mut high } else { &mut low }; + bucket.0 += 1; + bucket.1 += rows; + bucket.2 += dups; + lines.push(format!("era {n} R={} attempt {} dups {dups}", era.stride_rot, p.attempt)); + } + println!("CA4ERA per-load class over 16 drawn eras: R under 28: {} eras, {} rows, {} duplicate pairs; R 28 and up: {} eras, {} rows, {} duplicate pairs; per era {:?}", low.0, low.1, low.2, high.0, high.1, high.2, lines); + assert!(low.2 + high.2 <= 4, "duplicate reads beyond the chance floor across eras"); +}