# Weak-program census and the load-count rule Date: 3 October 2026. Ledger items M6 (weak programs) and M5 (load count varies 6x), `docs/fud-ledger.md`. Status: Measured. Every number below was produced on this machine on this date by the commands in section 10. Machine: Apple M5 Max (12 performance + 6 efficiency cores, 64 GB), macOS Darwin 25.6.0, rustc 1.99.0 (rustup), release build with LTO. Tool: `igneum-census/` (new crate, depends on `igneum-pow` as a path dependency and changes nothing in it). The census ran at `nice -n 15` on 8 threads while a devnet build shared the machine. ## 1. Summary | Question | Answer | |---|---| | Programs examined | 100,000 under the current generator (seeds `igneum-census-2026-10-03/0` to `/99999`), each on 4,096 nonces with the memory-hard 1 GiB dataset of day 2026-10-03; 100,000 more under the proposed generator on 2,048 nonces each; 100,000 again under the current generator with the closed-form dataset to check that the dynamic test does not depend on the day | | Weak programs under the current generator | 1.23 percent fail the register and bias tests (a dead or stuck register, a lane-constant load site, more than 1 percent of final registers saturated, or an output bit biased past 6 sigma); 2.43 percent have a register with no injecting write (a static property that produces most of the stuck registers) | | The bigger finding | 94.8 percent of programs re-read at least one dataset address inside every hash. 19.9 percent of all loads in the population are repeats of an address the same hash already read. The GPU does not pay for a repeat, so the static load count overstates the memory work by 20 percent on average and the hash rate tracks the distinct count, not the static count (section 5). This is the real cause of ledger M5 | | Rejection rule (M6) | Static: no cyclically redundant load, every register has at least one injecting write. Dynamic: 64 fixed warps on a seed-derived closed-form dataset, no constant register bit, no lane-constant load site, saturation under 1 percent, no output bit past 6 sigma, more than 120 distinct addresses per hash on average. Section 6 | | Rejection rate | Under the current generator the rule rejects 95.0 percent (the redundancy alone rejects 94.8 percent), which is why the fix is a generator rule, not a filter. Under the proposed generator it rejects 5.14 percent (3.93 static, 2.05 dynamic, overlapping), so 1.054 candidates per epoch on average. Section 7 | | Load-count rule (M5) | Exactly 16 load slots per program, drawn first from slots 1..63, and a load may only read a register that an earlier instruction of the program wrote and that no load has read since (the fresh-source rule). Every hash then does 128 distinct random reads; a 32-lane unit derives 4,096 items, the design bound of spec section 1.11. Sections 5 and 6 | | Hash-rate spread left by the rule | From the load count: none. Residual from program shape: 1.10x across the eight programs with measured Mac rates (3.51 to 3.85 G distinct loads/s), approximate; on the RTX 5090 the rule puts every hour at about 141 Mhash/s (18.0 G distinct loads/s over 128), against 118 to 321 Mhash/s for the 1st to 99th percentile program today. Section 8 | ## 2. Method Seeds. Program `i` is `generate(seed_words("igneum-census-2026-10-03/" + i))`, the production generator of `igneum-pow/src/generator.rs`, so any program here is reproducible with `igneum-pow hash --seed igneum-census-2026-10-03/i` or `igneum-census show --seed ...`. Interpreter. An instrumented copy of `verify.rs::interpret_warp` (same register-major loops, same op semantics, same `MemhardCpu::fetch` for the dataset words). On every program the first warp is also run through `igneum_pow::hash_warp` and the 32 outputs compared; the run aborts on any disagreement (none occurred). At start the tool checks `igneum-genesis` lanes 0 and 31 against the spec vectors (`1fb0b3bbc1ac8279`, `fa052263a854f3de`). Nonce sample. 128 warps per program (4,096 nonces) drawn as 64 pairs: base `b` uniform over the aligned 32-bit range, partner `b XOR (1 << k)` with `k` uniform in 5..31, from a SplitMix64 stream seeded by the program's seed string. Within a warp, lanes `l` and `l XOR m` for `m` in {1, 2, 4, 8, 16} differ in exactly one nonce bit, so a warp gives 80 single-bit-flip pairs for bits 0..4 and a pair of warps gives 32 for a bit in 5..31. Avalanche is measured over all of them (11,264 pairs per program). Metrics per program (one TSV line, 36 columns): | Metric | Definition | |---|---| | `loads`, `lph` | load instructions, loads per hash (x8) | | `dist_mean`, `dist_min` | distinct masked dataset addresses per hash, mean and minimum over the 4,096 hashes | | `redundant_frac` | `1 - dist_mean / lph`, the share of loads that repeat an address the same hash already read | | `dist_total_frac` | distinct addresses over the whole sample divided by total loads (birthday collisions in 2^28 words are under 0.2 percent of this) | | `s_redundant` | static: loads per iteration whose source register was not written since the previous load from it, counted cyclically over the 64 instructions | | `s_cancel` | static: of those, loads that also share the destination, unwritten in between; the pair is the identity | | `sites_lane_const` | load sites (iteration x instruction) whose 32 lanes read the same address in every warp | | `sites_nonce_const` | load sites whose address is the same in every lane of every warp | | `end_sat_frac`, `end_zero`, `end_ones` | share and counts of final register values (8 x 4,096) equal to 0 or 0xffffffff | | `or_sat_frac` | share of `or` executions whose result is 0xffffffff | | `const_bits_max`, `const_bits_sum`, `nonce_indep_regs` | per register, bits that are the same in every final value over the sample (AND of all values OR NOT the OR of all values): the maximum over registers, the total, and the registers at 32 (nonce-independent) | | `bias_max`, `bias_z_max` | largest deviation of an output bit's ones frequency from 0.5, and in units of `0.5 / sqrt(4096)` | | `aval_mean`, `aval_std`, `flip_min`, `flip_max`, `aval_hi_mean` | output bits flipped per single-bit nonce flip: mean and standard deviation, the least and most flipped output bit, and the mean over the high-bit (cross-warp) pairs alone | | `never_written`, `never_read`, `rotl_only`, `inj_missing`, `last_contract` | static register facts: never a destination; never a source; only `rotl` writes it; no write by an injecting op (`add sub xor mad shfl load`, the ops that are bijective in `dst` and bring another register in); last write of the program is a contraction (`or mul mulhi`) | | `depth`, `mlp` | longest chain of loads in series over the 8 iterations (a load's depth is its source's depth plus one, carried through every op), and `lph / depth` | | `dups` | duplicate 64-bit outputs in the sample; see the note in section 3 | Runtime: 2,797 s for the 100,000-program memory-hard census (8 threads, 128 warps each; 12.8 million warps, about 1.7 ms per warp per thread under load). ## 3. The population under the current generator Distributions (100,000 programs): | Metric | mean | min | p1 | p10 | p50 | p90 | p99 | p99.9 | max | |---|---|---|---|---|---|---|---|---|---| | lph (static loads per hash) | 127.93 | 24 | 64 | 96 | 128 | 160 | 192 | 216 | 256 | | dist_mean (distinct addresses per hash) | 102.44 | 23.96 | 56 | 73 | 104.00 | 128.74 | 152 | 169 | 200 | | redundant_frac | 0.191 | 0 | 0 | 0.069 | 0.189 | 0.316 | 0.421 | 0.500 | 0.632 | | s_redundant (per iteration) | 3.20 | 0 | 0 | 1 | 3 | 6 | 8 | 11 | 15 | | s_cancel (per iteration) | 0.26 | 0 | 0 | 0 | 0 | 1 | 2 | 3 | 4 | | ors | 2.57 | 0 | 0 | 1 | 2 | 5 | 7 | 8 | 12 | | depth (loads in series) | 41.3 | 9 | 24 | 32 | 40 | 56 | 66 | 80 | 104 | | mlp | 3.20 | 1.14 | 1.83 | 2.33 | 3.12 | 4.19 | 5.33 | 6.50 | 9.41 | | sites_lane_const | 0.002 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 29 | | end_sat_frac | 0.0005 | 0 | 0 | 0 | 0 | 0.0000 | 0.0021 | 0.122 | 0.500 | | or_sat_frac | 0.0020 | 0 | 0 | 0.0000 | 0.0001 | 0.0035 | 0.0305 | 0.251 | 0.808 | | const_bits_max | 0.038 | 0 | 0 | 0 | 0 | 0 | 0 | 9 | 32 | | nonce_indep_regs | 0.0004 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 4 | | bias_max | 0.0203 | 0.0107 | 0.0142 | 0.0164 | 0.0200 | 0.0244 | 0.0295 | 0.0342 | 0.111 | | bias_z_max | 2.59 | 1.38 | 1.81 | 2.09 | 2.56 | 3.13 | 3.78 | 4.38 | 14.25 | | aval_mean | 32.000 | 31.854 | 31.916 | 31.954 | 32.000 | 32.046 | 32.084 | 32.111 | 32.148 | | aval_std | 3.9996 | 3.890 | 3.940 | 3.967 | 4.000 | 4.032 | 4.058 | 4.078 | 4.104 | | flip_min | 0.4894 | 0.4727 | 0.4837 | 0.4867 | 0.4896 | 0.4919 | 0.4933 | 0.4943 | 0.4959 | | flip_max | 0.5106 | 0.5046 | 0.5067 | 0.5081 | 0.5104 | 0.5132 | 0.5163 | 0.5189 | 0.5233 | | aval_hi_mean | 32.000 | 31.636 | 31.795 | 31.887 | 32.000 | 32.114 | 32.206 | 32.269 | 32.385 | Load count (static), 100,000 programs. The count is binomial(64, 0.25) times 8: mean 127.9, standard deviation 27.7. Observed range 24 to 256 (ledger M5 quoted 40 to 232 over 10,000). | loads/hash | 24 to 56 | 64 | 72 | 80 | 88 | 96 | 104 | 112 | 120 | 128 | 136 | 144 | 152 | 160 | 168 | 176 | 184 | 192 | 200 to 256 | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | programs | 485 | 643 | 1461 | 2503 | 4372 | 6196 | 8353 | 9937 | 11334 | 11402 | 10903 | 9421 | 7460 | 5642 | 4054 | 2406 | 1550 | 939 | 939 | | cumulative % | 0.49 | 1.13 | 2.59 | 5.09 | 9.46 | 15.66 | 24.01 | 33.95 | 45.28 | 56.69 | 67.59 | 77.01 | 84.47 | 90.11 | 94.17 | 96.57 | 98.12 | 99.06 | 100 | Distinct addresses per hash (mean per program, rounded to the nearest 8): 24 to 48 in 0.6 percent, 56 to 72 in 9.8 percent, 80 to 120 in 74.6 percent, 128 to 144 in 13.3 percent, 152 to 200 in 1.7 percent. The population Summed over the 100,000 programs, one hash of each does 12,792,736 static loads and 10,243,731 distinct ones: 19.93 percent of all loads are repeats. Flags: | Flag | programs | share | |---|---|---| | s_redundant > 0 (a load re-reads an address inside the hash, static) | 94,774 | 94.77% | | dist_mean < lph (the same, measured) | 96,771 | 96.77% | | s_cancel > 0 (two loads cancel to the identity) | 22,530 | 22.53% | | inj_missing > 0 (a register with no injecting write) | 2,429 | 2.43% | | never_written > 0 | 143 | 0.14% | | rotl_only > 0 | 142 | 0.14% | | never_read > 0 | 160 | 0.16% | | last_contract > 0 (a register whose last write is or, mul or mulhi) | 79,833 | 79.83% | | end_sat_frac > 0.001 | 3,367 | 3.37% | | end_sat_frac > 0.01 | 680 | 0.68% | | or_sat_frac > 0.1 | 294 | 0.29% | | const_bits_max > 0 (a register with a nonce-independent bit) | 726 | 0.73% | | const_bits_max >= 4 | 172 | 0.17% | | nonce_indep_regs > 0 (a whole register nonce-independent) | 31 | 0.03% | | sites_lane_const > 0 (a load site all 32 lanes read at one address) | 30 | 0.03% | | sites_nonce_const > 0 | 30 | 0.03% | | bias_z_max > 4 | 409 | 0.41% (405 expected by chance over 64 bits x 100,000 programs) | | bias_z_max > 5 | 12 | 0.012% (3.7 expected by chance) | | bias_z_max > 6 | 6 | 0.006% (0.0 expected by chance) | | aval_mean outside 31..33, or any output bit flipping outside 0.45..0.55 | 0 | 0 | | dups > 0 | 6 | 0.006% (see below) | Avalanche is clean across the whole population: no program has a mean outside 31.85 to 32.15 or an output bit that flips outside 0.473 to 0.523, and the cross-warp (high nonce bit) means stay in 31.64 to 32.39. Bias and avalanche are not where this generator is weak. The six `dups` programs are sampler collisions (two of the 128 base nonces coincided; about 6 expected in 100,000 at 128 draws from 2^27), confirmed by `igneum-census show` ("sampler collision: base nonce ... drawn twice"), so `dups` is not used by any rule. Saturation rises with the number of `or` instructions, as expected, and is not confined to programs with many: | or instructions | programs | end_sat_frac > 0.01 | or_sat_frac > 0.1 | const_bits_max > 0 | |---|---|---|---|---| | 0 | 7,202 | 9 (0.12%) | 0 | 59 (0.82%) | | 1 | 19,449 | 38 (0.20%) | 14 (0.07%) | 120 (0.62%) | | 2 | 25,608 | 68 (0.27%) | 42 (0.16%) | 157 (0.61%) | | 3 | 22,341 | 130 (0.58%) | 79 (0.35%) | 150 (0.67%) | | 4 | 14,076 | 127 (0.90%) | 65 (0.46%) | 112 (0.80%) | | 5 | 7,040 | 133 (1.89%) | 34 (0.48%) | 73 (1.04%) | | 6 | 2,882 | 90 (3.12%) | 26 (0.90%) | 29 (1.01%) | | 7 | 1,009 | 45 (4.46%) | 19 (1.88%) | 14 (1.39%) | | 8 or more | 393 | 40 (10.18%) | 15 (3.82%) | 12 (3.05%) | Programs with no `or` at all still produce stuck registers (59 of 7,202), which is the `mulhi` mechanism below. ## 4. What the tails are Three mechanisms account for the dynamic tails. Each was read off `igneum-census show`, which prints the instruction list and, per register, the share of final values at 0 and at all-ones, the constant bits, and the list of ops that write it. 1. A register with no injecting write (`inj_missing`, 2.43 percent of programs). Program 5110: r3 is written only by `rotl@6 or@20 rotr@57`. `or` with a random word each iteration clears a zero bit with probability 1/2, so after 8 iterations r3 is 0xffffffff in every hash (`ones 1.0000, const bits 32`). r5, written by `mul@8` and three `xor`s whose sources are r3 and r7 (both stuck), ends at 0 in every hash. 29 of the program's load sites read from a stuck register, so all 32 lanes read one address (`sites_lane_const 29`): a GPU serves those loads from one cache line. 18 of the 30 lane-constant programs and 23 of the 31 nonce-independent-register programs have `inj_missing > 0`. 2. A zero-absorbing set of registers (no static signature). Program 52079: r2, r4, r6 and r7 end at 0 in every hash although each has loads and xors among its writes. `mulhi(a, b)` is bits 32..63 of the product, so for uniform inputs it loses about a bit of magnitude per application; a register fed mostly by `mulhi` and `mul` of its neighbours contracts to 0 within a few iterations, and once a set of registers is at 0 it stays there: `mulhi` and `mul` by a zero register give 0, `mad` with a zero factor leaves the destination alone, and a load through a zero register reads `dataset[0]`, a constant that the xor then cancels against the next read of it. The last write of each of the four registers is a `mulhi`, `mul` or `mad` whose other operand is in the set. Half the final state of this program is a constant, and `end_sat_frac 0.5000`. 3. Contraction as the last write (`last_contract`, 79.8 percent of programs, mostly harmless). Program 99109 has no stuck register and no redundant load, but r5, r6 and r7 are finished by `or@53 or@60 or@62` and `mul@59`, so the final r7 is all-ones 12 percent of the time and one output bit has ones frequency 0.611 (`bias_z_max 14.25`). Only 6 programs in 100,000 exceed 6 sigma on any output bit at 4,096 nonces, against 0.03 expected by chance, and all six are of this kind. A bias common to every miner is a difficulty distortion for one hour, not an unfairness, but a rule that costs nothing should still exclude it (section 6). ## 5. Redundant loads: the finding behind M5 A load is `dst ^= dataset[src AND MASK]`. If the next load from the same `src` comes before anything writes `src`, it reads the same address. The generator draws `src` uniformly from the seven registers other than `dst` and writes a uniformly drawn `dst` on every instruction, so after a load from register `s` the next instruction is a load from `s` with probability 1/28 and a write to `s` with probability 1/8; the next load from `s` is a repeat with probability about 0.22, and with 16 loads per program the expected number of repeats is about 3.2 per iteration, which is what the census measures (`s_redundant` mean 3.20, 94.8 percent of programs above 0). Of those, a repeat into the same destination with nothing in between cancels to the identity (`s_cancel`, 22.5 percent of programs): two loads that do nothing at all. The static count matches the dynamic one. `dist_mean` equals `lph - 8 x s_redundant` in nearly every program; the small shortfalls (for example `igneum-genesis/epoch1`: 79.47 against 80) are `rotr` by a lane whose amount is 0 mod 32, `or` into a saturated word and similar value-level coincidences. The GPU does not pay for a repeat. The kernel has no stores to global memory inside the loop, so a compiler is free to reuse the loaded value, and an L1 hit costs nothing against a DRAM miss even if it does not. The Mac rates measured in `proto-metal/README.md` and `MEMHARD.md` section 2.4 bear this out. "G loads/s" is Mhash/s times loads per hash; the distinct count is from this census (`igneum-census probe`, 64 warps per seed): | Seed | Static loads/hash | Distinct per hash | Repeats per iteration | Load critical path | M5 Max Mhash/s | G static loads/s | G distinct loads/s | RTX 5090 Mhash/s | G static | G distinct | |---|---|---|---|---|---|---|---|---|---|---| | igneum-genesis | 104 | 80.00 | 3 | 49 | 45.2 | 4.70 | 3.62 | 228.1 | 23.72 | 18.25 | | igneum-genesis/epoch1 | 104 | 79.47 | 3 | 32 | 48.4 | 5.03 | 3.85 | | | | | igneum-genesis/epoch2 | 112 | 88.00 | 3 | 40 | 40.0 | 4.48 | 3.52 | | | | | igneum-second-seed | 104 | 104.00 | 0 | 41 | 35.5 | 3.69 | 3.69 | | | | | igneum-second-seed/epoch1 | 144 | 104.00 | 5 | 48 | 35.4 | 5.10 | 3.68 | | | | | igneum-hourly | 128 | 96.00 | 4 | 40 | 36.6 | 4.68 | 3.51 | 185.3 | 23.72 | 17.79 | | igneum-hourly/epoch1 | 128 | 95.76 | 4 | 40 | 37.5 | 4.80 | 3.59 | | | | | igneum-hourly/epoch2 | 120 | 96.00 | 3 | 48 | 36.6 | 4.39 | 3.51 | | | | | spread over the 8 Mac rows | | | | | | 3.69 to 5.10, 1.38x | 3.51 to 3.85, 1.10x | | | | Two readings. `igneum-second-seed` (104 static, 104 distinct) and `igneum-second-seed/epoch1` (144 static, 104 distinct) hash at the same rate on the Mac, 35.5 and 35.4 Mhash/s: a 144-load program and a 104-load program do the same memory work because 40 of the 144 are repeats. And README observation 3, which called `igneum-second-seed` an outlier at the same 13 loads as `igneum-genesis`, is resolved: `igneum-genesis` does 80 real reads per hash, `igneum-second-seed` 104. Distinct loads per second vary 1.10x across the eight programs (coefficient of variation 3 percent); static loads per second vary 1.38x (9 percent). The load critical path (32 to 49 here) does not order the residual; the GPU hides that latency with occupancy. The two 5090 programs both carry about 24 percent repeats, so they cannot tell the two readings apart (static and distinct are each constant across them). One run decides it: `igneum-second-seed` on the 5090 is predicted at about 173 Mhash/s if the card is bound by distinct loads (18.0 G/s over 104) and 228 if by static loads (23.7 G/s over 104). That run is the next bench-log item for the Windows machine. Consequence for M5. Fixing the static load count does not fix the memory work per hash: a program with 16 load instructions does anywhere from 6 to 16 distinct reads per iteration today. The rule must fix the distinct count, which the fresh-source rule below does by construction. ## 6. The rules Two generator rules and one rejection rule. The generator rules change every program (the test vectors are re-cut when they are adopted, spec section 1.16 already schedules that); the rejection rule is what makes a node skip a seed. G1, exact load count. Draw the 16 load slots first (a uniform 16-subset of the 64 slots by a partial Fisher-Yates over the program stream), then draw the other 48 ops from the ten non-load weights. Every program has 16 load instructions and 128 loads per hash. G2, fresh source. On a load slot the source is drawn from the registers (other than `dst`) that an earlier instruction of this program has written and that no load has read since that write. Such a register holds a value produced in this iteration, so the load's address cannot repeat any earlier load's address in the hash, including across the iteration boundary. Nothing is eligible at instruction 0, so the 16 load slots are drawn from slots 1..63. If the eligible list is empty (a load early in the list whose few written registers are all `dst` or already read) the source is drawn as on an ALU slot and the program fails R-a below. A first form of G2 was measured too (section 7): eligible meant "not read by a load since its last write", with all eight registers eligible at instruction 0. It stops repeats inside the linear program but not across the wrap, and the wrap alone made R-a reject 36 percent of programs, which is why the definition above is the one proposed. R, rejection (deterministic, evaluated by every node on the candidate program before it is used): | Part | Test | What it catches | |---|---|---| | R-a (static) | no load whose source register is unwritten since the previous load from it, in cyclic order over the 64 instructions | the empty-list case of G2; under the current generator, 94.8 percent of programs | | R-b (static) | every register has at least one write by `add`, `sub`, `xor`, `mad`, `shfl` or `load` | the saturating registers of mechanism 1 (2.43 percent of programs today) | | R-c (dynamic) | over 64 fixed warps (base nonces from SplitMix64 seeded by the program seed, as the census draws them) on the closed-form dataset keyed by the program's own seed words: no register has a bit that is constant over all 2,048 final values; no load site reads one address in all 32 lanes of any warp; final values equal to 0 or 0xffffffff are under 1 percent; no output bit's ones frequency deviates from 0.5 by more than 6 sigma (0.066 at 2,048 nonces); the mean number of distinct addresses per hash exceeds 120 (fewer than one repeat per iteration) | mechanism 2 (zero-absorbing sets), the rest of mechanism 1, mechanism 3, and the value-level repeats of section 7.1 | R-c uses the closed-form dataset (`verify.rs::dataset_elem`, six integer ops) rather than the day's memory-hard dataset so that the test is a pure function of the program, costs under 2 ms on one core (the closed-form interpreter runs at 0.002 ms per warp, `igneum-pow` README), and does not have to wait for the day's cache. The closed-form words are as random-looking as the memory-hard ones for every property R-c measures; section 7.3 checks this on the whole population. A candidate that fails R is skipped and the next candidate is generated from `seed_words_from_bytes(program_seed || k_le32)` for attempt `k = 1, 2, ...` (attempt 0 is the bare seed, so every existing vector stands). Section 7 measures how often that happens. Not in the rule, and why: `last_contract` (80 percent of programs) is too common and R-c already catches the cases where it matters; `or` count caps the same; a cap on the load critical path (`depth`) is a hash-rate question, not a weakness, and section 5 shows the GPU does not care at this depth. Cross-check of the static half against the dynamic flags under the current generator (R-a or R-b, 94,889 programs rejected): of the 726 programs with a constant register bit it rejects 697, of the 30 with a lane-constant load site all 30, of the 31 with a nonce-independent register all 31, of the 680 with more than 1 percent saturation 663, of the 6 past 6 sigma 5. The register and bias conditions of R-c alone (1,232 programs) catch all of those by definition; the static half exists so that the common structural causes are excluded without running anything, and so that a reviewer can read the rule. ## 7. The proposed generator, measured The same 100,000 seeds were run through G1 + G2 (`--gen fixed16-fresh2`, the proposed form) and through G1 with the first form of G2 (`--gen fixed16-fresh`), 64 warps (2,048 nonces) per program, memory-hard dataset, same nonce sampler. The generator is implemented in `igneum-census/src/main.rs` (`generate_fixed16`), not in `igneum-pow`. ### 7.1 Three generators side by side | | Current generator | G1 + G2 first form (all registers eligible at instruction 0) | G1 + G2 proposed (eligible = written earlier in the program and not read by a load since) | |---|---|---|---| | Loads per hash | 24 to 256, mean 127.9 | 128 | 128 | | Distinct addresses per hash, mean (p1 / p50 / p99) | 102.4 (56 / 104 / 152) | 124.9 (113.7 / 127.8 / 128) | 127.7 (120 / 127.9990 / 128) | | Programs with a repeated load (static, cyclic) | 94.77% | 36.19% | 1.5710% | | Programs with a cancelling pair | 22.53% | 1.56% | 0.2390% | | inj_missing > 0 | 2.43% | 2.41% | 2.4050% | | Load critical path, mean (p1 / p99) | 41.3 (24 / 66) | 47.6 (32 / 72) | 51.6 (32 / 80) | | const_bits_max > 0 | 0.73% | 0.73% | 0.6920% | | sites_lane_const > 0 | 0.03% | 0.03% | 0.0390% | | end_sat_frac > 0.01 | 0.68% | 0.61% | 0.6000% | | bias_z_max > 6 | 6 programs | 2 programs | 2 programs | | R-a or R-b (static) rejects | 94.89% | 37.44% | 3.93% | | R-c register and bias conditions reject | 1.23% | 1.15% | 1.08% | | R-c distinct-count condition (mean distinct per hash above 120) rejects | 93.36% | 6.28% | 1.06% | | R-c (all) rejects | 93.44% | 7.19% | 2.05% | | R (all) rejects | 94.95% | 38.25% | 5.14% | | Expected candidates per epoch | 19.8 | 1.62 | 1.054 | Reading. The register-level weaknesses (stuck registers, saturation, lane-constant sites) are properties of the ALU mix and come out the same under every generator, at about 1.2 percent of programs; the generator rules do not touch them and R-c is what removes them. The first form of G2 removes repeats inside the linear program but leaves the wrap (a load from a register that was read by a load late in the previous iteration and not written since), which by itself fails 36 percent of programs. The proposed form is cyclically fresh by construction and fails R-a only when a load early in the list finds no written-and-unread register other than its destination. Under the proposed form a load's address can still coincide with an earlier one by value. Mostly this is noise at the 1-in-32 level (`rotr` by a lane amount that is 0 mod 32, `mad` with a zero product, `or` into a saturated word), different per lane, and the accepted population's `dist_mean` sits within a fraction of a load of 128 (section 7.2). In 1.06 percent of programs it is structural: `igneum-census show` on program 83504 lists 16 pairs of load sites that read the same address in every lane (two registers holding the same word when the second load executes; the cause was not isolated), so the hash does 112 reads where the kernel text says 128. That is why R-c carries the distinct-count condition: a program must average more than 120 distinct addresses per hash over the test warps, and the accepted population then has no program below 120.05. ### 7.2 The accepted population under the proposed generator Metrics over the 94,858 programs of the proposed generator that R accepts: | Metric | mean | min | p1 | p50 | p99 | max | |---|---|---|---|---|---|---| | lph | 128 | 128 | 128 | 128 | 128 | 128 | | dist_mean | 127.89 | 120.05 | 126.91 | 128.00 | 128 | 128 | | redundant_frac | 0.0009 | 0 | 0 | 0.0000 | 0.0085 | 0.0621 | | depth (loads in series) | 51.7 | 24 | 32 | 49 | 80 | 104 | | mlp | 2.58 | 1.23 | 1.60 | 2.61 | 4.00 | 5.33 | | end_sat_frac | 0.0000 | 0 | 0 | 0 | 0.0018 | 0.0094 | | or_sat_frac | 0.0012 | 0 | 0 | 0.0001 | 0.0243 | 0.2304 | | const_bits_max | 0 | 0 | 0 | 0 | 0 | 0 | | bias_z_max | 2.60 | 1.37 | 1.81 | 2.56 | 3.76 | 5.88 | | aval_mean | 32.000 | 31.787 | 31.882 | 32.000 | 32.119 | 32.223 | | flip_min | 0.4851 | 0.4671 | 0.4769 | 0.4854 | 0.4906 | 0.4941 | | flip_max | 0.5150 | 0.5062 | 0.5094 | 0.5146 | 0.5231 | 0.5326 | Every accepted program does 128 loads per hash of which at least 120 and typically 128 are distinct, has no constant register bit, keeps saturation under 1 percent, and shows the avalanche and bit-flip figures of an ideal 64-bit function to within sampling noise at 2,048 nonces. The load critical path is longer than under the current generator (median 49 against 40 loads in series) because a fresh source is often the register the previous load just wrote; section 5 found no rate effect of the critical path at this depth on the Mac, and the 5090 run of ten programs from this generator (section 10) is what settles it on NVIDIA. ### 7.3 The dynamic test does not depend on the dataset R-c is specified on the closed-form dataset so that it is a pure function of the program. To check that this sees the same programs as the memory-hard dataset, both censuses were repeated with `--closed-form` (same seeds, same warps, `dataset_elem` words instead of the cache-derived ones; 125.6 s and 59.4 s on 8 threads once the machine was quiet) and the R-c conditions compared program by program. Current generator, 100,000 programs, 4,096 nonces each, same nonce sample on both datasets: | Condition | memory-hard | closed-form | both | memory-hard only | closed-form only | |---|---|---|---|---|---| | const_bits_max > 0 | 726 | 729 | 708 | 18 | 21 | | const_bits_max >= 4 | 172 | 168 | 164 | 8 | 4 | | nonce_indep_regs > 0 | 31 | 36 | 22 | 9 | 14 | | sites_lane_const > 0 | 30 | 33 | 22 | 8 | 11 | | end_sat_frac > 0.01 | 680 | 679 | 679 | 1 | 0 | | bias_z_max > 6 | 6 | 8 | 5 | 1 | 3 | | dist_mean <= lph - 8 | 93362 | 93362 | 93362 | 0 | 0 | | R-c, register and bias conditions | 1232 | 1232 | 1216 | 16 | 16 | | R-c, all five conditions | 93437 | 93438 | 93437 | 0 | 1 | Per-program differences between the two datasets: `dist_mean` max 0.327 (mean 0.0021), `end_sat_frac` max 0.0110 (mean 0.000018), `or_sat_frac` max 0.0102, `bias_z_max` max 3.50 (mean 0.464, the sampling noise of two independent draws), `aval_mean` max 0.220. Proposed generator (G1 + G2), 100,000 programs, 2,048 nonces each, same nonce sample on both datasets: | Condition | memory-hard | closed-form | both | memory-hard only | closed-form only | |---|---|---|---|---|---| | const_bits_max > 0 | 692 | 681 | 635 | 57 | 46 | | const_bits_max >= 4 | 147 | 145 | 137 | 10 | 8 | | nonce_indep_regs > 0 | 31 | 34 | 24 | 7 | 10 | | sites_lane_const > 0 | 39 | 40 | 33 | 6 | 7 | | end_sat_frac > 0.01 | 600 | 600 | 599 | 1 | 1 | | bias_z_max > 6 | 2 | 2 | 2 | 0 | 0 | | dist_mean <= lph - 8 | 1057 | 1056 | 1056 | 1 | 0 | | R-c, register and bias conditions | 1078 | 1079 | 1059 | 19 | 20 | | R-c, all five conditions | 2054 | 2055 | 2035 | 19 | 20 | Per-program differences between the two datasets: `dist_mean` max 0.416 (mean 0.0042), `end_sat_frac` max 0.0123 (mean 0.000023), `or_sat_frac` max 0.0159, `bias_z_max` max 3.31 (mean 0.464, the sampling noise of two independent draws), `aval_mean` max 0.317. Reading. The per-program metrics are the same to three decimals on both datasets; saturation, the distinct count and the lane-constant sites are properties of the program, and the dataset words only need to look random for them to show. The verdicts that differ are at the edges of the thresholds: every `const_bits_max > 0` disagreement is a program with exactly one constant bit on one dataset and a nearly constant bit on the other (the strong cases, four or more constant bits, agree in 164 of 172 and 168 under the current generator and 137 of 147 and 145 under the proposed one), the `bias_z_max > 6` disagreements are programs between 5 and 7 sigma, and the single `end_sat_frac` disagreement sits at 1.0 percent. Under the proposed generator the full R-c verdict agrees on all but 39 of 100,000 programs (2054 rejected on the memory-hard dataset, 2055 on the closed-form one), all threshold-edge cases of the kinds just listed. A borderline program that one dataset accepts and the other rejects is a program with a nearly constant bit or a bias near 6 sigma; whichever side of the line it falls, R-c on the closed-form dataset is the definition, every node evaluates the same one, and the cost of accepting a borderline program is a near-constant bit in one register, not a shortcut. ## 8. Hash-rate spread, before and after The RTX 5090 is bound at about 23.7 G random loads per second past its L2 (`docs/bench-log.md`, RTX 5090 sweep), which is 18.0 G distinct loads per second once repeats are discounted (section 5, mean of the two programs). Rate per hour is that figure over the hour's loads per hash. | Population | Distinct loads per hash (p1 / p50 / p99, min / max) | RTX 5090 Mhash/s at p1 / p50 / p99 | Spread p1 to p99 | Spread min to max | |---|---|---|---|---| | Current generator, by static count (the ledger's reading) | 64 / 128 / 192, 24 / 256 | 370 / 185 / 123 (at 23.7 G static loads/s) | 3.0x | 10.7x | | Current generator, by distinct count (what the card does) | 56 / 104 / 152, 24 / 200 | 321 / 173 / 118 | 2.7x | 8.3x | | G1 + G2 + R | 128 / 128 / 128 | 141 / 141 / 141 | 1.0x from the load count | 1.0x | What remains after the rule is the program-shape residual: 1.10x across the eight programs with measured Mac rates (section 5), approximate, cause not isolated, to be re-measured on the 5090 with ten programs from the new generator. The Mac projection is 3.62 G distinct loads/s over 128, about 28 Mhash/s per hour on the M5 Max. Both projections assume the card stays random-access bound at 128 distinct loads, which the 5090 sweep showed for 104 and 128 static loads. Note the level: 141 Mhash/s on the 5090 is below today's 185 and 228 because today's measured programs do 80 and 96 real reads per hash, not 104 and 128. An alternative that keeps today's median memory work is 13 load slots (104 distinct, 173 Mhash/s, 3,328 items per unit). The recommendation is 16: it is the spec's own candidate, it matches the 25 percent load weight, it puts a unit at exactly the 4,096-item design bound of spec section 1.11, and the Rust verifier at 4,608 items measured 0.579 ms per unit, so 4,096 items is about 0.52 ms against the 10 ms gate. Hash rate is a number the difficulty absorbs; memory work per hash is the defence. ## 9. Proposed spec text For `docs/spec/01-lottery-hash.md`, replacing the "What fixes them" paragraph of section 1.4.2 and the op draw of section 1.4.3, and adding a section 1.4.6. The spec's owner applies it; nothing below has been written into the spec. > **1.4.2 Op weights and the load count.** Every program contains exactly 16 `load` instructions (`LOAD_SLOTS`, > Definition), so every hash performs 128 loads and a 32-lane unit derives at most 4,096 dataset items, the > bound of section 1.11. The other 48 instructions are drawn from the ten non-load families with the weights of > the table (12 10 8 8 8 7 6 6 6 4, sum 75). The load weight 25 of the earlier table is retired; it survives only > as the ratio 16 of 64. > > **1.4.3 Draw order.** From the program stream of 1.3.3, in this order. (1) Load slots: let `p[0..62] = 1..63` > (instruction 0 is never a load: nothing is fresh before it); for `i` in 0..15 draw `j = i + below(63 - i)` and > swap `p[i]` and `p[j]`; the load slots are `p[0..15]`. > (2) For each instruction `k` in 0..63, nine draws: `roll = below(75)` (on a load slot the roll is drawn and > ignored; otherwise the op is the first entry of the ten-family table whose cumulative weight exceeds `roll`); > `dst = below(8)`; on an ALU slot `a = below(7)` and `src = a + (a >= dst)`, on a load slot `a = below(|E|)` > and `src = E[a]` where `E` is the list, in register order, of registers other than `dst` that an earlier > instruction of this program has written and that no later `load` has used as its source (`E` is empty before > instruction 0), and if `E` is empty then `a = below(7)` and `src = a + (a >= dst)` as on an ALU slot; > `b = below(8)`; `imm = low32(next())`; `imm2 = low32(next())`; `rot = 1 + below(31)`; `bit = below(32)`; > `mask = 1 << below(5)`. 592 draws per program. A load's source holds a value written in the same iteration > that no earlier load has read, so no load repeats the address of an earlier load of the same hash; 1.4.6 (a) > excludes the programs where `E` was empty. > > **1.4.6 Program acceptance.** A candidate program is accepted only if all of the following hold, and every > conforming implementation MUST evaluate them identically. (a) For every `load`, some instruction between the > previous `load` from the same source register and this one, in cyclic order over the 64 instructions, writes > that register. (b) Every register `r0..r7` is the destination of at least one `add`, `sub`, `xor`, `mad`, > `shfl` or `load`. (c) The program is interpreted (section 1.7) for 64 units at base nonces > `low32(next()) AND NOT 31` from a SplitMix64 stream seeded with `FNV-1a-64("igneum-accept/" || seed words as > little-endian bytes)`, with init words `I` equal to the seed words and dataset words > `dataset_elem(idx, S[0], S[1])` of `verify.rs` in place of the memory-hard dataset, and over those 2,048 > evaluations: no register has a bit equal in every final value; no `load` site (iteration, instruction) reads > one address in all 32 lanes of any unit; the number of final register values equal to 0 or 2^32 - 1 is below > 164 (1 percent of 16,384); every output bit's ones count is within 6 x sqrt(2048) / 2 = 136 of 1,024; and the > number of distinct masked dataset addresses read by one lane in one evaluation, summed over the 2,048 > evaluations, exceeds 245,760 (a mean above 120 of the 128 loads). If > the candidate fails, attempt `k + 1` is generated from `seed_words_from_bytes(program_seed || k_le32)` for > `k = 1, 2, ...`, attempt 0 being `seed_words_from_bytes(program_seed)`; the first accepted candidate is the > program of the epoch. Measured rejection rate under this generator: see `docs/analysis/weak-program-census-2026-10-03.md` > section 7: 5.14 percent, so the probability that 15 consecutive candidates fail is below 2^-64, and an > implementation MAY treat 32 consecutive failures as a consensus fault. The 592-draw count: 16 slot draws plus 64 x 9. The test-vector consequence: every pack and every vector in sections 1.4.3, 1.15 and 1.17 is re-cut on adoption; `igneum-census show --gen fixed16-fresh --seed igneum-genesis` prints the first program of the new generator today (`--gen fixed16-fresh2`, op mix `load=16 add=8 shfl=8 xor=6 mad=5 mul=5 mulhi=5 sub=4 rotl=3 rotr=3 or=1`, load critical path 48). ## 10. What this does not show, and reproduction Not shown: any GPU number for the new generator (the 141 Mhash/s figure is a projection from the 5090's measured random-load rate; the next 5090 session should run ten programs from `--gen fixed16-fresh2` and `igneum-second-seed` through the existing pack path); the distinct-versus-static reading on NVIDIA (one run decides it, section 5); an adversary who grinds the epoch seed (the VDF of section 4 of the spec is the answer, and the rejection rule removes the programs a grinder would want); cryptographic strength of anything (ledger M7); whether 6 sigma at 2,048 nonces is the right bias threshold for the lottery (it is the loosest threshold that catches every bias the census found and triggers by chance about once in 10^6 programs). Commands (from `igneum-census/`, `~/.cargo/bin/cargo build --release` first; `S` is a scratch directory): ``` nice -n 15 ./target/release/igneum-census run --root igneum-census-2026-10-03 --count 100000 --warps 128 --threads 8 --gen default --out $S/census-default-100k.tsv nice -n 15 ./target/release/igneum-census run --root igneum-census-2026-10-03 --count 100000 --warps 64 --threads 8 --gen fixed16-fresh2 --out $S/census-fixed16-fresh2-100k.tsv nice -n 15 ./target/release/igneum-census run --root igneum-census-2026-10-03 --count 100000 --warps 64 --threads 8 --gen fixed16-fresh --out $S/census-fixed16-fresh-100k.tsv nice -n 15 ./target/release/igneum-census run --root igneum-census-2026-10-03 --count 100000 --warps 128 --threads 8 --gen default --closed-form --out $S/census-default-cf-100k.tsv ./target/release/igneum-census summarise --in $S/census-default-100k.tsv ./target/release/igneum-census probe --warps 64 --seed igneum-genesis --seed igneum-genesis/epoch1 ... (section 5 table) ./target/release/igneum-census show --warps 16 --seed igneum-census-2026-10-03/52079 (section 4 diagnostics) ``` The TSVs (about 25 MB each) are not checked in; every table is reproducible from the commands above.