Squashed from six commits (tag ca2-era-pre-squash) for one merge. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
20 KiB
Era layout: table layout and working set drawn per era and per program (Counter ASIC 2.0, layers 4 and 8)
5 October 2026, branch ca2-era, worker "ca2-era". Status: Designed and Implemented behind the class flag (LoadClass::era, not the lottery hash); nothing here changes the default generator, the pinned packs or any live program. Measured sections are marked as such; everything else is design.
Plan: docs/plans/counter-asic-2.md, layers 4 ("table layout drawn per era: item size, stride, interleave") and 8 ("working-set size drawn per program"). The era seed is E_n of spec 04 section 4.4. Confirmed on 5 October 2026 by grep over igneum-pow/src on branches master, readwidth and opencl-rdna4-telemetry: no era draw existed in code before this branch (igneum-era, EraParams, era_seed: no match).
1. What is drawn, and from what
Two streams, nothing else:
| Stream | Seeded from | Draws | Sets |
|---|---|---|---|
| Era stream | seed_words_from_bytes("igneum-era/" || E_n) words 0 and 1 (spec 01 section 1.13.1 wrote n_le64 || E_n; the index is dropped here because E_n already commits to n through the VDF input of section 4.4 step 2, and the node's seam hands the generator the era bytes alone: Epoch::from_chain_seeds(epoch, day, era, class, label), branch ca2-v3) |
7 per era, fixed | the load width W, the stride (M, R), the interleave pos[0..3] |
| Program stream | the epoch seed words as today (spec 01 section 1.3.3) | 11 per instruction instead of 9: the 9 of version 2, then 2 window draws (a class whose loads are not version 2's takes the read-width width roll between them, ca2-v3's rule) | per load site: the window shrink k_off and its offset o |
The era parameters change the class of every program of the era. The program stream does not see the era parameters (two eras with the same epoch seed draw the same instruction list and the same windows, and differ in width, stride and layout); this is what makes the six era packs below a controlled comparison.
1.1 Era draw (proposed spec text for section 1.13.1, replacing its parameter table)
One SplitMix64 stream S seeded with lo = words[0] | (words[1] << 32) of seed_words_from_bytes("igneum-era/" || E_n). Seven draws, in this order, whether or not a value is used:
W = allowed[below(|allowed|)]: the width in words of every dataset load of the era, drawn from the genesis-fixed ascending setallowed, a subset of {1, 4, 16} (4, 16 or 64 bytes). A set of one element pins the width; the draw is still consumed. The set is{1}(4 bytes, v2's load): the read-width decision of 5 October 2026 (docs/plans/read-width.md, "keep v2; w16 the only width that passes the rules and closes nothing") and the adoption rule of the same evening (a draw that changes the bytes per hash changes the rate; the six-era hash-rate spread must stay under 5 percent per card). The set is recorded in every pack (IGNEUM_ERA_ALLOWED_WIDTHS,program.jsonera.allowed_widths) and enters the program id. The code keeps the draw general so the set can be widened at genesis without a new derivation.M = low32(next()) OR 1: the stride multiplier, odd, sox -> x * Mis a bijection on 32-bit words.R = 1 + below(31): the stride rotation, in 1..31 (never 0: spec 01 section 1.14 item 3).- to 7.
r_i = next()foriin 0..3: the interleave draws. Letb = log2(W)(0, 2 or 4) andfree = 4 - b. Letc = [b, b + 1, ..., 15](16 - b candidates). Foriin0..free:j = i + (r_i mod (16 - b - i)), swapc[i]andc[j]. The interleave ispos = [0, ..., b - 1] ++ sort(c[0..free]), four ascending bit positions in 0..15. Drawsr_free..r_3are consumed and ignored.
The era parameters are (W, M, R, pos). Era 0 of the devnet packs is listed in section 5.
1.2 Dataset mapping with the interleave (replaces the last sentence of section 1.8.5)
Word w of the dataset holds word j(w) of item t(w), where j(w) is the 4-bit number whose bit i is bit pos[i] of w, and t(w) is w with bits pos[0..3] removed (the remaining bits in order). With pos = [0, 1, 2, 3] this is today's dataset[w] = item(w >> 4)[w AND 15], byte for byte.
Properties kept:
- An item has the same value at every dataset size (item derivation is untouched), and because every
pos[i] < 16,dataset[w]is the same at every dataset size of at least 2^16 words. The 1 GiB vectors of an era remain valid for words below 2^28 at any larger size, as today. - The low
b = log2(W)positions are 0..b-1, so theWwords of one aligned load lie in one item (tis the same for all of them andjrunsj0 .. j0 + W - 1). The verifier derives one item per lane per load, as today: the 4,096-item bound of section 1.11 holds (16 loads x 8 iterations x 32 lanes, whatever the width). - The dataset build writes 16 words of one item to 16 addresses
w(t, j)(a scatter of 4-byte writes instead of one 64-byte line whenpos != [0, 1, 2, 3]). This is the only GPU cost of the interleave and is paid once per day; section 6 measures it.
What the interleave does and does not buy. A chip that hard-wires today's layout (64-byte items, a 64-byte line per item) reads the wrong 15 words with every word once the era draws another layout; the layout changes every 180 days inside rules fixed at genesis. A chip whose address decoder can permute 28 address lines under firmware control pays nothing for it. The honest claim is the first sentence only. The stride below is the same kind of lever: two integer operations per load on a chip, nothing on a GPU.
1.3 Load address (replaces "a load reads one 4-byte word at src AND MASK" in section 1.5 for era programs)
For a load site with window draws (k_off, o) (section 1.4), a dataset of 2^D words (MASK = 2^D - 1), and register value x:
k = min(k_off, D - 26) (0 when D <= 26)
y = rotl(x * M, R)
idx = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK
base = idx AND NOT (W - 1) (W words from base are folded as verify::fold_words)
Uniformity: x * M with M odd and rotl are bijections of the 32-bit word, so y is uniform when x is; y AND (MASK >> k) is uniform on the window; the offset picks which of the 2^k aligned windows. Branch-free, integer only, three operations before the mask (multiply, rotate, and-or) against one today. The emitted text has one form per dialect, checkable by text search (section 1.14 item 2): CUDA and OpenCL ds[((rotl_imm(rN * 0x........u, Ru) & 0x........u) | 0x........u) & mask], Metal the same with dataset[ and & MASK].
1.4 Window draw per load site (layer 8; proposed text for section 1.4.3)
After the nine draws of version 2 (and the width roll, for a class whose loads are not version 2's), every instruction takes two more draws, used only on a load slot:
k_off = below(3) window = the dataset, a half or a quarter of it (2^(D - k_off) words)
o = low32(next()) AND (2^k_off - 1) which aligned window
Bounds: the window never goes below 2^26 words (256 MiB; k = min(k_off, D - 26) in 1.3), which exceeds the largest on-chip cache of any card in the benchmark (the RTX 5090's 96 MiB L2, the RX 9070 XT's 64 MB Infinity Cache, vendor figures), and never above the dataset. At the prototype dataset (2^28) the windows are 1 GiB, 512 MiB and 256 MiB; at the genesis dataset (2^29) 2 GiB, 1 GiB and 512 MiB. The dataset grows by the step schedule recommended to the project lead (spec 01 section 1.13.3 option (b), docs/analysis/card-lifetime-2026-10-05.md: power-of-two steps, 4 GiB at year 4, 8 GiB at year 12, 16 GiB at year 28, 32 GiB at year 60, every index AND MASK), so the window ceiling follows the steps and the floor stays the genesis constant 2^26 words; nothing in the address of 1.3 needs a range reduction. A program has 16 load sites and so up to 16 windows; the set a program reads is their union (section 7 computes its distribution). The verifier bound is unchanged (1.2).
Why per load site and not per program: a per-program window of a quarter of the dataset would hand a 256 MiB SRAM mirror a third of the hours at the prototype size. Sixteen sites with drawn offsets cover the dataset with high probability (section 7), so the mirror a chip would need is the whole dataset in every hour, and the hour-to-hour variation lands on the memory design (which quarter, which half, how many distinct windows), not on its size.
1.5 Program stream (replaces "592 draws per program" in section 1.4.3 for era programs)
16 slot draws, then 64 x 11 = 704: 720 draws per program (768 and 784 for a class with the width roll). On the chain an era program is a class v3 program (branch ca2-v3's seam: ProgramClass::V3, generator version 3): its id is program_id(3, seed words, attempt) as that branch defines it, and the era it was drawn under is identified beside the id by IGNEUM_ERA_SEED_HEX (packcheck::verify_pack_dir_chain refuses a pack whose era is not the job's), so the pair (id, era seed) names the program. The experiment classes that are not class v3 (--class <other> --era ...) carry the era inside the class id instead: FNV-1a-64("igneum-program-rw/" || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots || "era/" || allowed[3] || W || M_le32 || R || pos[4]). The class name is the base name with -era<first stream word as hex> (w4-era401998a5).
1.5.1 The seam (branch ca2-v3, kept as its signatures stand)
V3_CLASS is the base class of class v3 (16 loads of 4 bytes, the lottery hash's load; the era rides inside it), V3_ALLOWED = [1] its width set. generate_from_seed_bytes_program_class(label, seed, V3, Some(era)) draws LoadClass::era(V3_CLASS, era, &V3_ALLOWED) and stamps generator 3; without era bytes (a template before the era is known) the bare V3_CLASS stands. Epoch::chain_dataset(day, class) is unchanged: the layout of 1.2 is a property of the program (program.class.layout()), applied by the interpreter and by Epoch::dataset_word, so the day's cache is shared by every era of a day and the engine keys its caches on (day, class) as before.
1.6 Acceptance
The rule of section 1.4.6 is unchanged in its tests. Its interpreter mirrors 1.3 at the rule's constant D = 28 (idx as above with MASK = 0x0fffffff), as igneum-pow/src/accept.rs does. The distinct-address bound counts dataset loads as the read-width branch defines it.
2. The era seed on the devnet (stand-in for E_n)
Until the 1-hour VDF of section 4.4 is in the node, in the shape of the epoch seed's stand-in (docs/fork-divergence.md "Epoch seed"):
| Era | E_n |
|---|---|
| 0 | the genesis block hash (32 bytes) |
| n >= 1 | the hash of the last selected-chain block whose DAA score is below 15,552,000 n - 7,200 (the 2-hour lead of section 4.4 step 1) |
The era of a block is floor(DAA score / 15,552,000), a function of the header alone. E_n for n >= 1 is known 7,200 DAA seconds before the era starts, which covers the 1-hour VDF when it arrives and the kernel compile and dataset rebuild now. Test seeds for packs and tests: E_n = the 32 bytes (little-endian words) of seed_words_from_bytes("igneum-era-test/<n>") (igneum-pow ... --era igneum-era-test/<n>, the number only names the pack); raw bytes with --era <n>:<64 hex>.
3. Memory budget (the project lead, 5 October 2026: under 6 GB on an 8 GB card)
The era layout adds no resident memory: a window is a mask and an offset in the kernel text, the interleave is address arithmetic, the stride is two operations. The whole working set on a card, every item from this branch and the others:
| Item | Bytes | Who |
|---|---|---|
| Dataset (prototype) | 1 GiB | existing |
Cache (256 MiB, resident only while the day's dataset is built, then free; the decided reading, confirmed by the coordinator on 5 October 2026, docs/analysis/card-lifetime-2026-10-05.md carries the per-tier working set with the cache freed as the best case) |
256 MiB peak | existing |
| Layer 5 hot table | the ca2-cache worker's figure | not this branch |
| Scratch per resident warp (read-width variant 5, 32 or 128 KiB per warp) | 2,048 warps x 128 KiB = 256 MiB at most | readwidth branch |
| Output and read-back buffers | 2^24 nonces x 8 B = 128 MiB per dispatch | existing harness |
| Era windows, stride, interleave | 0 | this branch |
The era window never exceeds the dataset, so it never grows the footprint.
4. Implementation (behind the flag)
| Piece | Where | What |
|---|---|---|
EraParams, era_draw, LoadClass::era |
igneum-pow/src/generator.rs |
the 7-draw era stream of 1.1; the class carries era: Option<EraParams> beside mix, load_slots, scratch, scratch_kb; the two window draws per instruction (Instr::win, Instr::off); the program id of 1.5 |
Layout |
igneum-pow/src/memhard.rs |
split(w) -> (t, j), join(t, j) -> w, LINEAR = [0, 1, 2, 3]; MemhardCpu and DatasetSource carry it |
load_index |
igneum-pow/src/verify.rs |
the address of 1.3, shared by the interpreter and the acceptance mirror |
| Emitters | igneum-pow/src/emit.rs |
the one load form of 1.3 in Metal, CUDA and OpenCL C; mh_word, the three igneum_build kernels and the Metal build kernel with mh_t, mh_j, mh_addr when the layout is not linear; IGNEUM_ERA_* in program.h, an "era" object in program.json |
| CLI | igneum-pow/src/main.rs |
--era <igneum-era-test/n | n:hex> and --era-widths 4,16,64 (one width pins) on every command |
| Tests | igneum-pow/src/*.rs, igneum-pow/tests/packs.rs |
the draw is deterministic and within bounds; split/join are inverse and the dataset is a prefix at every size; six era programs pass the generator contract and the acceptance rule; vectors round-trip; the pinned packs are byte-identical; the six era packs match the emitters and every load has the form of 1.3 |
The default class is untouched: LoadClass::V2 has era: None, every emitter branch on era keeps today's text, and tests/packs.rs diffs the pinned packs (igneum-genesis-mh, igneum-devnet-v4-epoch0) against the emitters as before. Section 6 records the diff of a fresh export against the checked-in files.
5. The six era packs
proto-cuda/packs-ca2-era/era-<n>, n in 0..5: the devnet's 32-byte epoch seed edc4fa84...fb07 and day bytes igneum-day/20730 (the seeds of the pinned pack igneum-devnet-v4-epoch0, which is the v2 baseline with the same program seed), dataset 2^28 words, era seed igneum-era-test/<n>, width pinned at 4 bytes (--era-widths 4, the default). The program seed is held fixed so that the six packs differ in the era parameters only (section 1); each carries seeds.txt for the one-click workers. Every pack is attempt 1 (attempt 0 of this seed is rejected under the era class: 12 draws per instruction give a different stream from v2's). The drawn parameters (igneum-pow show --epoch-hex edc4... --era igneum-era-test/<n>, 5 October 2026):
| Pack | Class | Era seed E_n (first 16 hex) |
W (bytes) | M | R | pos |
|---|---|---|---|---|---|---|
| era-0 | w4-erab2ed8a89 | 5e0587f455a86e91 | 4 | 0x625e5ab3 | 19 | 0, 2, 10, 15 |
| era-1 | w4-era676a17fc | df57136f2ad5f410 | 4 | 0xb2a9d70d | 6 | 1, 3, 8, 13 |
| era-2 | w4-era843155d7 | 7f450623297a954f | 4 | 0x2b4a5b97 | 28 | 1, 3, 4, 8 |
| era-3 | w4-erad6367bfe | 8bffdd3366b9c3ff | 4 | 0x27ea7eff | 30 | 2, 3, 8, 13 |
| era-4 | w4-era4488f3ed | e593fc1d48475c88 | 4 | 0x4d38603d | 10 | 2, 9, 13, 15 |
| era-5 | w4-eraf897c84e | ff87ad96a1b53f36 | 4 | 0x03ac37ad | 22 | 0, 2, 10, 13 |
All six are class v3 packs (generator 3, IGNEUM_PROGRAM_CLASS "v3", IGNEUM_ERA_SEED_HEX), attempt 0, program id 73bcbfe8ccf988f1 in every pack (the seam's program_id(3, seed, attempt); the era seed beside it names the program), the era layout over version 2's item construction (mixer x1, the genesis cache) so that the v2 baseline pack is the same dataset; the chain's class v3 composes the same draw over LoadClass::MX4 (mixer x4, growth), and the integration re-exports these packs on it after the PC rows. The era-seed-to-pack assignment above is from program.h of each pack; the test-seed numbering is only the pack name.
The windows are a property of the program, so they are the same in all six packs (site:shrink:offset): 1:2:3 4:2:2 6:0:0 10:1:0 12:1:1 20:2:0 27:1:0 30:0:0 35:0:0 40:0:0 41:0:0 43:2:3 45:1:0 52:0:0 53:2:2 54:0:0: 7 sites read the whole dataset, 5 a half, 4 a quarter; the union is the whole dataset.
6. Measurements
Pending at the time of this commit; each table below says the machine, the date, the harness and the command when filled.
6.1 Byte-identical default path
6.2 Bit-exactness on the Mac (Metal, Apple OpenCL, CUDA emulation)
6.3 Hash rate per era on the M5 Max (Metal) and the CPU verifier
6.4 PCs (RTX 5090 CUDA, RX 9070 XT OpenCL): prepared, waiting for the go
7. The chip-model line per draw
From docs/analysis/m16-recompute-attacker-2026-10-05.md and the random-read ceilings of docs/bench-log.md ("the 9070 XT on the eGPU": 9070 XT 2.42 to 2.68 G loads/s at 1 GiB, 5090 16.4 to 18.0, M5 Max 3.41 to 3.49; every random 4-byte read costs AMD a 64-byte line):
| Quantity | Formula |
|---|---|
| Bytes read per hash | 128 loads x W bytes |
| Distinct 64-byte lines per hash | 128 (one line per load at every W up to 64 bytes; the census's 120 to 128 distinct addresses per hash) |
| SRAM a chip needs to mirror what the hash reads | the union of the program's 16 windows (a distribution over programs; section 7.1) |
| Latency-bound share | measured rate / (the card's 4-byte random-read ceiling / 128) |
The latency-bound share uses the 4-byte ceiling for every width because the 9070 XT line probe showed the same count per second for 4-byte and 64-byte random reads; the 5090's 16-byte and 64-byte ceilings are the read-width branch's measurement, cited when they land.
7.1 Union of windows per program
Filled from a CPU census over programs (section 6).
7.2 Found on the way (harness defects, both fixed on this branch)
| Where | Defect | Fix | Checked |
|---|---|---|---|
proto-cuda/nvrtc/packfile.h (the one-click workers) |
the seed words were re-derived from the bare epoch seed, so every pack of attempt 1 or higher was refused ("the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT"); 5.14 percent of epochs under v2, all six era packs, and the epoch 34 fleet outage of 18:23Z on 5 October 2026 (branch pack-loop af983a7, which this branch takes: pf_program_words) |
the pack-loop derivation merged over the readwidth packfile (class fields and string seeds kept) | the devnet pack (attempt 0) loads, the six era packs (attempt 1) load, a copy of era-0 with the attempt tampered to 0 is refused on the re-derivation (section 6) |
proto-cuda/host.cu, proto-opencl/host.c |
the host-side dataset word was mh_item(w >> 4)[w AND 15], the harness's own copy of the linear layout; under an interleaved layout the "64 random points vs host derivation" check failed while the Mac samples and the vectors passed |
host_ds_word calls the pack's mh_word (memhard.h), which carries the layout |
proto-cuda/emu/test-layout.sh: the CUDA emulation on era-1 (interleaved) and the devnet pack (linear) must pass the random-point check; it failed on era-1, era-3 and era-5 before the fix (section 6) |
8. What is unverified
- Everything in section 6 marked pending.
- The 1-hour VDF does not exist; the devnet stand-in of section 2 is a proposal.
- The interleave's value against a chip with a programmable address decoder is nil (1.2); the claim is limited to hard-wired layouts.
- The window floor of 2^26 words is set by the 5090's L2 (96 MiB) and the 9070 XT's Infinity Cache (64 MB, vendor figures); a future card with a larger cache moves the floor, which is a genesis constant.
- No cryptanalysis of the stride (a multiply and a rotate before the mask); it is a bijection, so the address distribution is that of the register value, as today.