igneum/docs/plans/era-layout.md

21 KiB

Era layout: table layout and working set drawn per era and per program (Counter ASIC 2.0, layers 4 and 8)

5 October 2026, branch ca2-era, worker "ca2-era". Status: Designed and Implemented behind the class flag (LoadClass::era, not the lottery hash); nothing here changes the default generator, the pinned packs or any live program. Measured sections are marked as such; everything else is design.

Plan: docs/plans/counter-asic-2.md, layers 4 ("table layout drawn per era: item size, stride, interleave") and 8 ("working-set size drawn per program"). The era seed is E_n of spec 04 section 4.4. Confirmed on 5 October 2026 by grep over igneum-pow/src on branches master, readwidth and opencl-rdna4-telemetry: no era draw existed in code before this branch (igneum-era, EraParams, era_seed: no match).

1. What is drawn, and from what

Two streams, nothing else:

Stream Seeded from Draws Sets
Era stream seed_words_from_bytes("igneum-era/" || E_n) words 0 and 1 (spec 01 section 1.13.1 wrote n_le64 || E_n; the index is dropped here because E_n already commits to n through the VDF input of section 4.4 step 2, and the node's seam hands the generator the era bytes alone: Epoch::from_chain_seeds(epoch, day, era, class, label), branch ca2-v3) 7 per era, fixed the load width W, the stride (M, R), the interleave pos[0..3]
Program stream the epoch seed words as today (spec 01 section 1.3.3) 11 per instruction instead of 9: the 9 of version 2, then 2 window draws (a class whose loads are not version 2's takes the read-width width roll between them, ca2-v3's rule) per load site: the window shrink k_off and its offset o

The era parameters change the class of every program of the era. The program stream does not see the era parameters (two eras with the same epoch seed draw the same instruction list and the same windows, and differ in width, stride and layout); this is what makes the six era packs below a controlled comparison.

1.1 Era draw (proposed spec text for section 1.13.1, replacing its parameter table)

One SplitMix64 stream S seeded with lo = words[0] | (words[1] << 32) of seed_words_from_bytes("igneum-era/" || E_n). Seven draws, in this order, whether or not a value is used:

  1. W = allowed[below(|allowed|)]: the width in words of every dataset load of the era, drawn from the genesis-fixed ascending set allowed, a subset of {1, 4, 16} (4, 16 or 64 bytes). A set of one element pins the width; the draw is still consumed. The set is {1} (4 bytes, v2's load): the read-width decision of 5 October 2026 (docs/plans/read-width.md, "keep v2; w16 the only width that passes the rules and closes nothing") and the adoption rule of the same evening (a draw that changes the bytes per hash changes the rate; the six-era hash-rate spread must stay under 5 percent per card). The set is recorded in every pack (IGNEUM_ERA_ALLOWED_WIDTHS, program.json era.allowed_widths) and enters the program id. The code keeps the draw general so the set can be widened at genesis without a new derivation.
  2. M = low32(next()) OR 1: the stride multiplier, odd, so x -> x * M is a bijection on 32-bit words.
  3. R = 1 + below(31): the stride rotation, in 1..31 (never 0: spec 01 section 1.14 item 3).
  4. to 7. r_i = next() for i in 0..3: the interleave draws. Let b = log2(W) (0, 2 or 4) and free = 4 - b. Let c = [b, b + 1, ..., 15] (16 - b candidates). For i in 0..free: j = i + (r_i mod (16 - b - i)), swap c[i] and c[j]. The interleave is pos = [0, ..., b - 1] ++ sort(c[0..free]), four ascending bit positions in 0..15. Draws r_free..r_3 are consumed and ignored.

The era parameters are (W, M, R, pos). Era 0 of the devnet packs is listed in section 5.

1.2 Dataset mapping with the interleave (replaces the last sentence of section 1.8.5)

Word w of the dataset holds word j(w) of item t(w), where j(w) is the 4-bit number whose bit i is bit pos[i] of w, and t(w) is w with bits pos[0..3] removed (the remaining bits in order). With pos = [0, 1, 2, 3] this is today's dataset[w] = item(w >> 4)[w AND 15], byte for byte.

Properties kept:

  • An item has the same value at every dataset size (item derivation is untouched), and because every pos[i] < 16, dataset[w] is the same at every dataset size of at least 2^16 words. The 1 GiB vectors of an era remain valid for words below 2^28 at any larger size, as today.
  • The low b = log2(W) positions are 0..b-1, so the W words of one aligned load lie in one item (t is the same for all of them and j runs j0 .. j0 + W - 1). The verifier derives one item per lane per load, as today: the 4,096-item bound of section 1.11 holds (16 loads x 8 iterations x 32 lanes, whatever the width).
  • The dataset build writes 16 words of one item to 16 addresses w(t, j) (a scatter of 4-byte writes instead of one 64-byte line when pos != [0, 1, 2, 3]). This is the only GPU cost of the interleave and is paid once per day; section 6 measures it.

What the interleave does and does not buy. A chip that hard-wires today's layout (64-byte items, a 64-byte line per item) reads the wrong 15 words with every word once the era draws another layout; the layout changes every 180 days inside rules fixed at genesis. A chip whose address decoder can permute 28 address lines under firmware control pays nothing for it. The honest claim is the first sentence only. The stride below is the same kind of lever: two integer operations per load on a chip, nothing on a GPU.

1.3 Load address (replaces "a load reads one 4-byte word at src AND MASK" in section 1.5 for era programs)

For a load site with window draws (k_off, o) (section 1.4), a dataset of 2^D words (MASK = 2^D - 1), and register value x:

k    = min(k_off, D - 26)          (0 when D <= 26)
y    = rotl(x * M, R)
idx  = ((y AND (MASK >> k)) OR ((o AND (2^k - 1)) << (D - k))) AND MASK
base = idx AND NOT (W - 1)         (W words from base are folded as verify::fold_words)

Uniformity: x * M with M odd and rotl are bijections of the 32-bit word, so y is uniform when x is; y AND (MASK >> k) is uniform on the window; the offset picks which of the 2^k aligned windows. Branch-free, integer only, three operations before the mask (multiply, rotate, and-or) against one today. The emitted text has one form per dialect, checkable by text search (section 1.14 item 2): CUDA and OpenCL ds[((rotl_imm(rN * 0x........u, Ru) & 0x........u) | 0x........u) & mask], Metal the same with dataset[ and & MASK].

1.4 Window draw per load site (layer 8; proposed text for section 1.4.3)

After the nine draws of version 2 (and the width roll, for a class whose loads are not version 2's), every instruction takes two more draws, used only on a load slot:

k_off = below(3)                      window = the dataset, a half or a quarter of it (2^(D - k_off) words)
o     = low32(next()) AND (2^k_off - 1)  which aligned window

Bounds: the window never goes below 2^26 words (256 MiB; k = min(k_off, D - 26) in 1.3), which exceeds the largest on-chip cache of any card in the benchmark (the RTX 5090's 96 MiB L2, the RX 9070 XT's 64 MB Infinity Cache, vendor figures), and never above the dataset. At the prototype dataset (2^28) the windows are 1 GiB, 512 MiB and 256 MiB; at the genesis dataset (2^29) 2 GiB, 1 GiB and 512 MiB. The dataset grows by the step schedule recommended to Josh (spec 01 section 1.13.3 option (b), docs/analysis/card-lifetime-2026-10-05.md: power-of-two steps, 4 GiB at year 4, 8 GiB at year 12, 16 GiB at year 28, 32 GiB at year 60, every index AND MASK), so the window ceiling follows the steps and the floor stays the genesis constant 2^26 words; nothing in the address of 1.3 needs a range reduction. A program has 16 load sites and so up to 16 windows; the set a program reads is their union (section 7 computes its distribution). The verifier bound is unchanged (1.2).

Why per load site and not per program: a per-program window of a quarter of the dataset would hand a 256 MiB SRAM mirror a third of the hours at the prototype size. Sixteen sites with drawn offsets cover the dataset with high probability (section 7), so the mirror a chip would need is the whole dataset in every hour, and the hour-to-hour variation lands on the memory design (which quarter, which half, how many distinct windows), not on its size.

1.5 Program stream (replaces "592 draws per program" in section 1.4.3 for era programs)

16 slot draws, then 64 x 11 = 704: 720 draws per program (768 and 784 for a class with the width roll). On the chain an era program is a class v3 program (branch ca2-v3's seam: ProgramClass::V3, generator version 3): its id is program_id(3, seed words, attempt) as that branch defines it, and the era it was drawn under is identified beside the id by IGNEUM_ERA_SEED_HEX (packcheck::verify_pack_dir_chain refuses a pack whose era is not the job's), so the pair (id, era seed) names the program. The experiment classes that are not class v3 (--class <other> --era ...) carry the era inside the class id instead: FNV-1a-64("igneum-program-rw/" || generator_le32 || seed words || attempt_le32 || mix[3] || load_slots || "era/" || allowed[3] || W || M_le32 || R || pos[4]). The class name is the base name with -era<first stream word as hex> (w4-era401998a5).

1.5.1 The seam (branch ca2-v3, kept as its signatures stand)

V3_CLASS is the base class of class v3 (16 loads of 4 bytes, the lottery hash's load; the era rides inside it), V3_ALLOWED = [1] its width set. generate_from_seed_bytes_program_class(label, seed, V3, Some(era)) draws LoadClass::era(V3_CLASS, era, &V3_ALLOWED) and stamps generator 3; without era bytes (a template before the era is known) the bare V3_CLASS stands. Epoch::chain_dataset(day, class) is unchanged: the layout of 1.2 is a property of the program (program.class.layout()), applied by the interpreter and by Epoch::dataset_word, so the day's cache is shared by every era of a day and the engine keys its caches on (day, class) as before.

1.6 Acceptance

The rule of section 1.4.6 is unchanged in its tests. Its interpreter mirrors 1.3 at the rule's constant D = 28 (idx as above with MASK = 0x0fffffff), as igneum-pow/src/accept.rs does. The distinct-address bound counts dataset loads as the read-width branch defines it.

2. The era seed on the devnet (stand-in for E_n)

Until the 1-hour VDF of section 4.4 is in the node, in the shape of the epoch seed's stand-in (docs/fork-divergence.md "Epoch seed"):

Era E_n
0 the genesis block hash (32 bytes)
n >= 1 the hash of the last selected-chain block whose DAA score is below 15,552,000 n - 7,200 (the 2-hour lead of section 4.4 step 1)

The era of a block is floor(DAA score / 15,552,000), a function of the header alone. E_n for n >= 1 is known 7,200 DAA seconds before the era starts, which covers the 1-hour VDF when it arrives and the kernel compile and dataset rebuild now. Test seeds for packs and tests: E_n = the 32 bytes (little-endian words) of seed_words_from_bytes("igneum-era-test/<n>") (igneum-pow ... --era igneum-era-test/<n>, the number only names the pack); raw bytes with --era <n>:<64 hex>.

3. Memory budget (Josh, 5 October 2026: under 6 GB on an 8 GB card)

The era layout adds no resident memory: a window is a mask and an offset in the kernel text, the interleave is address arithmetic, the stride is two operations. The whole working set on a card, every item from this branch and the others:

Item Bytes Who
Dataset (prototype) 1 GiB existing
Cache (256 MiB, resident only while the day's dataset is built, then free; the decided reading, confirmed by the coordinator on 5 October 2026, docs/analysis/card-lifetime-2026-10-05.md carries the per-tier working set with the cache freed as the best case) 256 MiB peak existing
Layer 5 hot table the ca2-cache worker's figure not this branch
Scratch per resident warp (read-width variant 5, 32 or 128 KiB per warp) 2,048 warps x 128 KiB = 256 MiB at most readwidth branch
Output and read-back buffers 2^24 nonces x 8 B = 128 MiB per dispatch existing harness
Era windows, stride, interleave 0 this branch

The era window never exceeds the dataset, so it never grows the footprint.

4. Implementation (behind the flag)

Piece Where What
EraParams, era_draw, LoadClass::era igneum-pow/src/generator.rs the 7-draw era stream of 1.1; the class carries era: Option<EraParams> beside mix, load_slots, scratch, scratch_kb; the two window draws per instruction (Instr::win, Instr::off); the program id of 1.5
Layout igneum-pow/src/memhard.rs split(w) -> (t, j), join(t, j) -> w, LINEAR = [0, 1, 2, 3]; MemhardCpu and DatasetSource carry it
load_index igneum-pow/src/verify.rs the address of 1.3, shared by the interpreter and the acceptance mirror
Emitters igneum-pow/src/emit.rs the one load form of 1.3 in Metal, CUDA and OpenCL C; mh_word, the three igneum_build kernels and the Metal build kernel with mh_t, mh_j, mh_addr when the layout is not linear; IGNEUM_ERA_* in program.h, an "era" object in program.json
CLI igneum-pow/src/main.rs --era <igneum-era-test/n | n:hex> and --era-widths 4,16,64 (one width pins) on every command
Tests igneum-pow/src/*.rs, igneum-pow/tests/packs.rs the draw is deterministic and within bounds; split/join are inverse and the dataset is a prefix at every size; six era programs pass the generator contract and the acceptance rule; vectors round-trip; the pinned packs are byte-identical; the six era packs match the emitters and every load has the form of 1.3

The default class is untouched: LoadClass::V2 has era: None, every emitter branch on era keeps today's text, and tests/packs.rs diffs the pinned packs (igneum-genesis-mh, igneum-devnet-v4-epoch0) against the emitters as before. Section 6 records the diff of a fresh export against the checked-in files.

5. The six era packs

proto-cuda/packs-ca2-era/era-<n>, n in 0..5: the devnet's 32-byte epoch seed edc4fa84...fb07 and day bytes igneum-day/20730 (the seeds of the pinned pack igneum-devnet-v4-epoch0, which is the v2 baseline with the same program seed), dataset 2^28 words, era seed igneum-era-test/<n>, width pinned at 4 bytes (--era-widths 4, the default). The program seed is held fixed so that the six packs differ in the era parameters only (section 1); each carries seeds.txt for the one-click workers. Every pack is attempt 1 (attempt 0 of this seed is rejected under the era class: 12 draws per instruction give a different stream from v2's). The drawn parameters (igneum-pow show --epoch-hex edc4... --era igneum-era-test/<n>, 5 October 2026):

Pack Class Era seed E_n (first 16 hex) W (bytes) M R pos
era-0 w4-erab2ed8a89 5e0587f455a86e91 4 0x625e5ab3 19 0, 2, 10, 15
era-1 w4-era676a17fc df57136f2ad5f410 4 0xb2a9d70d 6 1, 3, 8, 13
era-2 w4-era843155d7 7f450623297a954f 4 0x2b4a5b97 28 1, 3, 4, 8
era-3 w4-erad6367bfe 8bffdd3366b9c3ff 4 0x27ea7eff 30 2, 3, 8, 13
era-4 w4-era4488f3ed e593fc1d48475c88 4 0x4d38603d 10 2, 9, 13, 15
era-5 w4-eraf897c84e ff87ad96a1b53f36 4 0x03ac37ad 22 0, 2, 10, 13

All six are class v3 packs (generator 3, IGNEUM_PROGRAM_CLASS "v3", IGNEUM_ERA_SEED_HEX), attempt 0, program id 73bcbfe8ccf988f1 in every pack (the seam's program_id(3, seed, attempt); the era seed beside it names the program), the era layout over version 2's item construction (mixer x1, the genesis cache) so that the v2 baseline pack is the same dataset; the chain's class v3 composes the same draw over LoadClass::MX4 (mixer x4, growth), and the integration re-exports these packs on it after the PC rows. The era-seed-to-pack assignment above is from program.h of each pack; the test-seed numbering is only the pack name.

The windows are a property of the program, so they are the same in all six packs (site:shrink:offset): 1:2:3 4:2:2 6:0:0 10:1:0 12:1:1 20:2:0 27:1:0 30:0:0 35:0:0 40:0:0 41:0:0 43:2:3 45:1:0 52:0:0 53:2:2 54:0:0: 7 sites read the whole dataset, 5 a half, 4 a quarter; the union is the whole dataset.

6. Measurements

Pending at the time of this commit; each table below says the machine, the date, the harness and the command when filled.

6.1 Byte-identical default path

6.2 Bit-exactness on the Mac (Metal, Apple OpenCL, CUDA emulation)

6.3 Hash rate per era on the M5 Max (Metal) and the CPU verifier

6.4 PCs (RTX 5090 CUDA, RX 9070 XT OpenCL): prepared, waiting for the go

7. The chip-model line per draw

From docs/analysis/m16-recompute-attacker-2026-10-05.md and the random-read ceilings of docs/bench-log.md ("the 9070 XT on the eGPU": 9070 XT 2.42 to 2.68 G loads/s at 1 GiB, 5090 16.4 to 18.0, M5 Max 3.41 to 3.49; every random 4-byte read costs AMD a 64-byte line):

Quantity Formula
Bytes read per hash 128 loads x W bytes
Distinct 64-byte lines per hash 128 (one line per load at every W up to 64 bytes; the census's 120 to 128 distinct addresses per hash)
SRAM a chip needs to mirror what the hash reads the union of the program's 16 windows (a distribution over programs; section 7.1)
Latency-bound share measured rate / (the card's 4-byte random-read ceiling / 128)

The latency-bound share uses the 4-byte ceiling for every width because the 9070 XT line probe showed the same count per second for 4-byte and 64-byte random reads; the 5090's 16-byte and 64-byte ceilings are the read-width branch's measurement, cited when they land.

7.1 Union of windows per program

Filled from a CPU census over programs (section 6).

7.2 Found on the way (harness defects, both fixed on this branch)

Where Defect Fix Checked
proto-cuda/nvrtc/packfile.h (the one-click workers) the seed words were re-derived from the bare epoch seed, so every pack of attempt 1 or higher was refused ("the epoch seed bytes do not give the pack's IGNEUM_SEEDW_INIT"); 5.14 percent of epochs under v2, all six era packs, and the epoch 34 fleet outage of 18:23Z on 5 October 2026 (branch pack-loop af983a7, which this branch takes: pf_program_words) the pack-loop derivation merged over the readwidth packfile (class fields and string seeds kept) the devnet pack (attempt 0) loads, the six era packs (attempt 1) load, a copy of era-0 with the attempt tampered to 0 is refused on the re-derivation (section 6)
proto-cuda/host.cu, proto-opencl/host.c the host-side dataset word was mh_item(w >> 4)[w AND 15], the harness's own copy of the linear layout; under an interleaved layout the "64 random points vs host derivation" check failed while the Mac samples and the vectors passed host_ds_word calls the pack's mh_word (memhard.h), which carries the layout proto-cuda/emu/test-layout.sh: the CUDA emulation on era-1 (interleaved) and the devnet pack (linear) must pass the random-point check; it failed on era-1, era-3 and era-5 before the fix (section 6)

8. What is unverified

  • Everything in section 6 marked pending.
  • The 1-hour VDF: BUILT on 7 October 2026 (era VDF lane, after the attack pass's F7 row named it the gating dependency): kaspa_consensus_core::era_vdf (the class-group Wesolowski scheme on a fixed-width integer and the hash-chain fallback behind the genesis byte vdf_scheme), kaspa_consensus::processes::era_vdf (the cut rule, the day-of-blues input, the evaluator thread, the record store), behind Params::era_vdf_activation_daa (never on every network until Josh's word per network); the stand-in of section 2 stands below the activation and is what the VDF reads its input from above it. Verified: the F7 re-roll harness against the real era cut fires with the VDF off and is silent with it on across 6 cuts (tools/era-vdf/reroll.mjs, the record docs/analysis/era-vdf-2026-10-07.md section 3), the parameters and the measured prove and verify times are in spec 04 section 4.6. Still unverified: the P2P relay of a record to a syncing peer (spec 4.5, owed before era 1 of any network with the switch set), the binding of the cut to the certified checkpoint (left at the O-4.3 reading, one function to change), an external review of the class-group port (O-4.1), and the 2019-class-core verify time, which is measured on a proxy until a 2019 host is rented (record section 5).
  • The interleave's value against a chip with a programmable address decoder is nil (1.2); the claim is limited to hard-wired layouts.
  • The window floor of 2^26 words is set by the 5090's L2 (96 MiB) and the 9070 XT's Infinity Cache (64 MB, vendor figures); a future card with a larger cache moves the floor, which is a genesis constant.
  • No cryptanalysis of the stride (a multiply and a rotate before the mask); it is a bijection, so the address distribution is that of the register value, as today.