diff --git a/docs/analysis/class-v6/1p5x/x0/README.md b/docs/analysis/class-v6/1p5x/x0/README.md new file mode 100644 index 000000000..5463bcdc8 --- /dev/null +++ b/docs/analysis/class-v6/1p5x/x0/README.md @@ -0,0 +1,236 @@ +# X0: the latest model reconciled with the exact frozen software (the 1.5x programme, 9 October 2026, 10:0x UK) + +The adversary lane with the hash lane's pins (its line of 09:58 UK, every closing report staged with its sha256 at +build-1:/srv/artefacts/tas/x0-pins/.report.txt, SHA256SUMS.txt beside them; the recipe every X lane reads, +build-1:/srv/artefacts/1p5x/gpu-row-recipe.md sha cbab453c). The brief: the founder's +igneum-1p5x-experiment-brief.md (sha256 74ee073286d4a6e2…, landing under docs/plans/igneum-2.0-master/1p5x/). This +file reconciles every served and placed chip figure with the software and the measurements it rests on, names the +assumptions that differ between the served 1.9x to 2.1x hybrid line and the placed 2.4x line, and states the r0 the +programme uses with its range. Nothing here is a new measurement; no PASS is claimed. + +## 1. The frozen control, as pinned + +| Pin | Value | Read back | Label | +|---|---|---|---| +| The generator tree | class-v6 at 1a938abe4 (on branch class-v6, tip 755c2dbcf), generator 6, ProgramClass::V6, class string mx8-erad810f22d+sh256x27+state+reg64c+fold+rw, layer 8 ON, class_v6_family_flags 0xf, CLASS_V6_FLAG_NOWIN (1<<4) defined and off | the fingerprint recipe run on build-1 at 09:34 UK: 5f4d6dc6199294db89042171004e6420c1e5791e716d4b39b73d375818c10b6f, equal to the brief's | read back | +| The dataset policy | 70a6c703 (4 GiB at genesis, no later step) | F0's dataset row | pinned by F0 | +| The signing id | 2a1d6caab4c24564 at (epoch seed af89be5d, era edc4fa84, day 20730, state abb58003) | F0's freeze-candidate row | pinned by F0 | +| The research pairing (never the freeze) | hl-v6-all, id 0x4de7b836cc40a4ea (the genesis seed over node1's state) and the l8off kit's export 0x9d40978601a7df2a (era edc4fa84, day 20730, the node1 state) | the hash lane's line | pinned | + +## 2. The GPU side: every figure, its software, its boundary, its state + +Every figure is the three-card Windows rig's RTX 5090 (or 5080) alone through the runner's cards-off; power is +nvidia-smi board power.draw sampled at 1 Hz, the row's watts the sampler's mean from 8 s in to the end of the timed +dispatches, idle NOT subtracted (the microbench rows are the one exception, residency-subtracted); no wall meter +anywhere; fingerprints equal at every state. No GPU figure carries layer-8-off or a result compaction. + +| Figure (as served) | Job and report sha256 | Software by sha | Pack | State | Where it is used | +|---|---|---|---|---|---| +| The 5090 at its 1,300 MHz lock: 134.764 MH/s at 312.5 W, 2.319 microjoules per hash (served as 2.33) | run-ca3-pc1-v4-eff-5090-20261007-b (7 Oct 18:40 to 19:16Z), report ead72106… | the installed app's igneum-worker-cuda.exe d7a413c7… (0.3.20); kit igneum-ca3-v4-sub3-pc1-kit.zip 1d441f5a… (program.json a746c924…, program.h 18bca6be…, kernel.cu 80e92839…) | v4-devnet-epoch0, the Devnet 3 epoch 0 class v4 program, fingerprint e370fb2080b7dbb1; 491 batches of 2^24 | the knee | the numerator of every served chip ratio (the "served convention") | +| The class v5 control at the lock: v5-genesis 125.932 MH/s at 299.8 W against v4-genesis 125.924 at 294.0 (+2.0 percent watts) | run-ca3-pc1-v5lock-5090-20261008-b, report 0f444306… | the class v5 kit's worker | v5-genesis, v4-genesis | the knee | denominator.md's +2 percent (the 2.37 row) | +| The class v6 control on the frozen tree at the lock: hl-v6-all 70.306 MH/s at 282.0 W (4.011 microjoules); stock 70.526 at 412.0 (5.842) | run-ca3-pc1-l8off-5090-20261008 (9 Oct 02:52 to 03:00Z), report 05f00bd8… | the l8off kit igneum-ca3-l8off-kit-20261008.zip c20999d1…, CUDA worker 218ca414…; fingerprint 59e6708e46f1e87c; 250 batches | hl-v6-all 0x9d40978601a7df2a (the frozen class at the research pairing) | both | NOT yet in any served chip ratio (section 4) | +| The 5080 at its 1,100 lock: 71.201 MH/s at 146.6 W, 2.059 microjoules | run-ca3-pc1-v4-eff-5080-20261007-d (8 Oct 00:45 to 01:54Z), report 6acec71d… | the same worker d7a413c7…, the same pack | v4-devnet-epoch0, e370fb2080b7dbb1 | the knee | the 5080 column | +| The cohort rows (14 rented card classes, stock only) | the fleet lane's rows.jsonl; denominator.md sha a4ec182a… on master | per row | per row | stock | the cohort column (count-weighted about 3.6 microjoules, approximate) | +| Layer 8 off (register row 18, CLOSED 04:49 UK): hl-v6-all against hl-v6-all-nowin 0xec0c3757f8caef4a: stock 70.526 / 412.0 W against 70.209 / 417.2 (rate -0.4, joules +1.7 percent); lock 70.306 / 282.0 against 66.303 / 275.5 (rate -5.7, joules +3.6); the v5 pair the same way, +5.7 percent at the lock | the same job, report 05f00bd8…; the row document docs/analysis/class-v6/rows/pc1-l8off-ds2g-20261009.md sha fb2d5686… (master ac2f143c) | the l8off kit | the four packs | both | the D2 row: performed and closed, the layer stays ON | +| The microbench per-op rows (int_arx 11.3 pJ stock, 6.2 at the lock; the class v4 draw read as 10.3 pJ per op at the lock from the packs job's 10.8 / 6.4) | run-ca4-pc1-microbench-5090-20261008-c (8 Oct 05:15 to 05:56Z), report 62dbb138…; the 10.3 is counter-asic-4-research.md's reading (sha 82af4424…) | igneum-worker-cuda-ca4mb.exe 49aa60c4… | 20 probes x 60 s at full residency | both | the denominator of every per-family k | + +A caveat the hash lane names: since 14:2x UK on 8 October the rig's 5090 reads about 13 percent under the morning +(the enclosure swap), so pairs inside a job are the figures and a job's absolutes are not compared with 7 October's. +The served 2.33 microjoules is a 7 October absolute; the frozen-tree control is a 9 October job. Section 4 carries +what that does to r0. + +## 3. The chip side: every figure, its software, its boundary, its node + +| Figure | Software by sha | What it is | Boundary | Node scaling | +|---|---|---|---|---| +| The placed 8-lane 18-family core: 9.36 pJ per lane-op on the class v4 draw at ASAP7 (8.6 to 11.1), 6.55 at N5, 4.72 at N3 | class-v6-adversary 43c9f7cd8: tools/chip-model/mf (RTL core_mf.v, the flow, results/table.csv), docs/analysis/class-v6/multi-family-adversary.md (landed e3d0208d) | ORFS on ASAP7 (Yosys 0.68, OpenROAD, the orfs docker image), placed and routed with SPEF, OpenSTA power under a gate-level VCD of a random-input simulation; the register window a FakeRAM2.0 64x256 macro per 8 lanes whose dynamic energy is MODELLED (3.5 pJ per access, 2.0 to 7.0) | the core only: no memory controller, no PHY, no host; the machine terms are board.py's (section 5) | ASAP7 to N5 x0.70, to N3 x0.50 (TSMC's claimed per-node reductions at the same speed) | +| The placed 32-lane 18-family core: 9.54 pJ (8.9 to 11.0), 6.68 at N5 | the same; parasitics from the global route, calibrated within 1 percent of SPEF on the 32-lane genesis core; the detailed route did not converge (the fault, multi-family-adversary.md section 14) | | | the same | +| The program the chip rows execute | the testbench's draw with the class v4 WEIGHTS (add 12, xor 10, mul 8, mad 8, shfl 8, rotl 7, sub 6, mulhi 6, rotr 6, or 4), NOT the frozen tree's drawn program: the chip rows are per-op energies by family under the draw's weights; the frozen program's op counts (102,100 shadow ops, 512 base, 128 loads, W = 1) set the per-hash sum | | | | +| The memory system | chip-model-v3.md 5.3 and 5.5 (the GDDR7 board: 2.0 nJ per random read, 1.5 to 2.6; 21.3 G activates per second, the 5090's 82 percent as the sustained fraction; 20 W static, 15 W controller; USD 320 of devices), the N2 SRAM die (floor lane 3, 0.25 nJ per read at W = 1) | modelled on claimed device figures | the devices, controller and PHY | the memory does not scale with the core's node | +| The hybrid's hit rate | adv-cache-2's window-layer distribution on Devnet 3's programs (the hottest half of the items serves 0.7188 of reads: the p98 program), the hash lane's census reconcile (the mean over programs 0.581, coexistence-model.md line 537), census-packs.md's flat-control ratio 1.1 to 1.7x (sha 146b578c…) | measured on programs of the class v5 and v6 crates, the window model as the null | | | +| The complete machine | tools/chip-model/mf/flow/board.py: PSU 92 percent, VRM 90 percent, cooling 3 percent, one full node per 100 machines (85 W, USD 1,500), board and assembly USD 200, the core die sized to the memory's sustained rate | approximate allowances | the wall of the machine, less the host's share | | + +## 4. The reconciliation, line by line + +### 4.1 The served 1.9x to 2.1x hybrid against the placed 2.4x hybrid + +Both are the stored-half hybrid (1 GiB of SRAM beside the 16 GDDR7 devices) against the 5090 at its lock, same +node. They differ in four assumptions, each named: + +| Assumption | The served line (10.0r, "about 2.4x") | This lane's placed line (1.9x to 2.1x) | Which the programme uses | +|---|---|---|---| +| The hit rate of the hottest half | 0.72 (adv-cache-2's p98 program) | 0.581 (the hash lane's reconciled mean over programs) for the 1.93x; 0.72 for the 2.09x | the mean, 0.581: the p98 is the tail, carried beside it | +| The core's energy | the synthesised 8-lane core (4.61 pJ at N5) scaled by the first placed genesis core's factor | the placed 8-lane 18-family core itself, 6.55 pJ at N5 | the placed row | +| The machine terms | the board identity (2.33 over 0.466 + the shadow) with the hybrid's memory term, no power train | the complete machine: PSU, VRM, cooling, the host share, the sustained derate | the complete machine (the review's evidence standard) | +| The node | N5 node-for-node | N5 node-for-node | the same | + +So the 2.4x is the synthesised core at the tail hit rate on the board identity, and the 1.9x to 2.1x is the placed +core at the mean and the tail on the complete machine; the two agree once the three assumptions are aligned +(section 13 of multi-family-adversary.md: 1.93x at 0.581 placed, 2.09x at 0.72 placed; 2.69x at 0.72 synthesised +on the board identity, the 10.0q figure). The programme's hybrid r0 is the placed, complete-machine, mean-hit +figure. + +### 4.2 The window-conditioned 72 percent against the flat control, as the brief states it + +The brief reads: the hottest half serves 72 percent because the window model concentrates traffic, and a flat +distribution would give 50 percent. The record's reading (census-packs.md section 3): the 72 percent is the p98 +program on the window-conditioned null, the mean over programs is 0.581, and a flat control sits at 1.1 to 1.7x of +the window model's top-0.1-percent share (printed beside the census, not used). The layer-8-off experiment is +PERFORMED and CLOSED (register row 18): with the windows removed the hit of the hottest half is exactly 0.50 at +every program (multi-family-adversary.md section 13 (5)), and the honest side at the knee loses 5.7 percent of rate +and 3.6 to 5.7 percent of joules, so the layer stays ON. The chip-side arithmetic of the brief's conditional cache +calculation holds (misses 0.50 against 0.419 at the mean, +19 percent of DRAM reads; against 0.28 at the tail, +79 +percent) and the placed rows say what it buys: the hybrid at hit 0.50 reads 1.85x against 1.93x at the mean and +2.09x at the tail, 4 percent at the mean, 11 at the tail; the honest side pays more than that at the knee. X1 is the +new-consensus-version form of the same question (a class v7 shape without the fixed windows, everything else +frozen) and its traces replace this arithmetic on the hybrid rows when they land. + +### 4.3 Layer-8-off and compaction in any figure + +Layer 8 off is in NO served or placed figure (every GPU row and every chip row above carries the windows ON; the +hit-0.50 row is a labelled variant). Result compaction (X2) is in NO figure on either side. + +### 4.4 The wall-power boundary, both sides + +The GPU side's boundary is the card's board power (nvidia-smi, 1 Hz, idle not subtracted), the card alone through +cards-off: it excludes the host, the PSU's loss and the fans' share beyond the card's own. The chip side's boundary +is the complete machine (the core, the memory, the controller and PHY, PSU and VRM losses, cooling, the host's +share). The two are not the same boundary: the GPU's host, PSU loss and case fans are NOT in its 2.33, while the +chip carries about 21 percent of power-train loss, 3 percent of cooling and 0.85 W of host. The reconciliation the +programme should carry: either add the GPU's own machine terms (a mining rig's PSU at 92 percent and its host share, +about +10 percent on the card's figure, approximate) or read the chip at the board identity; the first is the honest +one and moves every served ratio up by about 0.1x. It is NOT applied in this file's r0 (the served convention is +kept so every row stays comparable); it is the first of the three assumptions below. + +### 4.5 The node scaling applied + +Every chip row is stated at ASAP7 and scaled by claimed factors: N5 (the 5090's class, "same node") x0.70, N3 ("a +node ahead") x0.50, N2 x0.36. The served sentence's "2.0x on the GPU's own node" is the N5 column; "2.2x to 2.4x" +the N3. The memory terms do not scale. The GPU side is measured at its own node and never scaled. + +## 5. The r0 the programme uses + +| Opponent (complete machine, placed core, same node N5) | r0 | Range | Source rows | +|---|---|---|---| +| The pure DRAM board | 1.5x | 1.3x to 1.7x (the SRAM access band and the memory band) | multi-family-adversary.md section 6 at the placed energy; the 32-lane core the same | +| The stored-half hybrid at the mean hit 0.581 | 1.93x | 1.65x to 2.12x | section 13 (placed) | +| The stored-half hybrid at the tail hit 0.72 (the 10.0q case) | 2.09x | 1.78x to 2.29x | section 13 (placed) | +| The N2 SRAM die | 2.3x | 2.0x to 2.7x | section 6 at the placed energy | + +**The reconciled r0 the programme uses is the worst credible opponent's: the pure DRAM board at 1.5x (1.3x to +1.7x) same-node; the hybrid at 1.93x (1.65x to 2.12x) is the X5 sweep's starting row and X1's target.** The +yardstick from 1.5x to 1.5x is nothing; from the hybrid's 1.93x, a 22 percent GPU cut or a 29 percent specialist +rise. A node ahead: the board 1.8x (1.5x to 2.0x), the hybrid 2.4x (2.05x to 2.63x). + +The three assumptions most likely to move r0: +1. The boundary mismatch (4.4): putting the GPU's own PSU loss and host share into its figure moves every ratio up + about 0.1x (the board to about 1.6x, the hybrid to about 2.1x), approximate. +2. The frozen-tree control against the class v4 numerator: the served 2.33 is class v4 on a 7 October job; the + frozen class v6 object reads 4.011 microjoules at the lock on the l8off kit's export at the research pairing (a + 9 October job on a card reading 13 percent low). The chip's shadow at class v6's op count is the same core on + 102,612 ops; if the honest card's class v6 energy per hash is 4.0 against the chip's 1.53, the board ratio reads + 2.6x and the hybrid 3.3x, which is the number the programme must reconcile first: either the class v6 object + costs the 5090 1.7x what class v4 did (then the chip's shadow op count for class v6 must be re-read from the + frozen program, not from class v4's 102,100) or the l8off job's absolute is the enclosure-swap artefact. The + hash lane's paired control (class v4 and class v6 in one session on one card, same kit) decides it; until it + lands r0 stays on the class v4 pair. +3. The SRAM macro's access energy (the one modelled chip term; 2.0 to 7.0 pJ per access gives the band's width) and + the hybrid's hit rate (0.50 with the windows off, 0.581 mean, 0.72 tail). + +## 5a. The decider in (the hash lane, 09:59 UK): assumption 2 is real, and what it does to r0 + +One session on a rented RTX 5090 (Vast 54992886, driver 580.105.08, the gen-6 Linux CUDA worker sha 804a6f7f…, 250 batches of +2^24 per pack, board power at 1 Hz, mean from 8 s in, idle not subtracted, stock only; evidence +build-1:/srv/artefacts/tas/x0-pins/x0-run-fk-5090c/, run-x0.log sha 9bc40685…, smi.csv 7ab9f16f…): class v4 v4-devnet-epoch0 +(0xa785001687d8688a, fingerprint e370fb2080b7dbb1) 141.963 MH/s at 500.0 W = 3,522 nJ per hash; the frozen class v6 hl-v6-all +(0x9d40978601a7df2a, 59e6708e46f1e87c) 70.900 at 445.1 W = 6,278 nJ; the signing object hl-v6-all-cs (0x2a1d6caab4c24564, +01f51b9d4805e5e6) 70.650 at 457.1 W = 6,470 nJ. Class v6 over class v4: 1.78x and 1.84x in joules per hash, 2.00x in rate; the +house 5090's own pair 1.68x at stock and 1.73x at the 1,300 MHz knee (2,319 against 4,011 nJ). So the served numerator (2.33 +microjoules, class v4) understates the frozen object's card cost by about 1.7x, and the only class v6 lock row is 4.01 +microjoules per hash on the house 5090. + +What it does to r0: the chip must be paired with the card on the same program. X4's read of the exported texts (10:0x UK) +gives the frozen object 256 loads per hash against class v4's 128, so the card's doubled cost per hash is doubled work, and +the chip's memory term doubles with it while its shadow (carried at class v4's 102,612 ALU ops until the hash lane's count lands) +does not. On the consistent class v6 pair (the card 4.011 microjoules at the knee, the chip at 256 loads) every ratio moves UP, +not down: X5's sweep (section 2, pair B) reads the DRAM board 1.85x (1.5x to 2.1x), the stored-half hybrid 2.53x, the N2 die 4.0x +(3.4x to 4.4x) same-node. The earlier worry (a 1.7x structural card cost the chip does not share) is not what the data say: the +cost is loads, which the chip pays too. The programme's r0 is therefore pair B's, with pair A (the served class v4 convention: +the board 1.5x, the hybrid 1.93x, the die 2.44x) kept as the served convention's row until the ALU count lands: + +| Opponent, same node | r0, pair A (class v4: the card 2.319, 128 loads) | r0, pair B (the frozen object: the card 4.011, 256 loads, the shadow at class v4's count) | Source | +|---|---|---|---| +| The pure DRAM board | 1.5x (1.3 to 1.7) | 1.85x (1.5 to 2.1) | X5's sweep on the placed core | +| The stored-half hybrid, mean hit | 1.93x (1.65 to 2.12) | 2.53x | the same | +| The N2 SRAM die (the worst credible design) | 2.44x (2.05 to 2.66) | 4.0x (3.4 to 4.4) | the same | + +The hash lane's per-family count (10:11 UK): the frozen object executes 62,120 ALU instructions per hash (loads 256, the base +1,024, the per-load address path 5,376, the shadow 55,296 unchanged, the reg64 init and fold 168); class v4 about 56,300; the card +halves its rate on the read count, which the chip pays too. That settles assumption 2 and opens the largest term of all: every +placed chip row charges the chip 102,612 ops per hash (the record's counted-op convention, from which the card's 0.652 +microjoules reads 6.4 pJ per counted op), but the chip executes one instruction per lane-op, so at the instruction count its +shadow energy is 0.36 to 0.41 microjoules instead of 0.67 and every ratio rises by 1.3x to 1.6x (X5 section 2: the board 2.05x +and 2.2x, the die 4.3x and 6.3x same-node on the two pairs). Both counts are carried until the hash lane defines the counted op +(asked 10:1x UK, default 10:45); the instruction count is the chip's physics and the default. + +## 5b. The amendment (the coordinator's order, 10:1x UK; every row WITNESS E_D): the frozen object's own numerator, the instruction count, the measured hits + +Two facts change every chip ratio. (1) The decider (5a): the frozen class v6 object does twice the random dataset reads per hash of +class v4 (256 against 128) with 9 to 10 percent more ALU (62,120 instructions against about 56,300; the hash lane's per-family +read of the signing object, 10:11 UK); the card pays 1.78x to 1.84x the joules per hash at stock on the rented 5090 and 1.68x at +stock / 1.73x at the knee on the house 5090. (2) The units fault (the hash lane, 10:13 UK, from the record's own units table, +attack-pass/f1-shadow.md section 1 and shadow-k.md section 1): the record's "counted op" is unit B, the GPU's micro-op tally (add +5, rotr 2, shfl 2, every other family 1: 1.83 per shadow instruction), so the 102,100 counted ops per class v4 hash are 55,296 +instructions; the chip executes one instruction per lane-op and its k was derived per instruction (algorithm.md 5.3), so every +placed row that multiplied 102,612 counted ops by a per-instruction chip energy (chip-model-v3's re-reads, the design document's +10.0r, this lane's sections 4, 6, 13 of multi-family-adversary.md and X5's first tables) overstated the chip's shadow term by +1.85x. The record's own gate unit is the instruction (unit A); the chip datapath count (unit C, rotates as wiring, the constants +hoisted) is lower still, about 50,100 per class v4 hash. + +Recomputed on the placed 8-lane 18-family core's per-family energies (ASAP7 placed with SPEF, N5 x0.70, N3 x0.50, claimed) +weighted by the frozen object's own per-family instruction counts: the chip's ALU energy per hash 0.312 microjoules at N5 (0.280 +to 0.388; 0.231 at unit C with the rotates as wiring), 0.225 at N3; 5.05 pJ per instruction at N5 (3.64 at N3). The GPU numerator +is the frozen object's own measured joules per hash: the house 5090 at the 1,300 MHz knee 4.011 microjoules (run-ca3-pc1-l8off-5090- +20261008, hl-v6-all, the only class v6 lock row; the card alone, nvidia-smi board power at 1 Hz, idle not subtracted); at stock +6.470 (the signing object on the rented 5090, x0-run-fk-5090c). The board-power boundary excludes the rig's PSU loss and host, +which the chip's machine carries; the wall-meter column adds 10 percent to the card's figure for them (approximate, unmeasured). +The hybrid's hit rate is the measured one for the hottest half: 0.56 on the signing object, 0.61 the mean over programs (the +quarter and three quarters scaled by the same ratio); 0.72 is the p98 tail and is not used. 256 loads per hash on both sides. +Nothing here is a chip measurement; every row is a modelled witness; no pass and no fail is claimed on these lines. + +| Opponent (WITNESS E_D, complete machine) | microjoules per hash, same node (band) | USD per MH/s | Knee, same node, hit 0.56 (signing) | Knee, same node, hit 0.61 (mean) | Knee + the wall-meter term (+10 percent on the card) | Stock, same node | Knee, a node ahead | Stock, a node ahead | +|---|---|---|---|---|---|---|---|---| +| The pure DRAM board | 1.707 (1.516 to 2.112) | 8.91 | 2.35x (1.90 to 2.65) | 2.35x (1.90 to 2.65) | 2.58x (2.09 to 2.91) | 3.79x (3.06 to 4.27) | 2.51x (2.02 to 2.83) | 4.05x (3.26 to 4.57) | +| The hybrid, the hottest quarter | 1.362 (1.222 to 1.679) | 7.02 | 2.86x (2.32 to 3.19) | 2.94x (2.39 to 3.28) | 3.24x (2.63 to 3.61) | 4.75x (3.85 to 5.29) | 3.20x (2.59 to 3.58) | 5.16x (4.18 to 5.77) | +| The hybrid, the hottest half | 1.077 (0.973 to 1.322) | 5.11 | 3.49x (2.84 to 3.86) | 3.73x (3.03 to 4.12) | 4.10x (3.34 to 4.54) | 6.01x (4.89 to 6.65) | 4.15x (3.36 to 4.60) | 6.69x (5.43 to 7.42) | +| The hybrid, the hottest three quarters | 0.894 (0.811 to 1.094) | 3.77 | 4.03x (3.29 to 4.45) | 4.49x (3.67 to 4.95) | 4.94x (4.03 to 5.44) | 7.24x (5.91 to 7.98) | 5.11x (4.16 to 5.64) | 8.25x (6.71 to 9.11) | +| The full SRAM die (one N2 reticle) at 1 kW | 0.513 (0.464 to 0.622) | 0.68 | 7.81x (6.45 to 8.65) | 7.81x (6.45 to 8.65) | 8.60x (7.10 to 9.51) | 12.60x (10.41 to 13.95) | 10.01x (8.22 to 11.13) | 16.15x (13.26 to 17.95) | +| Recomputation | 7.000 (6.459 to 8.238) | 7.74 | 0.51x (0.44 to 0.55) | 0.57x (0.49 to 0.62) | 0.63x (0.54 to 0.68) | 0.92x (0.79 to 1.00) | 0.77x (0.66 to 0.84) | 1.25x (1.06 to 1.35) | + +The one sentence: at the frozen object's own numerator and the instruction count, the pure DRAM board as a complete machine sits +ABOVE 1.5x by 0.85x (2.35x same-node at the knee, band 1.90x to 2.65x; 2.58x with the wall-meter term; 3.79x at stock; 2.51x a node +ahead), the stored-half hybrid at the measured hit 3.5x to 3.7x, and the N2 SRAM die, the worst credible design and the +programme's a (0.513 microjoules per hash, 0.464 to 0.622), 7.8x; the earlier 1.5x read was two errors in the honest side's +favour (class v4's numerator against a class v6 chip, and the GPU's micro-op tally charged to the chip), not the chip's cost. + +What this means for the yardstick: r_new = r0 x g / a from the board's 2.35x needs a 36 percent GPU cut at a = 1 (a 15 plus 20 +split reaches 1.66x); from the die's 7.8x no candidate within the brief's arithmetic reaches 1.5x on energy, which puts the +die's hold where the coexistence model put it (its dollars and its project), and the per-joule target at the board. X5's rows +and X6's blind run follow this line; the served energy sentence should not be re-worded from this file (the panel reads it). + +## 6. BLOCKED rows + +| Row | Why | Owner | +|---|---|---| +| The class v6 paired control at the lock (the rented pod refuses the clock lock; the stock pair is in, section 5a) | stock done 09:59 UK; the knee pair is the house 5090's 1.73x | the hash lane | +| The definition of the record's "counted op" | CLOSED 10:13 UK: unit B, the GPU's micro-op tally (1.83 per instruction); the chip's unit is the instruction (5b) | the hash lane | +| The chip's op count on the frozen program (the shadow's 55,296 writes in r0..r7, the base program's 512, the loads' 128 at W = 1: read from 1a938abe4's generator, not from class v4) | the chip rows use the class v4 draw's weights; the frozen program's per-family counts are owed | this lane, with the hash lane | +| The GPU's own machine terms (PSU, host) | not measured; approximate +10 percent | the hash lane (a wall meter on the rig) | + +## 7. Sources + +The hash lane's pin line (09:58 UK) and build-1:/srv/artefacts/tas/x0-pins/; the gpu-row-recipe.md (sha cbab453c…); +docs/analysis/class-v6/multi-family-adversary.md (master e3d0208d); docs/analysis/class-v6/coexistence-model.md (line +537); docs/analysis/class-v6/census-packs.md (sha 146b578c…); docs/design/class-v6-rotating-family.md 10.0h, 10.0i, +10.0q, 10.0r (sha 6eaf43bd…); docs/analysis/counter-asic-4-research.md 15.1a and 20.3 (sha 82af4424…); +docs/analysis/class-v6/floor/denominator.md (sha a4ec182a…); the register rows 18 and 105 to 126. diff --git a/docs/analysis/class-v6/1p5x/x5/README.md b/docs/analysis/class-v6/1p5x/x5/README.md new file mode 100644 index 000000000..61fa0fdd2 --- /dev/null +++ b/docs/analysis/class-v6/1p5x/x5/README.md @@ -0,0 +1,168 @@ +# X5: the co-optimised full opponent sweep (the 1.5x programme, 9 October 2026; first rows 09:4x UK on the frozen control) + +The adversary lane with the floor lanes under it. The sweep: the pure DRAM board, the hybrid at stored fractions 0.25, +0.5 and 0.75, the full SRAM store and recomputation, each as a complete machine (the memory devices, controller and +PHY, the core die sized to the memory's sustained rate, PSU and VRM losses, cooling, the host's share) at the +calibrated node, with its power, area and throughput; the worst credible design retained as the specialist side a. +The tool: `tools/1p5x/x5/sweep.py` (run on build-4 under the lease pool with a pid file; seconds). Every row is a +MODEL on the placed core's measured energies and the record's claimed memory figures; no chip is measured; the +vocabulary is NOT RUN / BLOCKED / FAIL / PASS and this lane never writes PASS. The registry batch for these rows reads +NOT RUN (`tools/ci/batches/adversary-20261009-x5.json`) until X6 runs the model blind. + +## 1. Sources by sha + +- The core: the placed 8-lane 18-family core, 6.55 pJ per lane-op at N5 on the class v4 draw (6.02 to 7.78), class-v6-adversary + f5a82bf6a `tools/chip-model/mf/results/table.csv` (ASAP7 placed with SPEF, gate-level VCD, the SRAM window a FakeRAM 64x256 macro per 8 + lanes with its access energy modelled at 3.5 pJ, 2.0 to 7.0); the node scaling N5 = ASAP7 x0.70, N3 = N5 x0.72 (claimed). +- The frozen control's counts: 102,100 shadow ops, 512 base, 128 loads at W = 1 (X0 section 3; the frozen program's own per-family + counts BLOCKED, X0 section 6). +- The honest card: the 5090 at its 1,300 lock, 2.319 microjoules per hash (X0 section 2, job run-ca3-pc1-v4-eff-5090-20261007-b). +- The memory: chip-model-v3.md 5.3 and 5.5 (GDDR7: 2.0 nJ per random read, 1.5 to 2.6; 21.3 G activates per second, 82 percent + sustained, the high case at the 5090's measured 17.5 G; 20 W static, 15 W controller; USD 320 of devices); the SRAM read at W = 1 + 0.25 nJ (0.20 to 0.35), floor lane 3 2.1; SRAM USD 250 per GiB and 15 W per GiB static (claimed, approximate); the N2 die USD 500. +- The hit rates (the window layer ON, the frozen control): the hottest half serves 0.581 of reads on the mean program (the hash lane's + reconcile, coexistence-model.md line 537), the quarter 0.341 and the three quarters 0.720 scaled from adv-cache-2's distribution by the + same ratio (approximate); the p98 program's 0.7188 is the tail. X1's traces replace these on the hybrid rows when they land. +- The machine terms: `tools/chip-model/mf/flow/board.py` (PSU 92 percent, VRM 90, cooling 3, one node per 100 machines). + +## 2. The rows on the frozen control: two consistent pairs (the first run 09:43 UK; the class v6 pair 10:1x UK, build-4) + +The chip side must be paired with the card side on the SAME program. Two consistent pairs exist today: (A) the class v4 pair, the +served convention (the card 2.319 microjoules per hash at the lock, the program 128 loads and 102,612 ALU ops: X0 section 2); +(B) the class v6 pair, the frozen object (the card 4.011 microjoules at the lock, the house 5090's only class v6 lock row, X0 5a; +the program 256 loads per hash, X4's read of the exported texts, with the ALU count carried at 102,612 until the hash lane's +line). Pair A is the first table (09:43 UK, 128 loads); pair B the second. The yardstick's r0 is pair B's, the frozen object's. + +Pair A (class v4, 128 loads, the card 2.319), same node (N5), the card over the machine; the band is the SRAM access band and +the memory band together; "credible" is a capital per MH/s under ten times the card's USD 15.9: + +| Variant | Sustained MH/s | Machine W | microjoules per hash (band) | Core lanes | Core mm^2 (N5) | Capex USD | USD per MH/s | Ratio (band) | Credible | Retained | +|---|---|---|---|---|---|---|---|---|---|---| +| The pure DRAM board (16 GDDR7 devices) | 136 | 209 | 1.528 (1.380 to 1.851) | 104,960 | 210 | 661 | 4.84 | 1.52x (1.25 to 1.68) | yes | | +| The hybrid, the hottest 25 percent in 0.5 GiB of SRAM (hit 0.341) | 207 | 283 | 1.367 (1.244 to 1.649) | 159,272 | 319 | 825 | 3.98 | 1.70x (1.41 to 1.86) | yes | | +| The hybrid, the hottest 50 percent in 1.0 GiB (hit 0.581) | 326 | 402 | 1.235 (1.128 to 1.483) | 250,502 | 501 | 1,015 | 3.12 | 1.88x (1.56 to 2.06) | yes | | +| The hybrid, the hottest 75 percent in 1.5 GiB (hit 0.720) | 487 | 561 | 1.151 (1.054 to 1.378) | 374,859 | 750 | 1,230 | 2.52 | 2.02x (1.68 to 2.20) | yes | | +| The full SRAM store (one N2 reticle, 2 GiB at W = 1) at a 1 kW budget | 1,419 | 1,350 | 0.951 (0.872 to 1.129) | 1,091,866 | 2,184 | 1,701 | 1.20 | 2.44x (2.05 to 2.66) | yes | the worst credible design on pair A | +| Recomputation: the hottest half in SRAM, the other 42 percent derived on the core, no DRAM | 251 | 1,359 | 5.411 (4.988 to 6.379) | 1,138,480 | 2,277 | 1,285 | 5.11 | 0.43x (0.36 to 0.46) | yes | loses at every point | + +Pair B (the frozen class v6 object, 256 loads, the card 4.011 at the knee), same node: + +| Variant | Sustained MH/s | Machine W | microjoules per hash (band) | Core lanes | Core mm^2 | Capex USD | USD per MH/s | Ratio, the card over the machine (band) | Credible (USD per MH/s under 10x the card) | Retained | +|---|---|---|---|---|---|---|---|---|---|---| +| the pure DRAM board (16 GDDR7 devices, 64 channels, 32 GB) | 68 | 148 | 2.172 (1.944 to 2.661) | 52,480 | 105 | 623 | 9.13 | 1.85x (1.51 to 2.06) | yes | | +| the hybrid, the hottest 25 percent of items in 0.5 GiB of SRAM beside the DRAM (hit 0.341); SRAM USD 125 claimed | 104 | 192 | 1.850 (1.671 to 2.257) | 79,636 | 159 | 767 | 7.41 | 2.17x (1.78 to 2.40) | yes | | +| the hybrid, the hottest 50 percent of items in 1.0 GiB of SRAM beside the DRAM (hit 0.581); SRAM USD 250 claimed | 163 | 258 | 1.585 (1.440 to 1.924) | 125,251 | 251 | 925 | 5.68 | 2.53x (2.08 to 2.79) | yes | | +| the hybrid, the hottest 75 percent of items in 1.5 GiB of SRAM beside the DRAM (hit 0.720); SRAM USD 375 claimed | 244 | 345 | 1.417 (1.292 to 1.715) | 187,429 | 375 | 1,095 | 4.49 | 2.83x (2.34 to 3.10) | yes | | +| the full SRAM store (one N2 reticle, 2 GiB at W = 1) at a 1 kW budget; the die USD 500 claimed; the reconciled ticket USD 1.0 per MH/s (0.5 to 1.6) | 1358 | 1347 | 0.992 (0.905 to 1.187) | 1,044,425 | 2089 | 1,667 | 1.23 | 4.04x (3.38 to 4.43) | yes | WORST CREDIBLE: a | +| recomputation: the hottest half in SRAM, the other 42% of items derived on the core (1,003,991 extra ops per hash), no DRAM; the core-bound rate at a 1 kW budget | 137 | 1359 | 9.907 (9.132 to 11.681) | 1,137,987 | 2276 | 1,284 | 9.36 | 0.40x (0.34 to 0.44) | yes | | + +The retained specialist side: the full SRAM store (one N2 reticle, 2 GiB at W = 1) at a 1 kW budget at 0.992 microjoules per hash (0.905 to 1.187); its ratio from the card's 4.011: 4.04x (3.38 to 4.43). + +A node ahead (N3) on pair B: 2.07x (1.68 to 2.32); 2.48x (2.03 to 2.75); 2.97x (2.44 to 3.27); 3.39x (2.79 to 3.72); 5.34x (4.46 to 5.87); 0.55x (0.47 to 0.60). + +The op count the chip executes (the hash lane's per-family read of the signing object, 10:11 UK): the frozen object is 62,120 +ALU instructions per hash (loads 256, the base 1,024, the per-load address path 5,376, the shadow 55,296 unchanged from class v4, +the reg64 init and output fold 168); class v4 about 56,300. The two tables above carry the record's COUNTED-OP convention +(102,612 per hash, the figure every served chip row rests on, from which the card's 0.652 microjoules reads 6.4 pJ per counted +op); the chip executes one instruction per lane-op, so at the INSTRUCTION count its shadow energy is 0.36 to 0.41 microjoules +per hash instead of 0.67, and every ratio rises. Both counts are carried until the hash lane says what a counted op is (asked +10:1x UK, default 10:45): the instruction count is the chip's physics and the default; the counted-op rows are the record's +convention. + +Both pairs at the INSTRUCTION count, same node, the card over the complete machine: + +| Variant | Pair A, instructions (56,300; the card 2.319): microjoules (band), ratio | Pair B, instructions (62,120; the card 4.011): microjoules (band), ratio | Pair B a node ahead | USD per MH/s (pair B) | +|---|---|---|---|---| +| The pure DRAM board | 1.129 (1.012 to 1.381), 2.05x (1.68 to 2.29) | 1.823 (1.622 to 2.250), 2.20x (1.78 to 2.47) | 2.39x (1.93 to 2.69) | 8.91 | +| The hybrid, the hottest quarter (hit 0.341) | 0.968 (0.875 to 1.179), 2.40x (1.97 to 2.65) | 1.501 (1.349 to 1.846), 2.67x (2.17 to 2.97) | 2.95x (2.39 to 3.29) | 7.19 | +| The hybrid, the hottest half (hit 0.581) | 0.835 (0.760 to 1.013), 2.78x (2.29 to 3.05) | 1.236 (1.118 to 1.513), 3.25x (2.65 to 3.59) | 3.67x (2.98 to 4.06) | 5.46 | +| The hybrid, the hottest three quarters (hit 0.720) | 0.752 (0.686 to 0.908), 3.09x (2.55 to 3.38) | 1.068 (0.971 to 1.304), 3.76x (3.08 to 4.13) | 4.33x (3.53 to 4.77) | 4.28 | +| The full SRAM store (one N2 reticle) at 1 kW | 0.540 (0.493 to 0.646), 4.29x (3.59 to 4.70) | 0.633 (0.574 to 0.764), 6.34x (5.25 to 6.99) | 8.24x (6.80 to 9.12) | 0.77 | +| Recomputation (the half stored, the rest derived) | 4.998 (4.607 to 5.893), 0.46x (0.39 to 0.50) | 9.547 (8.799 to 11.256), 0.42x (0.36 to 0.46) | 0.57x (0.49 to 0.62) | 9.02 | + +Reading. (1) At the instruction count the worst credible design is the die at 4.3x on pair A and 6.3x on pair B same-node (5.7x +and 8.2x a node ahead); the board 2.05x and 2.20x; the hybrids 2.4x to 3.8x. At the counted-op count (the tables above) the die +2.44x and 4.0x, the board 1.52x and 1.85x. The op-count definition is the largest single open term in the programme's r0 (a +factor of 1.3x to 1.6x on every chip ratio) and is the hash lane's to close. (2) On the frozen object the card pays 1.73x class +v4's joules for twice the loads, and the chip's memory term doubles while its ALU term grows 10 percent, so every pair-B ratio +sits above its pair-A row: the read count is the chip's cost too, but the card's per-read cost is 4x the chip's (8.7 nJ against +2.0). (3) Recomputation loses at every point and is never credible; the hybrids' ratio rises with the stored fraction and their +capital per MH/s falls. (4) The core's lanes are a tenth or less of every machine's capital except the die's, where the core is +3x the memory's silicon; the per-box placements (section 4) bracket the placement factor, which no single-port interleaving +changes. + +## 3. a, and the ratio from r0 with X1's g + +The programme's a is the worst credible design's energy per accepted hash on the frozen object (pair B), the N2 SRAM die: 0.992 +microjoules at the counted-op count (0.905 to 1.187), 0.633 at the instruction count (0.574 to 0.764); r0 for the die 4.0x to +6.3x same-node across the two counts (5.3x to 8.2x a node ahead); the DRAM board 1.85x to 2.2x. r_new = r0 x g / a: from the +die's 4.0x the honest card must cut its energy per hash by 63 percent at a = 1 for 1.5x and no 15 plus 20 split reaches it +(2.8x); from 6.3x, 76 percent. The X1 candidate's g (the card's class v7 energy over its class v6 energy, same card, same session) +and its traces re-run the hybrid rows and move nothing on the die (it serves every read whatever the distribution). BLOCKED until +X1 lands: g; the hybrid rows on the traces. BLOCKED on the hash lane: the counted-op definition (section 2). + +## 4. The co-optimised core (the per-variant hosts; PENDING with clocks) + +The 5-phase single-port slot buys the adversary energy (one macro access serves eight lanes; the units isolated) and pays in lanes: +105,000 lanes for the board, 1.1 million for the die. The co-optimised form interleaves four hash contexts per lane group on one +FakeRAM 256x256 macro (4 contexts x 64 registers x 8 lanes), so a lane group issues four ops per five cycles instead of one: the same +accesses per op, a 256-deep macro (its access about 4.5 pJ, 2.5 to 9.0, approximate), a quarter of the lanes and about a third of the +core's silicon and leakage. It changes the die's capital (USD 1.20 to about 0.6 per MH/s) and its leakage (the core's 55 W at 1.1 M +lanes to about 15), and so the die's a by a few percent; it changes the board's row under 1 percent. The RTL variant (`core_mf` with a +context index on the macro address and the phase counter) synthesises and places on the build boxes the coordinator assigns per +variant (the founder's order of 10:2x UK: build-2, 3, 5, 6, 7 and 8, one design each; no rented host); the converged SPEF row of the 32-lane core at 22 percent utilisation runs on build-4 (pid file +build-1:/srv/queue/pids/build-4-floorplan.pid, started 09:35 UK, about 13:30). Rows by 12:30 where the hosts land them; else PENDING +with the morning's clock. + +## 5. How X6 evaluates a (pack, seed) pair blind + +`python3 tools/1p5x/x5/sweep.py --rows tools/chip-model/mf/results/table.csv --node N5 --trace --json ` on any +box with python3, under a pid file. The trace JSON: `{"ops_per_hash": N, "loads_per_hash": L, "hit": {"0.25": h, "0.5": h, "0.75": +h}}`, the pair's own op count, load count and the fraction of its reads served by the hottest quarter, half and three quarters of +its items (the attack-f8 item histogram over the pair's nonces; X6 draws them). The retained design and its energy per accepted work +are the row marked retained; `--node N3` for a node ahead; `--gpu-uj` the card's energy for the same pair. The P03 cost lines (the +GPU-side cost of the candidate against the control) are the hash lane's measured rows, never this tool's. + +## 6. BLOCKED and NOT RUN + +| Row | State | Why | +|---|---|---| +| The hybrid rows on X1's traces | BLOCKED | X1's traces (11:30 UK) | +| g for r_new | BLOCKED | X1's paired card rows | +| The numerator on the frozen class v6 object | BLOCKED | the hash lane's paired class v4 against class v6 control (X0 assumption 2) | +| The co-optimised core's placed rows | PENDING | the assigned build boxes; by 12:30 or the morning's clock | +| The converged SPEF row (32-lane, 22 percent) | PENDING | build-4, about 13:30 | +| Every row of this file in the registry | NOT RUN | X6's blind run | + +## 7. B5 (the research plan's Track B item, the coordinator's ask of 10:0x UK): the full-SRAM design on fresh state and the near-memory design, as WITNESS rows + +The unit is Track B's proof-carrying, input-specific memory-work unit (an MTP / DRSample labelling of N labels of 32 bytes per +instance, the state built fresh per instance, in-degree 2, the proof search small beside the construction); the unit is PROVISIONAL +until the B2 lane defines the lottery-contribution and security-budget comparison, and NO ratio is written (no GPU side exists for +the unit until B3); the design goal for Track B is 1.25x with the guardband of the plan's section 3.6 (1.25 x 1.05 / 0.90 = 1.46) +beside the programme's 1.5x yardstick. Every row is a WITNESS (E_D), a modelled design the specialist could build, never a BOUND +(L), which is the B4 theory lane's. The tool: `tools/1p5x/x5/b5.py` (build-4, seconds; the label hash taken at 1,000 ALU ops on +the placed core, a provisional Blake2-class count the B1 reference replaces; the SRAM write 0.30 nJ per label, 0.25 to 0.45, the +read 0.25, 0.20 to 0.35, the N2 die's wire figures; the PIM in-bank access 0.35 nJ, 0.25 to 0.5, approximate from the HBM2 +breakdown's data-movement share; one hash unit per pseudo-channel, 32 per stack, the chain sequential per instance). + +| Design (WITNESS E_D) | State per instance | Labels N | Core energy per instance (microjoules) | Memory energy per instance | Energy per instance, device (band) | Instances per second per machine | Chains in flight | SRAM held (GiB) | Core lanes | Core mm^2 | Machine W | Energy per instance, whole machine | Capex USD | USD per (instance per second) | +|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| +| fresh-state full SRAM | 32 MiB | 1,048,576 | 6,868 | 576.7 | 7,445 (6,791 to 9,012) | 134,320.33 | 1,055,810 | 32,994.1 | 1,055,810 | 2,112 | 616,960 | 4,593,197 | 8,249,477 | 61 | +| near-memory (PIM) on one HBM3 stack | 32 MiB | 1,048,576 | 6,868 | 734.0 | 7,602 (6,843 to 9,222) | 4.07 | 32 | 0.0 | 256 | 1 | 5 | 1,235,621 | 550 | 135 | +| fresh-state full SRAM | 128 MiB | 4,194,304 | 27,473 | 2,306.9 | 29,780 (27,162 to 36,048) | 33,580.08 | 1,055,810 | 131,976.3 | 1,055,810 | 2,112 | 2,463,910 | 73,374,159 | 32,995,027 | 983 | +| near-memory (PIM) on one HBM3 stack | 128 MiB | 4,194,304 | 27,473 | 2,936.0 | 30,409 (27,372 to 36,887) | 1.02 | 32 | 0.0 | 256 | 1 | 5 | 4,942,483 | 550 | 541 | +| fresh-state full SRAM | 512 MiB | 16,777,216 | 109,891 | 9,227.5 | 119,118 (108,649 to 144,192) | 8,395.02 | 1,055,810 | 527,905.1 | 1,055,810 | 2,112 | 9,851,712 | 1,173,518,537 | 131,977,226 | 15,721 | +| near-memory (PIM) on one HBM3 stack | 512 MiB | 16,777,216 | 109,891 | 11,744.1 | 121,635 (109,488 to 147,547) | 0.25 | 32 | 0.0 | 256 | 1 | 5 | 19,769,931 | 550 | 2,162 | + +Reading. (1) On fresh state the full-SRAM design's cost per instance is the construction's core energy (N hashes at 1,000 ops: 6.9 +joules per 512 MiB instance, 0.43 per 32 MiB), with the SRAM's write and read energy under a percent of it, so the design does +not escape the labelling's cost by holding the state in SRAM: what the SRAM buys it is the sequential chain's speed (a label every +7.5 microseconds on one lane group) and the instances it can hold in flight, and what it pays is the SRAM held (the chains in flight +times the state, 2.4 to 38 GiB for the 32 to 512 MiB instances at 1 kW: the capital column). (2) The near-memory design is bound by +the chain: one instance per hash unit, 32 units per stack, so a 512 MiB instance's stack runs 32 chains and finishes one instance +every 250 ms per unit; its energy per instance is the same core energy (the hash is the cost, not the access) with the in-bank +access a third of the across-interface figure, which changes the device energy under a percent. (3) So for this unit the memory +technology is not where the specialist's energy goes; the hash's per-op energy on the core is, which puts the Track B question +where the plan's B4 puts it: the energy of accepted proofs under partial evaluation, shortcuts and amortisation, a bound, not a +witness. These rows stay the witness the B4 bound must beat. diff --git a/docs/analysis/class-v6/multi-family-adversary.md b/docs/analysis/class-v6/multi-family-adversary.md index f73c81dfc..7ccb061d7 100644 --- a/docs/analysis/class-v6/multi-family-adversary.md +++ b/docs/analysis/class-v6/multi-family-adversary.md @@ -922,4 +922,11 @@ per hash for the whole-window coupling, 1.2 percent of its shadow energy, and th from 1.5x to 1.49x same-node (not to 1.42x, which was the literal fold's cost); the card meanwhile keeps the fold at 6.20 nJ per hash on the 5090 and would pay 7.13 on the prefix form. The two shapes together: the chip's cost of the coupling is 8 nJ (connected-state shape, prefix form) to 56 nJ (shipped shape, literal form), 0.02x to 0.08x -of the bracket, and no design in the record charges the chip a 63-read cost it can avoid. +of the bracket, and no design in the record charges the chip a 63-read cost it can avoid. (6) X4's corrections (10:0x UK, +9 October; tools/1p5x/x4/chip_rows.py on 1p5x-x4, read from the exported class v6 texts): the class executes 256 loads +per hash (32 per iteration x 8), not 128, so the literal fold on the shipped shape is 16,128 ops and reads, 113 nJ at +the slot pricing (23 nJ at the datapath pricing); and the running total gives S but not the prefix P_s the identity +needs at every load, so the cheapest byte-identical form keeps 8 group sums (groups 1..7 moved on the base program's +writes to r8..r63, group 0 recomputed at each load: 7,560 ops, 2,784 extra reads, 36,864 flop bits, +32 B of state per +lane), 10 to 52 nJ per hash (the datapath to the slot pricing), which replaces the 7 nJ (5 to 10) row above; the +bracket moves 0.01x to 0.08x either way. diff --git a/docs/analysis/mhpow/b2/README.md b/docs/analysis/mhpow/b2/README.md new file mode 100644 index 000000000..9911614a2 --- /dev/null +++ b/docs/analysis/mhpow/b2/README.md @@ -0,0 +1,159 @@ +# B2: binding the economically expensive work (Track B, the 1.5x research programme) + +Lane: B2, the binding spec lane. Written 9 October 2026, 10:0x to 10:3x UK; re-cut at 10:3x to the B1 lane's encoding pins. Documents only. Every row here is NOT RUN in the +registry until the panel reads it; this lane writes no PASS. The status words in the table are this lane's reading of a paper +(ANSWERED), a written counter-example to a candidate rule (FAIL), or an open question with its owner (BLOCKED). + +## 0. Sources, by sha + +| Source | Identity | What was read | +|---|---|---| +| The founder's research plan, "IGNEUM - The 1.5x Research Programme" | sha256 269ceaa5824d09140a1bcc124dba59438db5a5e68fddc18960db7141fefc0305 | sections 2, 3, 4, 6 (B2 above all), 8, 9, 10 | +| Blocki and Smearsoll, "Provably Memory-Hard Proofs of Work With Memory-Easy Verification", TCC 2025, eprint 2025/1456 (the plan's R1) | the FULL PDF, sha256 13c3646e7d85c1aa58a92914582caab5798d90cf2a3cad33e58d81d49c1d31db, 608,919 bytes; fetched on build-9 from the Internet Archive capture of `https://eprint.iacr.org/2025/1456.pdf` at 2025-12-31 15:20:15Z (the eprint host answers a Cloudflare challenge, HTTP 403, from the box and from the Mac); the capture's CDX digest equals the 2025-08-12 capture's, so one version is on record | all 46 pages: sections 1.3, 2, 3, 4, 5, 6, 7 and Appendices A (Lemmas 7 to 10) and C (Merkle reveal and check) in full; Appendix B (the proofs of Theorems 9 and 10) skimmed. Copy at build-9:/srv/builds/b2/2025-1456.pdf for the B1 lane | +| The frozen class's header binding | igneum-pow `src/bind.rs` at the class-v6 freeze tree 1a938abe4 (generator fingerprint 5f4d6dc6...) | the init words, the pre-PoW hash rule, the interim day rule | +| The node's header hash | fork branch class-v6-node-review e8773ff5, `consensus/core/src/hashing/header.rs`, `hash_override_nonce_time` | every field the pre-PoW hash absorbs | +| The spec | `docs/spec/01-lottery-hash.md` on master bc6bfa75e: 1.0 (the frozen object), 1.6, 1.10, 1.12, 1.13.3 | the dataset policy digest in full, the epoch, day and era clocks | +| The signing pair | `igneum-pow/tests/composition.rs` at class-v6 755c2dbcf | epoch seed af89be5d..., era edc4fa84..., day 20730, program id 0x2a1d6caab4c24564 | +| The B1 lane's encoding pins (R15-05, a4490b2dfe9114a55, 10:2x UK) | unkeyed BLAKE2b-256, `lambda = 256`, one ASCII role byte per query kind, u64 LE integers, raw-label leaves, path bit `j` = bit `j - 1` of the leaf index (LSB at the root), no path sharing, the graph from `'G' || graph_seed || u64 v || u64 ctr` by rejection sampling (DRSample from ABH17 Algorithm 1) | adopted here whole; B1 owns the primitive and its fixtures | +| Figures | `tools/mhpow/b2/b2_figures.py` (sha256 7b8664ea44b770c9f6af246898aa576f3f0a63e65e7bf50a0b278c5727a1ee8b) | output `figures.txt` (sha256 27bb414c8534a438fdeae98cba2d4b9ece9c21d0487aa872673b0da465d9702a), run on build-9 under `/srv/builds/b2/b2-run.pid`: `python3 tools/mhpow/b2/b2_figures.py > docs/analysis/mhpow/b2/figures.txt` | + +## 1. What the paper proves, and what it does not + +The construction (paper section 3.2). `Prove^H(chi, N, k)` labels a DAG `G` on `N = 2^n` nodes: `l_1 = H(chi, 1)`, +`l_v = H(chi, v, l_v1, ..., l_vk)` over the parents in ascending order; commits to the labels with a Merkle tree salted by `chi`; +derives `k` challenges `c_i = H(chi, i, tau) mod N` from the root `tau`; opens `l_(c_i + 1)` and its parents. `Verify` recomputes +the challenges from `chi` and `tau`, checks every opening against `tau` and checks local consistency. The certificate is +`tau` plus the openings. + +| Paper statement | Content | Constants | +|---|---|---| +| Definition 5 | an MHPoW is `(Prove, Verify)` with prover efficiency `cmc <= N^2 lambda`, `O(N)` rounds, verifier efficiency, perfect completeness, and `(eps, C)`-soundness for every input `chi` in `{0,1}^lambda` | there is no difficulty parameter and no target: one certificate per input | +| Lemma 1 (oracle prediction) | a predictor with a hint of `|h|` bits outputs `|S|` fresh correct `lambda`-bit pairs with probability at most `|h| 2^(-lambda |S|)` | | +| Lemmas 7, 10 (Appendix A) | COLLISION at most `C(q,2) 2^-lambda`; BADORDER at most `q^2 n 2^-lambda`, `n` the query length in bits | | +| Lemma 2 | MISCOLOR at most `|V(G)| 2^-lambda` | | +| Theorem 4 | without COLLISION, BADORDER, MISCOLOR, the ex-post-facto pebbling is a legal pebbling of the ex-post-facto (green-only) graph `G'` | | +| Theorem 5, Lemma 3 | the extractor recovers `|P'_i|` oracle pairs from `sigma_i` and a hint of `(2 log2 q + log2 indeg) |P'_i|` bits; IE at most `t 2^(-lambda/2)` | precondition `lambda/4 >= 2 log2 q + log2 indeg(G')` | +| Theorem 6 | `cmc(trace) >= cc(G') lambda / 4` absent the bad events | the 1/4 is the rounding of `lambda - 2 log2 q - log2 indeg` under Lemma 3's precondition | +| Lemma 4 | LUCKYQUERY (a fresh root whose `k` challenges all land on green nodes while at least `beta N` nodes are red) at most `q (1 - beta)^k` | | +| Theorem 7, Corollary 1 | `cmc >= (lambda/4) min over |S| <= beta N of cc(G - S)` except with `2^-lambda C(q,2) + n 2^-lambda q^2 + 2^-lambda |V| + 2^(-lambda/2) t + q (1 - beta)^k` | `q` queries, `t` rounds | +| Lemma 5, Corollary 2 | for `(e, d)`-depth-robust `G`, `beta = e/(2N)`, `k = (2 N ln 2 / e) lambda`: `cmc >= (lambda/4)(e d / 2)` except with the sum above, last term `q 2^-lambda` | the statement writes `q 2^-lambda`, the proof's last line `q e^-lambda`; harmless, noted for B4 | +| Fact 8, Corollary 3 | DRSample: indeg 2, `(c3 N / log N, c4 N)`-depth-robust, so `cmc >= (lambda/4)(c3 c4 / 2) N^2 / log N`; `k = 2 lambda log N / c3` (the statement's `c1` is `c3/2`) | `c3`, `c4` exist (ABH17, ABP17); the paper gives no value | +| Section 7, Theorems 9, 10 | attacks on any MTP-style MHPoW: delete `e = N/k` nodes, pebble `G - S`; for constant indegree `k` must be `omega(log N / log log N)` | Attack 1: `cmc <= O(N d lambda + N lambda log N)`, success `(1 - e/N)^k` | + +Outside the paper: a lottery or difficulty target; more than one accepted certificate per input; amortisation over many inputs +(every statement fixes one `chi`); an input chosen after seeing the oracle; energy, bandwidth, SRAM against DRAM (the measure +is cumulative memory, bits times PROM rounds); quantum adversaries (the model is the classical parallel random oracle). + +## 2. The canonical instance input (work version 1, draft) + +Every field has a fixed width, integers are little-endian, 328 bytes in all (figures section 1, self-checked contiguous). +The VERIFIER builds the instance from the header and from the chain's own state. The certificate carries only `tau` and the +openings (paper section 3.2, step 4); no instance field is ever read from the certificate. + +| Offset | Width | Field | Value rule | Example (the signing pair) | +|---|---|---|---|---| +| 0 | 24 | tag | ASCII `igneum-mhpow/instance/1/` | | +| 24 | 32 | network_id | u8 length, ASCII, zero padded | `igneum-devnet-4` | +| 56 | 8 | chain_id | u64 | 4465 | +| 64 | 2 | header_version | the header's `version` (object byte in bits 8 to 13) | PLACEHOLDER | +| 66 | 2 | work_version | 1 for this draft | 1 | +| 68 | 4 | reserved_a | zero | | +| 72 | 32 | dataset_policy | the consensus digest arm over `class_v6_dataset_steps` and `class_v6_family_flags` (spec 1.0, 1.13.3) | 70a6c703787d5d75cdbc486b34acf3eb901f4188bae3e948c94274d5f4ac4bba | +| 104 | 8 | program_id | the class's program id for the epoch (the signing id) | 0x2a1d6caab4c24564 | +| 112 | 32 | epoch_seed | the epoch's program seed bytes (devnet: the epoch block hash, spec 1.12) | af89be5ddbadb6f6b4aee28ac8f249713be5d4c12621e3cea7f83ceada3c66b3 | +| 144 | 8 | epoch_start_daa | the epoch's start score `s = 86,400 d + L e` (spec 1.12) | PLACEHOLDER | +| 152 | 32 | era_seed | the era seed bytes | edc4fa844da9dc98d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07 | +| 184 | 8 | day_index | the day as the frozen class derives it (`bind.rs`: `timestamp_ms / 86,400,000`) | 20730 | +| 192 | 32 | state_root | the state stream root the class's dataset is keyed on | PLACEHOLDER (the node1 state file is sha256 abb5800350b02dd9...; the root bytes: BLOCKED B-N3) | +| 224 | 32 | header_prehash | `hash_override_nonce_time(header, 0, header.timestamp)` (`bind.rs`, the fork at e8773ff5) | PLACEHOLDER (`00...01`, the composition test's prehash) | +| 256 | 8 | daa_score | the header's | PLACEHOLDER | +| 264 | 8 | timestamp_ms | the header's | PLACEHOLDER (first millisecond of day 20730) | +| 272 | 4 | bits | the header's | PLACEHOLDER | +| 276 | 4 | reserved_b | zero | | +| 280 | 8 | trial_id | the header nonce, all 64 bits | 0 | +| 288 | 1 | log2_n | consensus parameter | PLACEHOLDER 24 | +| 289 | 1 | indeg | 2 (DRSample) | 2 | +| 290 | 2 | label_bytes | `lambda / 8` | PLACEHOLDER 32 (section 3, O-12) | +| 292 | 2 | k | challenges | PLACEHOLDER 0 (O-11) | +| 294 | 2 | reserved_c | zero | | +| 296 | 32 | graph_seed | consensus constant: the DRSample sampling seed | PLACEHOLDER zero | + +The header fields the brief names are all bound. `header_prehash` absorbs, at e8773ff5: `version`, `parents_by_level` (the +parent set, every level), `hash_merkle_root` (the coinbase commitment: the coinbase is in the transaction root), +`accepted_id_merkle_root`, `utxo_commitment`, `timestamp`, `bits`, the nonce as zero, `daa_score`, `blue_score`, `blue_work`, +`pruning_point`, `vote_key_hash`. The timestamp window is the node's acceptance rule; the instance binds the exact timestamp, so +each admissible timestamp is a different instance. `daa_score` and `timestamp_ms` also stand as explicit fields because the +epoch, day and era rules read them; the verifier checks each explicit field against the header and the chain. + +The oracle and the domain separation: + +| Use | Query | Paper form | +|---|---|---| +| Oracle | `H(role, x)` = unkeyed BLAKE2b, digest `lambda/8` = 32 bytes at B1's `lambda = 256`, input `role || x`; BLAKE2b stops at 512 bits (a larger `lambda` needs an XOF: B-T6) | one random oracle `H: {0,1}* -> {0,1}^lambda` | +| Instance digest | `chi = H('I', instance)`, 32 bytes | the paper's input `chi` in `{0,1}^lambda` (Definition 5) | +| Source label | `l_1 = H('L', chi || u64 1)` | `l_1 = H(chi, 1)` | +| Label | `l_v = H('L', chi || u64 v || l_p1 || l_p2)`, parents ascending; node 2 has the one parent 1; the parent count is fixed by the graph | `l_v = H(chi, v, l_v1, ..., l_vk)` | +| Merkle node | `tau_x = H('M', chi || tau_x0 || tau_x1)`; leaves are the raw labels, node `w` at leaf index `w - 1`; root `tau = tau_empty`; a reveal is the label and `n` siblings | section 4.1, Appendix C, salt `chi` | +| Challenge | `c_i = int_le(H('C', chi || u64 i || tau)) mod N`, `i = 1 .. k`; open node `c_i + 1` and each parent, each with its own path | `c_i = H(chi, i, tau) mod N` | +| Graph | DRSample, indeg 2, its draws from `H('G', graph_seed || u64 v || u64 ctr)` by rejection sampling; `graph_seed` a 32-byte public consensus constant, never `chi` | Fact 8; section 1.2 (the graph fixed a priori) | +| Lottery value | NONE in this draft (O-07) | the paper has none | + +Every query is fixed width once the role and the node are known, which keeps Theorem 5's parse unique (the extractor reads the +`h`-th parent at a fixed offset after `chi` and `v`). The role byte keeps the query families disjoint (B1's note: without it the challenge query `H(chi, i, tau)` has the shape of +an indegree-1 node's prelabel). The paper uses one oracle with structured inputs and no role byte; the role byte is a refinement that leaves every event the proofs bound +(COLLISION, BADORDER, MISCOLOR, LUCKYQUERY) defined as before. + +The example instance's bytes, its sha256 (7e24f382...), `chi` (12b3c35a... at trial 0, 3a02f2f6... at trial 1) and `l_1` are in +`figures.txt` section 2. They check the layout only: eleven fields are PLACEHOLDER until a live template (the node lane) and the +parameter choice (B3, B4) fill them. They are not conformance vectors; the B1 lane owns the primitive's known-answer fixtures. + +## 3. The obligations table + +| Id | Obligation | The construction's answer | Paper reference and constants | Reading | +|---|---|---|---|---| +| O-01 | The instance encoding is injective: two different (network, version, epoch, template, trial) tuples never give one `chi` | fixed widths, explicit lengths, a tag and a work version; `chi` from the same oracle | a `chi` collision is a COLLISION event: Lemma 7, `C(q,2) 2^-lambda` | ANSWERED | +| O-02 | Domain separation of the oracle's roles | the role byte (section 2; B1's pin, 'G' for graph sampling keyed by `graph_seed`, never `chi`) | Theorem 5's parse needs fixed offsets; Lemmas 7, 8, 10 hold for any query form | ANSWERED (the role byte is not in the paper; B-T7 asks B4 to confirm no step uses a cross-role coincidence) | +| O-03 | Reuse on another template: an accepted certificate replayed with another parent set, DAA score, coinbase or timestamp | every one of those changes `header_prehash`, so `chi`; every label, Merkle and challenge query carries `chi`, so the openings fail local consistency and the challenges move | Definition 5 soundness is per `chi`; Corollary 3 fixes `chi`; a replay needs `chi = chi'` (O-01) | ANSWERED | +| O-04 | Reuse on another epoch, day, era, network or protocol version | the same: those fields are in the instance, and the verifier recomputes each from the chain (never from the certificate) | as O-03 | ANSWERED, conditional on B-N1 (the verifier's recomputation rule and the live day rule) | +| O-05 | The graph is fixed before the prover acts (the Dinur and Nadler lesson) | DRSample sampled once from `graph_seed`, a consensus constant; never from `chi`, the epoch, the era or the template | the paper's own condition (section 1.2): MTP is sound "as long as the underlying graph is fixed a priori"; Fact 8 is an existence statement | BLOCKED B-T1: the probability that one sampled DRSample instance at the chosen `N` fails the `(c3 N / log N, c4 N)` depth-robustness, and whether a per-era re-seed is admissible | +| O-06 | Grinding the commitment to dodge the challenges (try roots until all `k` challenges land on honestly computed nodes) | `k` from Corollary 2 | Lemma 4: `q (1 - beta)^k`; Corollary 2: `beta = e/(2N)`, `k = (2 N ln 2 / e) lambda` gives `q 2^-lambda`; figures section 4 | ANSWERED (the constants of `k` wait on `c3`: B-T2) | +| O-07 | Grinding the lottery: what the prover can vary per draw, at what cost, once one labelling exists | no lottery rule is adopted. Two candidate rules fail: (a) `W = H('W' || chi || tau) <= target`: after one honest labelling the prover sets the sink label `l'_N` to fresh bytes (the sink turns red), rehashes its Merkle path and draws again, `log2 N + 1` calls a draw, and the certificate still verifies unless a challenge lands on the sink, probability `(1 - 1/N)^k` (0.99988 at `N = 2^24`, `k = 2,000`); honest work per trial is `2N` calls, so one labelling buys about `1.3 x 10^6` draws at `N = 2^24` (figures section 5). Varying any unchallenged node with few descendants does the same; a fixed always-checked set of the last `m` nodes moves the cost to about `m` labels a draw, and only full verification removes it. (b) `W = H(chi) <= target`: the prover screens trials for free and labels only the winner, so the memory work per block is one labelling whatever the difficulty | the paper's theorems lower-bound the cost of a trace that outputs AN accepted certificate (Theorem 7); they say nothing about how many distinct accepted certificates one trace outputs for one `chi`, and Definition 5 has no target. The re-roll is consistent with every theorem in the paper | (a) FAIL, (b) FAIL; the replacement BLOCKED B-T3 | +| O-08 | Grinding through the template: timestamp inside the window, coinbase extranonce, transaction set, parent choice, the nonce | each is a new `chi`, hence a full new labelling (no label query is shared across `chi`) | per `chi`: Corollary 3, `(lambda/4)(c3 c4 / 2) N^2 / log N` cumulative memory; across many `chi`: O-10 | ANSWERED per trial; the many-trial bound BLOCKED (O-10) | +| O-09 | Chosen instance: the prover picks a favourable `chi` | the graph is data-independent and fixed (O-05), so no `chi` changes the graph's structure; the only freedom is which `chi` to label | Corollary 3 fixes `chi` before the oracle; a `chi` chosen after querying the oracle needs a union over the candidates (at most `q`), a loss the paper does not state | BLOCKED B-T4 | +| O-10 | Amortisation: one labelling, or one shared state, serving many trials or many templates | none can share labels across `chi` (O-03); the lower bound for producing accepted certificates on `m` distinct inputs is the open part | not in the paper (every statement fixes one `chi`); the natural route is the disjoint union of `m` copies (cumulative pebbling cost adds over components), the extractor run on the union, LUCKYQUERY bounded per copy | BLOCKED B-T5 | +| O-11 | Partial evaluation: answering the challenges without the full labelling | the red-node budget `beta N` and the challenge count `k` | Theorem 7 (`beta`), Lemma 5 (`cc(G - S) >= (e - |S|) d`), Corollary 2 (`|S| <= e/2`, `cc(G') >= e d / 2`); Section 7: `k` must be `omega(log N / log log N)`, Theorem 10 attacks below that; certificate sizes in figures section 4: 19.5 MiB at `lambda = 256`, `N = 2^24`, `c3 = 1`; 194.9 MiB at `c3 = 0.1`; with a lottery-sized luck term (`q = 2^80`, `eps = 2^-40`) 9.1 MiB and 91.4 MiB | ANSWERED as formulas; the numbers wait on `c3` (B-T2) and the network budget (B6) | +| O-12 | The concrete constants: no hidden big-O in a numerical claim (plan section 8) | `lambda = 256` as B1 pins it; it must be chosen against the adversary's query count `q` | Lemma 3 needs `lambda/4 >= 2 log2 q + 1`: `lambda = 256` covers `q <= 2^31.5`, `512` covers `2^63.5`, `1024` covers `2^127.5` (figures section 3). Without the rounding the charge per pebble is `lambda - 2 log2 q - 1` bits: 127 of 256 at `q = 2^64` (0.496), 95 at `q = 2^80`. The proved ratio of the honest prover's cumulative memory (about `N^2 lambda / 2`) to the bound is `4 log N / (c3 c4)` | BLOCKED B-T2 (`c3`, `c4` for DRSample at the chosen `N`) and B-T6 (`lambda` for a network-scale `q`, and the XOF above 512 bits) | +| O-13 | The cost measure is the one the 1.5x target needs | none in the paper: the bound is cumulative memory in the parallel random oracle model, not joules, bandwidth or SRAM against DRAM | the plan's section 2 and B4: a theorem in bits times rounds is level 2 evidence; the physical translation (level 3) and the full-SRAM design (B5) are separate obligations | BLOCKED B-T8 | +| O-14 | Rejected trials and cancellation: a trial in flight when the template changes | the instance binds the parent set, so a new block on the network stales every trial in flight; the honest trial takes `O(N)` sequential rounds (Definition 5, item 1) | not a soundness matter; it is a feasibility bound: `N x t_seq <= rho x T_block`, with `T_block` 1 s on the devnet (spec 1.12) and `rho` the stale fraction the chain accepts. The bound caps `N`, and a small `N` fits one instance in SRAM, which is the B5 question | BLOCKED B-N2 (the node lane: `T_block` and `rho` on devnet-4) and B3 (`t_seq` measured, never assumed) | +| O-15 | The frozen class stays the chain's binding: the dataset policy 70a6c703 and the signing id 2a1d6caab4c24564 | the instance carries `dataset_policy` (all 32 bytes), `program_id`, `epoch_seed`, `epoch_start_daa`, `era_seed`, `day_index` and `state_root`; the verifier recomputes each from the chain and the header and rejects a mismatch, so no certificate carries across a class object, a policy or an epoch | by construction plus O-01, O-03. Two facts for the node lane: the policy digest sits outside the consensus digest until the class v6 floor is set (f0 manifest), so the instance is the only place a certificate binds it today; and the day rule differs between `bind.rs` (the interim `timestamp_ms / 86,400,000`) and spec 1.12 (DAA seconds) | ANSWERED by construction, conditional on B-N1 and B-N3 | +| O-16 | The epoch-seed lead and the day-pack: no precomputation across trials | the labelling needs `chi`, which needs the template's parents; the 600 DAA-second epoch-seed lead and the day key (known before the day) give no label in advance. Static per-day work (the v6 dataset) stays amortisable by design; the instance removes it from the MHPoW part only | O-03 | ANSWERED | + +## 4. The BLOCKED questions, with their owners + +| Id | Owner | The exact question | +|---|---|---| +| B-T1 | B4 theory lane (to be named) | For DRSample sampled from a fixed public seed at `N = 2^20` to `2^24`, what is the probability that the sampled graph is not `(c3 N / log N, c4 N)`-depth-robust, with the constants written out? Is a per-era re-seed from the era seed admissible, given that a party able to bias the era seed could search for a weak graph? | +| B-T2 | B4 | The values of `c3` and `c4` (Fact 8, from ABH17 or ABP17) for the sampled DRSample at the chosen `N`, proved, so that Corollary 3's bound and `k = 2 lambda log N / c3` become numbers. Without them every certificate size and every lower bound in this README is a sensitivity row. | +| B-T3 | B4 | The lottery. Give a lottery value `W` (or prove none exists) such that the expected number of distinct accepted pairs (`chi`, certificate) with `W` under the target, output by an adversary of cumulative memory `C`, is at most `C` divided by a constant fraction of Corollary 3's bound, with sampled verification of polylog cost. Known: any `W` that depends on labels the verifier samples can be re-rolled through an unchallenged node at the cost of that node's descendants (O-07 (a)); any `W` of `chi` alone is screened for free (O-07 (b)). Candidates for the panel: a succinct proof of the full labelling (a unique output, so Alwen and Serbinenko's bound per trial applies, at a prover cost the paper's own section 1 calls non-egalitarian); or a construction that is not MTP. | +| B-T4 | B4 | The adaptive-input version of Corollary 3: the adversary chooses `chi` among at most `q` candidates after querying `H`. State the loss in the bad-event sum. | +| B-T5 | B4 | The multi-instance version: producing accepted certificates for `m` distinct inputs costs at least `m (lambda/4)(c3 c4 / 2) N^2 / log N` except with what probability, and how the `q` and `t` terms scale with `m`. | +| B-T6 | B4 | The label width `lambda` for an adversary of `q = 2^64` to `2^96` queries: either `lambda >= 8 log2 q + 4` under Lemma 3 as stated (516 to 772 bits, so an XOF in place of BLAKE2b's 512-bit maximum), or a tighter statement of Theorem 6 with the charge `lambda - 2 log2 q - log2 indeg` per pebble at `lambda = 256`. | +| B-T7 | B4 | Confirm that no step in Theorems 4 to 7 or Lemmas 2 to 4 uses a coincidence between a label query and a Merkle or challenge query that the role byte would remove (the role byte only splits the domain; the check is that the extractor's hint and parse still hold). | +| B-T8 | B4 with the adversary lane | The translation from cumulative memory to energy and bandwidth (the plan's R2, R3), including the full-SRAM design at the `N` that O-14 allows. | +| B-N1 | node lane a283f5f0d364ceef0 | Which day rule is live on devnet-4 (`bind.rs`'s timestamp day or spec 1.12's DAA day), and the rule the verifier uses to recompute `epoch_seed`, `epoch_start_daa`, `era_seed`, `program_id` and the policy digest from a header and the chain. | +| B-N2 | node lane | The block interval on devnet-4, the timestamp acceptance window (past median and future bound), and the stale fraction the chain accepts, so O-14's bound on `N x t_seq` is a number. | +| B-N3 | node lane | The 32 bytes of the state stream root the class v6 dataset is keyed on for the signing pair (epoch af89be5d, era edc4fa84, day 20730), and whether that root is in the header's past by the epoch boundary. | + +## 5. What this means for Track B + +1. The input binding is complete for reuse: no certificate carries across a template, a trial, an epoch, a day, an era, a network + or a class object (O-01 to O-04, O-15, O-16). +2. The economic binding is not: the reference construction has no lottery, and the two direct ways to add one either buy about a + million draws per labelling at `N = 2^24` or screen trials for free (O-07). Until B-T3 is answered, an MTP certificate per + block adds memory work per block, never per lottery trial, and cannot move the 1.5x ratio. +3. Binding the template forces each trial to finish inside a fraction of the block interval (O-14), which caps `N` and pushes the + instance toward a size that fits in SRAM. B3 and B5 own that collision; it is recorded here because the binding causes it. +4. Even granted a lottery, the proved bound sits a factor `4 log N / (c3 c4)` under the honest prover's cumulative memory with + unknown `c3 c4` (O-12), and its unit is memory-time, not energy (O-13). + +No row here is a PASS. The registry batch is not run. diff --git a/docs/analysis/mhpow/b2/figures.txt b/docs/analysis/mhpow/b2/figures.txt new file mode 100644 index 000000000..652bcd7fa --- /dev/null +++ b/docs/analysis/mhpow/b2/figures.txt @@ -0,0 +1,96 @@ +== 1. canonical instance layout (work_version 1, draft) +offset width field encoding +0 24 tag ASCII 'igneum-mhpow/instance/1/' +24 32 network_id u8 length, then ASCII, zero padded +56 8 chain_id u64 LE +64 2 header_version u16 LE, the header's version field (object byte in bits 8 to 13) +66 2 work_version u16 LE, 1 for this draft +68 4 reserved_a zero +72 32 dataset_policy the consensus digest arm's 32 bytes +104 8 program_id u64 LE, the class's program (signing) id +112 32 epoch_seed the epoch's program seed bytes +144 8 epoch_start_daa u64 LE, the epoch's start DAA score +152 32 era_seed the era seed bytes +184 8 day_index u64 LE, the day as the frozen class derives it +192 32 state_root the state stream root the class's dataset is keyed on +224 32 header_prehash hash_override_nonce_time(header, 0, header.timestamp) +256 8 daa_score u64 LE +264 8 timestamp_ms u64 LE +272 4 bits u32 LE +276 4 reserved_b zero +280 8 trial_id u64 LE, the header nonce, all 64 bits +288 1 log2_n u8 +289 1 indeg u8, 2 for DRSample +290 2 label_bytes u16 LE, lambda / 8 +292 2 k u16 LE, challenges +294 2 reserved_c zero +296 32 graph_seed consensus constant: the DRSample sampling seed +total 328 + +== 2. example instance (layout check only; PLACEHOLDER fields named in the docstring) + 0 69676e65756d2d6d68706f772f696e7374616e63652f312f0f69676e65756d2d + 32 6465766e65742d34000000000000000000000000000000007111000000000000 + 64 010801000000000070a6c703787d5d75cdbc486b34acf3eb901f4188bae3e948 + 96 c94274d5f4ac4bba6445c2b4aa6c1d2aaf89be5ddbadb6f6b4aee28ac8f24971 + 128 3be5d4c12621e3cea7f83ceada3c66b30000000000000000edc4fa844da9dc98 + 160 d37e965176f6558a31560e40502ab3ae5491b21aaaabfb07fa50000000000000 + 192 0000000000000000000000000000000000000000000000000000000000000000 + 224 0000000000000000000000000000000000000000000000000000000000000001 + 256 000000000000000000d83504a101000000000000000000000000000000000000 + 288 1802200000000000000000000000000000000000000000000000000000000000 + 320 0000000000000000 +instance_sha256 7e24f382cef23fa1e70a401c9a1501aa8ffaf0f53876079ad9801e3eb7ecf446 +chi = H_256('I' || instance) 12b3c35a844467d4d290bc396c4a369d531fce8d33480d1f1148f488fb681943 +chi at trial_id 1 3a02f2f609d6d0216cf0a0e9294d6c96e9902ae04ce2420883eba287991a708b +l_1 = H_256('L' || chi || u64 1) 732d54bfc39303ce001cf5dc29e1434eba4ff01a01ce36cccdb0366be8bc5ae3 + +== 3. Lemma 3 precondition and the per-pebble charge (indeg 2) +lambda log2_q precondition lambda/4 >= 2 log2 q + 1 charge bits lambda - 2 log2 q - 1 charge / lambda max log2 q under the precondition +256 32 FAILS 191 0.746 31.5 +256 48 FAILS 159 0.621 31.5 +256 64 FAILS 127 0.496 31.5 +256 80 FAILS 95 0.371 31.5 +256 96 FAILS 63 0.246 31.5 +512 32 holds 447 0.873 63.5 +512 48 holds 415 0.811 63.5 +512 64 FAILS 383 0.748 63.5 +512 80 FAILS 351 0.686 63.5 +512 96 FAILS 319 0.623 63.5 +1024 32 holds 959 0.937 127.5 +1024 48 holds 927 0.905 127.5 +1024 64 holds 895 0.874 127.5 +1024 80 holds 863 0.843 127.5 +1024 96 holds 831 0.812 127.5 + +== 4. Corollary 2: k = (2 N ln 2 / e) lambda with e = c3 N / log2 N, so k = 2 ln 2 lambda log2 N / c3 + certificate bytes ~ k * (1 + indeg) * (lambda/8) * (log2 N + 1) (one label and its Merkle path per opened label; + no sharing of paths; the root and the encoding overhead omitted). c3 is DRSample's depth-robustness constant: + the paper (Fact 8) states it exists and gives no value; the rows are a sensitivity table, not a parameter. +lambda log2_N c3 k certificate_MiB +256 20 1.0 7098 13.6 +256 20 0.5 14196 27.3 +256 20 0.1 70979 136.5 +256 24 1.0 8518 19.5 +256 24 0.5 17035 39.0 +256 24 0.1 85174 194.9 + + the same luck term for a lottery target eps instead of 2^-lambda: k = ln(q / eps) / beta, beta = c3 / (2 log2 N) +log2_q log2_eps log2_N c3 k certificate_MiB (lambda 256) +64 -40 20 1.0 2884 5.5 +64 -40 20 0.1 28835 55.4 +64 -40 24 1.0 3461 7.9 +64 -40 24 0.1 34602 79.2 +80 -40 20 1.0 3328 6.4 +80 -40 20 0.1 33272 64.0 +80 -40 24 1.0 3993 9.1 +80 -40 24 0.1 39926 91.4 + +== 5. O-07 counter-example: one labelling, many lottery draws, when the lottery value is W = H('W' || chi || tau) + honest calls per trial: N labels + (N - 1) Merkle nodes + 1 lottery hash + re-roll per draw: set the sink label l'_N to fresh bytes (sink red), rehash its Merkle path (log2 N calls), + one lottery hash; the certificate passes iff no challenge lands on the sink: (1 - 1/N)^k +log2_N k honest_calls_per_trial reroll_calls_per_draw amortisation pass_probability +20 2000 2097152 21 9.99e+04 0.998094 +20 9000 2097152 21 9.99e+04 0.991454 +24 2000 33554432 25 1.34e+06 0.999881 +24 9000 33554432 25 1.34e+06 0.999464 diff --git a/docs/analysis/mhpow/b2/registry-batch-mhpow-b2-20261009-01.json b/docs/analysis/mhpow/b2/registry-batch-mhpow-b2-20261009-01.json new file mode 100644 index 000000000..1690e6867 --- /dev/null +++ b/docs/analysis/mhpow/b2/registry-batch-mhpow-b2-20261009-01.json @@ -0,0 +1,37 @@ +{ + "run_id": "mhpow-b2-20261009-01", + "manifest_sha": "13c3646e", + "evidence_dir": "docs/analysis/mhpow/b2", + "method": "model", + "note": "B2, the binding spec lane of the 1.5x programme's Track B (9 October 2026): the canonical instance input for an MTP/DRSample memory-hard proof of work in Igneum's setting (328 bytes, work version 1, draft) and the obligations table O-01 to O-16 read against Blocki and Smearsoll, eprint 2025/1456 (the full PDF, sha256 13c3646e...). Reuse across template, trial, epoch, day, era, network and class object: answered by the paper's per-input soundness plus the layout. The lottery: two candidate rules FAIL with written counter-examples (re-roll through a red sink, about 1.3 x 10^6 draws per labelling at N = 2^24; free screening on chi alone); the replacement and the multi-instance, adaptive-input, constants and energy questions BLOCKED to the B4 theory lane and the node lane. Figures by tools/mhpow/b2/b2_figures.py on build-9. NOT RUN until the panel reads it; this lane writes no PASS.", + "cells": [ + { + "cell": "model:mhpow-b2-binding", + "cases": [ + "POW-05", + "POW-06" + ], + "status": "NOT RUN", + "method": "model", + "evidence": "docs/analysis/mhpow/b2/README.md", + "note": "POW-05: obligations O-06 to O-11 and O-16 (grinding, chosen instance, amortisation, partial evaluation, precomputation); O-07 records two FAIL rules and B-T3 the open question. POW-06: O-11's certificate sizes as functions of DRSample's unproved constant c3 (13.6 to 194.9 MiB at lambda 256), a sensitivity table only; the verifier budget is B6's.", + "claim_impact": "none: no public figure moves; Track B's construction is not adopted and nothing activates", + "in_progress": true + } + ], + "map_cell_requested": { + "model:mhpow-b2-binding": { + "command": "python3 tools/mhpow/b2/b2_figures.py > docs/analysis/mhpow/b2/figures.txt (a build box under a pid file; byte-identical output, sha256 27bb414c...)", + "box_class": "build box, CPU only (build-9)", + "fixtures": [], + "cases": [ + "POW-05", + "POW-06" + ], + "coverage": { + "POW-05": "partial: the binding obligations for the Track B work unit only; the lottery binding is BLOCKED (B-T3) and the multi-instance bound BLOCKED (B-T5)", + "POW-06": "partial: certificate size as a formula; no verifier timing, no malformed-input cost" + } + } + } +} diff --git a/docs/analysis/mhpow/b6/README.md b/docs/analysis/mhpow/b6/README.md new file mode 100644 index 000000000..aff2b9035 --- /dev/null +++ b/docs/analysis/mhpow/b6/README.md @@ -0,0 +1,88 @@ +# B6 (R15-10): the verification budget of a proof-carrying memory-work unit against the node's limits + +Track B of the 1.5x Research Programme (the founder's second document, `docs/plans/igneum-2.0-master/1p5x/igneum-1p5x-research-plan.md`, section 6 B6; the package R15-10, section 10). Written by the node lane, 9 October 2026, 10:1x UK. Documents only: nothing here activates, nothing runs on a box of the fleet, no number below is a comparison by raw hashes per second (the Track B rule). Every row's verdict is NOT RUN until the panel reads it; every value the B1 lane's reference (R15-05, the Blocki and Smearsoll TCC 2025 MTP/DRSample shape, lane a4490b2dfe9114a55) has not yet produced is a BLOCKED row named by its symbol, with that lane as its owner. + +The question B6 asks (the plan, section 6): what a node, a pool and the network pay to verify, refuse and carry a work unit whose proof travels with the block, against the limits the node fork carries today (`vendor/igneum-node`, release-2.0.3-node-k7 27f54124; the shipped 2.0.2 line 5d53a591 carries the same limits unless a line below says otherwise). The plan's own warning stands: proof arrival rate times per-proof CPU time is a budget check, not a denial-of-service proof; the adversarial rows are bounds the node's rules impose, not a claim that nothing worse exists. + +## 1. The symbols + +| Symbol | Meaning | First value | Owner | +|---|---|---|---| +| `q` | challenges per proof (the paper's sampled positions) | BLOCKED: the B1 fixture's challenge count | B1 lane | +| `d` | Merkle depth of the commitment (log2 of the leaf count `N`) | BLOCKED: the B1 fixture's depth | B1 lane | +| `L` | leaf (label) size in bytes | BLOCKED: the B1 fixture's label width | B1 lane | +| `h` | inner node size: one hash, 32 bytes (the paper's hash; the node's `Hash` is 32 bytes, `crypto/hashes/src/lib.rs`) | 32 | node lane | +| `Δ` | openings per challenge beyond the challenged leaf itself (the DRSample parents the verifier re-derives a label from; MTP's δ) | BLOCKED: the B1 fixture's parent count | B1 lane | +| `B_p` | proof bytes per block (section 2) | formula below; BLOCKED as a number | node lane (formula), B1 lane (values) | +| `H_hdr` | the block header's wire bytes today (section 2) | about 222 + 32 × (parents by level); at most about 222 + 32 × 16 × levels | node lane | +| `r_64` | single-core SHA-256 rate on 64-byte inputs (one Merkle node) on the hub-class box | 6.8 × 10^6 per second (build-7, section 3) | node lane, measured | +| `r_L` | single-core SHA-256 bytes per second on `L`-byte inputs | 1.55 × 10^9 bytes/s at 1 KiB, 1.02 × 10^9 at 256 B (build-7, section 3) | node lane, measured | +| `T_v` | cold CPU verification time of one proof (section 3) | formula below; BLOCKED as a number | node lane (formula), B1 lane (values) | +| `bps` | blocks per second: devnet-4 1, testnet 1, mainnet 10 (`consensus/core/src/config/params.rs` 2342 via 2208 `BlockrateParams::new::<1>()`, 1928, 1789 `new::<10>()`) | 1 / 10 | node lane | + +## 2. Proof bytes per block and the header + +`B_p = h + q × (L + Δ × L + d × h)` — the commitment root (carried in the header or beside it, the B1 lane's call), then per challenge the challenged label, its `Δ` parent labels and one `d`-deep authentication path of 32-byte nodes (sibling paths shared between challenges shrink this; the formula is the upper bound the paper's non-shared layout gives). With the B1 fixture's `q`, `d`, `L`, `Δ` the number follows; until then the row is BLOCKED. + +The header today (`consensus/core/src/header.rs` 158–174: version u16, parents by level, hash merkle root, accepted id merkle root, utxo commitment, timestamp u64, bits u32, nonce u64, daa score u64, blue work u192, blue score u64, pruning point, Igneum's vote key hash): about 222 bytes of fixed fields plus 32 bytes per parent reference, with at most `max_block_parents` direct parents (10 at 1 bps, `consensus/core/src/config/bps.rs` 60–66: `ghostdag_k / 2` clamped to 10..16) per level. A proof-carrying unit adds its root to the header (32 bytes) and `B_p` to the block or beside it; which is the B1 lane's design call and a BLOCKED row here. + +| Row | Formula | First value | Node limit (source) | Verdict | +|---|---|---|---|---| +| Proof bytes per block `B_p` | `h + q(L + ΔL + dh)` | BLOCKED (q, d, L, Δ) | one P2P message is at most 256 MiB (`protocol/p2p/src/core/connection_handler.rs` 47, `P2P_MAX_MESSAGE_SIZE`); a block body's mass is bounded by the network's mass limits, which a proof would have to be exempt from or counted in (the B1 lane's call) | NOT RUN | +| Header growth | +32 bytes (the root) | 32 | the header is hashed whole for PoW; the class v5 bound hash reads the header's prehash (the engine's `hash_bound(prehash, nonce)`, `pool/src/verify.rs` 54): a 32-byte field is within every reader's prehash layout only if the layout is re-frozen (a digest move) | NOT RUN | +| Proof bytes per second at the block rate | `B_p × bps` | BLOCKED | a node relays every block it accepts to up to 128 inbound and 8 outbound peers (`kaspad/src/args.rs` 124–125): the uplink carries `B_p × bps × peers`; the plan's propagation row | NOT RUN | + +## 3. Cold verification time on the hub-class box and on the minimum node + +`T_v = q × (1 + Δ) × d / r_64 + q × (1 + Δ) × L / r_L + T_label`, where `T_label` is the cost of re-deriving one label from its `Δ` parents (the paper's per-node function; BLOCKED until the B1 lane names it) times `q`. Measured today, so the formula has a rate to multiply: on build-7 (AMD EPYC 9454, the hub-class box main named for this row), single-core SHA-256 via `openssl speed -evp sha256 -seconds 2`, read at 10:07 UK: 435.4 MB/s on 64-byte inputs = 6.8 × 10^6 hashes per second (one Merkle node per hash), 1,024.4 MB/s at 256 bytes, 1,550.9 MB/s at 1 KiB, 1,805.8 MB/s at 8 KiB. One hash of a 64-byte node therefore costs about 147 ns on one core; a path of `d = 20` costs about 2.9 µs; a proof with `q = 100` challenges and `Δ = 3` would cost about 1.2 ms of path hashing plus its labels — a shape, not the number, which waits on the fixture. The B1 lane's verifier, timed cold (no cache, a fresh process) on build-7, replaces the product the moment its fixtures exist. + +| Row | Formula | First value | Node limit (source) | Verdict | +|---|---|---|---|---| +| `T_v` on the hub-class box (build-7, one core) | above, with `r_64` = 6.8 M/s | BLOCKED (q, d, L, Δ, T_label) | the block time: 1,000 ms at 1 bps, 100 ms at 10 bps (`consensus/core/src/config/bps.rs` 52–56, `target_time_per_block = 1000 / BPS`); a node must verify at the admitted rate with headroom, so `T_v × bps` must sit well under one core-second per second on the smallest node | NOT RUN | +| `T_v` on the minimum node | the same formula with that node's measured `r_64` | BLOCKED: no minimum node is defined in any file the node lane holds (TV-04 carries the same gap for the hub count); the Mac entry serves 2.0.2 to 16 GB and 32 GB machines | the fd limit and the memory bound are the node's two hard floors today: `DESIRED_DAEMON_SOFT_FD_LIMIT` 1,048,576 (`kaspad/src/daemon.rs` 57, 2.0.3) and the executor's memory bound (the 2.0.3 rules' ring and record thinning; a fresh 2.0.2 sync read 62 to 86 GB on build-1 at 09:5x UK, main's separate read by 12:00) | NOT RUN | +| Verification at the maximum admitted rate (the plan's "cold verification, not only cached success") | `T_v × (bps + r_orphan)` where `r_orphan` is the orphan and side-block rate a node also verifies | BLOCKED | every block a node validates is PoW-checked before its body (`consensus/src/pipeline/header_processor`); the class v5 engine already costs a dataset read per header, so the proof's `T_v` adds to, never replaces, that read | NOT RUN | + +## 4. Invalid-proof amplification against the node's ban budget + +What an invalid proof costs a node is `T_v` (the verifier runs to the failing opening; a verifier that checks openings in order and stops at the first failure pays less on a random corruption and `T_v` on a last-opening corruption, the adversary's choice). What the node's rules then do: + +| Row | Mechanism | Value today (source) | Bound on the cost per peer | Verdict | +|---|---|---|---|---| +| The PoW strike guard | a header whose PoW fails is a strike; `DEFAULT_STRIKE_LIMIT` strikes inside `STRIKE_WINDOW` disconnect the peer and ban its IP for `BAN_DURATION` | 2 strikes per hour, window 3,600 s, ban 3,600 s (`protocol/flows/src/flowcontext/pow_guard.rs` 25, 31, 34; `IGNEUM_POW_STRIKES` overrides the limit) | an invalid proof that fails as PoW costs a peer at most 2 × `T_v` per hour per IP before a one-hour ban; a flood from `n` IPs costs `2n × T_v` per hour | NOT RUN | +| Ledger N6, the refused-block memory | a block refused by a rule is remembered for `RELAY_REFUSAL_WINDOW`; a peer that advertises it again is banned for `RELAY_REFUSAL_BAN` | window 3,600 s, ban 600 s (`protocol/flows/src/flow_context.rs` 106, 109) | the same invalid block costs `T_v` once per node per hour; a flood of distinct invalid blocks is bounded by the strike guard above, not by N6 | NOT RUN | +| The class v5 state wait (not a strike) | a block whose epoch state this node has not reached is held off without a strike (`protocol/flows/src/v10/blockrelay/flow.rs` 328–337; `ibd/flow.rs` 1282 `is_v5_state_wait`) | — | a proof whose verification needs the epoch's state (an input-specific unit, B2) inherits this: a node behind the cut cannot verify it and must hold it, which is the body-sync class the night's kits 2 to 7 worked through (fixes 2 to 7 on foreign-seed-capture-2 0156e250) | NOT RUN | +| The fd class (closed this morning) | accepted connections that never close held sockets in ESTAB until "Too many open files" at 524,288 on lp-4090-04 (05:49 UK); fixed by the transport close on router teardown and the server's HTTP/2 and TCP keepalive (76cc5a88 on release-2.0.3-node-k6/k7) | keepalive 10 s, unanswered 20 s (`protocol/p2p/src/core/connection_handler.rs` `server_keepalive`); 128 inbound (`kaspad/src/args.rs` 125) | a flood's connections cost sockets, not proofs: 128 inbound at once, each closed within 30 s of silence; a proof flood rides the inbound limit times the strike budget | NOT RUN | +| Amplification | `A = (bytes a peer sends) / (CPU-seconds the node spends)` per hour per IP | `A ≤ 2 × T_v` CPU-seconds per `2 × B_p` bytes per hour per IP under the strike guard | the plan's caution: batching and worst-case payloads need their own analysis; the rows above are the rules' bounds, not a proof | NOT RUN | + +## 5. The pool's share checks + +Today a share is one hash: the pool evaluates the member's lane on the CPU, `check(epoch, job, nonce, claimed)` (`pool/src/verify.rs` 52–63), one `epoch.hash_bound(prehash, nonce)` (the class v5 engine's bound hash, a dataset read) compared with the claimed hash and the share target (`share_target64`, `target64`: `pool/src/verify.rs` 34–38, `pool/src/protocol.rs` 85–101), the verdict's own `cost_ms` recorded, under a verification semaphore (`pool/src/pool.rs` 110 `verify_permits`); the refusal codes are `stale`, `duplicate`, `above_target`, `wrong_hash`, `unknown_job` (`verify.rs` 3). Shares are weighted by vardiff (`pool/src/server.rs` 7) and paid by PPLNS (`pool/src/pplns.rs` 76 `distribute`). + +| Row | Today | With a proof-carrying unit | Node limit (source) | Verdict | +|---|---|---|---|---| +| Work per share | one bound hash (`verify.rs` 54), `cost_ms` per verdict | a share must carry the unit's proof and the pool verifies `T_v` per share, or the share is a hash of a committed unit the pool verifies once per unit (the B1/B2 design call) | the pool's semaphore and its box's cores: `shares/s × T_v` core-seconds per second | NOT RUN | +| Bytes per share | nonce and hash (`protocol.rs` `ShareResult`) | `+ B_p` per share, or per unit | the pool protocol's line length (`protocol.rs` 178 `line`) | NOT RUN | +| A member's invalid share | `wrong_hash` / `above_target`, counted (`server.rs` 275, 293), no ban today | an invalid proof costs the pool `T_v`; the pool needs its own strike rule (the N6 shape) | none today in `pool/src` | NOT RUN | + +## 6. Propagation, the epoch wall and the body sync + +| Row | Formula | Value today (source) | Verdict | +|---|---|---|---| +| Bytes per block on the wire | `H_hdr + body + B_p` | `H_hdr` about 222 + 32 × parents; `B_p` BLOCKED | NOT RUN | +| Against the message limit | `H_hdr + body + B_p ≤ 256 MiB` | `P2P_MAX_MESSAGE_SIZE` 256 MiB (`connection_handler.rs` 47); an IBD batch carries `IBD_BATCH_SIZE` = 99 blocks per request (`protocol/flows/src/ibd/streams.rs` 25), so a batch carries `99 × B_p` of proof bytes | NOT RUN | +| Against the block time | the proof must reach every peer inside the block time with headroom: `(H_hdr + body + B_p) × peers / uplink < 1/bps` | 1,000 ms at 1 bps, 100 ms at 10 bps (`bps.rs` 52–56); up to 136 peers per node (`args.rs` 124–125) | NOT RUN | +| The class v5 epoch wall | every `POW_EPOCH_BLOCKS` = 3,600 blocks the engine's seed moves to the state after the cut block less `POW_EPOCH_LEAD` = 600 (`consensus/core/src/igneum.rs` 227, 233); the IBD catch-up waits up to `V5_STATE_WAIT` = 300 s per cut for the executor (`protocol/flows/src/ibd/flow.rs` 1278) | an input-specific proof (B2) keyed to the epoch's state is verifiable only by a node at or past that cut; the sync shape of the night (the per-cut body sync e189afb4, the heaviest anchor 94e4c6ba, the foreign-seed capture fixes 2 to 7) is the cost a joiner pays before it can verify a proof at all | NOT RUN | +| The body sync | a proof that travels in the body is fetched with the body (`RequestBlockBodies`, 99 per batch) after the header validated; a proof that travels in the header is fetched with the header | the headers-first IBD validates headers before bodies (`ibd/flow.rs` `sync_headers`), so a header-carried proof is verified at header time — `T_v` per header in IBD, at the IBD's rate, not the block rate | NOT RUN | +| Cancellation and state | a block refused after its proof verified has cost `T_v` for nothing; a proof kept for a block never accepted is memory | the proof record store (c90212e3) holds proof records under the block store with the lock from carried certificates; the executor's memory bound (the 2.0.3 rules) thins records — a proof-carrying unit's records are a new class in that store, BLOCKED until the B1 lane names the record's size | NOT RUN | + +## 7. What is BLOCKED and who unblocks it + +| Symbol or row | Owner | Clock | +|---|---|---| +| `q`, `d`, `L`, `Δ`, `T_label` (the fixture's parameters and the per-label function) | B1 lane (a4490b2dfe9114a55) | asked 10:08 UK; by 12:00, else these rows stay BLOCKED in the 13:00 landing | +| `T_v` measured cold on build-7 with the B1 verifier | node lane, the moment the B1 fixtures and verifier exist on a box | after the B1 lane's clock | +| the minimum node and its `r_64` | main (no file defines it; TV-04 carries the same gap for the hub count) | main's word | +| the proof's place (header, body, beside) and the record's size | B1/B2 lanes | their design call | +| the equivalent lottery-contribution and security-budget comparison the Track B rule requires before any number is compared | the coordinator (written and approved first) | before any row here reads anything but NOT RUN | + +Nothing in this document changes a rule, a parameter or a line of the node fork; the limits cited are the lines as they stand on 27f54124 and the parent's master, read 9 October 2026. diff --git a/docs/ops/build-queue.md b/docs/ops/build-queue.md index 623991804..65c4f00b5 100644 --- a/docs/ops/build-queue.md +++ b/docs/ops/build-queue.md @@ -13,6 +13,7 @@ build-1:/srv/queue/build-queue.md (the reader reads that one; this file is its s - An entry carries its owner lane, its box, its pid-file path and its clock (UK). A box with no live pid file from its queue and a load under 1.0 for a 30-minute read is a fault, reported by the reader to the coordinator (/srv/queue/faults.log). - A lane that finishes an entry replaces it with the next or hands the box back here with a line; an empty box is the fault. +- A box over 90 percent memory is a fault line from the reader to the coordinator (9 October 2026, 10:0x UK), with the box-1 rule above. - Main's rule (9 October 2026, 09:5x UK): an owner starts its matrix job on its box under /srv/queue/pids within 15 minutes of a read that found the box idle, or the reader marks the box unclaimed (open to any lane, written in status.json and faults.log); a box idle for 30 minutes is a red against its owner on the steward's board. A shared box names every owner and none ends another's pid. @@ -21,7 +22,7 @@ build-1:/srv/queue/build-queue.md (the reader reads that one; this file is its s | box | threads | owner lane | standing use | |---|---|---|---| -| build-1 | 96 | build-server lane | cuts and kits, the hands (observer-node, node1), the capacity fuzz slices, the workers page, the queue reader | +| build-1 | 96 | build-server lane | cuts and kits, the hands (observer-node, node1), the devnet-4 seed and hand, the light-reader, the capacity fuzz slices, the workers page, the queue reader. MEMORY RULE (the OOM of 09:54 to 10:04 UK, 9 Oct 2026: a fresh sync with a public listener reached 85 GB and the kernel killed nodes, Caddy, cron and rsyslogd): one fresh sync at a time, with its listener on loopback until within ten; one read node at most beside the hands, a second is a refusal; the queue reader faults any box over 90 percent memory | | build-2 | 96 | site lane and the adversary lane (shared; neither ends the other's pid) | the site gate (Playwright), the scene-parity suites; X5's dram design beside them | | build-3 | 32 | node lane | the long consensus fuzz and property suites (kaspa-consensus, kaspa-consensus-core) | | build-4 | 96 | adversary lane | the chip model (OpenROAD, kepler-formal): the 20 to 25 percent floorplan for the converged SPEF row | diff --git a/docs/plans/counter-asic-3-status.md b/docs/plans/counter-asic-3-status.md index 1c6e159e9..f0e70d8b3 100644 --- a/docs/plans/counter-asic-3-status.md +++ b/docs/plans/counter-asic-3-status.md @@ -879,6 +879,13 @@ X1's CENSUS, TRACES AND HIT CURVES (the research lane, 09:5x to 10:0x UK): (a) t X0's DECIDER (the hash lane, 09:59 UK, one session on the rented 5090 fk-x0-5090c, Vast 54992886, driver 580.105.08, the gen-6 Linux worker 804a6f7f, 250 batches of 2^24, nvidia-smi 1 Hz, stock only, the provider refusing -lgc): class v4 v4-devnet-epoch0 (0xa785001687d8688a, the 7 October numerator's pack) 141.963 MH/s / 500.0 W / 3,522 nJ per hash / e370fb2080b7dbb1 (the card at its 500 W limit); the frozen class v6 hl-v6-all (0x9d40978601a7df2a, the house knee rows' pack) 70.900 / 445.1 / 6,278 / 59e6708e46f1e87c; the signing object hl-v6-all-cs (0x2a1d6caab4c24564, generator 6) 70.650 / 457.1 / 6,470 / 01f51b9d4805e5e6 (its first CUDA fingerprint, self-test PASS); the ratio class v6 over class v4 in joules per hash 1.78x (the kit pack) and 1.84x (the signing object), the rate halving (2.00x); the house rig's same pair 1.68x at stock and 1.73x at the 1,300 MHz knee, so across three sessions 1.7 to 1.8x: class v6 costs the 5090 about 1.75x class v4's joules per hash and half its rate, the 7 October numerator (2.33 microjoules at the lock) understating the frozen object's card cost by about 1.7x; the evidence build-1:/srv/artefacts/tas/x0-pins/x0-run-fk-5090c/ (run-x0.log 9bc40685, smi.csv 7ab9f16f, SHA256SUMS.txt); the row document and registry batch 5 to land. X0 AMENDED at 25fe02109 (the adversary lane, 10:03 UK): what the decider does to r0 turns on one fact asked of the hash lane by 10:30: whether the class v6 program carries about twice the instructions per hash (then the chip's shadow scales with the card's and r0 holds: the board 1.5x, the hybrid 1.93x, the die 2.44x same node) or the same instructions at a structural card cost (then every ratio rises 1.7x: the board 2.6x, the hybrid 3.3x, the die 4.2x); X0 carries the band with the first reading as the default (the connected-state lane measured the 64-register window at under 1 percent of the card's energy, so the structural reading has no measured support yet); X5's six designs on their boxes (build-8's launcher relaunched at 10:01, the container's missing work-root variable). X3's FIRST PAIRED ROW (the Ember lane; the rented 4090 at stock, board power, the frozen control via the kit zip 4 bb66a546 pack hl-v6-all, worker f5f3846c, the anchor fingerprint e8f4f3289c6ee1fc equal to the CPU on the 4090 and the 5090; on an island, not a network): the control (base, batch 2^22) 31.252 MH/s at 250.0 W = 7,999 nJ per hash; tuned lb4-w4 at batch 2^26 31.302 MH/s at 251.6 W = 8,038 nJ (rate plus 0.16 percent, energy plus 0.49 percent per hash: no gain); 0 rejected shares over 12 shared 2^28-nonce windows (204 finds equal); the 4090 screen of all 17 race variants x batch 2^20 to 2^26 flat within 0.1 percent (the kernel memory-latency bound at 2,775 MHz SM, 250 of 450 W), so the in-container knobs carry no g; the only Ember lever with energy in it the clock/power ladder (the recipe's 5090 lock row: minus 31 percent nJ at minus 0.3 percent rate), which rented containers refuse: those rows NOT RUN; the rows to 11:00, the best setting's 30-minute stability run to 11:30. X4's FIRST PAIRED GPU ROW (the k lane, 10:0x UK; fk-x4-4090 at stock, worker 804a6f7f, 250 x 2^24, power.draw at 1 Hz from 8 s in): all four texts passing the pack self-test (96 of 96 vector lanes) with fingerprint e8f4f3289c6ee1fc, the GPU equivalence holding; pass 1 against the frozen hl-v6-all (35.080 MH/s, 285.4 W, 8,134 nJ per hash, 93 regs): direct 35.021 (minus 0.17 percent), 7,839 nJ (minus 3.6 percent), 182 regs; generated 35.053 (minus 0.08), 7,824 nJ (minus 3.8), 142 regs; prefix 35.056 (minus 0.07), 7,905 nJ (minus 2.8), 154 regs; pass 1 alone possibly the card warming up, the reversed pass 2 (D C B A, the frozen pack last) ending about 10:08 deciding it, g the mean of both passes; the knee NOT RUN (the container refusing -lgc). THE STALE-SINK CLASS CLOSED ON THE MOVING CHAIN AND BUILD-1's OOM (the node lane, 10:05 UK): kit 6's stale-sink kept datadir under 27f54124 reached the moving tip at 09:54:37 UK (26,324 / DAA 50,430 against lp-04's 26,312 / 50,413 read in the same second), one capture (lp-04's chain's seed, no refusal, no strike, no lost proof), five sink probes in all: PASS against a moving reference, the standing-peer probe closing the stale-sink class, no further flows change; the fresh 27f54124 beside it reached DAA 12,657 by 09:52 and then build-1 ran out of memory (the hands' two fresh kit-2 syncs from 09:49, the light-reader and the node lane's two nodes took the box to 120 GB used / 5 GB free at 09:53; both the node lane's nodes and the reader gone by 09:55 by the OOM killer's shape, dmesg root's), the fresh read VOID on memory, the fleet lane's kit canary on build-8 the fresh read of 27f54124 on the moving chain; build-1 at 101/23 GB after; the node lane's rule: one read node of its own on build-1 at a time and none while the hands sync; the build-server lane asked for the dmesg read, the light-reader's state (restarted through its unit only when the seed and the hand are within ten), the memory rule in the queue file. The node lane's queue entries up since 10:03:42 (build-7-node-flows.pid: the 2.0.3 flows suites on 27f54124 continuous; build-3-consensus-fuzz.pid; build-9-pow-fuzz.pid; driver loops on the Mac, ssh orchestration only, every build and run on the box); the HEAL lane's build-6-exec-p2p-fuzz.pid up at 10:05 (the exec, p2p-flows and p2p-lib suites each pass from the foreign-seed-capture-2 worktree 0156e250, build-6 shared with X5's hyb75); the D02 lane's build-3 flagged idle at 09:53 (its model under a second, its next run the TV-01 reconcile at 13:00; the box free beside its short runs); the steward recorded the D02 batch at 000ec2b2 (TV-01, TV-03, TV-04 NOT RUN in progress), the hashed identity list fix 780df1d7 (a path entry's login only), TV-04 on the map 7ad28f75, the same-work recorder class f5a560b5 (one record per case, cell and run; an older run kept as evidence). THE WINDOWS CANARY BOX (10:0x UK): no provider rents a Windows VM by API from the Mac (Hetzner Cloud: Linux images only, the Server 2022 install ISO 8637 needing a console hand; no Azure or AWS CLI or credential); the relay lane's route, a console install being a manual input under the founder's rule: Windows Server 2022 (4 vCPU, 8 GB, 80 GB, CPU-only) built unattended under KVM in a container on build-7 (/dev/kvm, Microsoft's public evaluation ISO, an autounattend with OpenSSH and auto-logon, the virtio drivers, ssh through a forwarded port on the box; the relay id win-canary with its secret minted and the bundle baked with the watchdog agent, the workers intake key issued, the bootstrap and the relay-task canary written; the pid line /srv/queue/pids/build-7-winvm.pid 2504745 since 08:58:41Z); the clocks: the VM answering ssh by 12:30 UK, the agent registered and the first workers report by 13:15, the first rule-33 record (2.0.2, the shipper's three changes for a CPU box) by 15:00; the shipper's cpx31-with-ISO order withdrawn (its correction once), the fleet lane released from fk-win-canary; the founder's folder: #1836 and the PC 2 note withdrawn, #1860 "PC 1 and PC 2: nothing to do". THE MINI's HOLD (the shipper, 09:57 UK): nothing armed to restart the mini (its launchd jobs: the app's own, the supervisor which never reboots, mini-tip, mini-collector, the relay agent; the snapshot script's go file removed); every future canary's reboot leg NOT RUN "founder's hold" unless he says otherwise; the mini at 09:57:14: app 2.0.2, the node synced at DAA 50,580 with 7 peers (the feed 50,573), the card between work, uptime 12 minutes. THE RESET AT 09:56 UK (the fleet lane): wave two done at 09:54:57, 40 of 40, no errors; lp-4090-43's stage running; of 92 prepared boxes at 09:55, 8 synced=true, 11 past DAA 46,800, the median DAA 17,991 (a fresh 4090 syncing in about 15 minutes), the island count high (28 tips among 35 mining boxes) until the syncs land; the hubs lp-04 DAA 50,495 (84 peers), lp-05 50,507, lp-06 50,520 still its own fork (the 09:50 reset did nothing: the forwarded scp of the 5 GB snapshot lost its connection and the sha check refused to touch the box; retried); the hub-1 archive stream restarted at 09:55 (13,039 files, 12.7 GB, hub-1 free 6.9 GB). THE 1.5x BRIEF LANDED (the landing hand, 09:56 UK): box master 37ee2667140411cd (the brief under docs/plans/igneum-2.0-master/1p5x/ with its README carrying the owners and clocks and the permanent rule, the ten BST and the one rig-name replacements with both sha256s, the master README line, the token-value README's ratification line citing d3686955); thirty-seven landings through it; X0's document gating. +THE CLOCK (the coordinator, 10:15:56 UK): read on the Mac: date = Fri Oct 9 10:15:56 BST 2026, date -u = 09:15:56 UTC, TZ=Europe/London date = 10:15:56 BST; the lanes' "UK" stamps match the Mac's date; main's clock reads one hour ahead of it; nothing converted; every clock in the record the Mac's Europe/London date. +RESET END (the recipe's line 10:11:33 UK; the tracker's first read 10:11:46): the reset pass complete (the hubs; wave one 53 of 54; wave two 40 of 40; lp-06's snapshot reset at 10:02; lp-4090-43 last at 10:08); the convergence read at the end 33 distinct tips among 35 mining boxes, a mid-sync figure; the within-ten tracker the measure: 22 of 101 boxes within ten of the reference (the feed 26,966) at 10:11:46, among them lp-04, lp-05 and lp-06 (all three hubs on ONE chain, DAA 51,424 / 51,436 / 51,449, lp-04 72 peers), lp-4090-02/09/21/29/36/37/38/44, lp-48-01/03/05, lc-3090-02/04, lp-3090-05, ln-node-02/04/05/06/08; the other 79 mid-sync, each one's time written as it lands, the counts every five minutes; lp-4090-43 at DAA 10,999 with four peers. The seeds: build-1 could not run its three fresh syncs at once: the OOM killer took the seed at 09:54:58 UK at 62 GB (its unit restarting it), the hand at 10:04:11 at 85.6 GB anon RSS (a fresh 5d53a591 sync from genesis grows to that in fifteen minutes, a 2.0.2 class for the node lane's stale-sink item) and with it cron, atd, agetty, rsyslogd and Caddy (pid 1054819), the hand's unit crash-looping 26 times; the kill count since 09:54 by command sshd 23, node 17, systemd 16, bash 16, sleep 16, igneumd 8, the fuzz binary 2, cron 3, atd 3, rsyslogd 2, caddy 1; build.igneum.network, git.igneum.network (the Forgejo container exited with Caddy), the intake route, the faucet host and /light dark from 09:55 to 10:13 UK (found by the empty edge read-back at 10:12); the build-server lane's restore at 10:13:41 (Caddy pid 451390, Forgejo at 10:14, timesyncd, polkit, mdmonitor, resolved, unattended-upgrades back; systemd-networkd failed, the links up, its restart waiting for main's word); the three nodes now in sequence under /run/dn4-sequence.pid: the seed alone first (restarted at 10:10 on a fresh datadir with its listener on loopback, the three hubs its only dial-outs, after its first partial datadir took island blocks from 93 inbound peers and stalled: EVM 4,443 at 5.9 GB two minutes in, the box at 8 percent), within ten about 10:50 UK, then explorer.igneum.network and /live pointed at the seed (about 10:55), then the hand with the indexer and the reader (about 11:35), then the light-reader with igneum-light-service restarted (about 12:20); the memory read every minute with a FAULT line over 90 percent; the queue file's build-1 entry carrying the memory rule. THE HUB-1 ARCHIVE DONE at 10:11:25 UK: 13,039 of 13,039 files sha256 equal on build-1:/srv/archive/hub-1-devnet-3/ (the sha list beside them; the direct rsync hub-1 to build-1 after the Mac route broke at 815 MB), the aside removed on hub-1 under the script's pid, free 6,874 MB before, 19,630 MB after. The 27f54124 rule-33 canary armed on a wave-two box (a fresh datadir, SEED the three hubs, the eight lines, a shard proved), verdict by 11:30 UK; the hive-entry canary on the tracker's end. The page live and clean at 10:14 UK (master e9efde24: the rig cards "agent not reporting since