81 lines
No EOL
28 KiB
HTML
81 lines
No EOL
28 KiB
HTML
<!doctype html>
|
|
<html lang="en">
|
|
<head>
|
|
<meta charset="utf-8">
|
|
<meta name="viewport" content="width=device-width, initial-scale=1, viewport-fit=cover">
|
|
<title>Engineering log</title>
|
|
<meta name="description" content="Every Igneum benchmark and test, with the commands that produced it.">
|
|
<meta name="theme-color" content="#0C0C0E">
|
|
<link rel="icon" href="data:image/svg+xml,%3Csvg xmlns='http://www.w3.org/2000/svg' viewBox='0 0 100 100'%3E%3Cpolygon points='50,4 74,34 67,58 80,54 61,96 39,96 20,54 33,58 26,34' fill='%23F2541B'/%3E%3C/svg%3E">
|
|
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Unbounded:wght@500;700;900&family=IBM+Plex+Sans:wght@400;500;600&family=IBM+Plex+Mono:wght@400;500&display=swap">
|
|
<style>
|
|
:root{--obsidian:#0C0C0E;--graphite:#16161A;--line:#2A2A30;--ember:#F2541B;--molten:#FFB35C;--bone:#F4F1EC;--ash:#9A9A9E;--ink-2:#C9C7C2}
|
|
*{box-sizing:border-box}body{margin:0;background:var(--obsidian);color:var(--bone);font-family:'IBM Plex Sans',system-ui,sans-serif;font-size:16px;line-height:1.6;padding-inline:clamp(16px,4vw,32px)}
|
|
a{color:var(--ember);text-decoration:none}a:hover{text-decoration:underline}
|
|
.wrap{max-width:1180px;margin:0 auto}
|
|
header{display:flex;flex-wrap:wrap;justify-content:space-between;align-items:center;gap:16px;padding-block:18px;border-bottom:1px solid var(--line)}
|
|
.brand{display:flex;align-items:center;gap:10px;color:var(--bone)}.brand b{font-family:'Unbounded',sans-serif;font-weight:900;letter-spacing:.06em;font-size:18px}
|
|
header nav{display:flex;flex-wrap:wrap;gap:18px;font-size:14px}header nav a{color:var(--ink-2)}
|
|
h1{font-family:'Unbounded',sans-serif;font-weight:900;font-size:clamp(28px,5vw,48px);line-height:1.05;margin:36px 0 10px}
|
|
.note{color:var(--ash);font-size:15px;max-width:72ch;margin-bottom:28px}
|
|
.layout{display:grid;grid-template-columns:minmax(0,1fr);gap:32px;padding-bottom:80px}
|
|
@media(min-width:960px){.layout{grid-template-columns:240px minmax(0,1fr);gap:56px}}
|
|
.toc{align-self:start;font-size:14px}@media(min-width:960px){.toc{position:sticky;top:24px}}
|
|
.toc ol{list-style:none;margin:0;padding:0;display:flex;flex-direction:column;gap:2px}.toc a{display:block;padding:6px 10px;border-radius:8px;color:var(--ink-2)}.toc a:hover{background:var(--graphite);text-decoration:none;color:var(--bone)}
|
|
article{max-width:78ch;min-width:0}
|
|
article h2{font-family:'Unbounded',sans-serif;font-weight:700;font-size:clamp(20px,2.6vw,26px);margin:44px 0 12px;padding-top:24px;border-top:1px solid var(--line)}
|
|
article h3{font-family:'Unbounded',sans-serif;font-weight:500;font-size:16px;margin:26px 0 8px}
|
|
article h4{font-size:15px;margin:20px 0 6px;color:var(--molten)}
|
|
article p{margin:0 0 14px}article ul,article ol{margin:0 0 14px;padding-left:22px}article li{margin-bottom:6px}
|
|
code{font-family:'IBM Plex Mono',monospace;font-size:.92em;background:var(--graphite);padding:1px 5px;border-radius:4px}
|
|
pre{margin:0 0 16px;padding:14px 16px;background:var(--graphite);border-radius:12px;font-family:'IBM Plex Mono',monospace;font-size:13px;line-height:1.55;overflow-x:auto}
|
|
.tbl{overflow-x:auto;margin:14px 0 20px;border:1px solid var(--line);border-radius:12px;background:var(--graphite)}
|
|
table{border-collapse:collapse;width:100%;min-width:560px;font-size:14px}th,td{padding:10px 12px;text-align:left;vertical-align:top;border-bottom:1px solid var(--line)}
|
|
th{font-family:'IBM Plex Mono',monospace;font-size:11px;letter-spacing:.12em;text-transform:uppercase;color:var(--ash);font-weight:500;background:var(--obsidian)}tr:last-child td{border-bottom:0}
|
|
footer{border-top:1px solid var(--line);padding-block:24px 48px;font-size:13px;color:var(--ash)}
|
|
</style>
|
|
</head>
|
|
<body>
|
|
<div class="wrap">
|
|
<header><a class="brand" href="/"><svg viewBox="0 0 100 100" width="26" height="26" aria-hidden="true"><polygon points="50,4 74,34 67,58 80,54 61,96 39,96 20,54 33,58 26,34" fill="#F2541B"></polygon><polygon points="50,42 59,58 50,82 41,58" fill="#0C0C0E"></polygon></svg><b>IGNEUM</b></a><nav><a href="/">Home</a><a href="/litepaper">Litepaper</a><a href="/bench">Engineering log</a><a href="/ledger">FUD ledger</a><a href="https://github.com/[second-owner-login]/igneum" rel="noopener">GitHub</a></nav></header>
|
|
<h1>Engineering log</h1>
|
|
<p class="note">Every measurement the project has made, newest at the bottom, written by the people and agents who ran it, with the commands and hardware. Prototype numbers are not mining numbers and say so.</p>
|
|
<div class="layout">
|
|
<nav class="toc" aria-label="Contents"><ol><li><a href="#2026-10-03-proto-metal-igneum-bench-first-run">2026-10-03 proto-metal / igneum-bench, first run</a></li><li><a href="#2026-10-03-proto-cuda-program-pack-export-mac-side-only-rtx-5090-run-pending">2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)</a></li><li><a href="#2026-10-03-sim-finality-sim-py-sustained-mining-finality-vote-weight-model-not-hardware">2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)</a></li><li><a href="#2026-10-03-proto-metal-hardening-tests-correctness-and-soundness-of-the-lottery-hash-metal-only">2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only)</a></li><li><a href="#3-october-2026-rtx-5090-first-run-[user]-s-pc-windows-cuda-12-8-runtime-driver-13-4-visual-studio-2026-with-the-14-30-toolset-selected-via-vcvarsall-vcvars-ver-14-30">3 October 2026, RTX 5090 first run (the project lead's PC, Windows, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30)</a></li><li><a href="#2026-10-03-sim-finality-v2-py-finality-rule-v2-with-latency-partitions-and-eclipses-model-not-hardware">2026-10-03 sim/finality_v2.py, finality rule V2 with latency, partitions and eclipses (model, not hardware)</a></li><li><a href="#2026-10-03-proto-metal-memory-hard-dataset-cache-8-dependent-reads-metal-only-cuda-pack-emulated">2026-10-03 proto-metal memory-hard dataset (cache + 8 dependent reads), Metal only; CUDA pack emulated</a></li><li><a href="#2026-10-03-rusty-kaspa-base-build-and-3-node-devnet-on-the-mac-consensus-engineer-pre-fork-proof">2026-10-03 rusty-kaspa base build and 3-node devnet on the Mac (consensus-engineer, pre-fork proof)</a></li><li><a href="#3-october-2026-rtx-5090-memory-hard-dataset-pack-igneum-genesis-mh">3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)</a></li></ol></nav>
|
|
<article><h1 id="igneum-bench-log">Igneum bench log</h1>
|
|
<p>Append-only. Every number here was measured on the machine named, on the date given.</p>
|
|
<h2 id="2026-10-03-proto-metal-igneum-bench-first-run">2026-10-03 proto-metal / igneum-bench, first run</h2>
|
|
<p>Machine: Apple M5 Max, 40 GPU cores, 64 GB unified memory, macOS Darwin 25.6.0, Swift 5.8.1, Metal 4. Build: <code>swiftc -O -o igneum-bench main.swift -framework Metal</code>. Source: <code>proto-metal/main.swift</code>. Setup: 1 GiB dataset (2^28 uint32), 64 instructions x 8 iterations, threadgroup 32 (threadExecutionWidth 32), 4 timed batches x 2^22 nonces after one warm-up batch. Verification: 3 warps per program, CPU interpreter vs GPU.</p>
|
|
<div class="tbl"><table><thead><tr><th>Seed</th><th>Loads/hash</th><th>Compile ms</th><th>Mhash/s</th><th>GB/s useful</th><th>CPU verify ms/warp</th><th>Verify</th></tr></thead><tbody><tr><td>igneum-genesis</td><td>104</td><td>49.6 (cold)</td><td>45.2</td><td>18.8</td><td>0.015</td><td>PASS</td></tr><tr><td>igneum-genesis/epoch1</td><td>104</td><td>20.3</td><td>48.4</td><td>20.1</td><td>0.019</td><td>PASS</td></tr><tr><td>igneum-second-seed</td><td>104</td><td>46.6 (cold)</td><td>35.5</td><td>14.8</td><td>0.016</td><td>PASS</td></tr><tr><td>igneum-second-seed/epoch1</td><td>144</td><td>18.7</td><td>35.4</td><td>20.4</td><td>0.017</td><td>PASS</td></tr><tr><td>igneum-hourly</td><td>128</td><td>52.0 (cold)</td><td>36.6</td><td>18.7</td><td>0.021</td><td>PASS</td></tr><tr><td>igneum-hourly/epoch1</td><td>128</td><td>21.6</td><td>37.5</td><td>19.2</td><td>0.017</td><td>PASS</td></tr><tr><td>igneum-hourly/epoch2</td><td>120</td><td>23.7</td><td>36.6</td><td>17.6</td><td>0.016</td><td>PASS</td></tr></tbody></table></div>
|
|
<p>Dataset size sweep (seed igneum-genesis): 4 MiB 569 Mhash/s, 64 MiB 183, 256 MiB 94, 512 MiB 69, 1 GiB 44. Dataset fill 1 GiB: 2.34 ms GPU time (427 GB/s) warm, 5.57 ms on first run of a process. Result: 21 warps, 672 hashes, zero mismatches. OVERALL PASS. Reading: memory bound at 1 GiB (12.8x drop from cache-resident), limited by random access rather than bandwidth, CPU verify roughly 250x under the 10 ms gate with a cheap dataset element. Apple silicon only. Details in <code>proto-metal/README.md</code>.</p>
|
|
<h2 id="2026-10-03-proto-cuda-program-pack-export-mac-side-only-rtx-5090-run-pending">2026-10-03 proto-cuda / program pack export (Mac side only; RTX 5090 run pending)</h2>
|
|
<p>Machine: the same Apple M5 Max. No CUDA toolchain exists on it, so nothing below is an NVIDIA measurement. Added <code>--export-pack <dir></code> to <code>proto-metal/main.swift</code>. It writes, per seed, the CUDA kernel (<code>kernel.cu</code>), C headers (<code>program.h</code>, <code>vectors.h</code>), JSON twins, and the Metal source, into <code>proto-cuda/packs/<seed>/</code>. Vectors: 3 warps (base nonces 0, 4096, 1000000), 96 x 64-bit outputs from the CPU interpreter, plus dataset words 0..15 and word [MASK]. The exporter runs the Metal kernel for the same warps and refuses to write unless all 96 match.</p>
|
|
<div class="tbl"><table><thead><tr><th>Pack</th><th>Loads/hash</th><th>Op mix</th><th>CPU interpreter vs Metal GPU (3 warps)</th><th>CUDA text in CPU emulation (clang, 32 threads/warp)</th></tr></thead><tbody><tr><td>igneum-genesis</td><td>104</td><td>load=13 xor=13 sub=7 shfl=6 add=5 mulhi=5 mad=4 rotr=4 mul=3 rotl=3 or=1</td><td>PASS 3/3</td><td>PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 and 2 warps/block</td></tr><tr><td>igneum-hourly</td><td>128</td><td>load=16 add=8 xor=7 mad=5 mul=5 mulhi=5 shfl=5 rotl=4 or=3 rotr=3 sub=3</td><td>PASS 3/3</td><td>PASS: dataset self-test, 3/3 warps standalone, warps 0 and 4096 in batch at 1 warp/block</td></tr></tbody></table></div>
|
|
<p>Sweep sizes 4, 64, 256, 512 MiB in the emulation: dataset self-test PASS at each (vectors only apply at 1 GiB). The regular Metal bench was re-run after the change: igneum-genesis 44.8 Mhash/s and igneum-hourly 36.3 Mhash/s at 1 GiB with 2 x 2^20 batches, PASS 3/3 warps each (consistent with the first-run table above, not a new figure). Harness: <code>proto-cuda/host.cu</code>, <code>build.sh</code>, <code>build.bat</code>, <code>README.md</code>, <code>CHECKLIST.md</code>, <code>emu/</code>. Pending: build and run on the project lead's RTX 5090 (CUDA 12.8 or newer, <code>-arch=sm_120</code>). No NVIDIA hash rate exists yet. Emulation rates are not recorded because they measure the Mac's CPU, not a GPU.</p>
|
|
<h2 id="2026-10-03-sim-finality-sim-py-sustained-mining-finality-vote-weight-model-not-hardware">2026-10-03 sim/finality_sim.py, sustained-mining finality vote weight (model, not hardware)</h2>
|
|
<p>Model: 1 day steps, 86,400 Poisson blocks/day, 1,000 Pareto honest keys (top key 17%), perfect retarget, every block blue, no latency, no VRF noise. Seed 7, seed 11 agrees. Rule: weight = 30-day sum of counted blocks, counted = min(actual, 2 x yesterday + f). Lock at 2/3 of total. Floors f in {1, 10, 100, 1000}. A: weight tracks hashrate at steady state (corr 1.00000); full weight from zero history on day 41 (f=1) to 32 (f=1000). B: 60% renter on one key takes 59.9% of rewards on day 1, crosses 1/3 of weight on day 20 to 26, 50% on day 27 to 34, never 2/3. 75% renter reaches 2/3 on day 29 to 34, day 27 with the cap removed. C: splitting defeats the cap. 10,000 fresh keys at f=1 move the 50% crossing from day 34 to 27 (no-cap figure 26); at f=10 they match no-cap exactly. The 30-day trickle buys 2 more days for 11.6% of the network. D: honest doubling is under-weighted 28 to 31 days; old miners lock alone for 20 to 23 days. E: 30% churn leaves 71% live, lock never lost. Threshold is 1/3: 35% stalls 1 day, 50% stalls 10 to 11 days. Active-24h total removes every stall. F: 51% patient owner holds 51.0% weight from day 45, vetoes from day 7 to 20, never locks alone. 67% owner locks alone from day 34 to 44. Recommend: f=1, cap 2x, window 30 (the window is the defence, the cap is worth 1 to 8 days), total = all keys with weight in the window. Details in sim/results.md.</p>
|
|
<h2 id="2026-10-03-proto-metal-hardening-tests-correctness-and-soundness-of-the-lottery-hash-metal-only">2026-10-03 proto-metal hardening tests (correctness and soundness of the lottery hash, Metal only)</h2>
|
|
<p>Machine: the same Apple M5 Max. Added <code>--fuzz</code>, <code>--edge</code>, <code>--stats</code>, <code>--determinism</code>, <code>--memcheck</code>, <code>--inline-dataset</code> to <code>proto-metal/main.swift</code>. Full tables and commands in <code>proto-metal/TESTS.md</code>. Fuzz: 10,200 random programs (200 + 10,000, cold compiles 13.6 to 48.5 ms), 4 random full-range warps each, dataset drawn from 64 MiB / 256 MiB / 1 GiB: 40,800 warps, 1,305,600 hashes, 0 mismatches, 0 compile failures, 0 static mask failures, generator contract (rotl 1..31, mask in {1,2,4,8,16}, src != dst) held on every instruction. 196 s for the 10,000 run. Edge: 14 hand-built cases (rotr by register 0 / 32 / -32 / 31 / 63, rotl 1 and 31, mulhi max operands, shfl masks 1..16, loads at index 0 and MASK via in-range and out-of-range registers, add/sub/mul/mad wraparound, zero loads, 64 loads) with operand values proven by a traced interpreter: 14/14 PASS, 128/128 lanes each. rotl by 0 (never generated) agreed too, recorded as informational only. Stats (3 seeds, 2^20 nonces each): bit frequency max deviation 2.90 sigma over 192 bit positions; avalanche 16,000 flips mean 31.99 to 32.04 (expect 32), std 3.98 to 4.01 (expect 4), every output bit flips with probability 0.490 to 0.508; chi-square on four 16-bit windows all within 2.3 sigma; 0 duplicates. Looks uniform. Not a security proof. Determinism: 5 runs and 3 compiles (one forced cold, 30 ms) of 2^20 hashes gave fingerprint 933787e8cfefccb7 every time; dataset fill deterministic (0b1a77899ee60493 twice) and 4,096 sampled words incl. 0 and MASK match the CPU closed form. Memcheck: every <code>dataset[</code> in the MSL is <code>dataset[rN & MASK]</code> (13/13 at 3 sizes), CUDA twin 13/13 plus one guarded fill write; 4 MiB run with nonces up to 0xffffffff completed and 4 wrapping warps matched the CPU; 416/416 load indices exceeded MASK before masking. Bench re-run after the changes: igneum-genesis 44.56 Mhash/s, epoch1 47.74 Mhash/s at 1 GiB, PASS 3/3 warps each (within 2 percent of the first-run table). <code>--export-pack igneum-genesis</code> re-run is byte-identical to the existing pack. SHORTCUT MEASURED: <code>--inline-dataset</code> replaces every load with the six-op closed form ds_elem and never reads memory: 4,888 Mhash/s wall (6,274 GPU time) vs 44.6 honest at 1 GiB, about 110x, and about 9x the cache-resident honest rate. With a closed-form dataset the hash is not memory-hard; an expensive dataset derivation is required, not optional. Not demonstrated: cryptographic strength, weak-program frequency and rejection, NVIDIA/AMD bit-exactness (CUDA run still pending), CPU verify gate with an expensive dataset element. Next three tests for the cryptographer are listed in TESTS.md section 8.</p>
|
|
<h2 id="3-october-2026-rtx-5090-first-run-[user]-s-pc-windows-cuda-12-8-runtime-driver-13-4-visual-studio-2026-with-the-14-30-toolset-selected-via-vcvarsall-vcvars-ver-14-30">3 October 2026, RTX 5090 first run (the project lead's PC, Windows, CUDA 12.8 runtime, driver 13.4, Visual Studio 2026 with the 14.30 toolset selected via vcvarsall -vcvars_ver=14.30)</h2>
|
|
<p>Pack igneum-genesis, dataset 1024 MiB, 5 batches x 2^24 hashes, 1 warp per block.</p>
|
|
<div class="tbl"><table><thead><tr><th>Card</th><th>Mhash/s at 1 GiB</th><th>GB/s useful</th><th>random loads/s</th><th>dataset fill</th><th>vectors</th></tr></thead><tbody><tr><td>NVIDIA RTX 5090 (170 SMs, 32 GB)</td><td>228.1</td><td>94.9</td><td>23.7 G</td><td>0.66 ms, 1638 GB/s</td><td>96/96 PASS, standalone and in batch</td></tr><tr><td>Apple M5 Max (40 GPU cores, same program, same day)</td><td>45.2</td><td>18.8</td><td>4.6 G</td><td>2.34 ms, 427 GB/s</td><td>96/96 PASS</td></tr></tbody></table></div>
|
|
<p>Result: the same hourly program, generated on the Mac, compiled by Apple's Metal and NVIDIA's CUDA, produced identical hashes on both vendors. Vendor independence of the lottery program is demonstrated for one program; igneum-hourly and the dataset sweep are the next runs. The ratio 5090 to M5 Max is about 5x on hashes and on random loads per second, approximate, consistent with a memory-bound program (random-access bound, not bandwidth bound: the 5090 writes the dataset at 1638 GB/s but hashes at 95 GB/s of useful 4-byte loads). Caveat unchanged: the prototype dataset is closed-form and not yet memory-hard (see TESTS.md), so these are prototype numbers, not mining numbers.</p>
|
|
<p>Build note for Windows: CUDA 12.8 crashes (cudafe++ access violation) under Visual Studio 2026's 14.51 toolset even with -allow-unsupported-compiler. Fix: install the MSVC v143 (14.30) component and open the environment with <code>"C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30</code>, then build normally.</p>
|
|
<h3 id="rtx-5090-dataset-sweep-and-second-program-same-session">RTX 5090, dataset sweep and second program (same session)</h3>
|
|
<div class="tbl"><table><thead><tr><th>dataset MiB</th><th>Mhash/s</th><th>GB/s useful</th><th>random loads/s (G)</th></tr></thead><tbody><tr><td>4</td><td>1339.8</td><td>557</td><td>139.3</td></tr><tr><td>64</td><td>1352.7</td><td>563</td><td>140.7</td></tr><tr><td>256</td><td>269.8</td><td>112</td><td>28.1</td></tr><tr><td>512</td><td>241.8</td><td>101</td><td>25.2</td></tr><tr><td>1024</td><td>228.7</td><td>95</td><td>23.8</td></tr></tbody></table></div>
|
|
<p>Second program igneum-hourly (128 loads per hash): 96/96 vectors PASS, 185.3 Mhash/s at 1 GiB, 23.7 G random loads/s.</p>
|
|
<p>Reading: the 5090 carries 96 MiB of L2. At 4 and 64 MiB the dataset sits inside it and the program runs about 5.8x faster than at 1 GiB. Past the L2 the rate settles at about 23.7 G random loads/s for both programs regardless of loads per hash (104 vs 128 loads gives 228 vs 185 Mhash/s, proportional), so the program is random-access bound once the dataset exceeds on-chip cache. Each 4-byte random load moves a 32-byte sector, so DRAM traffic is roughly 760 GB/s, approximate, against a quoted peak near 1.8 TB/s for this card. Design consequence: the dataset must stay well above any plausible on-chip cache, which the 2 GB genesis size and the growth schedule provide; a chip would need gigabytes of on-chip memory to escape the random-access limit. Still prototype numbers: dataset derivation remains closed-form until the 256 MB cache construction lands.</p>
|
|
<h2 id="2026-10-03-sim-finality-v2-py-finality-rule-v2-with-latency-partitions-and-eclipses-model-not-hardware">2026-10-03 sim/finality_v2.py, finality rule V2 with latency, partitions and eclipses (model, not hardware)</h2>
|
|
<p>Model: 30-s slots, Poisson(30) blocks/slot, 1,000 Pareto honest keys in 3 regions (45/35/20), 2-s inter-region delay (0.5 and 5 swept), uptime 97% (99.5% for pools over 1%), warm 30-day start for B to G. Seed 7, seed 11 agrees on B and E. Run time 5 min. Details in sim/results_v2.md. Rule: weight = flat 30-day blue blocks, dust 100, checkpoint per 30 blocks, lock at 2/3 of ACTIVE (participation over 240 checkpoints) vs TOTAL weight. A: weight = hashrate (corr 1.00000), full weight day 30, all keys over dust by day 20, lock median 2.5 s / p99 4.6 s at 2 s delay, 14 s max at 5 s, 0 stalls in 60 days except 17 at genesis. B: share(t) = (t/30) x a/(1+a) holds to 0.04 points; 1/3 crossed at day 20.0 / 15.0 / 12.5 / 11.1 and 2/3 at never / 30.0 / 25.0 / 22.2 for a = 1 / 2 / 4 / 9; dust hands a 9x renter 1.3 extra points. C: silent set that keeps mining: active recovers in 0 / 13 / 20 / 29 / 38 min at 34 / 40 / 45 / 50 / 55%; total never (silent weight never ages out). D: churn: active 2 min (35%) and 31 min (50%); total 41 h and 10.1 days. E: active FAILS the partition test: 50/50 honest split, no attacker, both sides lock after 60 min (30 with DAA retarget), 60/40 after 121 min; first-lock time = presence x (1 - 1.5 s)/s slots, confirmed. Total: 0 conflicts in every honest partition. 34% attacker breaks every variant at 50/50 (67% per side). F: delayed eclipse of a 20% pool is harmless (participation 0 after 2 h, back in 2 h, 0 conflicts); a 34% attacker poisoning that pool finalises a private fork in 49 min under active, never under total. Floor hybrid: active denominator never below 0.85 x total (lock needs 56.7% of total) gives 0 conflicts in every partition and eclipse, recovers in 0 / 13 min at 34 / 40% silent and 2 min at 35% churn; costs 4.1 days at 50% churn and liveness ends near 42% silent. Floor 0.80 does not stop the eclipse (54% > 53.3%). Recommend: active/cert + floor 0.85, presence 240, dust 100, quorum 2/3, grace at least 3x worst delay. Not modelled: real GHOSTDAG merge and post-heal fork choice, DAA lag, VRF aggregators, certificate revocation.</p>
|
|
<h2 id="2026-10-03-proto-metal-memory-hard-dataset-cache-8-dependent-reads-metal-only-cuda-pack-emulated">2026-10-03 proto-metal memory-hard dataset (cache + 8 dependent reads), Metal only; CUDA pack emulated</h2>
|
|
<p>Machine: the same Apple M5 Max (one performance core for the CPU figures). Construction, every table and the commands are in <code>proto-metal/MEMHARD.md</code>. Default dataset is now memory-hard; <code>--closed-form</code> keeps the original for comparison. Construction: 256 MiB cache = 2^22 lines of 64 B in 2^16 chains of 64 ChaCha12 blocks with feed-forward (in_j = prev ^ (sigma || K[8] || seg || j || tag)); item t = 16 words, 8 rounds of (seed-parameterised ARX-multiply mixer, read cache line s[0] & (2^22-1), xor) plus a final mixer; dataset[w] = item(w >> 4)[w & 15]. Hash kernel unchanged. Cache fill: 2.0 ms GPU (0.6 to 2.1 across runs), 185 ms one CPU core (Swift), 162 ms C++ host reference. Dataset build 1 GiB: 20.6 ms GPU (29.4 first in process), 814 M items/s, 6.5 G cache-line reads/s. GPU cache == CPU cache on all 2^26 words every run (FNV-1a 64 48c4f5bf24166b2e for day 2026-10-03). Shortcut ratio, seed igneum-genesis, 1 GiB: honest 45.2 Mhash/s in both constructions. Inline kernel (never reads the dataset): closed form 5,014 Mhash/s (111x FASTER than honest); memory-hard 9.49 Mhash/s (0.21 of honest, 4.8x SLOWER). At a 256 MiB dataset: honest 94.8, inline 9.48 (0.10). CPU verify per 32-lane warp (holds only the cache, derives every word on demand, 32 lanes interleaved): 0.649 / 0.631 / 0.701 ms for igneum-genesis, /epoch1, /epoch2 (104, 104, 112 loads; 3,328 to 3,584 items); 0.801 ms igneum-second-seed (104 loads); 1.205 ms igneum-second-seed/epoch1 (144 loads, 4,608 items). Cold single warps 1.16 to 2.11 ms. Closed form was 0.017 ms. 10 ms GATE MET, margin about 8x steady. Levers (implemented, measured, OFF by default; default generator unchanged): (a) --load-weight 17: 72 to 80 loads/hash, CPU 0.457 to 0.512 ms/warp, GPU 55.0 to 73.4 Mhash/s. (b) --wide-frac 50 (warp-coalesced 128 B loads): CPU 0.233 to 0.489 ms/warp, GPU 56.1 to 135.2 Mhash/s and useful bandwidth up to 56 GB/s, so (b) erodes the random-access bound. (a)+(b): CPU 0.223 to 0.276, GPU 106 to 139. Recommendation: no lever; (a) is the fallback if a slower verifier ever threatens the gate; (b) not recommended. Tests re-run on the new dataset: fuzz 200/200 (800 warps, 25,600 hashes, 0 mismatches, CPU interpreter 1.23 s), edge 14/14, determinism PASS (fingerprint 62a4f0eb018df273), memcheck PASS, stats PASS (3 seeds, no obvious bias). 3 warps x 3 seeds bit-exact in the bench run. CUDA: new pack proto-cuda/packs/igneum-genesis-mh (kernel.cu with cache-fill and build kernels, memhard.h shared by device and host, vectors incl. cache head/last/FNV and 64 sampled words). host.cu handles both modes; old packs unchanged (closed-form export re-run is byte-identical in kernel.cu and program.metal). clang emulation (emu/emu.sh igneum-genesis-mh): cache check PASS (all words, FNV == Mac), dataset self-test PASS at 1 GiB, 3/3 vectors standalone and 2/2 in batch at 2 warps/block. RTX 5090 and AMD runs of this pack PENDING; no NVIDIA figure for the memory-hard dataset exists. Not demonstrated: cross-vendor results for the new dataset; the shortcut ratio on a discrete GPU; time-memory trade-offs between the two measured points; cryptographic strength of the mixer and the chained cache; distinct-lines-per-hash census.</p>
|
|
<h2 id="2026-10-03-rusty-kaspa-base-build-and-3-node-devnet-on-the-mac-consensus-engineer-pre-fork-proof">2026-10-03 rusty-kaspa base build and 3-node devnet on the Mac (consensus-engineer, pre-fork proof)</h2>
|
|
<p>Machine: Apple M5 Max (18 CPU cores), 64 GB, macOS 26.6.2. Toolchain: Homebrew rust 1.69.0 was too old (repo needs 1.91.0), so rustup 1.29.1 was installed non-interactively and gives rustc 1.99.0 and cargo 1.99.0; protobuf 36.2 added via <code>brew install protobuf</code> (protoc was missing); Apple clang 14.0.3 already present. Nothing else was needed. Source: <code>vendor/rusty-kaspa</code> at commit <code>01b532e8b553523216471682649693af92f0fd16</code> (v2.1.0, 2026-09-22). <code>cargo build --release --bin kaspad</code>: 2 min 36 s cold, binary 35,405,104 bytes (34 MB), 131 compiler warnings, zero errors. Devnet: three <code>kaspad --devnet --nodnsseed --disable-upnp --enable-unsynced-mining --yes --loglevel=info</code> nodes, separate <code>--appdir</code>, P2P 16611/16621/16631, gRPC 16610/16620/16630, nodes 2 and 3 <code>--connect</code> to node 1 (node 3 to node 2 never came up because both started at once, so the topology was a star through node 1). Network params: 10 BPS (100 ms blocks), GHOSTDAG k 124, merge depth 36,000 blocks, finality depth 432,000, pruning depth 1,080,000, DAA window 661 samples x 40 blocks, genesis bits 0x1e21bc1c (about 248,663 hashes per block). Miner: kaspad ships none, so a 150-line CPU miner on <code>kaspa-pow::State</code> (real kHeavyHash, 16 threads, 300 ms template refresh) submitted to node 1 only: 27.2 MH/s sustained, 5,718 blocks in 180 s, 0 rejected. Blocks per second over the 180 s run: 31.76 on all three nodes (1,691 to 7,409 blocks each). Two phases: 56 to 62 blocks/s while difficulty sat at genesis (first 6,000 blocks, min window 150 samples), then the DAA raised difficulty to 1.12 M at block 6,018 and the rate fell to 14 to 15 blocks/s, still converging toward the 10 BPS target when the run ended. Propagation: block counts, DAA scores and sink hash were identical on all three nodes at 18 of 19 ten-second samples; the one miss was node 2 trailing by a single block for one sample. Tips stayed at 1 because a single serial miner never produced parallel blocks, so GHOSTDAG k was not exercised; a second miner is the next step for that. Earlier 30 s warm-up run: 1,690 blocks, 56.3 blocks/s on all three nodes, 28.1 MH/s. Fork points mapped with line numbers in <code>docs/fork-map.md</code> (hash, coinbase, DAA, header, depth constants, BPS and k). All nodes stopped at the end. Miner source kept outside the repo (scratchpad); re-create from <code>testing/integration/src/common/utils.rs:271</code> if needed.</p>
|
|
<h2 id="3-october-2026-rtx-5090-memory-hard-dataset-pack-igneum-genesis-mh">3 October 2026, RTX 5090, memory-hard dataset (pack igneum-genesis-mh)</h2>
|
|
<div class="tbl"><table><thead><tr><th>Check</th><th>Result</th></tr></thead><tbody><tr><td>256 MiB cache, GPU vs host, all 67,108,864 words</td><td>PASS, FNV-1a 48c4f5bf24166b2e matches the Mac</td></tr><tr><td>Cache fill</td><td>0.67 ms GPU, 223 ms one host thread</td></tr><tr><td>Dataset build from the cache, 1 GiB</td><td>13.4 ms, 1,253 M items/s</td></tr><tr><td>Vectors, 3 warps, standalone and in batch</td><td>96/96 PASS</td></tr><tr><td>Hash rate at 1 GiB</td><td>228.95 Mhash/s, 95.2 GB/s useful, 23.8 G random loads/s</td></tr></tbody></table></div>
|
|
<p>Reading: the memory-hard construction is now bit-exact across Apple Metal, NVIDIA CUDA and the CPU reference, cache and dataset included. Hash rate is unchanged from the closed-form dataset on both vendors, as expected, since the hash kernel only loads; what changed is that computing items on the fly is now slower than loading them (4.8x slower measured on Apple, not yet measured on NVIDIA). Still unmeasured: the inline shortcut ratio on NVIDIA, and AMD on any dataset.</p></article>
|
|
</div>
|
|
<footer>© 2026 Igneum. Generated from the repository at build time. Nothing on this page is an offer to sell anything.</footer>
|
|
</div>
|
|
</body>
|
|
</html> |