Commit graph

134 commits

Author SHA1 Message Date
igneum-labs
85c267971c Counter ASIC 3.0 gates (hash): the RTX 5090 G1 and G2 evidence from the collected job.log (eight fingerprints equal to the Mac's, 8 x 1,024 of 1,024); hash-gates.md and the bench-log entry closed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 16:14:08 +00:00
igneum-labs
fa857683be Counter ASIC 3.0 gates (hash): the gate directory (hash-gates.md, G1 Mac, G2 Mac, G3 and verifier JSON) and the bench-log entry; the RTX 5090 G1 rows pending the collected job.log
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 16:07:55 +00:00
igneum-labs
f4478a8ee1 Merge branch 'ca3-reserve' into ca3-coord 2026-10-06 08:45:48 +00:00
igneum-labs
1d0cc97721 Merge branch 'ca3-shadow' into ca3-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 08:45:48 +00:00
igneum-labs
697b8833df Counter ASIC 3.0 items 6 and 7: the RTX 5090 step-cost rows (PC 2 job run-ca3-family-pc2-20261006, every family bit-exact, mm8 included) in the bench-log and the reserve document 2026-10-06 08:45:02 +00:00
igneum-labs
9417a14ac7 Counter ASIC 3.0 item 8: the RTX 5090 rows, the chip side on the measured 11 pJ per op, the verdict
PC 2 job run-ca3-shadow-pc2-20261006 (card empty, every pack bit-exact against the Mac): the 5090 holds its rate to
150,800 ops per hash and loses 2.7 percent at 199,600 under the app's 431 W cap, which binds from 102,100 ops up and
takes the clock from 3,037 to 1,834 MHz (86 MH/s at 330,700 ops); 350 W at the control, 2.65 to 3.27 microjoules per
hash; marginal ALU energy 10 to 13 pJ per counted op. Clock rows OWED (nvidia-smi refused -lgc without rights). Chip
side at N = 100,000 and k = 1: 2.1x over the 5090 on GDDR7, 0.9x over the M5 Max. GO at mx8+sh256x27.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:44:57 +00:00
igneum-labs
aab9858068 Merge branch 'ca3-derive' into ca3-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 08:31:10 +00:00
igneum-labs
c6e03662a6 Counter ASIC 3.0 item 2: the RTX 5090 rows from PC 2 (bit-exact on CUDA, +1.1 s NVRTC per pack, the card-off key finding)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:30:15 +00:00
igneum-labs
e3ee3049bd Merge branch 'ca3-reserve' into ca3-coord
# Conflicts:
#	docs/bench-log.md
#	tools/observer/observer.mjs
2026-10-06 08:24:29 +00:00
igneum-labs
3959d66e55 Merge ca3-shadow (item 8's Mac rows) into ca3-coord: LoadClass carries derive_len and shadow; igneum-pow suite 96 passed 2026-10-06 08:11:03 +00:00
igneum-labs
4605314315 Counter ASIC 3.0 item 8: the Mac rows
The M5 Max ladder (Metal, packbench, IOReport GPU and DRAM watts without root): latency-bound to about 100,000 ops per
hash, the 5 percent point about 130,000, 11 to 27 W GPU at 100,000 ops, 0.78 to 1.40 microjoules per hash; the
verifier's law 2.06 ms + 3.2 us per 1,000 shadow instructions per warp on one core; every pack bit-exact. The
analysis file with the knob, the chip side (k = 1, 1.5, 0.3), the gates and the consequences; the bench-log entry;
the 5090 rows pending the PC 2 job (the playbook now carries the core-clock rows and the sh256x40 rung).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:08:26 +00:00
igneum-labs
bcc2db992e Counter ASIC 3.0 item 2: the design, the Mac measurements, the chip-model row and the PC 2 job
docs/plans/counter-asic-3-derivation.md (the design, the acceptance test, the interpreter, the allowance argument,
the measurements, the PROPOSED reserve entry R0 for 1.13.2, what is owed), docs/analysis/chip-model-v3.md section 6
(the per-day derivation rows at 1.0x to 3x allowances), the bench-log entry, relay/playbooks/ca3-derive-pc2.ps1
(one PC 2 job: self-fetched packs zip, the installed worker through NVRTC, the card off only under test with its
key from settings.json). Verifier 4.875 / 4.944 ms per unit on one M5 Max core under the measure lock against
x8's 2.061 / 2.063; Metal build 29 ms against 22; hash rate equal; bit-exact on Metal and Apple OpenCL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:50:01 +00:00
igneum-labs
ea141f2a94 Counter ASIC 3.0 items 6 and 7: family step-cost probes (Metal, CUDA), the PC 2 playbook, the M5 Max rows in the bench-log 2026-10-06 07:36:12 +00:00
igneum-labs
c5bc0f6bc8 Merge ca2-coord (6bb8124) into release-0.3.11: the Counter ASIC 2.0 docs commits (N4 rows, gate cells, digests, rules, status; the five cited documents already in the tree)
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:06:38 +00:00
igneum-labs
b9f0158872 Merge branch 'ca2-epoch' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:04:46 +00:00
igneum-labs
dae3593578 Merge branch 'ca2-soundness' into ca2-coord
# Conflicts:
#	docs/bench-log.md
#	proto-metal/packbench.swift
2026-10-05 23:04:45 +00:00
igneum-labs
22b495aa1d Merge branch 'ca2-analysis' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:04:14 +00:00
igneum-labs
41889fcaea Merge ca2-epoch (6cb846b) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:03:46 +00:00
igneum-labs
bfdc632a14 Merge ca2-analysis (c9ebe7b) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:03:45 +00:00
igneum-labs
75a56182b2 Merge commit 'a2f08e3' into release-0.3.11
# Conflicts:
#	docs/bench-log.md
#	docs/evidence.md
#	site/litepaper.html
2026-10-05 22:39:34 +00:00
igneum-labs
629157b388 Merge branch 'ca2-coord' into release-0.3.11
# Conflicts:
#	docs/bench-log.md
2026-10-05 22:39:12 +00:00
igneum-labs
19b30e31f0 Merge remote-tracking branch 'origin/master' into ca2-v3
# Conflicts:
#	docs/bench-log.md
#	proto-opencl/host.c
2026-10-05 22:34:07 +00:00
igneum-labs
c877e11ff5 C29 and D11: the verifier line on the fixed crate, the bounty struck until escrowed, the final-class rates in the public tables; status 22:25 2026-10-05 22:25:02 +00:00
igneum-labs
a59e782159 Bench log: Counter ASIC 2.0, the numbers (the level 3 page section); litepaper anchor 2026-10-05 22:20:05 +00:00
igneum-labs
9597374c4d verifier regression of e08909f fixed (derive_items out of line, one instance per cache size with the line mask a constant: v2 0.609 ms per unit against readwidth's 0.607, was 1.33); x8 into class v3 under the delegated rule (V3_CLASS = MX8; mx8-genesis and mx8-devnet-epoch0 re-exported through the seam, the devnet one with the era inside; the mx4 packs kept as the x4 record, generator 2); tests/packs.rs and tests/mixer.rs on the x8 class; mixer-x4.md 6.2a, 6.4a, 6.5 decision, 6.6 the regression; chip-model-v3.md headline x8 (0.92x with the factor); bench-log addendum with the PC 1 build rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:10:55 +00:00
igneum-labs
4f0b9f5475 bench-log: the root-socket recurrence of 21:25Z and the fix job of 22:01Z (the prover proved again at 22:02Z)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:03:06 +00:00
igneum-labs
829687b1ab mixer x4: the measure session (v2 / x4 / x8 verifier 1.33 / 1.94 / 2.79 ms per unit on a loaded core, 1.45x and 2.1x; the 256 MiB fill 172 to 175 ms; the Metal 1 GiB build flat at 21 ms, latency-bound), the verification-throughput consequences (C19), what is unverified and what is owed; the bench-log entry; the PC 1 two-card playbook; the measured verifier row in the chip model
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:02:28 +00:00
igneum-labs
39ecd7c52b hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.

Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:57:39 +00:00
igneum-labs
1558971521 era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook
Squashed from six commits (tag ca2-era-pre-squash) for one merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:41:33 +00:00
igneum-labs
6cb846b84c ca2-epoch: the Mac compile-ahead measurement (10 fresh programs 15.9 / 17.7 / 20.4 ms min / median / max, the devnet pack 79 ms then 1 ms from the shader cache) in docs/bench-log.md and epoch-length.md section 6; the per-card table and the 600-s floor filled
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:19:33 +00:00
igneum-labs
594790b354 Merge branch 'amd-prove' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-05 21:03:58 +00:00
igneum-labs
ffa2633845 docs/analysis/amd-proving.md: no zkVM proves on AMD (SP1, RISC Zero, Jolt, OpenVM, ICICLE cited), the SP1 CPU prover measured on PC 1 beside the miners (282 s a shard at any size, 30 GB RSS: no CPU tier), the tier consequences and the public line; PC 1 job scripts with the bash -n gate; bench-log entry
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:02:59 +00:00
igneum-labs
46bad306d8 Proving v1: the harness passes on the final fork tree (N = 8, 21 checks in 244 s)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:57:33 +00:00
igneum-labs
4ddc9fcf4c Proving v1: the miner-on curve's 2^25 rows (the prototype shard 24.7 GB and 38.8 s beside the miner)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:56:32 +00:00
igneum-labs
1aff0b79d1 Proving v1: the miner-on curve (the adopted shard 22.2 GB and 13.2 s beside the miner; the prototype 30.1 GB): the 24 GB tier measured, the public lines updated
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:55:54 +00:00
igneum-labs
1eaeefd770 Proving v1: the S_p curve's re-plan points (28.4 GB at 20 M cycles: the peak is flat from 20 M to 60 M)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:49:26 +00:00
igneum-labs
03313281ba Proving v1: the S_p curve, card alone (13.9 GB floor, 20.4 GB at the v1 shard, 28.3 GB at the prototype shard): the 12 and 16 GB tiers cannot prove on SP1 6.8.1's GPU server, 24 GB proves v1 shards, 32 GB the prototype
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:43:22 +00:00
igneum-labs
344cba8e8c Proving v1: the memory sweep and the miner-on peaks, the root-socket class fix (cleanup lines, tools/ci/prover-socket-check.sh in CI), the host's --budget re-plan and the S_p curve job, the RAM and aggregation-card gates, N = 8 in the fast-time file and spec 7.4
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:35:13 +00:00
igneum-labs
c9ebe7bd14 Counter ASIC 2.0 layer 7: RTX 5090 and RX 9070 XT dot4 numbers from PC 1 (job run-dot4-20261005): dp4a 7.45 T/s on the 5090 via inline PTX, v_dot4_i32_iu8 0.66 T/s on the 9070 XT via the clang builtin in Adrenalin OpenCL C, signed emulation 7.1x / 1.46x / 4.7x the ALU step on NVIDIA / AMD / Apple; no PC platform lists cl_khr_integer_dot_product
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:31:33 +00:00
igneum-labs
73358ca92d Proving v1: the 12 GB memory sweep on the 5090 (13.9 GB floor on an empty shard, 28.3 GB on a full one, no knob moves the floor)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:27:45 +00:00
igneum-labs
2d04a1c553 read-width: bench-log entry and docs/plans/read-width.md (three cards, probes, widths, the per-load mix, the scratch at 32 and 128 KiB, chip model, recommendation: keep v2; w16 the only width that passes the rules and closes nothing); test multiply made wrapping (dev-profile overflow check)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:24:40 +00:00
igneum-labs
3f09a552eb scratch-soundness.md: the five findings, the recompute-versus-store arithmetic at 64, 256 and 2,048 slots, the named on-die-cache chip row per variant (2.4x at every share under the cap) beside the M16 mixer lever, the host tag contract, the vector requirements, the Metal results (228 of 228 pass, 3 of 3 built-in failures caught); bench-log entry with the commands and counts
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:49 +00:00
igneum-labs
d4848453ae Proving v1: the chain of 8 on the 5090 measured (N = 2, 4, 8: 32.6, 66.8, 135.6 s; chained aggregation 9.7 s a block on a mining card), the fleet table re-cut on the measured rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:14 +00:00
igneum-labs
e9508926dc Proving v1: the chain of 2 on the 5090 (31.8 s, the chained aggregation 9.5 s), the job aborted by the 0.3.10 restart at block 3
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:14 +00:00
igneum-labs
251ea75b23 Proving v1: step 1 measured with the fleet mining (the prover costs the 5090 4.0%, GPU peak 15.6 GB with the miner), the fleet coverage window (2.4%, p50 44 s)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:14 +00:00
igneum-labs
ae3b681b3d Proving v1: the fleet-size table from the measured inputs, the numbers table in the plan
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:14 +00:00
igneum-labs
3fcf4f705c Proving v1: the harness passes (21 checks), the bench-log entry with the step 1 GPU memory, RAM and SM-target numbers, the CPU chain of 2 and verify-segment
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:14 +00:00
igneum-labs
eeef9cd056 Counter ASIC 2.0 layer 7: integer matrix family design (vendor primitives cited: PTX dp4a and mma .u8/.s8, AMD v_dot4_i32_iu8 and WMMA iu8 on RDNA 3 and 4 via LLVM and GPUOpen, CDNA 3 MFMA i8, Metal 4 matmul2d char x char -> int found in MSL 4.1 table 7.3; dot4 and mm8 semantics, reserve entry R1, emulation rule); M5 Max dot4 probe numbers in the bench log; PC playbook dot4-probe.ps1 prepared, not published
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:12:53 +00:00
igneum-labs
566f60b7a2 Merge c4-fix tip (d898d36) into release-0.3.10: the signer-pipe-check CI script and the per-job build-inputs zip names, committed 2026-10-05 18:29:19 +00:00
igneum-labs
d898d36308 C4 fix: the rule v3 harness row measured (140-s split, B certified 11 and 12, n0 adopted 12 by certificate 3 s after the heal, every node on the certified chain, 0 conflicting, 0 disagreeing); bench-log, ledger and spec 3.11.7 updated
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 18:28:29 +00:00