Commit graph

134 commits

Author SHA1 Message Date
igneum-josh
1ecb730462 Counter ASIC 3.0 gates (hash): the RTX 5090 G1 and G2 evidence from the collected job.log (eight fingerprints equal to the Mac's, 8 x 1,024 of 1,024); hash-gates.md and the bench-log entry closed
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 17:14:08 +01:00
igneum-josh
5e2fbdd0c6 Counter ASIC 3.0 gates (hash): the gate directory (hash-gates.md, G1 Mac, G2 Mac, G3 and verifier JSON) and the bench-log entry; the RTX 5090 G1 rows pending the collected job.log
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 17:07:55 +01:00
igneum-josh
820f4d4eed Merge branch 'ca3-reserve' into ca3-coord 2026-10-06 09:45:48 +01:00
igneum-josh
12d8a9302a Merge branch 'ca3-shadow' into ca3-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 09:45:48 +01:00
igneum-josh
b85f0f18ed Counter ASIC 3.0 items 6 and 7: the RTX 5090 step-cost rows (PC 2 job run-ca3-family-pc2-20261006, every family bit-exact, mm8 included) in the bench-log and the reserve document 2026-10-06 09:45:02 +01:00
igneum-josh
70eaf06456 Counter ASIC 3.0 item 8: the RTX 5090 rows, the chip side on the measured 11 pJ per op, the verdict
PC 2 job run-ca3-shadow-pc2-20261006 (card empty, every pack bit-exact against the Mac): the 5090 holds its rate to
150,800 ops per hash and loses 2.7 percent at 199,600 under the app's 431 W cap, which binds from 102,100 ops up and
takes the clock from 3,037 to 1,834 MHz (86 MH/s at 330,700 ops); 350 W at the control, 2.65 to 3.27 microjoules per
hash; marginal ALU energy 10 to 13 pJ per counted op. Clock rows OWED (nvidia-smi refused -lgc without rights). Chip
side at N = 100,000 and k = 1: 2.1x over the 5090 on GDDR7, 0.9x over the M5 Max. GO at mx8+sh256x27.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 09:44:57 +01:00
igneum-josh
bc148a85ec Merge branch 'ca3-derive' into ca3-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 09:31:10 +01:00
igneum-josh
54bc188a5d Counter ASIC 3.0 item 2: the RTX 5090 rows from PC 2 (bit-exact on CUDA, +1.1 s NVRTC per pack, the card-off key finding)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 09:30:15 +01:00
igneum-josh
fcce185eeb Merge branch 'ca3-reserve' into ca3-coord
# Conflicts:
#	docs/bench-log.md
#	tools/observer/observer.mjs
2026-10-06 09:24:29 +01:00
igneum-josh
45f30196b9 Merge ca3-shadow (item 8's Mac rows) into ca3-coord: LoadClass carries derive_len and shadow; igneum-pow suite 96 passed 2026-10-06 09:11:03 +01:00
igneum-josh
a050a54c65 Counter ASIC 3.0 item 8: the Mac rows
The M5 Max ladder (Metal, packbench, IOReport GPU and DRAM watts without root): latency-bound to about 100,000 ops per
hash, the 5 percent point about 130,000, 11 to 27 W GPU at 100,000 ops, 0.78 to 1.40 microjoules per hash; the
verifier's law 2.06 ms + 3.2 us per 1,000 shadow instructions per warp on one core; every pack bit-exact. The
analysis file with the knob, the chip side (k = 1, 1.5, 0.3), the gates and the consequences; the bench-log entry;
the 5090 rows pending the PC 2 job (the playbook now carries the core-clock rows and the sh256x40 rung).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 09:08:26 +01:00
igneum-josh
dd5041b989 Counter ASIC 3.0 item 2: the design, the Mac measurements, the chip-model row and the PC 2 job
docs/plans/counter-asic-3-derivation.md (the design, the acceptance test, the interpreter, the allowance argument,
the measurements, the PROPOSED reserve entry R0 for 1.13.2, what is owed), docs/analysis/chip-model-v3.md section 6
(the per-day derivation rows at 1.0x to 3x allowances), the bench-log entry, relay/playbooks/ca3-derive-pc2.ps1
(one PC 2 job: self-fetched packs zip, the installed worker through NVRTC, the card off only under test with its
key from settings.json). Verifier 4.875 / 4.944 ms per unit on one M5 Max core under the measure lock against
x8's 2.061 / 2.063; Metal build 29 ms against 22; hash rate equal; bit-exact on Metal and Apple OpenCL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 08:50:01 +01:00
igneum-josh
192a683d08 Counter ASIC 3.0 items 6 and 7: family step-cost probes (Metal, CUDA), the PC 2 playbook, the M5 Max rows in the bench-log 2026-10-06 08:36:12 +01:00
igneum-josh
5debb36478 Merge ca2-coord (8bef299) into release-0.3.11: the Counter ASIC 2.0 docs commits (N4 rows, gate cells, digests, rules, status; the five cited documents already in the tree)
# Conflicts:
#	docs/bench-log.md
2026-10-06 00:06:38 +01:00
igneum-josh
f940101ae2 Merge branch 'ca2-epoch' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 00:04:46 +01:00
igneum-josh
271cd63116 Merge branch 'ca2-soundness' into ca2-coord
# Conflicts:
#	docs/bench-log.md
#	proto-metal/packbench.swift
2026-10-06 00:04:45 +01:00
igneum-josh
8c00f84c1e Merge branch 'ca2-analysis' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-06 00:04:14 +01:00
igneum-josh
54419244a4 Merge ca2-epoch (e95e8b5) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
# Conflicts:
#	docs/bench-log.md
2026-10-06 00:03:46 +01:00
igneum-josh
fe7df4f841 Merge ca2-analysis (ee42d7c) into release-0.3.11: docs and standalone probe sources the evidence rows cite (C38); nothing a built artefact reads
# Conflicts:
#	docs/bench-log.md
2026-10-06 00:03:45 +01:00
igneum-josh
5cbb796051 Merge commit '22c2363' into release-0.3.11
# Conflicts:
#	docs/bench-log.md
#	docs/evidence.md
#	site/litepaper.html
2026-10-05 23:39:34 +01:00
igneum-josh
18605f425c Merge branch 'ca2-coord' into release-0.3.11
# Conflicts:
#	docs/bench-log.md
2026-10-05 23:39:12 +01:00
igneum-josh
49c7e78307 Merge remote-tracking branch 'origin/master' into ca2-v3
# Conflicts:
#	docs/bench-log.md
#	proto-opencl/host.c
2026-10-05 22:34:07 +00:00
igneum-josh
7e6f77c8f9 C29 and D11: the verifier line on the fixed crate, the bounty struck until escrowed, the final-class rates in the public tables; status 22:25 2026-10-05 22:25:02 +00:00
igneum-josh
939ded01c8 Bench log: Counter ASIC 2.0, the numbers (the level 3 page section); litepaper anchor 2026-10-05 22:20:05 +00:00
igneum-josh
1ab8b2113e verifier regression of 0fc0ad1 fixed (derive_items out of line, one instance per cache size with the line mask a constant: v2 0.609 ms per unit against readwidth's 0.607, was 1.33); x8 into class v3 under the delegated rule (V3_CLASS = MX8; mx8-genesis and mx8-devnet-epoch0 re-exported through the seam, the devnet one with the era inside; the mx4 packs kept as the x4 record, generator 2); tests/packs.rs and tests/mixer.rs on the x8 class; mixer-x4.md 6.2a, 6.4a, 6.5 decision, 6.6 the regression; chip-model-v3.md headline x8 (0.92x with the factor); bench-log addendum with the PC 1 build rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:10:55 +00:00
igneum-josh
1d7897930c bench-log: the root-socket recurrence of 21:25Z and the fix job of 22:01Z (the prover proved again at 22:02Z)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 23:03:06 +01:00
igneum-josh
16dfd1eacc mixer x4: the measure session (v2 / x4 / x8 verifier 1.33 / 1.94 / 2.79 ms per unit on a loaded core, 1.45x and 2.1x; the 256 MiB fill 172 to 175 ms; the Metal 1 GiB build flat at 21 ms, latency-bound), the verification-throughput consequences (C19), what is unverified and what is owed; the bench-log entry; the PC 1 two-card playbook; the measured verifier row in the chip model
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 23:02:28 +01:00
igneum-josh
1950661c11 hot table (Counter ASIC 2.0 layer 5), measured and not adopted, on the ca2-v3 composed class (squash of tag ca2-cache-history-2026-10-05)
LoadClass::hot (Option<HotClass { mb, k, added }>) beside mix, slots, scratch, mixer_mult, growth and era; V3_CLASS = { era: None, hot: None, ..LoadClass::MX4 } (hot stays None: the hot code is behind the flag, measured and not adopted). Op::Hot, the hot slots drawn after the scratch slots, no width roll (v2_loads allows the added form's extra slots), the id suffix hot/<S><k>[added], the hot parsing inside parse_loads, the era branch first in name(). HotTable under seed_words("igneum-hot/" || epoch seed) with the cache chain and tag umHT, read at H[mulhi(src, HOT_WORDS)]; DatasetSource::hot attached by new_class_day and from_seed_bytes_class; the acceptance stand-in dataset_elem(idx, S[2], S[3]); the three emitters (hot argument after the init words, ht_segment and igneum_hot_fill beside the layout-aware cores); packfile.h hot fields beside class, era, attempt and mixerMult; OpenCL host, Metal packbench and NVRTC worker fill H on the device and self-test it. Eight packs under proto-cuda/packs-ca2-hot (replaced hot32k4 hot64k4 hot96k4 hot64k2 hot64k8, added hot32k4a hot64k4a hot96k4a) re-exported on the merged crate: vectors unchanged, program.h and program.json carry the mixer fields. docs/plans/hot-table.md (design, spec text, per-tier budget, chip model, Mac and PC measurements, the decision: layer 5 out of v3, the 3.0 note); bench-log entry and addenda with job ids and worker sha256s; the two PC playbooks.

Checks on this commit: cargo test --release 53 + 19 pass (the pinned v2, mx4, era, readwidth and hot packs); the pinned packs under proto-cuda/packs, packs-ca2-mixer, packs-ca2-era and packs-readwidth untouched; Metal packbench and Apple OpenCL --bench-pack on all eight hot packs (run lock, 2^20 at base 0): 96/96 lanes, hot table head, last line and FNV PASS, one fingerprint per pack on both harnesses, equal to the fingerprints before the rebase (hot32k4 679e5e83378d3790, hot64k4 d4c9e456b039fdef, hot96k4 7c98eceffee9fd73, hot64k2 f43b10a95879b8e5, hot64k8 c11309d743be9392, hot32k4a afb700b2d997c847, hot64k4a ba214baa9c1a9e85, hot96k4a 29e1916aed6deff5). No era or mixer behaviour changed: every resolution kept the ca2-v3 side and appended the hot branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:57:39 +01:00
igneum-josh
b105a5533c era layout (Counter ASIC 2.0 layers 4 and 8) behind the class flag, on the ca2-v3 seam: the 7-draw era stream from E_n (stride, interleave, width pinned at 4 B), per-site window draws (dataset, half, quarter at a 256 MiB floor), the strided windowed load address in the interpreter, the acceptance mirror and the three emitters, the interleaved dataset layout riding with the program (memhard::Layout, mh_t/mh_j/mh_addr, Epoch::dataset_word), V3_CLASS with the era drawn inside by generate_from_seed_bytes_program_class, --era / --era-widths on the CLI, six class v3 era packs (packs-ca2-era), tests, the design doc, host.cu/host.c deriving host words through the pack's mh_word, emu/test-layout.sh, the PC 1 playbook
Squashed from six commits (tag ca2-era-pre-squash) for one merge.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 22:41:33 +01:00
igneum-josh
e95e8b547d ca2-epoch: the Mac compile-ahead measurement (10 fresh programs 15.9 / 17.7 / 20.4 ms min / median / max, the devnet pack 79 ms then 1 ms from the shader cache) in docs/bench-log.md and epoch-length.md section 6; the per-card table and the 600-s floor filled
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:19:33 +00:00
igneum-josh
69ab0fb95d Merge branch 'amd-prove' into ca2-coord
# Conflicts:
#	docs/bench-log.md
2026-10-05 22:03:58 +01:00
igneum-josh
f1d7a7d32a docs/analysis/amd-proving.md: no zkVM proves on AMD (SP1, RISC Zero, Jolt, OpenVM, ICICLE cited), the SP1 CPU prover measured on PC 1 beside the miners (282 s a shard at any size, 30 GB RSS: no CPU tier), the tier consequences and the public line; PC 1 job scripts with the bash -n gate; bench-log entry
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:02:59 +00:00
igneum-josh
90d3299966 Proving v1: the harness passes on the final fork tree (N = 8, 21 checks in 244 s)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:57:33 +01:00
igneum-josh
f6e2308179 Proving v1: the miner-on curve's 2^25 rows (the prototype shard 24.7 GB and 38.8 s beside the miner)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:56:32 +01:00
igneum-josh
219517fc17 Proving v1: the miner-on curve (the adopted shard 22.2 GB and 13.2 s beside the miner; the prototype 30.1 GB): the 24 GB tier measured, the public lines updated
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:55:54 +01:00
igneum-josh
0294264afd Proving v1: the S_p curve's re-plan points (28.4 GB at 20 M cycles: the peak is flat from 20 M to 60 M)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:49:26 +01:00
igneum-josh
98014a375a Proving v1: the S_p curve, card alone (13.9 GB floor, 20.4 GB at the v1 shard, 28.3 GB at the prototype shard): the 12 and 16 GB tiers cannot prove on SP1 6.8.1's GPU server, 24 GB proves v1 shards, 32 GB the prototype
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:43:22 +01:00
igneum-josh
c2544be3bf Proving v1: the memory sweep and the miner-on peaks, the root-socket class fix (cleanup lines, tools/ci/prover-socket-check.sh in CI), the host's --budget re-plan and the S_p curve job, the RAM and aggregation-card gates, N = 8 in the fast-time file and spec 7.4
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:35:13 +01:00
igneum-josh
ee42d7c5d4 Counter ASIC 2.0 layer 7: RTX 5090 and RX 9070 XT dot4 numbers from PC 1 (job run-dot4-20261005): dp4a 7.45 T/s on the 5090 via inline PTX, v_dot4_i32_iu8 0.66 T/s on the 9070 XT via the clang builtin in Adrenalin OpenCL C, signed emulation 7.1x / 1.46x / 4.7x the ALU step on NVIDIA / AMD / Apple; no PC platform lists cl_khr_integer_dot_product
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:31:33 +00:00
igneum-josh
cde9c47b37 Proving v1: the 12 GB memory sweep on the 5090 (13.9 GB floor on an empty shard, 28.3 GB on a full one, no knob moves the floor)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:27:45 +01:00
igneum-josh
e752fc7d45 read-width: bench-log entry and docs/plans/read-width.md (three cards, probes, widths, the per-load mix, the scratch at 32 and 128 KiB, chip model, recommendation: keep v2; w16 the only width that passes the rules and closes nothing); test multiply made wrapping (dev-profile overflow check)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:24:40 +01:00
igneum-josh
a4658816e0 scratch-soundness.md: the five findings, the recompute-versus-store arithmetic at 64, 256 and 2,048 slots, the named on-die-cache chip row per variant (2.4x at every share under the cap) beside the M16 mixer lever, the host tag contract, the vector requirements, the Metal results (228 of 228 pass, 3 of 3 built-in failures caught); bench-log entry with the commands and counts
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:21:49 +00:00
igneum-josh
b3997089ff Proving v1: the chain of 8 on the 5090 measured (N = 2, 4, 8: 32.6, 66.8, 135.6 s; chained aggregation 9.7 s a block on a mining card), the fleet table re-cut on the measured rows
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:21:14 +01:00
igneum-josh
5d07095d3a Proving v1: the chain of 2 on the 5090 (31.8 s, the chained aggregation 9.5 s), the job aborted by the 0.3.10 restart at block 3
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:21:14 +01:00
igneum-josh
e5bb296c2e Proving v1: step 1 measured with the fleet mining (the prover costs the 5090 4.0%, GPU peak 15.6 GB with the miner), the fleet coverage window (2.4%, p50 44 s)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:21:14 +01:00
igneum-josh
b5e27037f1 Proving v1: the fleet-size table from the measured inputs, the numbers table in the plan
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:21:14 +01:00
igneum-josh
13dc9ce24e Proving v1: the harness passes (21 checks), the bench-log entry with the step 1 GPU memory, RAM and SM-target numbers, the CPU chain of 2 and verify-segment
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 21:21:14 +01:00
igneum-josh
f59708dff7 Counter ASIC 2.0 layer 7: integer matrix family design (vendor primitives cited: PTX dp4a and mma .u8/.s8, AMD v_dot4_i32_iu8 and WMMA iu8 on RDNA 3 and 4 via LLVM and GPUOpen, CDNA 3 MFMA i8, Metal 4 matmul2d char x char -> int found in MSL 4.1 table 7.3; dot4 and mm8 semantics, reserve entry R1, emulation rule); M5 Max dot4 probe numbers in the bench log; PC playbook dot4-probe.ps1 prepared, not published
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 20:12:53 +00:00
igneum-josh
9a51e322c1 Merge c4-fix tip (68f14b9) into release-0.3.10: the signer-pipe-check CI script and the per-job build-inputs zip names, committed 2026-10-05 19:29:19 +01:00
igneum-josh
68f14b90be C4 fix: the rule v3 harness row measured (140-s split, B certified 11 and 12, n0 adopted 12 by certificate 3 s after the heal, every node on the certified chain, 0 conflicting, 0 disagreeing); bench-log, ledger and spec 3.11.7 updated
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-05 19:28:29 +01:00