From 5334b4a5ecfc655d14707e88c0c71794e3915922 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 08:49:04 +0000 Subject: [PATCH] morning summary: Counter ASIC 3.0 closed, the class v4 candidate and its gate decision Co-Authored-By: Claude Fable 5.1 --- docs/plans/morning-2026-10-06.md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/plans/morning-2026-10-06.md b/docs/plans/morning-2026-10-06.md index 94281c42..b431517c 100644 --- a/docs/plans/morning-2026-10-06.md +++ b/docs/plans/morning-2026-10-06.md @@ -43,12 +43,15 @@ The finding that matters: the chip that wins is not the clever recompute chip 2. Why: at the hash the 5090 spends about 55 W on memory and the rest keeping a GPU alive at 0.15 percent of its integer budget. The lever is RandomX's lever: make the hash use the rest of the chip. The 5090 can hide about 330,000 operations per hash behind its 128 reads; today it hides 512. Item 8 measures that fill: on the Mac it costs 1.5 percent of rate at 100,000 ops, the verifier barely notices, and the chip's edge drops from 5x to 2.7x. The 5090's rows are queued on PC 2. The deciding number is the chip core's energy per op against a GPU's ALU, which item 8 is pricing. +**Closed at 08:50.** One class v4 candidate, measured on the hash's own numbers and ready for its six-gate run on your word: mixer x8 plus 100,000 operations of program work per hash. The 5090 loses 0.2 percent of rate and the M5 Max 1.5 percent; the verifier adds 0.17 ms per warp; bit-exact on Metal, CUDA, Apple OpenCL and the CPU emulation. The chip must then carry a 14,000-lane ALU array, and its per-joule edge over the 5090 falls from 5.6x to 2.1x at a core as efficient as the GPU's, 1.5x at a realistic one. Your test as a number: the chip crosses 2x only if its datapath spends under half the energy per op that a GPU does. The cost per tier: a 5090 draws 431 W instead of 350 for the same blocks (a rig pays about 23 percent more electricity), the M5 Max 37 W instead of 21, a pool user sees nothing. The 9070 XT and 4060-class rows are owed, the AMD ones because PC 1 was left alone. + Other items: item 2 (per-day random derivation) works bit-exact at no hash cost and drops the recompute chip to 0.29x to 0.43x, but at full size its CPU verifier is over the 10 ms gate on an old core; it goes in as a reserve, the half-size draw passes, and the class v4 candidate is the pairing of derivation class and program length under one verifier gate. Item 3: cryptanalysis brief and budget line, USD 80,000 to 160,000, one firm plus one academic group, verdict GO to commission, nobody contacted. Items 4 and 5: share-pattern detector in the observer (fires on a fabricated fixed design, quiet on the devnet), issuance trigger at USD 20,000 a day, FPGA soft overlay 0.3x to 0.4x a 5090 per watt, layer 9 ranked above layer 7. Items 6 and 7 in flight. ## Decisions for you The full list with recommendations is in `consequences-decisions.md` (14) and the 3.0 status file. The ones that bite first: +0. **Run the six gates on the class v4 candidate**, or wait for the 9070 XT and 4060-class rows first (3.0 status, decision 2). 1. **The public claim "under 2x".** True of the recompute chip per chip, false of the stored-dataset chip per joule. Two re-wordings are drafted in the 3.0 status file; nothing on the site changed. Pick one before any public push. 2. **The segment rule.** Answered at 08:50: it is a consensus rule (as shipped a fresh segment record is valid for an 8-second window). The fix is on the node fork behind a new switch `proving_v1_fresh_rule_daa`, so 0.3.12 carries the node and goes out as a two-manifest publish with the switch at tip + 14,400. Measured on PC 2 beside the miner: 9 whole segments in 30 minutes, 72 of 72 shard records paid, 11 percent of hash rate. 3. **0.3.12 go.** The state-reply fix, the update catch-up, the card order, Ember Tune and its guards, plus the node fork with the segment switch (two-manifest publish, switch at tip + 14,400). No prompt on any machine.