Class v6, the rotating family: the outline, section 0 per tier and the four layers' table with their chip rows' shape and the open numbers (the hash lane's three rows due 16:00 UK; the document complete by 18:00 UK)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 10:00:21 +00:00
parent 93a5e401b1
commit 2fc3377cd0
2 changed files with 69 additions and 0 deletions

View file

@ -0,0 +1,67 @@
# Class v6, the rotating family: four layers that change the hash on a schedule no release carries
8 October 2026, 10:0x to 17:xx UTC (11:0x to 18:xx UK), branch `counter-asic-4`, the Counter ASIC 4.0 research lane under the Counter ASIC coordinator, the hash lane for the measured per-tier rows. Opened on the founder's question of 11:0x UK: "can we add in any more layers? class rotating? things that would render an ASIC useless as soon as it dropped." Status: a design document and a gate plan; no consensus code this week; nothing here touches the devnet, the testnet object or any served number. Every figure carries a label: **measured** (a card on a named job), **modelled** (arithmetic on the chip model's cited figures, `docs/analysis/chip-model-v3.md`), **claimed** (a vendor's or an author's figure), **approximate** (from memory or a scaling). Outline and layer table at 10:5x UTC; the hash lane's rows due 16:00 UK; the document complete by 18:00 UK; at each clock the default lands with the gaps named.
## 0. For the founder tonight: what each layer does to a chip on its release day, and what it costs a 5090, a 5070 Ti and an M5 Max
The honest frame first, from last night's measured close (`docs/analysis/counter-asic-4-research.md` sections 2 and 15.1a). Two chips exist in the model. A **fixed-function chip** wires one hash: the mixer's rounds, the op mix's lane ratios, the read width, the program length and the block shape are silicon. A **GPU-like chip** stores the dataset in commodity DRAM and runs the hour's program on a programmable integer core (the `f = 1` chip of chip-model-v3 section 5, the chip anyone builds); for it every drawn parameter is firmware. The four layers below render the FIRST kind useless on the day a parameter leaves its wired value, and move nothing for the second kind except the size of the core it must carry and the capex that forces. What the second kind keeps is the identity of section 2: at zero shadow premium the card's whole-card energy over the chip's memory energy (3.6x on a 5090 at its knee, measured card, modelled chip), and with the class v4 shadow 2.1x at a core as good as a GPU lane. No rotation changes that; the layers change which chip can be built and how long its tape-out lives.
| Layer | A fixed-function chip on its release day | A GPU-like chip on its release day | RTX 5090 (measured where stated) | RTX 5070 Ti (DEFAULT until the hash lane's row: the rented 5070's row scaled, approximate) | Apple M5 Max (measured where stated) |
|---|---|---|---|---|---|
| 1. Per-era draws of the parameters now fixed by release (mixer round count, op-mix weights, read width, program length, shadow block shape), from chain state | dead at the first era whose draw leaves its wired value: a wired x8 mixer at an x16 era recomputes nothing right; a 4-byte-granularity controller at a 16-byte era moves 4x the bytes; a core sized for 100,000 ops at a 200,000 era runs at half rate; the tape-out lives one era (180 days, layer 3) | firmware; the core sized for the band's top (200,000 ops: about 60 mm^2 of N5 instead of 30, USD 25 to 40 more per chip, modelled); `k` unchanged; capex per MH/s +10 to 20 percent (modelled) | rate: 0 across the band while latency-bound (the ladder's rungs 0 to 2: 0 and -2.7 percent at the 431 W cap, measured); watts: the premium follows N and the mix (6.2 to 11.3 pJ per counted op measured; a shuffle-heavy draw up to 5x per op, so the band excludes it); the verifier +0.2 to +1.5 ms per era draw (measured ladder, estimated mixer) | pending the row; by scaling: holds rate to about 200,000 ops at its cap (the 4070 did), the premium 20 to 30 W | rate -3.3 to -10 points across the N band (measured ladder rungs 0 to 2), 0 for the mixer and the width; the Apple tier sets the band's top (130,000 ops at the 5 percent rule) |
| 2. The state-derived dataset's size tracks chain-state growth with a floor (class v5's leaves scaled by the state, never below the 1.13.3 schedule) | a chip with fixed memory ages out when the dataset passes it: one HBM3 stack 24 GB, the 5090's board 32 GB; the time-memory curve (chip-model 5.4) says the excess must be recomputed at 6.3 nJ per item against 2.0 per read, so its energy per hash rises with the overflow | the same memory limit; a chip buys DRAM a card cannot (24 GB stacks at USD 200, modelled), so it ages out LAST: the 8, 12 and 16 GB card tiers go first | fine to 32 GB: the 2 GiB genesis dataset plus the state's leaves; at a used chain's 5 GB state (class-v5 section 3) the dataset is capped at the sample size, 2 GiB | 16 GB: fine to the 8 GiB step (year 12 on the schedule); the state floor does not move it | 36 to 128 GB unified: fine to the 16 GiB step; the daily build grows with the size (13 to 30 ms measured at 1 GiB) |
| 3. Scheduled family epochs by height, every 180 days by default, no release (the reserve R0 to R8 of spec 1.13.2 unlocking by height, then rotating) | a chip without the family's datapath loses its weight of the mix at the unlock (4 points of 79) or emulates it at the vendor penalty (1.5x to 2.4x per op, measured on the cards); a chip taped out against one family set is a GPU-like chip or dead | pre-wires every family for about USD 4 of N5 silicon (algorithm.md 5.2, modelled); moves the per-joule edge under 10 percent per family | measured family step costs: shfla 1.53x, perm 1.30, mm8 2.43 the add-xor-rotate step; under 1 percent of rate at 4 points | pending; the same families on the same silicon generation | measured: shfla 1.91x, perm 1.13 emulated, mm8 emulated at 1.6x per dot4; under 1 percent of rate at 4 points |
| 4. The acceptance floor (c''') and the F8 uniformity test generalised to every era's draw, with a redraw on failure (plus the per-site largest-bucket bound and the value-level bias test as the next class's two tests) | nothing on the chip; it is what makes layers 1 and 3 safe without per-era cryptanalysis | nothing | the generator's attempts per seed (today about 30 at 2.4 percent rejection under (c'''); a redraw costs nothing on a card) | the same | the same |
What a fully general chip still gets, in one line: the stored-dataset chip with a programmable core at the top of the band keeps 3.6x at zero premium and 2.1x at `k = 1` on a 5090 at its knee (section 2 of the research file), and the four layers cost it about USD 30 to 60 per chip of extra silicon and a project that must be N5-class from the first tape-out (the USD 100 M break-even cap of the mission lane's model).
## 1. Where the four layers sit in the design as it stands
| Item | Today (spec 01, the class system) | What class v6 adds |
|---|---|---|
| The era draw (1.13.1) | one SplitMix64 stream from the 1-hour VDF's `E_n`; draws the op-weight perturbation (B = 2 points, proposed), the fold rotations, the table layout and the working-set window (layers 4 and 8 of Counter ASIC 2.0, IN); `epoch_len` and the ladder's draws consumed and not used (set by signal) | layer 1: five more draws of the same stream, in a fixed order, each within a genesis-fixed band, each consumed whether used or not |
| The mixer (1.8.5, `mixer_mult` 8 under class v3) | fixed at genesis "so the verify budget holds" | layer 1 draws it in {4, 8, 16} (the x4 and x8 rows measured, x16 the hash lane's row or the estimate), the verifier's budget the bound |
| The read width (1.13.1 draw 1, `allowed = {1}`) | pinned to 4 bytes by the 5 October decision (w64 made the 5090 bandwidth-bound) | layer 1 draws from {1, 4} words (w16 within 2.7 percent on the 5090 and the 9070 XT, measured); never 16 words |
| The shadow (class v4: 256 x 27, rung 0 of the ladder) | N moves by miner signal within the ladder | layer 1 draws the block shape (64 to 256 instructions, measured; never 1,024) and the op-mix weights within the band; N stays the ladder's (signal, not draw) |
| The dataset (1.13.3) | 2 GiB plus 0.5 GiB a year; class v5's leaves (64 B per state record) | layer 2: the leaf count tracks the state with the schedule as the floor; the sample cap stays at the dataset size |
| The reserve (1.13.2) | R1..R8 unlock one per era, by height | layer 3: the unlock cadence a genesis constant (180 days) and the set rotating after the reserve is exhausted (the retired family returns in a fixed cycle) |
| The acceptance rule (1.4.6; sub-version 3's (a') (c') (c'''); F8's census as a gate) | judged on the genesis parameters; the census run per class by hand | layer 4: the rule and the census parameterised by the era draw; the generator redraws a candidate that fails any of the four tests |
## 2. Layer 1: per-era draws of the released parameters
PENDING the hash lane's rows (16:00 UK): the x16 mixer verifier and build; the two re-weighted shadow packs. The skeleton:
| Parameter | Band (genesis) | Why that band (the measured rows that set it) | Chip rows: fixed-function / GPU-like | Per tier: 5090 / 5070 Ti / M5 Max | Verifier | Open number |
|---|---|---|---|---|---|---|
| Mixer applications per round `m` | {4, 8, 16} | x4 and x8 measured (mixer-x4.md 6.4: 1.92 and 2.79 ms per unit on a loaded M5 Max core; the build unmoved on every discrete card, latency-bound); x16 about 3.7 ms per unit on the M5 Max core, 9 ms on a 2019-class core (chip-model-v3 section 3 item 1, estimate) | the `f = 0` recompute chip's rate halves per doubling (0.31x bare at x8, 0.16x at x16, modelled); the `f = 1` chip unmoved (it recomputes nothing) | rate 0 / 0 / 0 (latency-bound, measured at x4 and x8); the daily build 42 / pending / 29 ms at x8 (measured) | +0.6 to +1.5 ms per doubling (measured x4, x8; x16 the row) | the x16 row; adv-mixer-3's margin (every statistic clean from k = 1, the SAT ladder k = 1 solved, k = 2..4 timeout): x4 keeps the margin by that report's reading, x16 widens it |
| Op-mix weights (the ten non-load families) | each weight within B points of table 1.4.2, B = 2 today; the proposal: B = 4 with the shuffle and mulhi weights capped at their class v4 values | the microbench (15.1a): shfl 55.8 pJ per op, mulhi 39.6, prmt 22.3, lop3 24.1, mul 13.9, arx 11.3 at the stock clock; a draw that raises shfl from 8 to 14 of 75 raises the premium per instruction by about 35 percent (modelled on those rows) | a chip that specialised its lane ratio loses the ratio; a general core nothing | the premium per instruction moves with the mix: pending the two packs | +0 ms | the two re-weighted packs' rows |
| Read width `W` | {1, 4} words | w4 and w16 measured 5 October: the 5090 139.8 against 136.1 MH/s, the 9070 XT 17.90 against 18.15; w64 bandwidth-bound (71.9 MH/s on the 5090) | a 4-byte-granularity controller moves 4x the bytes at a w16 era (the Ren-Devadas lever reversed); the `f = 1` chip's energy per read rises toward the card's | 0 / pending / 0 (measured w16 on the M5 Max: within 1 percent) | 0 | none |
| Shadow block shape | 64 to 256 instructions per block, the pass count the ladder's | 64-instruction blocks ran 2.5 to 3.5 percent FASTER than 256 on the 5090 and the M5 Max (measured 6 October); 1,024 cost the M5 Max 17 percent | nothing for any chip (the work is the same) | +2.5 to 0 percent / pending / +2.5 to 0 | 0 | none |
| Program length N | not drawn: the ladder's signal (latency-ladder.md) | an unconditional draw retires the Apple tier at 200,000 (measured -10 percent) | the core sized for the ladder's admissible top (rung 2, 199,600 ops) | the ladder's rows | the ladder's rows | none |
## 3. Layer 2: the dataset's size tracks the chain state with a floor
PENDING the write-up; the numbers exist: the time-memory curve (chip-model-v3 5.4: every memory system's cheapest point is `f = 1`; a recomputed item costs 6.3 nJ and 9,360 ops against a stored one's 1.2 to 2.0 nJ), the DRAM read cost per the microbench (10.9 nJ per dependent read on the 5090 unlocked, 8.7 at the lock, the whole card's marginal; the memory system's own 2.0 modelled), the state sizes (class-v5-stored-state.md section 3: 93 records today, 61 M on a used chain, the leaves capped at the dataset size), the card-lifetime table (card-lifetime-2026-10-05.md: the 4 GiB step retires 4 GB cards; 8 GB mines to year 12).
## 4. Layer 3: scheduled family epochs by height
PENDING the write-up; the numbers exist: the reserve's chip and card rows (algorithm.md 5.2; counter-asic-3-reserve.md), the family step costs (item 6, measured on the M5 Max, the 5090 and the 9070 XT), the emulation bound of 1.13.2 (8x per op, 5 percent of rate).
## 5. Layer 4: the acceptance rule and the census as a function of the era draw
PENDING the write-up; the pieces: (c''') `MIN_DISTINCT_RATIO_V5 = 0.995` (class-v5 section 14: 2.435 percent of candidates under it, attempts +3.4 percent); F8's largest-bucket 6-sigma test on 64-line segments over 2^28 derivations (f8-uniform.md section 7, PASS at +4.84 sigma against the control's +4.18); the per-site bucket bound from this morning's tail attribution (AP-F8-1: p4, p8, p10, p34 as per-site bucket concentration at a narrow-window site); the value-level bias test (adv-cache-2: a product's biased low bits at address bit R; the research file's 20.2b); the duplicate-lane test (the research file's 20.2a, the per-load class's record).
## 6. The gate plan: the family analysed as a family
PENDING: the attack board's shape (F1 to F10) and the in-house pass's nine lanes run over the parameter bands, not per era: each row's harness takes the era draw as an input and the census walks the band's corners and 64 random eras; the testnet period as the window; the known-failed test per layer (layer 1: a wired-mixer stand-in at an x16 era reads 0 of 32 lanes; layer 2: a stale-state hasher at a larger state reads 0 of 32; layer 3: a kernel without the live family refuses at packcheck; layer 4: the per-load class's candidate 0 is refused by the generalised rule, and a passing era's census equals the hand census).
## 7. What a fully general chip still gets
PENDING: the identity's numbers per layer (the research file's section 2), the capex rows (section 16), the honest line.
## 8. Unverified and owed
- The 5070 Ti tier: no card owned; the hash lane's rented row or the 5070's row scaled (approximate).
- The x16 mixer: the verifier and build rows, or the estimate.
- The op-mix band: the two re-weighted packs, or the microbench arithmetic.
- Every chip figure is the model's; no chip has been measured.

View file

@ -17,3 +17,5 @@ docs/analysis/ci-failures-2026-10-06.md
docs/analysis/mission
# 7 October 2026 (night): the Counter ASIC 4.0 research record (an internal research document: the founder's words, box names, lane records)
docs/analysis/counter-asic-4-research.md
# 8 October 2026: the class v6 design record (an internal research document: the founder's words, lane records)
docs/design/class-v6-rotating-family.md