Class v6 10.0k: the connected-state lane's first rows (63 of 64 registers necessary; the window +0.6 percent on a 5090 at stock; censuses as well as v5; the narrow per-step chain names the two-level-file re-optimisation)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
7ae4890164
commit
11bb46ed28
1 changed files with 15 additions and 0 deletions
|
|
@ -475,6 +475,21 @@ The multi-family adversary lane (a1a9876a88f5a72fc; synthesis only, ASAP7 TC, a
|
|||
|
||||
The disagreement, carried as two models until the placed rows settle it: the k lane's SRAM-banked time-multiplexed window (10.0i) is modelled at 8.5 to 10.5 pJ per lane-op, NOT cheaper than its gated flop window (6.2), with the RTL form in synthesis on a pod as a check; the multi-family lane's macro window reads 5.9 (band 5.1 to 7.7) for the whole core. The two differ in the macro's energy model and the access pattern (one 256-bit access per eight lanes against a per-lane port), and in whether the window's whole state is live across every wait (10.0i: 61 of 64 values necessary, so no lane can skip the access). Until the k lane's placed gated rows (17:30 UK) and the multi-family lane's placed 8-lane core (18:30 UK) land, the served window line is read with this caveat: **the window's measured cost to the card is under 5 percent and its defence against a flop-file adversary is about 0.13 of k; against a macro-file adversary it may be close to nothing, which is why the next programme's multi-family adversary is the lane that decides whether the window stays a defence or only a cost.** The 18-family variant (+43 percent of cells on the same microarchitecture), the live-family mixes and the 32-lane cores follow by 18:30 with the board and lifetime rows.
|
||||
|
||||
#### 10.0k Amendment (16:56 UK): the connected-state lane's first rows (for class v7; branch class-v6-connected, `docs/analysis/class-v6/connected-state.md` by 17:45 UK)
|
||||
|
||||
The variant cs64s27x16 (a research class): a 64-register window per lane; per step one 4-byte load from a window register, then 27 distinct instructions run 16 passes, the next address written by the block's last instruction from a fresh spine (add, sub, xor, shfl only), lossy ops feeding the next injection, every register covered by an injecting write, the result folding all 64 registers; 128 loads and 55,296 ALU per hash (class v5 128 and 55,680), text 448 per iteration (v5 320); no long program, select tree, wide read or scratchpad.
|
||||
|
||||
| Reading | The numbers | Label |
|
||||
|---|---|---|
|
||||
| Liveness (seed igneum-v6c/0) | 63 of 64 registers necessary for a later address at every address until the last iteration's tail; the result reads 64; an address depends on 1 to 15 registers (mean 11.2) of the previous address point; a block touches 16 to 24 registers; reads per register 384 to 3,208 per hash | measured |
|
||||
| What that means | the window is live state a specialist must hold, but the per-step chain is narrow (about 11 of 64), so the obvious re-optimisation is a two-level file, about 20 hot per step | the lane's reading |
|
||||
| Census (the sub-version 3 harness, (c''') on) | no era 256 of 256, 0 exhausted, 0.28 rejections per candidate, mean attempt 0.39 (the control mx8+sh256x27: 0.67, 2.0); eras 0 to 7 with the window-bit refusal 256 of 256, 0 exhausted, mean attempt 1.7 (the control 4.9), 306 window-bit refusals (the control 218); the F8 form at 16 x 2^20 0.999 to 1.002 of uniform (the control the same) | measured |
|
||||
| What that means | it censuses at least as well as v5; it carries v5's product-bit class, which layer 1's index fold removes for both | the lane's reading |
|
||||
| The card at stock (a rented 5090 at the 575 W cap, the class v5 nvcc harness, 250 x 2^24, the same card for both) | cs64 64.93 MH/s at 574.8 W against v5-genesis 65.30 at 574.8 W: +0.6 percent of energy per hash for the window; ptxas 80 registers per thread (v5 48), 0 spills, the same occupancy | measured |
|
||||
| What that means | the window is nearly free on the card at stock (the budget 10 percent); the 4090 row by 17:15, the PC 1 lock row owed (PC 1 booked to 17:40); the chip score waits on the k lane's re-optimised pJ and k; KEEP or KILL at 18:00 | owed |
|
||||
|
||||
For the window line of the close: this is a second measured card row (after 10.0e's) that the window costs the card under 1 to 5 percent, and a second liveness read (63 of 64) that a specialist must hold the state; the narrow per-step chain (about 11 of 64) is the first sign of the re-optimisation 10.0i and 10.0j are arguing over, a two-level file, and the chip score at 18:00 is what the window's defence is worth against it.
|
||||
|
||||
#### 10.0d The rotation schedule the close adopts (the rotation lane, `docs/design/class-rotation-four-layers.md` on class-v6-rotation at bd43f808, build-3, gate green, 14:4x UK; one line per layer; both of this document's constraints held: the 180-day family epoch not shorter, W = 4 not drawn)
|
||||
|
||||
| Layer | Boundaries a year | What it draws, from where | Exposure per boundary (this document's units) | Chip | Label |
|
||||
|
|
|
|||
Loading…
Reference in a new issue