Counter ASIC 2.0 status 20:55: layers 6 and 7, node agent, 0.3.11 scope with proving v1
This commit is contained in:
parent
b681a23914
commit
65f47ba38b
2 changed files with 16 additions and 2 deletions
|
|
@ -69,6 +69,12 @@ If any gate fails: stop at that gate, write why in `docs/plans/counter-asic-2-st
|
|||
|
||||
0.3.11 through the shipper's pipeline (`tools/ship-app.mjs`, the plan shape of `docs/plans/release-0.3.10.md`). The 0.3.10 shipper finishes 0.3.10 first, then takes the 0.3.11 tree, or the coordinator ships 0.3.11 with the same runbook if the shipper has stopped. The publish order is that of `docs/plans/finality-v3-devnet-publish.md`: the override field, the digest handshake, hand nodes and the seed first, then the manifest with `--activation-height` and the deadline note, then the apps, then the digest sweep, then the HiveOS package republish.
|
||||
|
||||
## 8a. Proving v1 rides with it
|
||||
|
||||
the project lead delegated the proving v1 decisions to its agent (acd4f36bc2c07a4e2): fork branch proving-v1 (from a24ab01a: segment records, the pool split, five RPCs, p2p message 72, parameters proving_v1_activation_daa, segment_blocks, unproven_daa, aggregator_share_bps, in the digest only once set) and its app branch, rebased onto the 0.3.10 tips. 0.3.11 carries program_class_v3 AND proving v1 together: one override object, one digest, one publish, the same gates for each half (its fast-time harness green on the final tree, suites on PC 2); both activation heights set at publish by the same rule (tip + 14,400, checked >= 10,800). If one half is not ready when the other is, the ready half ships as 0.3.11 and the other as 0.3.12; the status file says which.
|
||||
|
||||
Next-cut list (not consensus, not in 0.3.11 unless a one-file app change with tests): ota-k2 (second update key with revocation), rig-install, pool-v0 (its own service, no app change), repro-bench; fork-side pack-loop 05ef0fa3 is merged into the v3 node branch because v3 touches the same miner paths.
|
||||
|
||||
## 9. After the publish: Counter ASIC 3.0
|
||||
|
||||
The ASIC-history agent sends its ranked additions; they are measured the same way, folded into the class as v4 behind its own activation, same gates, same rollout; `docs/plans/counter-asic-3.md`.
|
||||
|
|
@ -81,6 +87,6 @@ The ASIC-history agent sends its ranked additions; they are measured the same wa
|
|||
| The scratch share and size (layer 3) | 0, 12.5, 25, 50% RMW at 32 or 128 KB per warp | `<from the readwidth table and docs/analysis/scratch-soundness.md>` | the same |
|
||||
| The hot table size (layer 5) | 32, 64, 96 MB | `<from docs/plans/hot-table.md>` | |
|
||||
| The era draw bounds (layers 4 and 8) | the draw as written in `docs/plans/era-layout.md` | | |
|
||||
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | `<from docs/analysis/sram-mirror.md>` | |
|
||||
| Layer 7 | reserved family, unlock by era height or 90% signal | `<from docs/analysis/int8-matrix-family.md>` | |
|
||||
| The cache schedule (layer 6) | flat 256 MiB or a growth schedule | Recommended C: the cache doubles when the dataset doubles (512 MiB at year 4); not delegated, a gate 1 decision for the project lead | `docs/analysis/sram-mirror.md`: a 256 MiB mirror is 54 mm^2 at N2 (0.0175 um^2 cell, array factor 0.70), about $26 per good die, approximate |
|
||||
| Layer 7 | reserved family, unlock by era height or 90% signal | Reserve R1 = mm8 (uint8 8x16 by 16x8 tile per unit), W_new 4, unlock era 4 or 90% signal; switched off, no consensus effect tonight | `docs/analysis/int8-matrix-family.md`: native on PTX mma.sync, AMD WMMA iu8, Metal 4 matmul2d; dot4 emulation on Apple 1.6x (unsigned) |
|
||||
| The activation height N4 | the rule of section 3 | | |
|
||||
|
|
|
|||
|
|
@ -65,3 +65,11 @@ Probe ceilings from round 1 (dependent reads per second at 1024 MiB, device time
|
|||
the project lead's width rule applied to the ceilings alone (the measured v3 rates will replace this when the table lands): a 128-load hash at 64 B reads 8,192 B; at the 5090's 64 B ceiling that is 71 MH/s and 584 GB/s, 37% of its stream bandwidth, over the one-third margin the rule sets; at 16 B it is 2,048 B per hash, 141 MH/s and 288 GB/s, 18%, inside the margin; 4 B is 9%. On the 9070 XT every width costs the same line fetch (2.4 G/s), so 16 B is where AMD gains 4x the bytes per hash at no cost and the 5090 stays latency-bound with margin. Provisional width under the rule: 16 B (w16), pending the measured rates and the latency-bound share per card.
|
||||
|
||||
Added deliverable (20:40): the public description in four levels, `docs/plans/counter-asic-2-public.md` (aec53bb): levels 1 and 2 are written as copy; level 3 carries the bench table with `[owed]` markers for every number not yet measured; level 4 lists the documents. The integration branch applies levels 1 to 3 to `site/index.html`, `site/litepaper.html` and `site/bench.html` with the final numbers.
|
||||
|
||||
## 20:55 layers 6 and 7 landed; the node-fork agent started; 0.3.11 scope
|
||||
|
||||
Branch ca2-analysis (5d5ba15, f59708d). Layer 6: no cache growth rule exists in the spec; cited bit cells N7 0.027, N5 / N3E / Intel 18A 0.021, N3B 0.0199, N2 0.0175 um^2, array factor 0.70; the 256 MiB mirror is 83 / 64 / 54 mm^2 at N7 / N5 / N2 (74 at N2 with a 96 MB hot table), $13 to $26 of silicon per good die (approximate). The mirror was never unaffordable; the cache's job is to stay above GPU L2 (5090 96 MB, GB202 128 MB). Recommendation C for the project lead (gate 1): cache doubles when the dataset doubles. Layer 7: Metal 4 matmul2d has int8 x int8 -> int32, so a unit-level mm8 tile is native on all three vendors; per-lane dot4 is emulation on Apple (M5 Max: 548 G unsigned dot4/s against an 880 G ALU chain, 1.6x; signed 4.7x). Reserve R1 = mm8, W_new 4, unlock era 4 or 90% signal. Owed: the 5090 and 9070 XT dot4 probe (job prepared: relay/playbooks/dot4-probe.ps1, exe sha256 5adaeb1a...6416f4; publishes when a PC frees).
|
||||
|
||||
Node-fork agent a3f505a9d981300cd started 20:50: ca2-v3 (igneum-pow seam: ProgramClass, Epoch::from_chain_seeds, generator 3 in the program id, class in the pack) and ca2-v3-node (vendor/igneum-node-ca2 under the ca2-v3 worktree, from 21d4c73c, with pack-loop 05ef0fa3 merged): the field in Params, OverrideParams and the digest, the epoch-boundary rounding, the era stand-in, the job line, the fast-time gate script.
|
||||
|
||||
0.3.11 scope (coordinator, 20:52): carries program_class_v3 and proving v1 together (one override object, one digest, one publish; each half under the same gates); the ready half ships as 0.3.11 and the other as 0.3.12 if one lags. Next-cut list recorded in the rollout plan section 8a.
|
||||
|
|
|
|||
Loading…
Reference in a new issue