Counter ASIC 3.0 status: item 2 interim (GO as reserve, NO-GO genesis-live at dr736), PC 2 clear about 08:35Z
This commit is contained in:
parent
1c08438a12
commit
f2e79316fe
1 changed files with 12 additions and 1 deletions
|
|
@ -17,7 +17,7 @@ The test every result is judged against (the project lead, 6 October): a chip ma
|
|||
| # | Item | Worker branch | State | Close |
|
||||
|---|---|---|---|---|
|
||||
| 1 | Partial-store chip and the time-memory curve | ca3-analysis | CLOSED, merged (71df794) | verdict OVER 2x: the f = 1 chip (the dataset stored in DRAM, nothing recomputed) is 5.1x per joule on GDDR7 and 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent, $2.8 per MH/s against the 5090's $14.7; the curve is monotone toward f = 1, so the partial-store chip is never built; the mixer and item 2 do not touch it; chip-model-v3.md section 5 |
|
||||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | running (Mac first; 5090 job after PC 2 clears) | |
|
||||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | INTERIM merged (acb96ee, dd5041b); 5090 job queued on PC 2 | GO as reserve R0; NO-GO for genesis-live at dr736 (verifier 4.9 ms per unit on one M5 Max core, about 12 ms on a 2019-class core by the 2.5x rule, over the gate; dr368 2.69 ms passes both); urgency LOW after item 1 (the stored-dataset chip derives no item) |
|
||||
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | CLOSED, merged (c0642af, merge ed06814) | detector.mjs + 7 of 7 tests + one observer hook, dry run quiet on the devnet (max correlation 0.53 against the 0.8 edge; excess spread 0 to 5.2 percent); funding.md rule 5: bounty escrowed and benchmark live before daily issuance crosses USD 20,000 a day; epoch-length.md sections 11 and 12: signal trigger M = 6 windows; FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the bank-bound ceiling (unmeasured); layer 9 ranks above layer 7. The live observer is NOT restarted yet: the write path is untested; one restart after item 7's hook merges, then the first live detector row is recorded here |
|
||||
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
|
||||
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | running (Mac first; 5090 job after PC 2 clears) | |
|
||||
|
|
@ -39,6 +39,17 @@ The whole case for the f = 1 chip: the 5090's memory system draws about 55 W of
|
|||
|
||||
What moves the f = 1 rows (chip-model-v3.md section 5.7): not the dataset size (one HBM3 stack holds 24 GB, the schedule reaches 4 GiB at year 4), not the read width (the decision to stay at 4 B stands), not the chain length; only (a) the honest card's watts at the hash (a 5090 holding 136 MH/s at a 250 W cap reads 3.9x on GDDR7, at 200 W 3.2x) and (b) program work in the latency shadow (the card hides 512 ops per hash behind 128 reads and could hide about 330,000 before compute binds; at N = 330,000 the model reads 1.85x at a chip core as efficient as the GPU's ALU, 2.5x at 1.5x worse), which is the design item this run adds (item 8 below) and the one that answers the project lead's test directly: a chip must carry the memory system AND the ALU budget.
|
||||
|
||||
### Item 2, the Mac rows (interim, ca3-derive dd5041b; measure lock, load average 4.5 to 5.3; the 5090 rows queued on PC 2; the 9070 XT OWED)
|
||||
|
||||
| Class | Ops per item (chip, RC + rk hoisted / GPU as written / multiplies) | Verifier ms per unit, one M5 Max core (worst cold) | 2019-class core, 2.5x rule, approximate | Metal 1 GiB build | Hash rate, M5 Max | Bit-exact | Chip row (bare / at a 1.2x / 1.5x allowance; the 3x no longer applies) |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| v2 (x1) | 1,152 / 1,296 / 144 | 0.598 | 1.5 | | | | 2.45x |
|
||||
| v3 (x8), live | 9,216 / 10,368 / 1,152 (the acceptance floors) | 2.061 to 2.063 | 5.2 | 22.1 ms | 27.1 MH/s | yes | 0.31x / 0.92x with the 3x factor |
|
||||
| dr368 (9 x 368 instructions) | about half of dr736 | 2.69 | 6.7, passes | | | | |
|
||||
| dr736 (9 x 736 instructions, the genesis-day draw: 9,992 / 10,659 / 1,461) | | 4.875 to 4.944 (5.241) | 12, over the gate | 29.0 ms | 27.1 MH/s (equal) | yes (Metal and Apple OpenCL, fingerprint 50e3eaa779da4f1e) | 0.29x / 0.34x / 0.43x |
|
||||
|
||||
Consequence: the derivation costs no hash rate on any card and 7 ms a day of build on the Mac (29 against 22 ms; the 5090 and the 9070 XT rows owed, the loaded-iGPU tier is the one to watch); it costs the verifier, and the verifier budget is shared with item 8's shadow ops under the one 10 ms gate, so the class v4 candidate is the pairing that fits, not either lever alone (both workers told).
|
||||
|
||||
### Item 6, the Mac rows (interim, ca3-reserve 192a683; the 5090 job waits on PC 2)
|
||||
|
||||
Step cost per family on the M5 Max as a ratio to the add-xor-rotate chain (881 G steps/s; best of 3; load average 7.64; all bit-exact): shl 0.85, shr 0.86, bfe 0.77, andn 0.75, byte permute 1.13 (emulated on Apple), popcount 0.87, clz 1.01, select 0.76, indexed shuffle 1.91 (2.2x the xor shuffle), live rotr 1.13, dot4 unsigned 1.60, dot4 signed 4.73. Every 32-bit datapath family costs an Apple lane under 1.2x a step, inside the 8x emulation bound of 1.13.2 with room; the matrix family is the only one past 1.6x. The rows feed item 8's ALU pricing. The full table with consequences lands at item 6's close.
|
||||
|
|
|
|||
Loading…
Reference in a new issue