Counter ASIC 3.0 status: item 1 closed (verdict over 2x, the f = 1 chip), item 8 (latency-shadow program work) opened
This commit is contained in:
parent
7b32a4e80b
commit
216a4b90c3
1 changed files with 27 additions and 3 deletions
|
|
@ -16,15 +16,39 @@ The test every result is judged against (the project lead, 6 October): a chip ma
|
||||||
|
|
||||||
| # | Item | Worker branch | State | Close |
|
| # | Item | Worker branch | State | Close |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| 1 | Partial-store chip and the time-memory curve | ca3-analysis | running | |
|
| 1 | Partial-store chip and the time-memory curve | ca3-analysis | CLOSED, merged (71df794) | verdict OVER 2x: the f = 1 chip (the dataset stored in DRAM, nothing recomputed) is 5.1x per joule on GDDR7 and 7.5x to 9.2x on HBM3 in the model, 2.1x to 4.8x by the Ethash precedent, $2.8 per MH/s against the 5090's $14.7; the curve is monotone toward f = 1, so the partial-store chip is never built; the mixer and item 2 do not touch it; chip-model-v3.md section 5 |
|
||||||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | running (Mac first; 5090 job after PC 2 clears) | |
|
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | running (Mac first; 5090 job after PC 2 clears) | |
|
||||||
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | running | |
|
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | running | |
|
||||||
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
|
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
|
||||||
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | running (Mac first; 5090 job after PC 2 clears) | |
|
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | running (Mac first; 5090 job after PC 2 clears) | |
|
||||||
|
| 8 (added by item 1's finding, coordinator 08:xx UTC) | Program work in the latency shadow: the hash rate, watts and verifier cost at N = 50,000, 100,000, 200,000 ops per hash on the M5 Max and the 5090; the 5090 power-cap rows | ca3-shadow | running (Mac first; 5090 job after PC 2 clears) | |
|
||||||
|
|
||||||
## 3. Measured numbers
|
## 3. Measured numbers
|
||||||
|
|
||||||
(filled at each close)
|
### Item 1 (analysis, no new measurement; every chip figure is arithmetic on cited memory figures, approximate where marked)
|
||||||
|
|
||||||
|
| Row | Rate against the 5090's 136.1 MH/s | Per joule against 2.40 microjoules per hash | $ per MH/s | Binding |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| f = 0 on-die recompute chip (sections 1 to 3; 9,360 ops per item hoisted, 10,512 unhoisted) | 0.31x bare, 0.92x with the 3x factor (0.27x / 0.82x unhoisted) | 1.86x (1.75x unhoisted; 1.3x to 2.4x over the on-die read energy) | $16.8 | compute |
|
||||||
|
| f = 1 on GDDR7 (the 5090's own memory system, 21.3 G reads/s activate ceiling) | 1.22x | 5.1x | $2.8 | memory |
|
||||||
|
| f = 1 on one HBM3 stack | 0.61x | 7.5x | $6.6 | memory |
|
||||||
|
| f = 1 on eight HBM3 stacks | 4.9x | 9.2x | $4.0 | memory |
|
||||||
|
| f = 0.25 to 0.75, both memories | between the ends and worse than both on $ per MH/s | | | |
|
||||||
|
|
||||||
|
The whole case for the f = 1 chip: the 5090's memory system draws about 55 W of its 326 at the hash (17 percent, approximate); the rest is the GPU spinning on loads at 0.15 percent of its integer budget. The op count per mixer application, counted from `memhard.rs`: 144 as written, 128 with the RC and rk adds hoisted; 72 x 128 + 144 = 9,360 per item, which is the spec's "about 130" x 72 exactly; the chip model's rows stand at 9,360 and are given at 10,512 beside them.
|
||||||
|
|
||||||
|
What moves the f = 1 rows (chip-model-v3.md section 5.7): not the dataset size (one HBM3 stack holds 24 GB, the schedule reaches 4 GiB at year 4), not the read width (the decision to stay at 4 B stands), not the chain length; only (a) the honest card's watts at the hash (a 5090 holding 136 MH/s at a 250 W cap reads 3.9x on GDDR7, at 200 W 3.2x) and (b) program work in the latency shadow (the card hides 512 ops per hash behind 128 reads and could hide about 330,000 before compute binds; at N = 330,000 the model reads 1.85x at a chip core as efficient as the GPU's ALU, 2.5x at 1.5x worse), which is the design item this run adds (item 8 below) and the one that answers the project lead's test directly: a chip must carry the memory system AND the ALU budget.
|
||||||
|
|
||||||
|
### Consequences per tier (item 1)
|
||||||
|
|
||||||
|
| Tier | Meaning | Being done |
|
||||||
|
|---|---|---|
|
||||||
|
| Home miner, 8 or 12 GB card | Nothing today (no chip exists; a 28 nm controller project is $5M to $30M and about 32 months by the Ethash precedent); when one lands it runs 0.3 to 0.5 microjoules per hash against 10 to 20 for this tier (approximate), the first tier out | the power-cap rows and the item 8 measurement; the detector (item 4) is what tells this miner a chip has arrived |
|
||||||
|
| 16 GB AMD (9070 XT) | 7x worse per joule than the 5090 and 30x worse than the chip (approximate); a chip ends AMD home mining first | vendor-share metric (item 7); nothing in the hash fixes AMD's 2.4 G dependent reads/s |
|
||||||
|
| 24 or 32 GB (5090, M5 Max) | the honest best at 2.40 microjoules; 5x to 9x behind the chip in the model, 2x to 5x by the precedent | a 200 W cap on the 5090, if the rate holds, halves the gap |
|
||||||
|
| Rig | per joule it is its cards; at $2.8 against $14.7 per MH/s a chip fleet is cheaper per dollar too | the issuance trigger: bounty and benchmark live before daily issuance crosses about $50K (item 4b) |
|
||||||
|
| Pool user | a chip fleet is a few operators; Monero's was found at 85 percent by its share pattern | the detector, before the public testnet |
|
||||||
|
| The public claim "under 2x" | held for the recompute chip at the op budget; per joule and against the stored-dataset chip the model reads over 2x on both memories and the precedent reads 2.1x to 4.8x; NOT SAFE TO PUBLISH as worded | decision for the project lead (section 6); nothing on the site or the devnet changes from this run |
|
||||||
|
|
||||||
## 3a. Corrections found by the run
|
## 3a. Corrections found by the run
|
||||||
|
|
||||||
|
|
@ -43,7 +67,7 @@ The test every result is judged against (the project lead, 6 October): a chip ma
|
||||||
|
|
||||||
## 6. Decisions for the project lead
|
## 6. Decisions for the project lead
|
||||||
|
|
||||||
(filled at the close)
|
1. The public claim. "Under 2x" is true of the recompute chip per chip at the op budget and false of the stored-dataset chip per joule (item 1). Re-word to the measured fact (the honest card runs at 1/128 of its dependent-read ceiling; a chip must out-read it per watt) until the item 8 rows land, or keep the claim with "per chip, against the recompute chip" stated. Nothing changes on the site until the project lead's word.
|
||||||
|
|
||||||
## 7. Unverified and owed
|
## 7. Unverified and owed
|
||||||
|
|
||||||
|
|
|
||||||
Loading…
Reference in a new issue