Counter ASIC 3.0 status: the close (items 6, 7 and 8 closed; the chip model before and after; the class v4 candidate mx8+sh256x27 ready for the gate run; nine decisions for the project lead; the owed list)
This commit is contained in:
parent
f4478a8ee1
commit
4bdf9af4c3
1 changed files with 74 additions and 7 deletions
|
|
@ -1,6 +1,6 @@
|
|||
# Counter ASIC 3.0: status
|
||||
|
||||
Started 6 October 2026, 07:15 UTC, on the project lead's order "run it all". Coordinator worktree `igneum-wt-ca3-coord`, branch `ca3-coord` from `50df751` (ca2-coord `9206aea` merged with master `ddfcdac`). The plan: `docs/plans/counter-asic-3.md`. The audit behind it: `docs/analysis/asic-resistance-history.md`. The 2.0 record: `docs/plans/counter-asic-2-status.md`, `docs/bench-log.md` "Counter ASIC 2.0, the numbers", `docs/analysis/chip-model-v3.md`. Class v3 is live on the devnet since DAA 154,800 (crossing 03:51:42Z, verdict PASS 05:21:48Z). Nothing in this run publishes to the devnet; what passes is a class v4 candidate behind `program_class_v4_activation_daa` under the six gates and the two-publish rollout of 2.0.
|
||||
Started 6 October 2026, 07:15 UTC, on the project lead's order "run it all"; closed 09:5x UTC with every item measured on the Mac and the 5090 and every AMD row owed (PC 1 not released). Final tree: ca3-coord, 45 commits over the base, `no-conflict-markers.sh` clean, observer tests 10 of 10, `igneum-pow` 96 of 96 with the pinned v2 and v3 packs byte for byte. Coordinator worktree `igneum-wt-ca3-coord`, branch `ca3-coord` from `50df751` (ca2-coord `9206aea` merged with master `ddfcdac`). The plan: `docs/plans/counter-asic-3.md`. The audit behind it: `docs/analysis/asic-resistance-history.md`. The 2.0 record: `docs/plans/counter-asic-2-status.md`, `docs/bench-log.md` "Counter ASIC 2.0, the numbers", `docs/analysis/chip-model-v3.md`. Class v3 is live on the devnet since DAA 154,800 (crossing 03:51:42Z, verdict PASS 05:21:48Z). Nothing in this run publishes to the devnet; what passes is a class v4 candidate behind `program_class_v4_activation_daa` under the six gates and the two-publish rollout of 2.0.
|
||||
|
||||
The test every result is judged against (the project lead, 6 October): a chip maker must have to build a better GPU than NVIDIA to launch an ASIC, the way Bitmain's Antminer X5 reached CPU parity per joule only at many times the price. The binding rule from 2.0: the hash stays bound by dependent random memory reads.
|
||||
|
||||
|
|
@ -9,7 +9,7 @@ The test every result is judged against (the project lead, 6 October): a chip ma
|
|||
| Machine | State | Rule |
|
||||
|---|---|---|
|
||||
| Mac M5 Max | mining paused; Metal worker free | measurements under `with-lock.sh measure` only |
|
||||
| PC 2 (1ccfe586, RTX 5090) | CLEAR at 08:24:27Z (the clear file carries the proving agent's constraints: prover left ON, no quit or restart, /opt/igneum, /opt/igneum-segal and settings.json untouched); the three 3.0 jobs run one at a time through the mkdir lock (items 8, 2, 6); "PC 2 released" to the proving agent after the last; the prover-floor agent next. Before that: the proving agent's jobs `segments-pc2-pv1` runs b and c (run b claimed nothing on a PowerShell key bug; run c from about 07:53Z); "PC 2 clear" expected about 08:35Z; then three 3.0 jobs one at a time (items 2, 6, 8); the prover-floor agent queues after "PC 2 released" | nothing published until "PC 2 clear"; one job at a time (`/tmp/igneum-devnet/pc2-ca3.lock`); released by message when done |
|
||||
| PC 2 (1ccfe586, RTX 5090) | RELEASED at 08:43:24Z after the three jobs (derive 08:26 to 08:29Z, shadow 08:29 to 08:41Z, family 08:41 to 08:43Z, all exit 0; prover ON, app untouched); then the prover-floor agent, then the 0.3.12 release engineer's build. Before that: CLEAR at 08:24:27Z (the clear file carries the proving agent's constraints: prover left ON, no quit or restart, /opt/igneum, /opt/igneum-segal and settings.json untouched); the three 3.0 jobs run one at a time through the mkdir lock (items 8, 2, 6); "PC 2 released" to the proving agent after the last; the prover-floor agent next. Before that: the proving agent's jobs `segments-pc2-pv1` runs b and c (run b claimed nothing on a PowerShell key bug; run c from about 07:53Z); "PC 2 clear" expected about 08:35Z; then three 3.0 jobs one at a time (items 2, 6, 8); the prover-floor agent queues after "PC 2 released" | nothing published until "PC 2 clear"; one job at a time (`/tmp/igneum-devnet/pc2-ca3.lock`); released by message when done |
|
||||
| PC 1 (ae432dc7, RTX 5090 + RX 9070 XT) | the project lead's desk | not used today; every AMD row OWED |
|
||||
|
||||
## 2. The items
|
||||
|
|
@ -20,8 +20,8 @@ The test every result is judged against (the project lead, 6 October): a chip ma
|
|||
| 2 | Per-day item-derivation program (reserve entry, verifier gate, daily build) | ca3-derive | CLOSED, merged (acb96ee, dd5041b, bdc07d3, 54bc188); 5090 job run-ca3-derive-pc2-20261006 exit 0 in 126 s | GO as reserve R0; NO-GO for genesis-live at dr736 (verifier 4.9 ms per unit on one M5 Max core, about 12 ms on a 2019-class core by the 2.5x rule, over the gate; dr368 2.69 ms passes both); urgency LOW after item 1 (the stored-dataset chip derives no item) |
|
||||
| 4 + 5 | Share-pattern detector, trigger rules, FPGA lane, layer 9 against 7 | ca3-detector | CLOSED, merged (c0642af, merge ed06814) | detector.mjs + 7 of 7 tests + one observer hook, dry run quiet on the devnet (max correlation 0.53 against the 0.8 edge; excess spread 0 to 5.2 percent); funding.md rule 5: bounty escrowed and benchmark live before daily issuance crosses USD 20,000 a day; epoch-length.md sections 11 and 12: signal trigger M = 6 windows; FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the bank-bound ceiling (unmeasured); layer 9 ranks above layer 7. The live observer is NOT restarted yet: the write path is untested; one restart after item 7's hook merges, then the first live detector row is recorded here |
|
||||
| 3 | Cryptanalysis brief in funding.md | ca3-crypto-brief | CLOSED, merged (43c3ead) | funding.md line item USD 80k to 160k, reviewer shortlist, ranked break list; verdict GO to commission (no outreach, no spend) |
|
||||
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | merged to 6422b6f (merge fcce185); the 5090 family job queued on PC 2 | proposed reserve order R1 byte permute, R2 popcount and clz, R3 indexed shuffle, R4 bit-field extract, R5 variable shifts, R6 select, R7 andn, R8 mm8 (`docs/plans/counter-asic-3-reserve.md`, a decision for the project lead, not in the spec); vendor-share.mjs with tests and one observer hook; repro.md carried from repro-bench (dda9fa3) with section 8: today's devnet NVIDIA 0.986 of blue blocks, Intel 0.014, coverage 1.00 at 135.8 MH/s with PC 1 off |
|
||||
| 8 (added by item 1's finding, coordinator 08:xx UTC) | Program work in the latency shadow: the hash rate, watts and verifier cost at N = 50,000, 100,000, 200,000 ops per hash on the M5 Max and the 5090; the 5090 power-cap rows | ca3-shadow | running (Mac first; 5090 job after PC 2 clears) | |
|
||||
| 6 + 7 | Reserve order with step costs, vendor-share metric | ca3-reserve | CLOSED, merged (b85f0f1, merge 820f4d4); 5090 job run-ca3-family-pc2-20261006 exit 0 in 36 s, the card to itself | proposed reserve order R1 byte permute, R2 popcount and clz, R3 indexed shuffle, R4 bit-field extract, R5 variable shifts, R6 select, R7 andn, R8 mm8 (`docs/plans/counter-asic-3-reserve.md`, a decision for the project lead, not in the spec); vendor-share.mjs with tests and one observer hook; repro.md carried from repro-bench (dda9fa3) with section 8: today's devnet NVIDIA 0.986 of blue blocks, Intel 0.014, coverage 1.00 at 135.8 MH/s with PC 1 off |
|
||||
| 8 (added by item 1's finding, coordinator 08:xx UTC) | Program work in the latency shadow: the hash rate, watts and verifier cost at N = 50,000, 100,000, 200,000 ops per hash on the M5 Max and the 5090; the 5090 power-cap rows | ca3-shadow | CLOSED, merged (70eaf06, merge 12d8a93); 5090 job run-ca3-shadow-pc2-20261006 (lock 08:29:22 to 08:41:03Z), the card empty | GO as a class v4 candidate at `mx8+sh256x27` (100,000 ops per hash): the chip's per-joule edge over the 5090 falls from 5.6x to 2.1x at k = 1 for 0.2 percent of the 5090's rate and 1.5 percent of the Mac's; NO-GO above 130,000 ops or a block over 256 instructions; the chip question is decided by k, the chip core's energy per op against the 5090's measured 11 pJ: over 2x only at k under 0.5 |
|
||||
|
||||
## 3. Measured numbers
|
||||
|
||||
|
|
@ -69,6 +69,41 @@ The verifier is not the constraint: 3.2 microseconds per 1,000 shadow instructio
|
|||
|
||||
Consequence per tier (interim): the M5 Max at 21 W is the per-joule best honest miner we own (0.78 against the 5090's 2.34 microjoules), so against Apple silicon the stored-dataset chip of item 1 reads 1.7x per joule on GDDR7 (0.47 against 0.78) and 2.4x on one HBM3 stack, inside the "1 to 2x" band on GDDR7; per pound the Mac stays the worst tier (list price against 27 MH/s). Filling the shadow to about 100,000 ops costs an Apple miner 16 W more for 1.5 percent of rate (0.78 to 1.40 microjoules) and buys 1.4x against a chip at k = 1; the N the project can pick without any card we own losing over 5 percent is set by the Mac at about 130,000 until the 5090 and 9070 XT rows land. An NVIDIA owner's number is the PC 2 job; an AMD owner's is owed.
|
||||
|
||||
### Item 8, the 5090 rows and the chip side (ca3-shadow 70eaf06; job run-ca3-shadow-pc2-20261006, the card confirmed empty by nvidia-smi compute-apps and the process list, the installed worker's `--bench`, nvidia-smi at 1 Hz, the app's 431 W cap; every pack bit-exact on Metal, CUDA, clang emulation and Apple OpenCL)
|
||||
|
||||
| N ops per hash | M5 Max MH/s (delta) | M5 Max W, microjoules | RTX 5090 MH/s (delta) | 5090 W, SM MHz, microjoules | Verifier ms per warp, one M5 Max core |
|
||||
|---|---|---|---|---|---|
|
||||
| 930 (class v3 today: 512 instructions, 384 ALU, about 1.83 counted ops each) | 27.08 | 21.0, 0.78 | 132.2 | 350, 3,040, 2.65 | 2.06 |
|
||||
| 49,700 | 26.75 (-1.2%; a 64-instruction block +2.5%) | 31.1, 1.16 | 132.3 (+0.1%; 64-block +3.5%) | 425, 3,034, 3.21 | 2.21 |
|
||||
| 102,100 (`sh256x27`, the candidate) | 26.67 (-1.5%) | 37.2, 1.40 | 131.95 (-0.2%) | 431 cap, 2,824, 3.27 | 2.23 |
|
||||
| 150,800 | 25.10 (-7.3%) | 36.7, 1.46 | 131.75 (-0.3%) | 431, 2,427, 3.27 | 2.33 |
|
||||
| 199,600 | 24.25 (-10.4%) | 38.2, 1.58 | 128.67 (-2.7%) | 431, 1,753, 3.35 | 2.43 |
|
||||
| 330,700 | 21.39 (-21%) | 40.1, 1.88 | 86.39 (-35%, compute-bound at the capped clock, 28.6 T op/s) | 431, 1,834, 4.99 | 2.62 (worst cold 2.77) |
|
||||
|
||||
Where each card leaves the latency bound (the 5 percent rule): M5 Max about 130,000 ops; RTX 5090 about 210,000 at its 431 W cap (the cap binds from 102,100 ops and the governor lowers the clock); RX 9070 XT OWED. The verifier's law: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp, so 330,700 ops add 0.56 ms (about 1.4 ms on a 2019-class core): it fits every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none), and the node never binds before the cards. Marginal ALU energy per counted op: the 5090 10 to 13 pJ at its shipping clock (twice the 5.5 pJ the item 1 model assumed), the M5 Max 6.9 pJ. Power caps: `-pl` below 400 W cannot be set, the 5 October sweep rows re-used (302 to 316 W at every cap); the clock rows are OWED (nvidia-smi refused `-lgc` without rights; the job did not ask; an elevated job is a decision for the project lead, section 6). Item 1's denominator: the 5090 at the hash is 290 W in the app and 350 W in the bench, not 326; the f = 1 rows move 2 to 11 percent.
|
||||
|
||||
Chip side (approximate: the f = 1 stored-dataset chip of item 1 plus an ALU core at k times the 5090's measured 11 pJ per op): at N = 100,000 its per-joule edge over the 5090 falls from 5.6x (shadow empty, 350 W) to 2.1x on GDDR7 and 2.3x on one HBM3 stack at k = 1, 1.5x at k = 1.5, 3.2x at k = 0.5, 4.1x at k = 0.3; over the M5 Max from 1.6x to 0.9x at k = 1. The chip needs a 14,000-lane ALU array (about 30 mm^2 at N5, 150 W at k = 1) at 100,000 ops and 22,600 lanes (60 to 100 mm^2, 500 W) at 331,000; at 28 nm the array is reticle-class. That is the project lead's test as a number: with the shadow filled the chip must carry the memory system and a GPU-class datapath, and its edge is the ratio k of its datapath's energy per op to the GPU's.
|
||||
|
||||
Consequences per tier at N = 100,000: an Apple miner loses 1.5 percent of rate and pays 16 W more (0.56x per watt, per pound unchanged); a 5090 loses 0.2 percent and goes from 350 to 431 W (0.81x per watt), so a 5090 rig pays about 23 percent more electricity for the same blocks; a pool user sees nothing; a small NVIDIA card (4060 class) binds near 100,000 by its ALU budget (model, owed); AMD holds by its budget (owed); every verifier tier is untouched (+0.17 ms per warp).
|
||||
|
||||
### Item 6, the 5090 rows (ca3-reserve b85f0f1; job run-ca3-family-pc2-20261006, 08:42:27 to 08:43:03Z, exit 0, the card quiet, prover off and back on; nvcc 12.8 sm_120; best of 3 runs; every row bit-exact)
|
||||
|
||||
| Family | M5 Max, Metal (ratio to the alu chain, 881 G steps/s) | RTX 5090, CUDA (ratio to the alu chain, 7,941 G steps/s) | RX 9070 XT | Native or emulated |
|
||||
|---|---|---|---|---|
|
||||
| rotr (live) | 1.13 | 1.32 | OWED | native everywhere |
|
||||
| shflx (live shuffle) | 0.86 | 1.49 | OWED | native |
|
||||
| variable shifts shl / shr | 0.85 / 0.86 | 1.27 / 1.28 | OWED | native |
|
||||
| bfe (bit-field extract) | 0.77 | 1.54 (a two-instruction sequence on NVIDIA) | OWED | native on Apple and AMD, sequence on NVIDIA |
|
||||
| andn | 0.75 | 1.26 | OWED | native |
|
||||
| perm (byte permute) | 1.13, EMULATED | 1.30 (`prmt`) | OWED | emulated on Apple |
|
||||
| popc / clz | 0.87 / 1.01 | 1.50 / 1.63 | OWED | native |
|
||||
| sel | 0.76 | 1.32 | OWED | native |
|
||||
| shfla (lane + delta) | 1.91 (2.2x the xor shuffle) | 1.53 | OWED | native; the 32-lane crossbar is the chip's cost (about 6x the xor butterfly, approximate) |
|
||||
| dot4 (comparison) | 1.60 unsigned / 4.73 signed, emulated | 1.16 | 1.06 (5 October) | |
|
||||
| mm8 (comparison, R8) | OWED (Metal 4 matmul2d) | 2.43 (`mma.m8n8k16.u8`, bit-exact) | OWED | the licensable block |
|
||||
|
||||
Reading: against the live rotr every candidate is 0.95x to 1.23x on NVIDIA; on Apple only perm (emulated) and shfla cost more than rotr; the 8x emulation bound of 1.13.2 holds everywhere by 4x or more; the 5 percent hash-rate bound at W_new = 4 is argued from the step cost (under 1 percent of ALU time on a read-bound hash), not measured, since no reserve family is live. Proposed order R1 perm, R2 popc and clz, R3 shfla, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8, W_new = 4 each, family n at era n (mm8's era-4 unlock kept as a named exception or moved to era 8: the project lead's call); the full proposed 1.13.2 text with edge vectors per family is `docs/plans/counter-asic-3-reserve.md` section 6. Consequence per tier: an Apple miner pays the emulated perm at 1.13x a step and shfla at 1.91x, under 1 percent of its hash rate at W_new = 4 (argued); an NVIDIA miner pays nothing measurable; an AMD miner's row is owed and its `ds_bpermute_b32` cost is the one number that could move R3; a chip pays a barrel shifter, a byte crossbar, a popcount tree and a 32-lane crossbar per lane, which is the point.
|
||||
|
||||
### Item 6, the Mac rows (interim, ca3-reserve 192a683; the 5090 job waits on PC 2)
|
||||
|
||||
Step cost per family on the M5 Max as a ratio to the add-xor-rotate chain (881 G steps/s; best of 3; load average 7.64; all bit-exact): shl 0.85, shr 0.86, bfe 0.77, andn 0.75, byte permute 1.13 (emulated on Apple), popcount 0.87, clz 1.01, select 0.76, indexed shuffle 1.91 (2.2x the xor shuffle), live rotr 1.13, dot4 unsigned 1.60, dot4 signed 4.73. Every 32-bit datapath family costs an Apple lane under 1.2x a step, inside the 8x emulation bound of 1.13.2 with room; the matrix family is the only one past 1.6x. The rows feed item 8's ALU pricing. The full table with consequences lands at item 6's close.
|
||||
|
|
@ -105,11 +140,29 @@ The live observer runs the shared checkout, which autosync fast-forwards from or
|
|||
|
||||
## 4. The chip model, before and after
|
||||
|
||||
(filled at the close)
|
||||
Every chip figure is arithmetic on cited memory and logic figures and is approximate; every GPU figure is measured and names its entry. "Per chip" is rate per chip against the 5090's rate; "per joule" is energy per hash, the Ethash chips' metric.
|
||||
|
||||
| Chip | Before 3.0 (the 2.0 record, 5 October) | After 3.0 (6 October) | Source |
|
||||
|---|---|---|---|
|
||||
| On-die 256 MiB cache recompute chip (f = 0), class v3 | 0.31x bare, 0.92x with the 3x fixed-function factor per chip; per joule not priced; the public claim "under 2x" rested on this row | per chip unchanged; per joule 1.86x (1.3x to 2.4x over the on-die read energy); with the per-day derivation (item 2, dr736) the 3x factor goes: 0.29x bare, 0.34x at a 1.2x allowance, 0.43x at 1.5x | chip-model-v3.md sections 2, 5.4, 6 |
|
||||
| Stored-dataset memory-controller chip (f = 1), the Ethash class | not priced (O-1.6 open, the curve never drawn) | per chip 1.22x on GDDR7, 0.61x on one HBM3 stack, 4.9x on eight; per joule 5.1x (GDDR7) to 9.2x (eight HBM3 stacks) at the 326 W denominator, 5.6x at the measured 350 W bench control; $2.8 per MH/s against the 5090's $14.7; the Ethash precedent for this class 2.1x to 4.8x per joule; the curve is monotone toward f = 1 so the partial-store chip is never built; the mixer, item 2 and x16 do not touch it | chip-model-v3.md section 5; history rows 3 and 4 |
|
||||
| The same chip with the latency shadow filled (item 8, N = 100,000 ops per hash, class v4 candidate `mx8+sh256x27`) | not a lever anyone had priced | per joule over the 5090 2.1x (GDDR7) and 2.3x (HBM3) at k = 1, 1.5x at k = 1.5, 3.2x at k = 0.5; over the M5 Max 0.9x at k = 1; the chip needs a 14,000-lane ALU array, about 30 mm^2 at N5 and 150 W at k = 1, reticle-class at 28 nm; k, the chip core's energy per op against the 5090's measured 11 pJ, decides it, and 2x is crossed only at k under 0.5 | latency-shadow-2026-10-06.md sections 6 and 8 |
|
||||
| The honest denominators | the 5090 at 136.1 MH/s and 326 W (a peak with the prover on) | the 5090 at 290 W in the app and 350 W in the bench (2.34 to 2.65 microjoules per hash), cap floor 400 W so no power cap binds; the M5 Max at 21 W GPU plus DRAM (0.78 microjoules), three times better per joule than the 5090 and the honest best we own | item 8; miner-eff's 4 October log |
|
||||
|
||||
What 3.0 did to the model in one line: the chip that matters is not the one 2.0 priced; it is the Ethash-class memory-controller chip, over 2x per joule today; the lever that answers it is program work in the latency shadow, which makes the chip carry a GPU-class datapath, and the measured candidate takes its edge from 5.6x to 2.1x at a chip core no better than the GPU's, for 0.2 percent of the 5090's rate.
|
||||
|
||||
## 5. What passes as class v4
|
||||
|
||||
(filled at the close)
|
||||
Judged on measurements against the six gates of `docs/plans/counter-asic-2-rollout.md` section 7. Nothing is published; nothing is cut.
|
||||
|
||||
| Candidate | Measured | The six gates | Verdict |
|
||||
|---|---|---|---|
|
||||
| `mx8+sh256x27`: class v3 plus a 256-instruction ALU block run 27 times per iteration, about 100,000 ops per hash (item 8) | hash rate -0.2 percent on the 5090, -1.5 percent on the M5 Max, under the 5 percent rule; verifier +0.17 ms per warp (2.23 ms, gate 10 ms, 2019-class core about 5.6 ms); bit-exact on Metal, CUDA, clang emulation and Apple OpenCL; the 5090 at 431 W (its cap) from 350; the 9070 XT owed | G1 bit-exact: NVIDIA and Apple GREEN, the AMD vendor NOT RUN (gfx1036 or the 9070 XT); G2 (1,000 random hashes per card re-hashed on the CPU): NOT RUN; G3 (crate suite green: 96 of 96; the Metal fuzz, edge and stats runs on the class): NOT RUN; G4 (the fast-time 3-node network across a v4 activation): NOT RUN; G5, G6: NOT RUN | READY FOR THE GATE RUN as THE class v4 candidate, on the project lead's word; not ready for a cut. Two design decisions ride with it: the acceptance rule does not yet interpret the block (cost stated in the analysis), and the block stays at 64 to 256 instructions |
|
||||
| `dr368` or `dr736`: the per-day derivation (item 2) | dr736 verifier 4.9 ms per unit on the M5 Max core, about 12 ms on a 2019-class core (over the gate); dr368 2.69 ms, passes both; bit-exact on Metal, Apple OpenCL and CUDA; build +7 ms a day on the Mac, +2 ms on the 5090; NVRTC +1.1 s per pack (a once-a-day module is a requirement) | not a v4 candidate: a reserve entry | GO as reserve R0 (proposed text, counter-asic-3-derivation.md section 6); NO-GO genesis-live at dr736 until O-1.14; urgency LOW (the f = 1 chip derives no item) |
|
||||
| The reserve order R1 to R8 (item 6) | step costs on two cards, every family under 1.91x a step, bit-exact | a spec ordering, no activation | GO for the order as proposed; a decision for the project lead |
|
||||
| Layer 9 (epoch length) ranked above layer 7 (mm8) (item 5) | FPGA soft overlay 0.30x to 0.39x per watt on the measured basis, 0.7x to 1.9x at the unmeasured bank-bound ceiling | reserve ranking | GO for the ranking; the rented FPGA hour owed |
|
||||
|
||||
So: one class v4 candidate, `mx8+sh256x27`, defined and measured on the hash's own numbers, with gates G1 (AMD), G2, G3 (Metal runs), G4, G5 and G6 still to run before any publish, by the two-publish rollout of 2.0 and a `program_class_v4_activation_daa` at tip + 14,400.
|
||||
|
||||
## 6. Decisions for the project lead
|
||||
|
||||
|
|
@ -129,6 +182,15 @@ Draft (b), the measured fact only: "An RTX 5090 mines this hash at 136 MH/s and
|
|||
|
||||
Either replaces "under 2x" on the site and in the litepaper once the project lead chooses; until then the claim stays as it is and this file records that it is not safe as worded.
|
||||
|
||||
2. The class v4 candidate: run the six gates on `mx8+sh256x27` (100,000 ops per hash) and, if green, cut it by the 2.0 rollout shape. Cost to the tiers: the 5090 goes from 350 to 431 W for the same blocks (a rig pays about 23 percent more electricity), the M5 Max from 21 to 37 W for 1.5 percent of rate; the verifier +0.17 ms per warp. Gain: the stored-dataset chip's per-joule edge over the 5090 falls from 5.6x to 2.1x at a chip core equal to the GPU's. Alternative: wait for the 9070 XT and the 4060-class rows (both owed) before the gates, since a small card binds near 100,000 by its ALU budget (model).
|
||||
3. The reserve order R1 perm, R2 popc and clz, R3 shfla, R4 bfe, R5 shifts, R6 sel, R7 andn, R8 mm8 (W_new = 4 each; mm8's era-4 unlock kept as the named exception or moved to era 8), and R0 the per-day derivation as a reserve family (dr368 the safe draw; dr736 after O-1.14).
|
||||
4. Commission the mixer cryptanalysis: USD 80,000 to 160,000, the brief and the reviewer shortlist in funding.md (item 3); it should name random ARX programs (item 2's class) beside M_r. No outreach has been made.
|
||||
5. The bounty trigger: escrowed and the benchmark live before daily issuance crosses USD 20,000 a day (funding.md rule 5); the detector is the clock (quiet on the devnet, max pairwise r 0.10 to 0.53 against the 0.8 edge).
|
||||
6. The live observer: the detector and the vendor-share hooks reach it only through a push to master (autosync); the project lead's word on the push, then one restart and the first live rows into this file.
|
||||
7. The 5090 clock rows (`-lgc` at 2,781 / 2,472 / 2,163 / 1,854 MHz): an elevated job on PC 2 would ask for administrator rights at the keyboard (the 5 October prompt class); not run without the project lead's word.
|
||||
8. PC 1: the 9070 XT rows for items 2, 6 and 8, the detector's band and the FPGA ranking's AMD line are all owed on PC 1's release.
|
||||
9. The job tooling: one card-key form everywhere (the settings.json key carries the device index) and a CI check that fails a job script posting a key without it (the class fix for item 2's loaded-card run).
|
||||
|
||||
## 7. Unverified and owed
|
||||
|
||||
| Item | Owed | Why |
|
||||
|
|
@ -137,4 +199,9 @@ Either replaces "under 2x" on the site and in the litepaper once the project lea
|
|||
| 4 | the 9070 XT rate band for the detector; the coinbase card-model tag (0.3.12) so the chain-attributed band needs no fleet log; the public testnet's first-week baseline (history check 2) | PC 1 not released; a miner change |
|
||||
| 5 | a rented FPGA hour (the soft-overlay reads-in-flight number is a model on cited HBM figures) | no FPGA in the fleet |
|
||||
| 6 | the 9070 XT step costs | PC 1 not released |
|
||||
| 4 + 7 | the live observer restart with both hooks, and the first live detector and vendor-share rows | after item 7 merges |
|
||||
| 4 + 7 | the live observer restart with both hooks, and the first live detector and vendor-share rows | needs a push to master (the project lead's word) |
|
||||
| 2 | the 2019-class core (O-1.14), which decides dr736 against dr368; the once-a-day NVRTC module for the item function (required before any activation); cryptanalysis of random ARX programs; the loaded-iGPU tier's build with the day program; the 5090 absolutes re-run with the card quiet (ratios stand) | unmeasured; unimplemented |
|
||||
| 8 | the 9070 XT and 4060-class rows (where a small card binds); the 5090 clock rows (an elevated job); the 5090 at a 575 W cap (model only); the Mac package watts (IOReport gives GPU + DRAM, Ember's 38 W approximate); the chip side's k, lane area and 28 nm scaling; the Metal fuzz, edge and stats runs and gates G2, G4 to G6 on the class; the acceptance rule's reading of the block | PC 1 not released; no elevated job; the gates are the next step on the project lead's word |
|
||||
| 6 | mm8 as a chain on Apple (Metal 4 matmul2d; the Mac's Swift toolchain has no tensor API); the 5 percent rule per family with the family live (argued only); the RDNA ISA guides unread (mnemonics from LLVM's tables) | |
|
||||
| 1 | the DRAM energy figures are streaming figures applied to random 32-byte reads; the GDDR7 burst and HBM3 tFAW are behind the JEDEC paywall; no chip has been built or torn down | |
|
||||
| all | every AMD RDNA 4 number in this run | PC 1 is the project lead's desk today |
|
||||
|
|
|
|||
Loading…
Reference in a new issue