igneum/docs/analysis/chip-model-v3.md
igneum-labs bcc2db992e Counter ASIC 3.0 item 2: the design, the Mac measurements, the chip-model row and the PC 2 job
docs/plans/counter-asic-3-derivation.md (the design, the acceptance test, the interpreter, the allowance argument,
the measurements, the PROPOSED reserve entry R0 for 1.13.2, what is owed), docs/analysis/chip-model-v3.md section 6
(the per-day derivation rows at 1.0x to 3x allowances), the bench-log entry, relay/playbooks/ca3-derive-pc2.ps1
(one PC 2 job: self-fetched packs zip, the installed worker through NVRTC, the card off only under test with its
key from settings.json). Verifier 4.875 / 4.944 ms per unit on one M5 Max core under the measure lock against
x8's 2.061 / 2.063; Metal build 29 ms against 22; hash rate equal; bit-exact on Metal and Apple OpenCL.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-10-06 07:50:01 +00:00

14 KiB

The on-die-cache recompute chip against the RTX 5090, class v2 and class v3, everything combined

5 October 2026 (night), Counter ASIC 2.0, worker ca2-mixer. The model is M16's (docs/analysis/m16-recompute-attacker-2026-10-05.md): the strongest chip the plan has priced holds the whole cache in SRAM and derives every dataset item instead of reading it, so its cost per hash is item derivations, and its rate at a 50 T op/s integer budget (an RTX 5090's, approximate) is 50 T / (ops per hash). Nothing here is a measurement of a chip; every GPU figure says where it was measured. "Approximate" marks a figure from memory.

1. Inputs

Input Value Source
Items per hash 128 (one item per load, 128 loads per hash, median 128.00 distinct) spec 01 sections 1.4.2 and 1.8.5; the 20,000-program census
Integer operations per mixer application about 130 spec 01 section 1.8.4
Mixer applications per item 9 under v2; 36 under v3 (m = 4, docs/plans/mixer-x4.md) memhard::Shape::mixers_per_item
Integer operations per item 1,170 (v2); 4,680 (v3) 9 x 130; 36 x 130
Integer operations per hash 149,760 (v2, "150,000"); 599,040 (v3, "600,000") 128 x the above
Chip integer budget 50 T op/s (approximate: 21,760 ALUs at about 2.4 GHz, one 32-bit operation each per clock) M16 section 3
Fixed-function factor 3x (approximate, from memory: 2x to 5x is the usual credit for a pipeline with no scheduling or divergence) M16 section 3
RTX 5090, version 2 programs, measured 136.1 MH/s (readwidth, tonight, docs/plans/read-width.md, pack w4 on PC 2); 139.7 MH/s (M11, 4 October, docs/bench-log.md) this analysis uses tonight's 136.1 as the denominator and quotes both
RTX 5090 at w16 (16-byte loads), measured 139.8 MH/s readwidth table, tonight (the width stays 4 B: w16 closes nothing)
Cache mirror, 256 MiB, N5 headline density 128 mm^2, $46 per good die (64 mm^2, $21 at the bit-cell lower bound) docs/analysis/sram-mirror.md revision 2, sections 4 and 5 (ca2-analysis e6085c6)
Cache mirror plus a 96 MB hot table, N5 headline 175 mm^2, $68 same, so a hot table costs 0.49 mm^2 and $0.23 per MB (linear, approximate)
512 MiB and 1 GiB mirrors, N5 headline 255 mm^2 and 510 mm^2; $111 to $306 same, section 4 (the growth rule's cache at years 4 and 12, priced at today's node)
GPU-class die 750 mm^2 (the equal-silicon comparison) M16 section 3
CPU verifier, one M5 Max core (loaded, load average 5.6; ratios are the measurement) v2 1.31 to 1.36 ms per unit, x4 1.92 to 1.96 (1.45x), x8 2.79 (2.1x); worst cold 1.58 / 2.04 / 2.94 ms docs/plans/mixer-x4.md section 6.4, 5 October 2026 21:40 UTC

2. The rows

Chip rate = 50 T op/s / ops per hash. "Bare" = chip rate / 136.1 MH/s. "With the factor" = bare x 3. "Equal silicon" = bare x (750 - SRAM) / 750 x 3: the SRAM takes die area the logic does not get, the M16 convention ("minus the area the SRAM takes"). SRAM in mm^2 and dollars at the N5 headline density.

Row Mixer Ops per hash Chip rate at 50 T op/s SRAM the chip holds mm^2 / $ (N5 headline) Bare gain against 136.1 MH/s With the 3x factor Equal silicon, SRAM deducted, with the factor
v2 as shipped (the M16 and scratch-soundness row) x1 149,760 334 MH/s 256 MiB 128 / $46 2.45x (2.39x against 139.7) 7.4x 6.1x
v2 at w16 (not adopted; the chip's cost is items, not bytes: unchanged) x1 149,760 334 256 MiB 128 / $46 2.39x against 139.8 7.2x 5.9x
x4 (the candidate measured beside v3; not v3) x4 599,040 83.5 MH/s 256 MiB 128 / $46 0.61x 1.84x 1.53x
MEASURED, NOT ADOPTED (layer 5 decided out of v3 on the PC rows, coordinator 21:40 UTC): v3 plus a 32 MiB hot table, added form (16 dataset loads and k hot loads): the honest card pays the hot loads, this chip pays SRAM only x4 599,040 (a hot load is one SRAM read, no item) 83.5 288 MiB 144 / $53 0.66x at the Mac's g = 0.93 (126.6 MH/s); 0.70x at the 5090's g = 0.87 (118.4); the 9070 XT's g 0.84 1.98x (Mac g), 2.11x (5090 g) 1.60x, 1.71x
MEASURED, NOT ADOPTED: v3 plus a 64 MiB hot table, added form x4 599,040 83.5 320 MiB 160 / $61 0.71x at the Mac's g = 0.87 (118.4 MH/s); 0.73x at the 5090's g = 0.84 (114.3); the 9070 XT's g 0.80 2.12x (Mac g), 2.19x (5090 g) 1.67x, 1.73x
v3 at year 4 (cache 512 MiB under option C, dataset 4 GiB), no hot table x4 599,040 83.5 512 MiB 255 / $111 0.61x 1.84x 1.21x
v3 at year 12 (cache 1 GiB, dataset 8 GiB) x4 599,040 83.5 1 GiB 510 / $306 0.61x 1.84x 0.59x
v3: mixer x8 (decided 22:05 UTC under the delegated rule: verify 2.1 ms per unit on one Mac core against the 10 ms gate, the daily 1 GiB build 23 to 77 ms on the 5090 and the 9070 XT) x8 1,198,080 41.7 256 MiB 128 / $46 0.31x 0.92x 0.76x
x8 at year 4 x8 1,198,080 41.7 512 MiB 255 / $111 0.31x 0.92x 0.61x

The era draws of spec 1.13.1 cost the chip nothing in this model: the mixer round count is not drawn, the op weights and fold rotations change the program, not the item derivation, so the chip's ops per hash stand. The width rule (4-byte loads kept) changes nothing either: w16 would have moved the honest denominator by 2.7% and the chip's cost not at all.

Arithmetic, row v3: 36 x 130 = 4,680 ops per item; x 128 = 599,040 per hash; 50 x 10^12 / 599,040 = 83.5 x 10^6 hashes per second; 83.5 / 136.1 = 0.613; x 3 = 1.84; equal silicon (750 - 128) / 750 = 0.829, x 1.84 = 1.53. Hot table rows: 32 MiB x 0.49 mm^2 per MB = 16 mm^2, 64 MiB = 32 mm^2 (the 96 MB column of sram-mirror.md scaled linearly); (750 - 144) / 750 = 0.808 and (750 - 160) / 750 = 0.787. The honest denominator in the added form is the v2 rate times g, the card's measured ratio with the hot loads added: on the M5 Max tonight g = 0.93 / 0.87 / 0.83 at 32 / 64 / 96 MiB (the cache agent, relayed by the coordinator at 21:23 UTC; docs/plans/hot-table.md carries the runs); the 5090's and the 9070 XT's g are the PC rows, owed, and until they land the row carries the Mac's g against the 5090's rate, which is a mixed figure and is marked so. Year 4 and 12 rows: the mirror of sram-mirror.md section 4 at N5 for 512 MiB and 1 GiB plus the 64 MiB table, at today's density (the node of those years is denser by about 1.8x at year 10 on the trend the same file cites; the row is a floor on the area, not a forecast).

3. The margin, plainly

The combined headline row is the mixer row alone (layer 5 is out: the added form costs the 5090 13 to 16 percent and the 9070 XT 16 to 20 percent against the 0.97 bar, coordinator 21:40 UTC; the width stays 4 bytes; the era draws and the cache growth cost this chip nothing at year 0), and class v3 is x8 (decided 22:05 UTC). The headline: the on-die-cache recompute chip at 50 T op/s reaches 41.7 MH/s against the 5090's 136.1, 0.31x bare, 0.92x with the 3x fixed-function factor, 0.76x with the mirror's area deducted: under 1x with the factor, 0.92x, a margin of 8 percent on the factor (a 3.3x factor reads 1.0x) and of 9 percent on the budget (55 T op/s reads 1.0x). The x4 candidate, measured beside it, read 1.84x and 1.53x. The hot-table rows above are kept as measured, not adopted: against THIS chip an added hot table is a cost to the honest card and none to the chip, so it would have moved the row the wrong way by the card's own g. The margin, plainly:

  • the 3x fixed-function factor is approximate and from memory; at 3.3x the equal-budget row reads 2.0x;
  • the denominator is one card's measured rate on one night (136.1 against 139.7 the night before: 2.6% apart);
  • the 50 T op/s budget is approximate; a chip at 55 T op/s reads 2.0x;
  • the hot table in the added form lowers the honest denominator by whatever the hot loads cost the GPU (owed from the PC rows), which raises the chip's gain by the same share, 1.84x or more if the hot loads are free, higher if not; the hot table's only cost to this chip is 16 to 32 mm^2 of die.

What keeps it under 1x is the mixer, and nothing else in Counter ASIC 2.0 moves this chip (the scratch at any share gave 2.4x, docs/analysis/scratch-soundness.md section 3.4; the hot table taxes the DRAM-only chip, not this one; the cache growth taxes it only in die area, which is cheap at year 0 and real at year 12). The next levers, in order:

  1. Mixer x16 (the next step of the same lever): 0.16x bare and 0.46x with the factor against 136.1; the verifier by the measured increments (+0.63 ms at x4, +1.46 at x8 on the M5 Max core: about +3.1 ms at x16, 3.7 ms per unit, 9 ms on a 2.5x slower laptop core, approximate) is at the edge of the 10 ms gate, so a 2019-class laptop core measurement (O-1.14) decides it, not this model.
  2. The hot table: adopted or not on the PC rows (docs/plans/hot-table.md); in the added form it costs the GPU 7 to 17 percent on the Mac and the chip die area only, so against this chip it is a lever in the wrong direction and against a DRAM-only chip the first lever; if it is adopted, the mixer must carry the extra 1/g (x8 at g = 0.87 reads 1.06x at the equal budget, 0.84x with the SRAM deducted).

4. What this does not settle

The items of M16 section 5 stand: the inline kernel on NVIDIA with a 64 MiB cache inside L2 (a measured point under the "50 T op/s" row) is a PC job not yet run; the time-memory curve (O-1.6) is not drawn; the mixer has had no cryptanalysis, and a shortcut inside it cuts the 4,680 directly; no chip has been priced beyond its SRAM.

6. The per-day derivation (item 2)

6 October 2026, Counter ASIC 3.0 item 2, worker derive (docs/plans/counter-asic-3-derivation.md; everything PROPOSED, a prototype behind load class dr736). The fixed-shape mixer of section 2's rows is replaced by nine straight-line programs of 736 instructions per item drawn from the day key stream (twelve two-register forms, the chain rule, an acceptance test with the x8 mixer's counts as floors). The chip's cost per hash is still item derivations; what changes is the fixed-function factor, because the chip must now execute an arbitrary program of the day from a 12-form set over 16 registers (a sequencer: instruction store, register file, operand muxes, a 32-bit ALU with a multiplier and a rotator) instead of a wired pipeline of 72 mixer stages with the day's constants in the wires. The counts are from the code (memhard::mixer: 144 ops per application as written, 128 with the round constants hoisted, 16 multiplies; the x8 item is 10,368 / 9,216 / 1,152), not the 130 of section 1; the day program's floor is those counts, so the bare row cannot fall below x8's.

Row Derivation Chip ops per hash Chip rate at 50 T op/s SRAM the chip holds mm^2 / $ (N5 headline) Bare gain against 136.1 MH/s Allowance 1.2x (ProgPoW's claimed range, history 2.4 [S67] [S70]) Allowance 1.5x (cautious upper bound, approximate) The old 3x (the fixed shape's; does not apply) Equal silicon, SRAM deducted, at 1.2x / 1.5x
x8 as shipped (section 2's v3 row, re-counted from the code with constants hoisted) fixed mixer, 72 x 128 1,179,648 42.4 MH/s 256 MiB 128 / $46 0.31x 0.37x 0.47x 0.93x 0.31x / 0.39x
dr736, the genesis day's draw (9,992 chip ops, 1,461 multiplies per item; the floor is x8's 9,216) the day program, 9 x 736 instructions 1,278,976 39.1 MH/s 256 MiB 128 / $46 0.29x 0.34x 0.43x 0.86x 0.29x / 0.36x
dr736 at the floor (a day whose draw sits exactly on the acceptance floor) the day program 1,179,648 42.4 256 MiB 128 / $46 0.31x 0.37x 0.47x 0.93x 0.31x / 0.39x
dr368, the fallback (the x4-equivalent count: 5,004 chip ops per item on the genesis day) the day program, 9 x 368 640,512 78.1 256 MiB 128 / $46 0.57x 0.69x 0.86x 1.72x 0.57x / 0.71x
dr736 at year 4 (cache 512 MiB) the day program 1,278,976 39.1 512 MiB 255 / $111 0.29x 0.34x 0.43x 0.86x 0.23x / 0.28x

Arithmetic, row dr736: 9,992 x 128 = 1,278,976; 50 x 10^12 / 1,278,976 = 39.1 x 10^6; 39.1 / 136.1 = 0.287; x 1.2 = 0.345; x 1.5 = 0.431; x 3 = 0.862; equal silicon (750 - 128) / 750 = 0.829, x 0.345 = 0.286, x 0.431 = 0.357. The allowance argument, plainly: the 3x of section 1 was the credit for "a pipeline with no scheduling or divergence", which a fixed dataflow earns because the chip wires the 72 applications and bakes the constants in; with a program that changes daily the chip keeps no divergence (the GPU has none here either: the item function is straight-line), the constants folded into an instruction store, and no warp scheduler or operand collector, and it loses the wiring. That residual is what ProgPoW's audits priced at 1.1x to 1.2x for a conventional compute chip (Rao: "conventional compute chips gain little on ProgPoW", history section 2.4); 1.5x is a cautious upper bound of mine (approximate) for a chip that also drops the GPU's float and graphics area. The chain rule (every instruction reads the register the previous one wrote) adds a cost the row does not credit: with no intra-item parallelism a single engine completes one dependent instruction per cycle at best and must interleave items to keep its multiplier busy, which is a register file per item in flight (RandomX's light-mode argument, history 2.4). The measured costs that buy this: the verifier 4.88 ms per unit on one M5 Max core against x8's 2.06 (the derivation document's section 5.1), the Mac's daily build 29 ms against 22, the hash rate unchanged; the 5090's build and compile are the PC 2 job, the 9070 XT's OWED.

What this does not settle: the rows are the same 50 T op/s budget and the same denominator as section 2 (their margins apply); no chip has been priced for its instruction store or its register files per item in flight; the random ARX programs have had no cryptanalysis (the item 3 brief should name them beside M_r); the 2019-class core measurement (O-1.14) decides whether 736 or 368 is the length, and the derivation document's section 0 carries that verdict.