Counter ASIC 3.0 status: item 6's Mac step costs (interim)
This commit is contained in:
parent
4ab969d182
commit
7f76308ff8
1 changed files with 4 additions and 0 deletions
|
|
@ -39,6 +39,10 @@ The whole case for the f = 1 chip: the 5090's memory system draws about 55 W of
|
|||
|
||||
What moves the f = 1 rows (chip-model-v3.md section 5.7): not the dataset size (one HBM3 stack holds 24 GB, the schedule reaches 4 GiB at year 4), not the read width (the decision to stay at 4 B stands), not the chain length; only (a) the honest card's watts at the hash (a 5090 holding 136 MH/s at a 250 W cap reads 3.9x on GDDR7, at 200 W 3.2x) and (b) program work in the latency shadow (the card hides 512 ops per hash behind 128 reads and could hide about 330,000 before compute binds; at N = 330,000 the model reads 1.85x at a chip core as efficient as the GPU's ALU, 2.5x at 1.5x worse), which is the design item this run adds (item 8 below) and the one that answers the project lead's test directly: a chip must carry the memory system AND the ALU budget.
|
||||
|
||||
### Item 6, the Mac rows (interim, ca3-reserve 192a683; the 5090 job waits on PC 2)
|
||||
|
||||
Step cost per family on the M5 Max as a ratio to the add-xor-rotate chain (881 G steps/s; best of 3; load average 7.64; all bit-exact): shl 0.85, shr 0.86, bfe 0.77, andn 0.75, byte permute 1.13 (emulated on Apple), popcount 0.87, clz 1.01, select 0.76, indexed shuffle 1.91 (2.2x the xor shuffle), live rotr 1.13, dot4 unsigned 1.60, dot4 signed 4.73. Every 32-bit datapath family costs an Apple lane under 1.2x a step, inside the 8x emulation bound of 1.13.2 with room; the matrix family is the only one past 1.6x. The rows feed item 8's ALU pricing. The full table with consequences lands at item 6's close.
|
||||
|
||||
### Consequences per tier (item 1)
|
||||
|
||||
| Tier | Meaning | Being done |
|
||||
|
|
|
|||
Loading…
Reference in a new issue