Counter ASIC 3.0 status: the verifier on a server core (igneum-build-1): rung 2 admissible, rung 3 over the gate loaded; the ladder's flags
This commit is contained in:
parent
ca16dde4e9
commit
7b61f14811
1 changed files with 3 additions and 1 deletions
|
|
@ -84,7 +84,9 @@ Consequence per tier (interim): the M5 Max at 21 W is the per-joule best honest
|
|||
| 199,600 | 24.25 (-10.4%) | 38.2, 1.58 | 128.67 (-2.7%) | 431, 1,753, 3.35 | 2.43 |
|
||||
| 330,700 | 21.39 (-21%) | 40.1, 1.88 | 86.39 (-35%, compute-bound at the capped clock, 28.6 T op/s) | 431, 1,834, 4.99 | 2.62 (worst cold 2.77) |
|
||||
|
||||
Where each card leaves the latency bound (the 5 percent rule): M5 Max about 130,000 ops; RTX 5090 about 210,000 at its 431 W cap (the cap binds from 102,100 ops and the governor lowers the clock); RX 9070 XT OWED. The verifier's law: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp, so 330,700 ops add 0.56 ms (about 1.4 ms on a 2019-class core): it fits every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none), and the node never binds before the cards. Marginal ALU energy per counted op: the 5090 10 to 13 pJ at its shipping clock (twice the 5.5 pJ the item 1 model assumed), the M5 Max 6.9 pJ. Power caps: `-pl` below 400 W cannot be set, the 5 October sweep rows re-used (302 to 316 W at every cap); the clock rows are OWED (nvidia-smi refused `-lgc` without rights; the job did not ask; an elevated job is a decision for the project lead, section 6). Item 1's denominator: the 5090 at the hash is 290 W in the app and 350 W in the bench, not 326; the f = 1 rows move 2 to 11 percent.
|
||||
Where each card leaves the latency bound (the 5 percent rule): M5 Max about 130,000 ops; RTX 5090 about 210,000 at its 431 W cap (the cap binds from 102,100 ops and the governor lowers the clock); RX 9070 XT OWED. The verifier's law: 2.06 ms + 3.2 microseconds per 1,000 shadow instructions per warp, so 330,700 ops add 0.56 ms (about 1.4 ms on a 2019-class core): it fits every pairing's headroom (x8 7.8 / 4.6 ms, dr368 7.1 / 2.8, dr736 5.1 / none), and the node never binds before the cards. The verifier on a SLOWER core, measured 7 October 08:46 UK by the node lane on igneum-build-1 (an EPYC 9454P core at 3.66 GHz under schedutil, the measure hold, load 6.1; the ladder worktree's igneum-pow 59ae70cf; 50 warps averaged, the cold warp beside it) for the testnet's latency ladder (its rungs are item 8's shadow packs): rung 2 = 199,566 counted ops (the sh256x53 pack) 8.99 ms loaded (9.25 cold), ADMISSIBLE under the 10 ms gate; rung 3 = 330,740 ops (sh256x88) 5.95 ms alone (9.04 cold), 10.11 ms with the sibling loaded (10.85 cold), OVER the gate by 0.85 ms, so the testnet freezes that rung FALSE. Reading against the M5 Max's law (2.06 ms + 3.2 microseconds per 1,000 shadow instructions: 2.62 ms at 330,700): the server core is 2.3x slower alone and 3.9x loaded, so the 2.5x rule's "about 6.5 ms" for a slower core was right alone and short under load; the candidate at 100,000 ops sits well inside the gate on this core by the same ratios (about 5.3 ms loaded, approximate, not measured), and the 2019-class laptop core (O-1.14) stays the owed row. Consequence per tier: a pool core or a node on a server-class CPU verifies the candidate class inside the gate with room; the ladder's top rung is out for any core slower than an M5 Max under load, which is why the testnet ships it false.
|
||||
|
||||
Marginal ALU energy per counted op: the 5090 10 to 13 pJ at its shipping clock (twice the 5.5 pJ the item 1 model assumed), the M5 Max 6.9 pJ. Power caps: `-pl` below 400 W cannot be set, the 5 October sweep rows re-used (302 to 316 W at every cap); the clock rows are OWED (nvidia-smi refused `-lgc` without rights; the job did not ask; an elevated job is a decision for the project lead, section 6). Item 1's denominator: the 5090 at the hash is 290 W in the app and 350 W in the bench, not 326; the f = 1 rows move 2 to 11 percent.
|
||||
|
||||
Chip side (approximate: the f = 1 stored-dataset chip of item 1 plus an ALU core at k times the 5090's measured 11 pJ per op): at N = 100,000 its per-joule edge over the 5090 falls from 5.6x (shadow empty, 350 W) to 2.1x on GDDR7 and 2.3x on one HBM3 stack at k = 1, 1.5x at k = 1.5, 3.2x at k = 0.5, 4.1x at k = 0.3; over the M5 Max from 1.6x to 0.9x at k = 1. The chip needs a 14,000-lane ALU array (about 30 mm^2 at N5, 150 W at k = 1) at 100,000 ops and 22,600 lanes (60 to 100 mm^2, 500 W) at 331,000; at 28 nm the array is reticle-class. That is the project lead's test as a number: with the shadow filled the chip must carry the memory system and a GPU-class datapath, and its edge is the ratio k of its datapath's energy per op to the GPU's.
|
||||
|
||||
|
|
|
|||
Loading…
Reference in a new issue