Counter ASIC 2.0 status 23:10: layer 7 PC rows, PC 1 free

This commit is contained in:
igneum-labs 2026-10-05 20:38:08 +00:00
parent 3f33c6fd00
commit 4ce31bbbe0

View file

@ -157,3 +157,18 @@ The attempt-0 packfile.h bug (the era agent's find) is the one that took the fle
ca2-epoch (a32a3ece66c02417a): the epoch length as an era parameter (600 to 7,200 DAA s, base 3,600; draw or 90% signal; the VDF rule; the difficulty-window constraint; the FPGA threat with citations), Mac compile-ahead measured now, the 5090 and 9070 XT compile times cited from the bench log, docs/plans/epoch-length.md. The dot4 probe job is running on PC 1 (ca2-analysis tip ee42d7c carries the playbook; the 5090 confirmation is the agent's watch). PC 1 queue after it: the era six-pack job, then the hot-table added-form job. PC 2: the proving agent's, then the ca2 node suites.
Agents running: ca2-era, ca2-cache, ca2-node, ca2-mixer, ca2-epoch; ca2-analysis watching its PC job. Done: ca2-soundness.
## 23:10 layer 7 complete on all three cards (ca2-analysis ee42d7c); PC 1 free
dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s; both cards restored and mining; 5090 SM clock 2,505 MHz before and after):
| Device | ALU chain, G steps/s | signed dot4 emulation, G dot4/s | dot4 instruction, G dot4/s | emulation vs instruction |
|---|---|---|---|---|
| RTX 5090 (NVIDIA OpenCL 3.0, driver 617.14) | 8,753.5 | 1,239.1 (7.1x an ALU step) | 7,453.6 (inline PTX dp4a.s32.s32, 1.17x) | 6.0x |
| RX 9070 XT gfx1201 (AMD-APP 3683.0) | 701.4 | 480.8 (1.46x) | 664.3 (__builtin_amdgcn_sudot4, 1.06x) | 1.38x |
| gfx1036 (RDNA 2 iGPU) | 40.6 | 15.8 (2.6x) | sudot4 does not build (needs dot8-insts) | |
| M5 Max (Metal) | 879.8 | 188.2 (4.7x); unsigned 548.2 (1.6x) | none exists | |
All bit-exact against the CPU reference. One dp4a costs about one ALU step on NVIDIA and AMD; the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on dot4 (the family does not widen the AMD gap), 10x the M5 Max on the chain and 13.6x on dot4 (Apple's emulation widens its gap 1.4x). At W_new 4 the family adds about 21 ops per hash per lane; hash-rate losses expected under 5% on every card (to be measured with the family live). Owed: the CUDA __dp4a cross-check (needs nvcc), Metal 4 matmul2d int8 on the M5, sdot4 on RDNA 2. cl_khr_integer_dot_product is listed by no driver we own.
PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job.