Counter ASIC 2.0 status 23:20: the PC 1 queue

This commit is contained in:
igneum-labs 2026-10-05 20:39:03 +00:00
parent 4ce31bbbe0
commit 9dcb9697cf

View file

@ -172,3 +172,18 @@ dot4 probe on PC 1 (jobs fetch-dot4-20261005, run-dot4-20261005, exit 0, 101 s;
All bit-exact against the CPU reference. One dp4a costs about one ALU step on NVIDIA and AMD; the 5090 is 12.5x the 9070 XT on the ALU chain and 11.2x on dot4 (the family does not widen the AMD gap), 10x the M5 Max on the chain and 13.6x on dot4 (Apple's emulation widens its gap 1.4x). At W_new 4 the family adds about 21 ops per hash per lane; hash-rate losses expected under 5% on every card (to be measured with the family live). Owed: the CUDA __dp4a cross-check (needs nvcc), Metal 4 matmul2d int8 on the M5, sdot4 on RDNA 2. cl_khr_integer_dot_product is listed by no driver we own.
PC 1 is free: the era six-pack job goes next when its package arrives, then the hot-table added-form job.
## 23:20 PC 1 scheduler (the coordinator's role from now): the queue
Rule: one PC 1 job at a time; an agent asks by message before publishing, gets a "go PC 1" from this coordinator, and reports when its RESULT lines are in and both cards are restored; a hash-rate or power number taken while another job holds a card is not a number. The CPU-only job runs only in a slot where no measurement overlaps it.
| # | Job | Agent | Cards | Length | State |
|---|---|---|---|---|---|
| 1 | Era six-pack (layers 4 and 8), 5090 and 9070 XT | ca2-era a452664c512c73b9b | one card at a time | about 15 to 20 min | waiting for the package (packs on the pack-loop packfile.h) |
| 2 | Hot table, added form, probe 32/64/96 MiB plus packs | ca2-cache a5271cf269757b118 | one card at a time | about 15 min | waiting for the re-measured Mac rows and the rebuilt zip |
| 3 | Reproducible benchmark run | a0b9f574775ef1693 | one card at a time, 120 s per card | about 5 min | queued |
| 4 | AMD sweep on the 9070 XT (core clock and power steps) | a01dcb34ae16d867c | 9070 XT only; the 5090 keeps mining | about 20 min | queued |
| 5 | Ember Tune end to end, both cards | a855dcc4bd05e0615 | both | to be stated | queued after the AMD sweep |
| 6 | AMD-proving CPU fallback (CPU-only SP1 run, both cards mining) | a39db54d4de4af51e | none; loads the CPU | to be stated | last, or in a gap where no measurement runs for its whole length |
If job 1's package is more than 15 minutes away when job 2 is ready, job 2 goes first; the short job 3 fills any gap of under 10 minutes between packages.