ca2-epoch: the Mac compile-ahead measurement (10 fresh programs 15.9 / 17.7 / 20.4 ms min / median / max, the devnet pack 79 ms then 1 ms from the shader cache) in docs/bench-log.md and epoch-length.md section 6; the per-card table and the 600-s floor filled

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-05 21:19:33 +00:00
parent 8482a6f4d3
commit 0db67deeb3
2 changed files with 41 additions and 2 deletions

View file

@ -1571,3 +1571,29 @@ The 9070 XT rows are 2,048 persistent warps (4,096 within 1 percent), arena 64 M
**Readings.** (1) Same count, wider: the vendor gap does not move at 16 B (7.8x) because on the 9070 XT a 4-byte read already costs a 64-byte line and on the 5090 a 16-byte read costs one 32-byte sector, the same as 4 bytes: the memory systems do identical work, only the fold's input grows. At 64 B the gap closes to 4.1x, entirely by the 5090 losing half its rate (its share falls to 0.58 and its DRAM traffic reaches 589 GB/s, 37 percent of the stream: bandwidth, not latency, bounds it), while the 9070 XT and the M5 Max do not move. (2) Fewer, wider (w64x4): 3.7x, but every card runs 4x faster because the dependent chain is 32 loads long instead of 128; the 5090 sits at a 0.56 share (bandwidth), so a chip with more bandwidth per dollar than a GPU gains, which is the Ethash shape the design avoids. (3) The mix: the hour-to-hour spread is 7 to 22 percent of the median per card (the 5090 the widest, because its 64-byte loads are the expensive ones and their count per program runs 2 to 8 of 16); the programs with many 64-byte loads (mixA-3, mixA-5, mixB-2) are the slow hours on the 5090 and the fast ones nowhere. (4) The scratch: on the 5090 every RMW share costs 12 to 48 percent against the persistent control, the 32 KiB arena less than the 128 KiB one (the smaller arena, 64 MiB over 2,048 warps, sits inside the 96 MB L2); on the M5 Max the 32 KiB rows are FASTER than the control (+12 and +74 percent at 25 and 50 percent), because the arena (128 MiB over 4,096 warps) lives in the chip's caches and a scratch op is cheaper than a dataset read, so replacing dataset loads raises the rate: the scratch at these sizes is not memory work on Apple and is partly cached on NVIDIA. The chip row for these variants comes from the ca2-soundness branch; what this entry gives is the GPU cost and the share. (5) Latency-bound shares: v2 0.87 to 1.01 on the three cards, w16 0.84 to 1.03, w64 0.58 (5090) and 0.78 (9070 XT); the Mac's shares above 1 are an Apple OpenCL probe under load against a Metal rate.
Jobs: `run-readwidth-5090-20261005` and `run-readwidth-9070-20261005` (probes; the packs refused for their string seeds, fixed in a9e002c), `run-readwidth-5090-20261005c`, `run-readwidth-9070-20261005c` (benches), `run-readwidth-9070-scratch-20261005d` (the scratch packs after the `__local` fix d0018cf, AMD's compiler requires the exchange buffer at the kernel's outermost scope); read back with `node tools/jobs.mjs <id> --all`. Mac commands and logs: `docs/plans/read-width.md` section 3. The worker exes for the jobs: `proto-cuda/nvrtc/build-windows.sh` on this branch (mingw), sha256 of the CUDA one `6f46336f...defe1`.
## 5 October 2026 (night), epoch length as an era parameter (Counter ASIC 2.0, layer 9): the Mac's compile-ahead per program
Branch `ca2-epoch`, worker "ca2-epoch"; design and the per-card table in `docs/plans/epoch-length.md`. Question (the project lead: "what about faster program changes?"): what a card spends per epoch between receiving the next seed and swapping, which sets the floor of the epoch-length ladder (600 to 7,200 DAA s). Machine: Apple M5 Max (Darwin 25.6.0, 64 GiB), 21:18 UTC, load average 11 to 14 from other agents' builds and runs; the measure lock held for the 3-s run (`tools/lock/with-lock.sh measure bash scratchpad/epoch-measure.sh`). `proto-metal/igneum-bench` built from this branch with `swiftc -O -target arm64-apple-macos11 -o igneum-bench main.swift -framework Metal` under a build slot.
Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/epoch9`; version 2 generator, 128 loads per hash), each generated and compiled at run time (`makeLibrary` from source plus `makeComputePipelineState`), dataset 2^28 words, one 2^20 batch and one verify warp per program:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Program | Compile ms (library + pipeline) | Mhash/s (GPU) | Verify |
|---|---|---|---|
| epoch0 | 18.8 (9.3 + 9.5) | 27.2 | PASS |
| epoch1 | 17.8 (8.8 + 9.1) | 27.9 | PASS |
| epoch2 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch3 | 16.1 (8.0 + 8.2) | 27.7 | PASS |
| epoch4 | 18.6 (9.0 + 9.6) | 27.5 | PASS |
| epoch5 | 15.9 (7.7 + 8.2) | 28.4 | PASS |
| epoch6 | 17.6 (8.6 + 9.1) | 27.8 | PASS |
| epoch7 | 17.6 (8.7 + 9.0) | 30.1 | PASS |
| epoch8 | 20.4 (9.8 + 10.6) | 28.4 | PASS |
| epoch9 | 18.2 (8.7 + 9.5) | 29.3 | PASS |
| min / median / max | 15.9 / 17.7 / 20.4 | | 10 of 10 |
Cache fill 1.95 ms GPU (192.4 ms one core), dataset build 20.8 ms GPU for 1 GiB. The devnet pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (the pack's two libraries, `memhard.metal` and `program.metal`): compile 79 ms, 1 ms, 1 ms (the system shader cache answers the identical source from the second run); cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU.
Reading: a fresh program compiles in about 18 ms on this card with the Metal compiler service warm, 79 ms for a pack with its dataset kernels, up to 1.8 s cold (the variant-racing entry's first seed), 0 to 444 ms at the fleet's live boundaries (M11). The hot table fill of layer 5 is 0.07 to 0.22 ms (ca2-cache). So the Mac's per-epoch compile-ahead is under 2 s without the race and about 38 s with it (M11: 34.0 / 34.9 / 37.8 s), and the race is the only item visible against the 600-s window in which the program is known (lead 1,200 s minus the 600-s VDF, fixed at every epoch length). PC cards, cited in the plan: RTX 5090 NVRTC 151 to 180 ms, prepare 0.5 to 1.0 s without the dataset (M11), race one round about 37 s; RX 9070 XT OpenCL compile NOT MEASURED at the current worker (owed: `host.c` times `clBuildProgram` only in the `prepare` path and no `prepared` line from gfx1201 is in any upload); Intel UHD build 3.0 to 6.4 s (M11). Floor by the rule (slowest compile-ahead under 10% of the epoch and inside the window, dataset excluded): 600 DAA s, carried by the race at 6.3% of 600 s; with the race off (M11 found base wins on both the 5090 and the Mac) the slowest measured row is the Intel iGPU at 1.1%. Consequences per tier and the difficulty-settle constraint (24% of a 600-s epoch in settle at the measured 144 s) are in the plan.

View file

@ -131,7 +131,7 @@ What a card does between receiving the next seed and swapping (`app/igneum-app/s
| Card, compiler | Program (generate + compile) | Hot table fill (layer 5, `docs/plans/hot-table.md`, ca2-cache) | Variant race | Compile-ahead total, race off | Total, race on | Source |
|---|---|---|---|---|---|---|
| Apple M5 Max, Metal | 82 ms (hot-swap entry, 4 October); 58 ms (serve check); 0 to 444 ms over 8 live boundaries (M11 fleet row); 1,798 ms with the Metal compiler cold (variant-racing entry); tonight's 10-compile sample in 6.2 | 0.07 to 0.22 ms on the GPU | 34.0 / 34.9 / 37.8 s min / median / max over 8 boundaries (M11); 39.8 s one round in the serve check | 0.5 s (1.8 s cold) | 38 s | `docs/bench-log.md`: "first hourly program swap on the live devnet" (4 October), "miner performance: variant racing" (4 October), M11 table (5 October); section 6.2 |
| Apple M5 Max, Metal | 82 ms (hot-swap entry, 4 October); 58 ms (serve check); 0 to 444 ms over 8 live boundaries (M11 fleet row); 1,798 ms with the Metal compiler cold (variant-racing entry); tonight: 15.9 / 17.7 / 20.4 ms min / median / max over 10 fresh programs, 79 ms for the devnet pack's two libraries (6.2) | 0.07 to 0.22 ms on the GPU | 34.0 / 34.9 / 37.8 s min / median / max over 8 boundaries (M11); 39.8 s one round in the serve check | 0.5 s (1.8 s cold) | 38 s | `docs/bench-log.md`: "first hourly program swap on the live devnet" (4 October), "miner performance: variant racing" (4 October), M11 table (5 October); section 6.2 |
| NVIDIA RTX 5090, NVRTC 12.8 | NVRTC 151 to 180 ms (M11; 151 ms at the PC 2 14:20 boundary); prepare total 580 to 1,074 ms including cache 68, dataset 113 and self-test 511 ms, so about 0.5 to 1.0 s without the dataset; 1,285 ms on the nvcc path of 4 October | 0.67 ms per 256 MiB cache fill on the 5090 scaled to `S / 256`, under 1 ms | 232 to 300 ms to compile 17 variants, 112 s of timing for 3 rounds (M11 race rows, PC 1); one round about 37 s (112 / 3, approximate) | 1.0 s | 38 s | `docs/bench-log.md` M11 table and race rows (5 October), hot-swap entry (4 October), "PC 2 at the 14:20 boundary" (4 October) |
| AMD RX 9070 XT, OpenCL (gfx1201, Adrenalin 26.9.2) | NOT MEASURED at the current worker: `proto-opencl/host.c` times `clBuildProgram` only in the `prepare` path (`buildMs` in the `prepared` line) and no `prepared` line from this card is in any upload (PC 1 app logs 17:26, 18:10, 19:02 UTC; jobs `rdna4-serve-4`, `rdna4-bench-1`, `run-readwidth-9070-20261005c` print cache, dataset and check only) | not measured on AMD | none: the OpenCL worker has no race (`docs/design/miner-tuning.md`: 17 names on NVIDIA, 14 on Metal) | 0.31 s plus the compile (cache 9 ms, self-test 300 ms measured, `rdna4-serve-4`) | the same | OWED (section 9) |
| AMD Radeon integrated gfx1036 (PC 2), OpenCL | prepare total 6.9 / 9.4 / 11.7 s with the 1 GiB dataset build on the iGPU inside; the compile is not separated | | none | under 11.7 s | the same | M11 table; the compile share OWED |
@ -140,7 +140,20 @@ What a card does between receiving the next seed and swapping (`app/igneum-app/s
### 6.2 The Mac, tonight (Measured)
`<filled after the measurement: machine, date, command, 10 compiles, median and worst, the identical-source re-compile>`
Apple M5 Max (Darwin 25.6.0, 64 GiB), 5 October 2026 21:18 UTC, load average 11 to 14 (other agents' builds and runs; the measure lock held for the 3-s run, `with-lock.sh measure bash scratchpad/epoch-measure.sh`), `proto-metal/igneum-bench` built from this branch (e752fc7 + this document) with `swiftc -O -target arm64-apple-macos11`. Ten distinct programs (seed strings `igneum-devnet-v4-epoch0`, `/epoch1` .. `/epoch9`, version 2 generator, 128 loads per hash), each generated and compiled at run time (`makeLibrary` from source plus `makeComputePipelineState`), dataset 2^28 words, one 2^20 batch and one verify warp per program so the run is the compiles:
./igneum-bench --seed igneum-devnet-v4-epoch0 --hours 10 --dataset-log2 28 --batch-log2 20 --batches 1 --verify-warps 1
| Compile (library + pipeline), ms | Values over the 10 programs |
|---|---|
| each | 18.8, 17.8, 17.6, 16.1, 18.6, 15.9, 17.6, 17.6, 20.4, 18.2 |
| min / median / max | 15.9 / 17.7 / 20.4 |
| library share | 7.7 to 9.8 ms; pipeline 8.2 to 10.6 ms |
| hash rate during the run (GPU time, loaded Mac) | 27.2 to 30.1 MH/s, all 10 verify warps PASS |
The same pack three times through `packbench --pack ../proto-cuda/packs/igneum-devnet-v4-epoch0 --batches 1 --batch-log2 20 --group 256` (two libraries, `memhard.metal` and `program.metal`): compile 79 ms, then 1 ms and 1 ms, the system shader cache answering the identical source; cache fill 0.6 to 0.7 ms GPU, dataset build 20.7 to 20.8 ms GPU for 1 GiB.
Reading: a fresh program costs this card about 18 ms to compile with the Metal compiler service warm, 79 ms for a pack with the dataset kernels, and up to 1.8 s when the compiler is cold (the variant-racing entry's first seed). The fleet's 0 to 444 ms per boundary (M11) sits between those, so the live figure is the app's cold-start states, not the compile itself. None of it is visible against a 600-s window: the Mac's compile-ahead is the race or nothing.
### 6.3 Shares and the floor