RTX 5090: dataset sweep shows the 96 MiB L2 cliff; igneum-hourly 96/96 PASS

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-josh 2026-10-03 16:33:18 +01:00
parent aba248d4a3
commit f2a1a64fa1

View file

@ -84,3 +84,17 @@ Result: the same hourly program, generated on the Mac, compiled by Apple's Metal
Build note for Windows: CUDA 12.8 crashes (cudafe++ access violation) under Visual Studio 2026's 14.51 toolset even with -allow-unsupported-compiler. Fix: install the MSVC v143 (14.30) component and open the environment with
`"C:\Program Files\Microsoft Visual Studio\18\Community\VC\Auxiliary\Build\vcvarsall.bat" x64 -vcvars_ver=14.30`, then build normally.
### RTX 5090, dataset sweep and second program (same session)
| dataset MiB | Mhash/s | GB/s useful | random loads/s (G) |
|---|---|---|---|
| 4 | 1339.8 | 557 | 139.3 |
| 64 | 1352.7 | 563 | 140.7 |
| 256 | 269.8 | 112 | 28.1 |
| 512 | 241.8 | 101 | 25.2 |
| 1024 | 228.7 | 95 | 23.8 |
Second program igneum-hourly (128 loads per hash): 96/96 vectors PASS, 185.3 Mhash/s at 1 GiB, 23.7 G random loads/s.
Reading: the 5090 carries 96 MiB of L2. At 4 and 64 MiB the dataset sits inside it and the program runs about 5.8x faster than at 1 GiB. Past the L2 the rate settles at about 23.7 G random loads/s for both programs regardless of loads per hash (104 vs 128 loads gives 228 vs 185 Mhash/s, proportional), so the program is random-access bound once the dataset exceeds on-chip cache. Each 4-byte random load moves a 32-byte sector, so DRAM traffic is roughly 760 GB/s, approximate, against a quoted peak near 1.8 TB/s for this card. Design consequence: the dataset must stay well above any plausible on-chip cache, which the 2 GB genesis size and the growth schedule provide; a chip would need gigabytes of on-chip memory to escape the random-access limit. Still prototype numbers: dataset derivation remains closed-form until the 256 MB cache construction lands.