Class v6 3.3: lane 3's reading of the measured 5.5 GiB step (the mapping safe at the v6 epoch; 4 percent at stock, about 9 at the knee; the 8 GB tier's RX 7600 row owed)
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
a41aa5a0cc
commit
731caa1f18
1 changed files with 1 additions and 1 deletions
|
|
@ -89,7 +89,7 @@ Reading: layer 2 does not move the chip anyone builds, because a chip buys DRAM
|
|||
|
||||
### 3.3 Per tier
|
||||
|
||||
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. **The genesis floor itself, 5.5 GiB, MEASURED on a rented 5090 at stock (the fleet hand, RunPod secure, driver 570.195, 15:19 to 15:23 UK, the ca3-ds55 kit's own worker, program 73bcbfe8 in both packs, 250 and 500 x 2^24, every row PASS with 96 of 96 vector lanes): the 1 GiB control 141.48 MH/s at 325.6 W (2.30 microjoules; memory.used peak 1,914 MiB) against ds55 136.56 at 305.3 W busy, 327.6 steady (2.24 busy, 2.40 on the steady watts; peak 6,522 MiB: the 5.9 GB dataset plus the 268 MB cache plus the context), fingerprints stable across both passes (ds55 23ced07a4d28b465 becomes the pin). So the 5.5 GiB floor costs 3.5 percent of the rate at the same watts, about 4 percent more energy per hash, and the non-power-of-two mapping (1,476,395,008 words, 92,274,688 items, loads as (src x words) >> 32) is not a cliff on sm_120. Caveat: this host's 5090 plateaued at 328 W on both packs (a host power cap; another host's 5090 pulled 443 W on class v5 genesis this afternoon), so the microjoules are capped-card numbers and the rate and fingerprints are the row.** So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
|
||||
The hash lane's VRAM rows (12:0x UK, modelled from the measured 0.4 GiB working set plus about 0.5 GiB of driver and app): the dataset needs 3.2, 5.4 and 9.9 GiB of device memory at the floor, 2x and 4x; a 12 GB card falls off at about 9.5 GiB (year 15 on the 1.13.3 schedule), a 16 GB GPU at about 13.5 GiB (year 23), a 16 GB unified Mac at about 8 GiB (year 12), the 5090 at about 29 GiB (year 54). **The DRAM-read cost per hash on the NVIDIA cards is NOT size-independent at the knee, MEASURED (the hash lane's kit b, PC 1, 12:49 to 12:58 UK, the pinned class v3 program 73bcbfe8 at 2, 4 and 8 GiB against the 1 GiB control 137.65 MH/s at 312.2 W unlocked and 127.39 at 212.6 W at the 1,300 MHz lock; 250 batches per row, fingerprints PASS): unlocked 133.86 at 315.7 W (-2.8 percent), 132.42 at 317.2 (-3.8), 131.75 at 319.2 (-4.3); at the lock 121.14 at 210.3 W (-4.9 percent), 113.06 at 204.5 (-11.2), 109.39 at 201.6 (-14.1); MH/W at the lock 0.599, 0.576, 0.553, 0.543, which is 4 / 8 / 10 percent more energy per hash at 2 / 4 / 8 GiB.** The lane's earlier reading (2 MiB pages keep the TLB's reach past 8 GiB) holds unlocked, where the card hides most of the page-walk term in its slack; the latency-bound regime at the lock exposes it. **The genesis floor itself, 5.5 GiB, MEASURED on a rented 5090 at stock (the fleet hand, RunPod secure, driver 570.195, 15:19 to 15:23 UK, the ca3-ds55 kit's own worker, program 73bcbfe8 in both packs, 250 and 500 x 2^24, every row PASS with 96 of 96 vector lanes): the 1 GiB control 141.48 MH/s at 325.6 W (2.30 microjoules; memory.used peak 1,914 MiB) against ds55 136.56 at 305.3 W busy, 327.6 steady (2.24 busy, 2.40 on the steady watts; peak 6,522 MiB: the 5.9 GB dataset plus the 268 MB cache plus the context), fingerprints stable across both passes (ds55 23ced07a4d28b465 becomes the pin). So the 5.5 GiB floor costs 3.5 percent of the rate at the same watts, about 4 percent more energy per hash, and the non-power-of-two mapping (1,476,395,008 words, 92,274,688 items, loads as (src x words) >> 32) is not a cliff on sm_120. Caveat: this host's 5090 plateaued at 328 W on both packs (a host power cap; another host's 5090 pulled 443 W on class v5 genesis this afternoon), so the microjoules are capped-card numbers and the rate and fingerprints are the row.** Lane 3's reading of it for the schedule (2a61cb46, 15:25 UK): the step lands between the 2 and 4 GiB stock rows (2.8 and 3.8 percent), so the non-power-of-two floors of 5.5 / 8.5 / 11.5 cost nothing beyond the size and the multiply-shift mapping is safe to adopt at the v6 epoch; per tier the 5.5 GiB step costs a 5090 4 percent per hash at stock (measured) and about 9 at its knee (interpolated from the 2, 4, 8 GiB knee rows); the 8 GB tier's fate at that step is the RX 7600 reading from PC 1 (about 16:00 to 16:30 UK, an amendment; its 1.8 GB of headroom is the question), the 5.5 GiB knee row with it; the SRAM store at that step is 3 reticles, USD 1,500, its joules unmoved. So each step of the schedule costs a tuned 5090 about 4 to 5 percent per hash while the chip's joules do not move (floor lane 3: its ticket goes USD 1,500 / 2,500 / 3,000 at 5.5 / 8.5 / 11.5 GiB), and every chip edge against a card at its knee rises by 4 to 11 percent across the schedule; the honest sentence for the schedule decision is USD 1,000 of chip ticket per step for about 1 to 5 percent of the tuned 5090's energy and about a quarter of today's measured cards by count. **On the M5 Max it is not size-independent, measured by this lane at 10:40 UTC under the Mac's measure lock (Metal packbench, the hash lane's class v3 packs at 2^28 to 2^31 words, the same seed and era, 3 batches of 2^24, vectors 3 of 3 and fingerprints per pack): 26.48 MH/s at 1 GiB (footprint 1,664 MiB, the build 32 ms), 23.26 at 2 GiB (-12.2 percent; 2,688 MiB; 54 ms), 21.31 at 4 GiB (-19.5 percent; 4,736 MiB; 94 ms), 20.61 at 8 GiB (-22.2 percent; 8,832 MiB; 193 ms).** The Apple GPU's dependent random read costs more time as the working set grows past its page reach (approximate reading: a TLB-reach effect on unified LPDDR5X; the power channels were not sampled this run, so the joules per hash move by at least the rate's share), which is a real per-tier cost of layer 2 that the NVIDIA model does not show: at an 8 GiB floor the Apple tier mines 22 percent slower per card than at 1 GiB, before any memory limit. The table below carries it.
|
||||
|
||||
| Tier | At the floor (today to year 4) | At a 4 GiB state-driven step | At 16 GiB | Label |
|
||||
|---|---|---|---|---|
|
||||
|
|
|
|||
Loading…
Reference in a new issue