Class v6 floor lane 5 (invention): the 5090 W = 16 stock row measured on a secure-cloud pod: the hinted form at sp170-w4 holds the 4-byte rate (145.7 against 142.6 MH/s) at 401 W against 316, +25 percent energy per hash, the second sector 4.7 nJ on the card (k about 0.24); the kill stands at stock; all five pods destroyed, USD 9 spent
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Documents-only replay of0a8c7883e(0a8c7883e9) for the box mirror master
This commit is contained in:
parent
9083c305ad
commit
ca381a45f7
1 changed files with 7 additions and 8 deletions
|
|
@ -1,12 +1,12 @@
|
|||
# Class v6 floor lane 5, invention: what none of the four floor lanes covers, checked against the identity
|
||||
|
||||
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK; corrected 13:1x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item; the rented-card rows of 13:2x to 13:4x UK carried in section 3.1, which reverse rank 1 at stock). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
|
||||
8 October 2026, 12:48 to 19:xx UK, branch `class-v6-floor-invention` from `counter-asic-4` (base 171618c5 after the fast-forward; first cut at 13:0x UK; corrected 13:1x UK for the read-width naming: `w16` is 16 bytes, `w64` is the 64-byte item; the rented-card rows of 13:2x to 13:4x UK carried in section 3.1, which reverse rank 1 at stock; the 5090 row at 13:4x UK). The open-ended lane of the floor research: four other lanes work the known levers (SM-sparse gating, the shadow's `k` from RTL synthesis, the SRAM full-store chip and the dataset floor, the honest denominator per tier); this file finds what they do not cover. Every candidate is checked against the identity first and killed or kept with a number. Every number carries a label: **measured** (a card on a named job, the file named), **modelled** (arithmetic on the chip model's cited inputs, `docs/analysis/chip-model-v3.md` section 5), **claimed** (a vendor's or an author's figure, URL given), **approximate** (from memory or a scaling). Nothing here changes a consensus object, a pack or a served number. Research and reasoning ran on this Mac; no build, no benchmark and no miner ran anywhere for this file.
|
||||
|
||||
## 0. One page
|
||||
|
||||
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes; at 64 bytes the chip pays a second sector, and this lane's question was whether the honest card pays it at the DRAM's price or at its own.
|
||||
|
||||
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at about 8.6 nJ (the 4090: +71 W at 8.3 G reads per second for no rate), the cost of a whole dependent read through its L2 and crossbar, not the DRAM's 1.15 nJ: the 4090's energy per hash rises 34 percent, the H100's 33 percent in its best form, the 3090's 25 percent, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 on GDDR7 (1.15 over 8.6), under the ALU shadow's 0.3 to 0.8, and the edge at zero shadow moves 0 on Ada and 0.9x in the chip's favour on Hopper. **KILL at stock on measured rows.** The one row that could revive it is the honest card's knee (a smaller fixed share and a lower fabric voltage; a rented pod refuses `-lgc`), owed as a single PC 1 row and not blocking anything: W = 16 stays out of class v6.
|
||||
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS) and on the RTX 5090 at a sparse shape (145.7 against 142.6 MH/s, sp170-w4, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at its own fabric's price, not the DRAM's 1.15 nJ: 8.6 nJ on the 4090 (+71 W at 8.3 G reads per second for no rate) and 4.7 nJ on the 5090 (+88 W at 18.6 G), so the energy per hash rises 34 percent on the 4090, 25 on the 5090, 33 on the H100's best form and 25 on the 3090, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 (4090) to 0.24 (5090) on GDDR7, under the ALU shadow's 0.3 to 0.8; the edge at zero shadow moves 0 on Ada, 7 percent on Blackwell (4.7x to 4.4x at stock) and 0.9x in the chip's favour on Hopper, and 0 with the shadow on (3.3x at `k` 0.5 either way). **KILL at stock on measured rows.** The one row that could revive it is the honest card's knee (a smaller fixed share and a lower fabric voltage; a rented pod refuses `-lgc`), owed as a single PC 1 row and not blocking anything: W = 16 stays out of class v6.
|
||||
|
||||
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash), and two items per read (W = 32) dies with W = 16.
|
||||
|
||||
|
|
@ -194,19 +194,19 @@ Reading: nothing published since 2023 offers a lever outside the identity; every
|
|||
| The custom HBM4E base die (lane B) | 0.9 to 1.0 | about 1.5 | 0.18 to about 0.27 (+50 percent) | unchanged | modelled, approximate |
|
||||
| The N2 SRAM full store (floor lane 3's rows) | 0.25 nJ (80 bits of wire) | 0.88 nJ (560 bits) | 0.036 to 0.126 (+250 percent): 66x to 19x against the 5090's stock point | power-bound: 8.3 to 2.4 GH/s per die | modelled (lane 3) |
|
||||
|
||||
**The card side, measured today** (8 October 2026, 13:2x to 13:4x UK; RunPod one-shot pods under the fleet's `oneshot.py`, the image `nvidia/cuda:12.8.1-devel-ubuntu24.04`; the worker built on each pod from this branch's `proto-cuda/nvrtc/worker.cpp` with `g++ -O2 -std=c++17 ... -ldl` through its Linux `dlopen` path, sha256 prefix 30eea48b; stock clocks, no power cap change; the packs `w4` (the lottery hash), `w64` (the item) and `w64-l2` (the `w64` kernel text with the first `uint4` load of each item replaced by inline PTX `ld.global.L2::64B.v4.u32`, so the L2 fetches the 64-byte line as one request; nothing else changed); every row's self-test PASS (96 of 96 vector lanes) and the 2^24 fingerprint `836e56e7d496e980` on every `w64` and `w64-l2` row, `25f96e7dce90bd4e` on every `w4` row, so the hinted text is bit-exact; rows of 2^24 nonces x 60 batches with `nvidia-smi` at 1 Hz over the row, the watts the mean of the upper half of the row's samples; a first pass of 5-batch rows for the shape sweep, a second of 60-batch rows for the watts; the sparse shapes are the worker's `sp<N>-w<W>` persistent grids, N = the card's SM count, W warps per block, so lanes in flight = N x W x 32):
|
||||
**The card side, measured today** (8 October 2026, 13:2x to 13:4x UK; four cards on five pods; RunPod one-shot pods under the fleet's `oneshot.py`, the image `nvidia/cuda:12.8.1-devel-ubuntu24.04`; the worker built on each pod from this branch's `proto-cuda/nvrtc/worker.cpp` with `g++ -O2 -std=c++17 ... -ldl` through its Linux `dlopen` path, sha256 prefix 30eea48b; stock clocks, no power cap change; the packs `w4` (the lottery hash), `w64` (the item) and `w64-l2` (the `w64` kernel text with the first `uint4` load of each item replaced by inline PTX `ld.global.L2::64B.v4.u32`, so the L2 fetches the 64-byte line as one request; nothing else changed); every row's self-test PASS (96 of 96 vector lanes) and the 2^24 fingerprint `836e56e7d496e980` on every `w64` and `w64-l2` row, `25f96e7dce90bd4e` on every `w4` row, so the hinted text is bit-exact; rows of 2^24 nonces x 60 batches with `nvidia-smi` at 1 Hz over the row, the watts the mean of the upper half of the row's samples; a first pass of 5-batch rows for the shape sweep, a second of 60-batch rows for the watts; the sparse shapes are the worker's `sp<N>-w<W>` persistent grids, N = the card's SM count, W warps per block, so lanes in flight = N x W x 32):
|
||||
|
||||
| Card (architecture, memory) | `w4`: MH/s at W | `w64` exported, full occupancy | `w64` best sparse shape | `w64-l2` best form | On the 95 percent line | Energy per hash, `w4` to the best `w64` form | Label |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| RTX 4090 (Ada, GDDR6X, 128 SMs; driver 595.71) | 64.73 at 225 W (3.48 microjoules); idle 29 W | 33.75 (-48 percent) at 242 W | sp128-w1 (4,096 lanes) 53.98 (-17) at 258 W | **full occupancy 63.65 (-1.7 percent) at 296 W**; sp128-w1 58.5 (-10) at 264 | **PASS on rate** | 3.48 to 4.66: **+34 percent** | measured |
|
||||
| H100 SXM (Hopper, HBM3, 132 SMs; driver 580.126) | 253.4 at 442 W (1.74); idle 70 W | 171.4 (-32) at 469 W | sp132-w8 200.3 (-21) at 463 W | 171.4 at full occupancy, 200.3 at sp132-w8: **the hint does nothing on Hopper** | FAIL (79 percent) | 1.74 to 2.31 (best form): +33 percent; the full-occupancy form 2.74, +57 | measured |
|
||||
| RTX 3090 (Ampere, GDDR6X, 82 SMs; driver 580.159) | 60.00 at 321 W (5.35) | 23.95 (-60) at 330 W | sp82-w2 42.19 (-30) at 334 W | sp82-w2 52.34 (-13) at 349 W; sp82-w4 51.0 (5-batch row) | FAIL (87 percent) | 5.35 to 6.67: +25 percent | measured |
|
||||
| RTX 5090 (Blackwell, GDDR7, 170 SMs) | the 5 October rows: 136.1 MH/s; `w64` 71.9 (-47) at full occupancy; the probe -13 at its best lane count | | | pending: RunPod's one community 5090 host (216.249.100.66, bus 85:00.0, driver 590.48) answered `cuInit` 999 on three pods in a row (destroyed); a secure-cloud pod is the fourth try, its row amended here when it lands | | | measured 5 October (rate only) |
|
||||
| RTX 5090 (Blackwell, GDDR7, 170 SMs; a secure-cloud pod, driver 580.126, bus 02:00.0; the community host 216.249.100.66 answered `cuInit` 999 on three pods, all destroyed) | 142.5 to 142.7 at 311 to 319 W (2.21 microjoules; the record's 2.26 at 311 W); idle not settled on the pod | 75.3 (-47 percent) at 344 W (the 5 October row, 71.9, reproduced) | sp170-w1 (5,440 lanes) 101.5 (-29) at 343 W; sp170-w8 88.2 | full occupancy 110.4 (-22.5) at 392 W; **sp170-w4 (21,760 lanes) 145.7 (+2.3 percent) at 401 W**; sp170-w8 141.9 (-0.4) at 408 | **PASS on rate at the sparse shape** | 2.21 to 2.75: **+25 percent** | measured |
|
||||
| RX 9070 XT, Apple M5 Max | 5 October: -3 percent and 0 at 64 bytes (a 64-byte line per read already) | | | not re-run | free | 0 on the memory side | measured 5 October |
|
||||
|
||||
The rows are the record: the 4090 and H100 pods were destroyed before their raw logs were copied home (the copy step failed quietly on a shell expansion), so their figures above are the extracted rows read off the pods in the session; the 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*` (result.log, result2.log, power.csv, power2.csv). Spend: about USD 6 of the USD 150 allowed.
|
||||
|
||||
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**, against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the card and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The one row that could change the sign is the knee: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row is one PC 1 run (the `w4` and `w64-l2` packs at `--block-warps 1`, 60 batches, at the lock) and is owed to the hash lane as a formality, not a gate.
|
||||
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**; on the 5090 at sp170-w4, 88 W more at 145.7 MH/s: **4.7 nJ**; against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the 4090 (0.6 on the 5090) and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13 on the 4090, 0.24 on the 5090**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The one row that could change the sign is the knee: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row is one PC 1 run (the `w4` and `w64-l2` packs at `--block-warps 1`, 60 batches, at the lock) and is owed to the hash lane as a formality, not a gate.
|
||||
|
||||
**Verdict: KILL at stock on measured rows; W = 16 words stays out of class v6; the design's "never 16" stands, for the measured reason (the card's fabric price of a sector) rather than the plan's (the pins).** What is kept: the hinted one-request load as a kernel fact for Ada (a 64-byte dependent read at the 4-byte rate), which a future width decision can use; the measured 8.6 nJ per sector as the card-side figure the chip model's width rows lacked; and the per-architecture rate table above.
|
||||
|
||||
|
|
@ -256,7 +256,7 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
|
|||
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | nothing ships; had W = 16 shipped, this tier would have paid about a third more energy per hash for no rate (Ada) and an eighth of its rate too (Ampere) | nothing; the knee row on PC 1 is a formality |
|
||||
| One 12 GB card (4070, 5070) | the same | the same |
|
||||
| One 16 GB card (RX 9070 XT) | nothing (its 64-byte line is free at W = 16, measured 5 October); its row against the chip is unchanged at 22.7x on class v3 | the vendor-share metric |
|
||||
| One 24 or 32 GB card (5090 class) | nothing ships; the 5090's own stock row is pending a working pod and amends this file; the record's 3.6x at the knee and 2.1x at `k = 1` stand | the secure 5090 pod; the knee row on PC 1 (the hash lane, one row) |
|
||||
| One 24 or 32 GB card (5090 class) | nothing ships; had W = 16 shipped, the 5090 would have held its rate at a sparse shape and paid +88 W (+25 percent per hash) at stock; the record's 3.6x at the knee and 2.1x at `k = 1` stand | the knee row on PC 1 (the hash lane, one row, a formality) |
|
||||
| Apple (M5 Max and the unified tiers) | nothing (free at 64 bytes, measured 5 October); the honest best per joule is unchanged | nothing |
|
||||
| A rig | nothing ships; the two model corrections raise the chip's project floor (USD 20 M to 75 M), which is the rig's real wall, by 4x to 15x over the record's cheapest row | the chip model's rows corrected (the Counter lane) |
|
||||
| A pool user | nothing changes | |
|
||||
|
|
@ -266,9 +266,8 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
|
|||
|
||||
## 6. Unverified and owed
|
||||
|
||||
- The 5090's own stock rows (`w4`, `w64`, `w64-l2` at full occupancy and the sparse shapes, with watts): pending the secure-cloud pod after three community pods on the same host failed `cuInit` with 999; amended into 3.1 when they land.
|
||||
- The knee row (the 5090 at the 1,300 MHz lock, `w4` against `w64-l2`, 60 batches, watts): a rented pod refuses the lock; one PC 1 row with the hash lane; the lever lives only under a 20 percent energy rise there, which the stock rows make unlikely.
|
||||
- The 4090's and H100's raw logs were not copied before their pods were destroyed; the rows in 3.1 are the figures read off the pods during the session. The 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*`.
|
||||
- The 4090's and H100's raw logs were not copied before their pods were destroyed; the rows in 3.1 are the figures read off the pods during the session. The 5090's and the 3090's raw logs are at `~/igneum-fleet/fl5-results/5090-*` and `3090-*` (result.log, result2.log, result3.log, power*.csv).
|
||||
- The watts are `nvidia-smi` board power at 1 Hz, the upper half of each row's samples averaged; idle on the rented pods read 29 W (4090) and 70 W (H100); the 3090's idle row was taken straight after a run and is not a settled idle.
|
||||
- The chip side of 3.1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).
|
||||
- The tensor chip figures are vendor specifications at TDP (claimed), the M4 ANE and the H800 loop measured; the RT figures are claims and one measured RTX 2080 row; the PHY node facts are vendor IP pages read today.
|
||||
|
|
|
|||
Loading…
Reference in a new issue