Class v6 floor lane 5 (invention): the knee row from PC 1 (the hash lane, kit d): at the 1,300 lock the hinted w64 form reads 37.6 against the 4-byte row's 100.6 MH/s at 2.3x the energy per hash, so W = 16 is dead at the knee as at stock; every row this lane owed on W = 16 is closed

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

Documents-only replay of 58635a97b (58635a97b7) for the box mirror master
This commit is contained in:
igneum-labs 2026-10-08 14:04:48 +00:00
parent af5744a888
commit b259cbc31d

View file

@ -6,7 +6,7 @@
**The identity, as the record states it** (`docs/analysis/counter-asic-4-research.md` section 2): against the chip that stores the dataset, `edge = (E_card + F) / (E_mem + k F)`; at zero premium the edge is `E_card / E_mem` (3.6x on a 5090 at its 1,300 MHz knee against the GDDR7 board, measured card, modelled chip), and a shadow op with `k` under 1 never fully closes it. The one thing with `k = 1` by construction is the memory itself: the same devices on both sides. The record forced one DRAM operation per read (the activate and one 32-byte sector) and called every wider read "nothing to force: the chip already pays the sector". That is true up to 32 bytes; at 64 bytes the chip pays a second sector, and this lane's question was whether the honest card pays it at the DRAM's price or at its own.
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS) and on the RTX 5090 at a sparse shape (145.7 against 142.6 MH/s, sp170-w4, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at its own fabric's price, not the DRAM's 1.15 nJ: 8.6 nJ on the 4090 (+71 W at 8.3 G reads per second for no rate) and 4.7 nJ on the 5090 (+88 W at 18.6 G), so the energy per hash rises 34 percent on the 4090, 25 on the 5090, 33 on the H100's best form and 25 on the 3090, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 (4090) to 0.24 (5090) on GDDR7, under the ALU shadow's 0.3 to 0.8; the edge at zero shadow moves 0 on Ada, 7 percent on Blackwell (4.7x to 4.4x at stock) and 0.9x in the chip's favour on Hopper, and 0 with the shadow on (3.3x at `k` 0.5 either way). **KILL at stock on measured rows.** The one row that could revive it is the honest card's knee (a smaller fixed share and a lower fabric voltage; a rented pod refuses `-lgc`), owed as a single PC 1 row and not blocking anything: W = 16 stays out of class v6.
**The candidate this lane found, measured today, and where it stands.** The dataset item is 64 bytes (16 words; spec 1.5). Reading the whole item (W = 16 words, the read-width branch's class `w64`, bit-exact on four runtimes, fingerprint `836e56e7d496e980`, the verifier +4 percent) forces a second 32-byte sector per read in the same open row: on the chip model's own inputs (909 pJ of activation plus 1,150 pJ of movement per sector, Micron's 4.5 pJ per bit, claimed) the GDDR7 chip's energy per hash rises from 0.466 to 0.62 microjoules (+33 percent, its rate unchanged), one HBM3 stack's from 0.321 to 0.40 (+24), and on floor lane 3's wire figure the N2 SRAM die's from 0.036 to 0.126 (66x to 19x). The design's "never 16 words" rested on one row, the 5090's 71.9 MH/s (-47 percent) in the exported kernel at full occupancy, while the same card's own 64-byte probe read -13 percent; so the lane rented cards and measured the card side at stock (section 3.1): **the rate question is settled per architecture, and the energy question kills the lever at stock.** A one-request form of the same kernel text (PTX `ld.global.L2::64B` on the first load of each item; bit-exact, the same fingerprint) holds the 4-byte rate on the RTX 4090 at full occupancy (63.65 against 64.73 MH/s, 98 percent, PASS) and on the RTX 5090 at a sparse shape (145.7 against 142.6 MH/s, sp170-w4, PASS); on the H100 the best form is -21 percent and the hint does nothing (FAIL); on the RTX 3090 the best form is -13 percent (FAIL). But the card pays the second sector at its own fabric's price, not the DRAM's 1.15 nJ: 8.6 nJ on the 4090 (+71 W at 8.3 G reads per second for no rate) and 4.7 nJ on the 5090 (+88 W at 18.6 G), so the energy per hash rises 34 percent on the 4090, 25 on the 5090, 33 on the H100's best form and 25 on the 3090, against the modelled chip's 33 (GDDR7) and 50 (HBM3). The sector's `k` is about 0.13 (4090) to 0.24 (5090) on GDDR7, under the ALU shadow's 0.3 to 0.8; the edge at zero shadow moves 0 on Ada, 7 percent on Blackwell (4.7x to 4.4x at stock) and 0.9x in the chip's favour on Hopper, and 0 with the shadow on (3.3x at `k` 0.5 either way). **KILL at stock on measured rows.** The knee row, measured by the hash lane on PC 1 at 15:0x UK, closes it from the other side: at the 1,300 MHz lock the hinted form reads 37.6 against the 4-byte row's 100.6 MH/s at 2.3x its energy per hash (section 6). W = 16 is out of class v6 at stock and at the knee.
**Everything else on the brief is killed, with the number.** The tensor core: the record's `k` 0.03 to 0.3 is not defensible at its floor (the 0.04 pJ figure is INT4 at 0.46 V on a test chip with no memory system); shipping merchant silicon at nominal voltage reads 0.30 to 0.56 pJ per INT8 MAC (Meta MTIA v2 0.51, Qualcomm AI 100 Ultra 0.34, AMD MI355X 0.56, NVIDIA B200 0.44, Apple M4 ANE 0.30 measured; all claimed unless marked), and a Hopper tensor core under a pure MMA loop measured 0.34, so the honest band against the tile the class runs (1.5 pJ per MAC at the knee, measured) is `k` 0.2 to 0.4 and against the dense wide tile (0.83 at the knee) 0.4 to 0.7: equal to or under the ALU shadow's 0.3 to 0.8, never above it, and the Apple cost (-35 percent of rate at 1,024 tiles per hash, measured) kills it as content regardless. The RT core: traversal results are implementation-defined on every API (Vulkan "no ordering guarantee", DXR "no defined order", NVIDIA "could change depending on which driver, which GPU"), the acceleration structure is opaque, and a dedicated unit reads 4 to 33 nJ per ray against 290 to 750 nJ measured board-level on an RTX 2080 (`k` 0.01 to 0.2): dead on bit-exactness and on `k`. Memory-level parallelism per dollar: no asymmetry per channel, but a correction to the chip model: no 28 nm GDDR7 PHY exists (the shipped ones are N3 and FinFET; the oldest GDDR6 PHYs are 12 nm), so the record's cheapest chip (a USD 5 M project, a USD 17 M break-even cap) is not buildable; the floor is a 12 nm GDDR6 part at about USD 20 M to 30 M (cap USD 70 M to 100 M) or an N5-class GDDR7 part at USD 50 M to 75 M (cap 170 M to 250 M), modelled. Proof of latency: the chain checks values, never time; the block is 1 s and a DRAM round trip 0.1 to 0.4 microseconds; the farm's node is on its own LAN; a sequential per-block prefix favours the lower-latency side, which is the chip (about 4x). The time dimension on the address map: a bit permutation is firmware (a few hundred gates), no prefetch exists for a dependent chain, and the literature has no such scheme. The literature since 2023 holds nothing outside the identity; what it adds is a higher Ethash precedent (iPollo V2H, 0.14 J per MH, about 13x a 5090, vendor claim), the measured PAM3 I/O energies (SK hynix, ISSCC 2024) and the PHY-node fact above. Of this lane's own ideas, four more die on the identity (video decode, row-straddling, scratch in DRAM, independent chains per hash), and two items per read (W = 32) dies with W = 16.
@ -218,7 +218,7 @@ Reading: RandomX promised parity and got roughly it (1.0x to 1.5x after 46 month
The rows are the record: the 4090 and H100 pods were destroyed before their raw logs were copied home (the copy step failed quietly on a shell expansion), so their figures above are the extracted rows read off the pods in the session; the 3090's raw logs are at `~/igneum-fleet/fl5-results/3090-*` (result.log, result2.log, power.csv, power2.csv). Spend: about USD 6 of the USD 150 allowed.
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**; on the 5090 at sp170-w4, 88 W more at 145.7 MH/s: **4.7 nJ**; against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the 4090 (0.6 on the 5090) and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13 on the 4090, 0.24 on the 5090**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The one row that could change the sign is the knee: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row is one PC 1 run (the `w4` and `w64-l2` packs at `--block-warps 1`, 60 batches, at the lock) and is owed to the hash lane as a formality, not a gate.
**What the rows say, on the identity.** (1) The rate question: the exported form's -47 to -60 percent is a request-shape and occupancy effect, as the 5090's probe said; one PTX qualifier on the first load makes the 64-byte item one L2 request and holds the 4-byte rate on Ada at full occupancy (98 percent), while on Hopper the hint changes nothing (the H100's loss is elsewhere: its HBM3 pseudo-channel moves a 64-byte access as two bursts and its best shape is 79 percent) and on Ampere the best shape is 87 percent. (2) The energy question, which decides it: the card pays the second sector at the cost of a whole dependent read, not at the DRAM's movement term. On the 4090 the hinted form draws 71 W more at the same 64.7 MH/s: 71 W over 8.3 G reads per second is **8.6 nJ per second sector**; on the 5090 at sp170-w4, 88 W more at 145.7 MH/s: **4.7 nJ**; against the 1.15 nJ the DRAM device charges for it (the 4090's GDDR6X at 6 pJ per bit would be 1.5), so the other 7 nJ is the card's own L2, crossbar, L1 and the SMs' longer wait, the same anatomy as the record's 8.7 nJ whole-card marginal per dependent read (15.1a). In the identity's terms `F` is 128 x 8.6 nJ = 1.1 microjoules per hash on the 4090 (0.6 on the 5090) and the chip's cost for the same work is 128 x 1.15 = 0.15: **`k` about 0.13 on the 4090, 0.24 on the 5090**, under the ALU shadow's 0.3 to 0.8 and no better than the L2 hot table's. The energy per hash rises 34 percent on the 4090 and 33 percent on the H100's best form against the modelled chip's 33 (GDDR7) and 50 (HBM3): the edge at zero shadow moves 0 on Ada against the GDDR7 chip and 0.9x in the chip's favour against HBM3 on Hopper; against the SRAM die the chip's 66x to 19x becomes about 25x against the card's new energy, not 19x. (3) The knee, which could have changed the sign, did not (the hash lane's PC 1 row, section 6: the hinted form at the 1,300 lock reads 37.6 MH/s against 100.6 at 2.3x the energy per hash; the hint recovers nothing in the latency-bound regime on Blackwell). The reasoning that made it the one open row, for the record: at the 1,300 MHz lock the 5090's fixed share is smaller and its fabric runs at a lower voltage, so the sector's 8.6 nJ would fall with the measured 10.9 to 8.7 nJ per read (15.1a) and the card's rise would be a larger fraction of a smaller base; the lever lives only if the card's energy per hash rises under about 20 percent there, and a rented pod cannot read it (`nvidia-smi -lgc` refused: "does not have permission to change clocks"). That row ran on PC 1 at 15:0x UK and is in section 6.
**Verdict: KILL at stock on measured rows; W = 16 words stays out of class v6; the design's "never 16" stands, for the measured reason (the card's fabric price of a sector) rather than the plan's (the pins).** What is kept: the hinted one-request load as a kernel fact for Ada (a 64-byte dependent read at the 4-byte rate), which a future width decision can use; the measured 8.6 nJ per sector as the card-side figure the chip model's width rows lacked; and the per-architecture rate table above.
@ -268,7 +268,7 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
| Home miner, one 8 GB NVIDIA card (32-byte sectors) | nothing ships; had W = 16 shipped, this tier would have paid about a third more energy per hash for no rate (Ada) and an eighth of its rate too (Ampere) | nothing; the knee row on PC 1 is a formality |
| One 12 GB card (4070, 5070) | the same | the same |
| One 16 GB card (RX 9070 XT) | nothing (its 64-byte line is free at W = 16, measured 5 October); its row against the chip is unchanged at 22.7x on class v3 | the vendor-share metric |
| One 24 or 32 GB card (5090 class) | nothing ships; had W = 16 shipped, the 5090 would have held its rate at a sparse shape and paid +88 W (+25 percent per hash) at stock; the record's 3.6x at the knee and 2.1x at `k = 1` stand | the knee row on PC 1 (the hash lane, one row, a formality) |
| One 24 or 32 GB card (5090 class) | nothing ships; had W = 16 shipped, the 5090 would have held its rate at a sparse shape and paid +88 W (+25 percent per hash) at stock, and at its knee lost 63 percent of its rate at 2.3x the energy (measured, PC 1); the record's 3.6x at the knee and 2.1x at `k = 1` stand | nothing |
| Apple (M5 Max and the unified tiers) | nothing (free at 64 bytes, measured 5 October); the honest best per joule is unchanged | nothing |
| A rig | nothing ships; the two model corrections raise the chip's project floor (USD 20 M to 75 M), which is the rig's real wall, by 4x to 15x over the record's cheapest row | the chip model's rows corrected (the Counter lane) |
| A pool user | nothing changes | |
@ -278,7 +278,7 @@ Both sides are capex-dominated (the record's 16.1: 7x to 10x their electricity p
## 6. Unverified and owed
- The knee row (the 5090 at the 1,300 MHz lock, `w4` against `w64-l2`, 60 batches, watts): a rented pod refuses the lock; one PC 1 row with the hash lane; the lever lives only under a 20 percent energy rise there, which the stock rows make unlikely.
- The knee row: DONE by the hash lane on PC 1's 5090 (kit d, 14:51 to 15:0x UK, the v5 kit's CUDA worker, `--block-warps 1`, 250 batches of 2^24, nvidia-smi 1 Hz, fingerprint 836e56e7d496e980 PASS on every `w64` row). Unlocked: `w4` 119.95 MH/s at 308.2 W (2.57 microjoules); `w64` 63.35 at 332.9 W; `w64-l2` 92.35 at 369.1 W (4.00 microjoules): the hint recovers 46 percent of the plain form's loss at full occupancy and stays 23 percent under the 4-byte rate at 1.56x its energy per hash. At the 1,300 MHz lock: `w4` 100.55 at 190.5 W (1.89 microjoules); `w64` 37.35 at 177.8 W; `w64-l2` 37.59 at 164.1 W (4.37 microjoules): **at the knee the hint gives nothing and the card's energy per hash is 2.3x the 4-byte row's**, far over the 20 percent line. The lever is dead at the knee as at stock; the sparse shape (sp170-w4) was not in the kit's worker and was not asked for again, because at the lock the card's fixed share is smaller and the sector's share larger, so no shape can bring a 2.3x rise under 1.2x. This closes every row this lane owed on W = 16.
- The 4090's and H100's raw logs were not copied before their pods were destroyed; the rows in 3.1 are the figures read off the pods during the session. The 5090's and the 3090's raw logs are at `~/igneum-fleet/fl5-results/5090-*` and `3090-*` (result.log, result2.log, result3.log, power*.csv).
- The watts are `nvidia-smi` board power at 1 Hz, the upper half of each row's samples averaged; idle on the rented pods read 29 W (4090) and 70 W (H100); the 3090's idle row was taken straight after a run and is not a settled idle.
- The chip side of 3.1 is the model's own data-movement term applied twice (4.5 pJ per bit, Micron, claimed, streaming; O'Connor's 3.48 for HBM2) and lane 3's wire figure for the SRAM die; no one has measured a second-sector column read's energy on GDDR7 (section 2.6: the public record has no random-read joule on any current part).