From af0f82798ee2c084e55096e2ef013f2e00cad0d2 Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Thu, 8 Oct 2026 12:27:08 +0000 Subject: [PATCH] Class v6: the coordinator's 13:30 UK clocks (floor rows by 15:30, the synthesis on master by 16:30; the decision stays the founder's) Co-Authored-By: Claude Fable 5.1 Documents-only replay of 6bbb33364 (counter-asic-4) for the box mirror master --- docs/design/class-v6-rotating-family.md | 18 +++++++++--------- 1 file changed, 9 insertions(+), 9 deletions(-) diff --git a/docs/design/class-v6-rotating-family.md b/docs/design/class-v6-rotating-family.md index 1b598524a..1c0f70bfa 100644 --- a/docs/design/class-v6-rotating-family.md +++ b/docs/design/class-v6-rotating-family.md @@ -261,7 +261,7 @@ The served sentence (`docs/plans/counter-asic-3-public-text-2026-10-07.md`, the | The convention behind every "x per joule with the shadow" figure | the record's convention (chip-model-v3 and the served text): `k` is the chip core's op cost as a fraction of the GPU's at the same operating point, so when the 5090 locks to 6.2 pJ the chip's core is priced at 6.2 k pJ too; the absolute convention: the chip's core costs its own pJ per op (the k lane's RTL figure), the GPU's its measured pJ at its point | measured GPU side; the chip side claimed until the k lane's RTL rows | the record's convention flatters every chip row at the lock by up to 1.6x (floor lane 3, 13:06 UK: the SRAM die 6.1x against 3.8x at k 0.5, 3.3x against 2.0x at k 1); the served line's figures are at the knee, so they are in the record's convention and need re-reading in the absolute one before main's word; the close carries both, absolute first | | The schedule's cost beside the SRAM store's edge (main's order) | section 3: the floor and ceiling per tier; main's candidate schedule (6, 10, 14 GiB) priced by lane A by 17:00; the Apple rate cost measured today (-12 / -20 / -22 percent at 2 / 4 / 8 GiB) | measured and modelled | lands at 17:00 | -The wording this lane proposes for main's word, if the SRAM-store reading stands at 20:00: "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word. +The wording this lane proposes for main's word, if the SRAM-store reading stands at the 15:45 UK close (the synthesis on the mirror's master by 16:30; the decision itself stays the founder's): "At launch the strongest chip we can price today, a memory-controller chip that stores the dataset, reaches 2.1x per joule against an RTX 5090 at its knee with a core as good as a GPU lane and 3.4x with one three times better, under class v4 from the first block; a 5090 locked at its knee pays 82 W for that shadow work. The strongest chip we can model for 2027 to 2028, a 2 GiB SRAM store on one 2 nm die, would reach 3x to 5x with that same shadow and 17x without it, at an N2 project of USD 100 M to 500 M and 18 months or more; the dataset's size is the one lever on it, and the schedule says how it grows. Class v5 makes the dataset the chain's own state, so a chip that does not follow the chain is wrong on every item. The model and every measurement are public." Every number in it is labelled in this document and the chip model; nothing served moves on this lane's word. ## 7c. Beyond the four: the invention lane's layers 5 to 7 and its rejected list (lane C, `docs/analysis/class-v6/invention.md` on master at a9f03598, 10:48 UTC; the census script `tools/attack/v6-invention/v6inv-census.sh`) @@ -274,11 +274,11 @@ The wording this lane proposes for main's word, if the SRAM-store reading stands Rejected by the invention lane with the number (its section 2): per-lane data-dependent branches, reads tied to the shard proof per block, randomised memory topology, the VRAM ratchet as a lever (layer 2's floor stands as written), pool-sampled witnesses, prover-gated eligibility (1.6 kW of proving network-wide at any hash rate: 0.7 percent of the hash at 100 GH/s), time-locked commitments beyond the era VDF, derived reads in the hash, row-straddling reads, scratch draws, the refresh per block. No candidate moves the per-joule identity. -## 10. The floor: the stored-dataset chip's per-joule floor and everything behind it (the five floor lanes of 12:55 UK; their rows by 16:00 to 16:30 UK, their documents under `docs/analysis/class-v6/floor/` by 19:30; this section closes at 15:45 UK on every row in by 15:30, each missing row's default stated and the row owed; later rows are amendments with their own time) +## 10. The floor: the stored-dataset chip's per-joule floor and everything behind it (the five floor lanes of 12:55 UK; their rows by 15:30 UK (the coordinator's clock, 13:30 UK; the earlier 19:30 defaults became 15:30), their documents under `docs/analysis/class-v6/floor/` by 19:30; this section closes at 15:45 UK on every row in by 15:30, each missing row's default stated and the row owed; later rows are amendments with their own time) The question the founder set: the stored-dataset chip pays the DRAM's own energy per random read and nothing else; the honest card pays that plus everything its silicon spends around the read; the floor is the ratio, and every term in it is a lever. The identity (the research file's section 2): at zero shadow `edge = E_card / E_mem`; with the shadow `(E_card + F) / (E_mem + k F)`. The lanes each take one term. -| Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 19:30 | The row owed | +| Lane (branch) | The term | What this document already holds (measured) | The default if the lane's rows are not in by 15:30 UK | The row owed | |---|---|---|---|---| | SM-sparse (`class-v6-floor-sm`) | `E_card` above `E_mem`: the honest card's watts toward the DRAM's own at the activate ceiling, on rented 5090, 4090 and H100 and PC 1 through the hash lane's jobs, with a worker patch | the research file's 20.3b (PC 1, 7 October night, the fourth exe): a quarter of the SMs holds 98.2 percent of the class v4 rate at the SAME draw (460 against 451 W) and 99.8 percent of class v3 at 4 W less; watts minus idle per MH/s never falls below base; the draw follows the work, not the SM count; at the 1,300 lock the sparse shapes collapse. The candidate was closed on that row | the 20.3b reading stands: the honest card's premium-free floor is its idle plus its memory system plus whatever the SMs spend waiting (99 W of 228 at the lock, measured as the residual, not reached by idling SMs); the 5090 at its knee 1.66 microjoules against the GDDR7 chip's 0.466: 3.6x | a measured breakdown of the 99 W that an idle SM does not save (clock tree, L2, fabric), and whether a different occupancy shape (fewer warps per SM at full SM count) moves it | | The shadow's k (`class-v6-floor-k`) | `k` from RTL synthesis (Yosys plus OpenROAD on build-4) replacing every claimed chip-side k; the op mix that maximises k | the GPU side measured (the research file's 15.1a): ARX 11.3 / 6.2 pJ per counted op, mul 13.9 / 8.3, mulhi 39.6 / 21.0, prmt 22.3 / 11.5, lop3 24.1 / 13.0, shfl 55.8 / 29.4, fp32 FMA 9.2 / 5.2, the int8 tile 1.4 to 4.1 per MAC; the chip side was claimed only (a 5 nm SIMD array 2 to 5 pJ per op, approximate) until the lane's first rows (10.2, 13:3x UK): synthesised on ASAP7, k 0.18 at the lock in the absolute convention at N3 (0.35 unscaled), the ARX families 0.17 to 0.18, mad 0.20, mulhi 0.03 | the default stays the record's rows at k 0.5 (the DRAM chip 2.9x at the knee) with the unit floors as the lower bound: the GDDR7 chip with the class v4 shadow at 4.0x at the lock and 5.8x at stock (absolute, floor k) is the worst case the served line must survive, not the reading; the sequencer-core row with the M5 Max column is the headline when it lands; the op mix held at class v4's (15.1b: the shuffle at 4.9x the add on the GPU side; mulhi the worst lever on the card by 6x) | the synthesised pJ per op per family for a 32-lane SIMD core at the chosen node, the resulting k per family, and the mix that maximises k at a fixed GPU premium | @@ -286,7 +286,7 @@ The question the founder set: the stored-dataset chip pays the DRAM's own energy | The honest denominator (`class-v6-floor-denominator`) | `E_card` per tier with every software knob (the core lock, the undervolt, the memory clock, the occupancy, the block shape); the Ember tier table; FIRST TABLE IN (10.4, 13:4x UK): the honest NVIDIA floor is the 16 GB Blackwell card at its knee (the 5080 2.06 measured), the only lever the operating point | the knee rows (the research file's 20.3a and 20.3b; the cost rows): the 5090 at 1,300 MHz 1.66 to 1.69 microjoules, the 4070 at its tune 2.57, the 5070 Ti stock 1.79, the M5 Max 0.78 (GPU and DRAM channels); the operating point is rank 1 of the research file | those rows stand as the per-tier floor; the Ember knob ships as ordered for 0.3.24 | the per-tier best point with every knob, the Ember table, and the AMD and Apple lines (ADLX or nothing; no lever) | | Invention beyond the four (`class-v6-floor-invention`) | the tensor core's k re-read, the RT core, controller plus PHY cost, proof of latency, era-driven address mapping, the literature since 2023 | the research file: the tensor tile's measured 1.4 to 4.1 pJ per MAC against a 5 nm array's claimed 0.04 to 0.4 (k 0.03 to 0.3, the worse lever); the L2 hit at 1.4 to 2.4 nJ against a chip's SRAM (k 0.1 to 0.3); the texture interpolator excluded (not bit-exact across vendors); the literature table (sections 5 and 8) | those readings stand; no new lever | anything that reads k above 1 on measured rows, which nothing has | -The close (20:00): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land. +The close (15:45 UK; the synthesis on the mirror's master by 16:30): the honest floor per tier against each chip row, the recommended class v6 changes (the op mix, the SM-sparse default, the dataset schedule), the served chip line as a measured range, and what stays for v7, written here from the lanes' rows as they land. ### 10.0 The close at 15:45 UK (the founder's standing rule, 13:3x UK: anything doable in 30 minutes to 2 hours gets a clock inside that window; so 15:45 is THE close, not a provisional: the floor lands on every row in by 15:30, and the lock rows and the k lane's remaining unit rows go in below as amendments, each with its own time; there is no 20:00 event. Lane 3's 7bc9de4b capex numbers are final for the close; the SRAM die headlines the k lane's synthesised sequencer-core row the moment it lands, by 15:00, with its 6x to 10x at the full shadow kept beside it marked as the floor-k worst case) @@ -309,7 +309,7 @@ The served line in one sentence, on the defaults: a chip that stores the dataset The three class v6 changes it implies: (1) the op mix stays class v4's with the lossy families capped at their base (lane D: 0 of 3,000 eras exhausted under the band; the k lane: mulhi is the card's worst lever by 6x, so no re-weight helps the card); (2) no SM-sparse default (dead on the measured 5090 rows; the per-watt gift of the hot table is the chip's, 20.3c); (3) the dataset schedule 5.5 / 8.5 / 11.5 GiB with the read width pinned at 4 words (the chip's ticket USD 1,500 / 2,500 / 3,000; a tuned 5090 pays 1 to 5 percent per step; about a quarter of today's cards by count per step). -### 10.1 SM-sparse on the honest card (floor lane 1, `docs/analysis/class-v6/floor/sm-sparse.md` on class-v6-floor-sm at 9e478daf, first rows 13:25 UK; the knee row by 15:00, the file by 19:30) +### 10.1 SM-sparse on the honest card (floor lane 1, `docs/analysis/class-v6/floor/sm-sparse.md` on class-v6-floor-sm at 9e478daf, first rows 13:25 UK; the knee row by 15:00, the file by 15:30) **The 20.3b reading holds on three more cards, and the residual over the DRAM's own energy is the clock domain, not the SMs.** Rented pods at stock (the worker's `--bench`, 250 x 2^24, nvidia-smi at 1 Hz, bit-exact on every row); microjoules per hash, the edge card over chip (GDDR7 0.466, HBM3 0.321, SRAM at lane B's 0.14): @@ -338,7 +338,7 @@ Host e (idle 12 W): base 142.0 at 352.7 W = 2.485; 43 SMs 141.2 at 337.4 = 2.389 The decomposition for section 2's term (the H100 microbench, full residency at the boost clock, measured): idle 80 W; resident SMs issuing nothing 136.6 W (+57 W, the SM clock domain); the ALU probe 437.9; the dependent DRAM chase 450.3 W at 32.2 G reads a second (11.5 nJ per read whole-card); the hash 450.4 W at 32.6 G reads a second. The memory path on the GPU side is 313 of the 370 W over idle (9.6 nJ per read above the HBM's modelled 1.2). The 4090 and 5090 probes land by 15:00; PC 1's memory-clock ladder at the 1,300 lock is queued (`run-ca4-pc1-memclk-5090-20261008`). **The fraction a chip cannot strip** (the memory side's own on chip-model 5.3's rows: DRAM plus static plus a controller, about 0.40 microjoules per hash on GDDR7 at 128 reads, 0.47 on the 4090's GDDR6X at its rate, 0.23 on HBM3): the 5090 at stock 16 to 19 percent, at the 1,300 lock 25 percent, the 4090 13 percent, the H100 13 percent. Everything else is the card's silicon around the read and a chip strips it; the honest floor's real number per tier is that fraction, and the only honest-side lever on the rest is the clock (rank 1). -### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 19:30 UK) +### 10.2 The shadow's k from RTL (floor lane 2, `docs/analysis/class-v6/floor/shadow-k.md` on class-v6-floor-k, build-4 and build-3, first synthesised rows 13:3x UK; RTL and flow under `tools/chip-model/rtl`; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30)) **The first synthesised k is 0.18 at the 5090's lock in the absolute convention at N3, and nothing reads inside the claimed 0.3 to 0.8 band except the unscaled ASAP7 figure at the lock (0.35). These are PER-UNIT FLOORS: no fetch, no decode, no register file beyond an 8-entry window, which is the chip rotation kills (section 1). The coordinator's order (13:5x UK): the headline k for the close is the programmable sequencer-core row (fetch, decode, a 32-register file, the per-era registers, the drawn program), asked of the lane with the M5 Max column; until it lands the default is the record's shadowed rows at k 0.5 with these unit floors beside them as the lower bound, and the GDDR7 board at 4.0x at the lock and 5.8x at stock on the floor k is the WORST CASE the served line must survive, not the reading.** Method: a minimal lane (instruction register, an 8 x 32-bit register window of flops, read muxes, the unit, one write port; no fetch or decode) in Yosys 0.68 plus OpenROAD on ASAP7, routed, SPEF, power from a random-input gate-level VCD with every pin annotated, the TC corner at 0.70 V; so every chip figure is a FLOOR and every k a floor. The node scaling is claimed from TSMC's headline per-node power reductions (N5 x0.70, N3 x0.50, N2 x0.36 of ASAP7; approximate). The GPU side is the research file's 15.1a, measured. @@ -361,7 +361,7 @@ The ranking by k, highest first (hardest for a chip; absolute, N3 against the lo What it does to the GDDR7 stored-dataset board (E_mem 0.466 microjoules, modelled) with the class v4 shadow (102,100 ops per hash): absolute, the chip's shadow is 102,100 x about 1.1 pJ = 0.11 microjoules at N3, E_chip about 0.58, so the edge reads **4.0x at the lock (the card 2.33) and 5.8x at stock (3.36)**, against the served 2.1x at k = 1 and 3.4x at k 0.33. In the record's convention it is k 0.18 at the lock, 2.33 / (0.466 + 0.18 x 0.652) = 4.0x, the same number there, and at stock k 0.10, 3.36 / (0.466 + 0.10 x 1.10) = 5.8x. **The shadow buys the card about 0.5x to 1x of edge out of the 5.2x and 3.6x zero-shadow figures, not the 1.5x the served k = 1 row implies.** This is the row the served line must survive (section 9): the served "2.1x" is the k = 1 reading; the unit floor says a core's units are five to six times cheaper than that at N3, and the sequencer core that a rotating family forces (fetch, decode, register file, the drawn program's control) sits between the two, where the lane's next row puts it; until then the default reading is k 0.5 (the DRAM chip 2.9x at the knee in the record's convention) with 4.0x as the floor-k worst case. Pending from the lane by 19:30 UK: the 32-lane shuffle butterfly and the general crossbar (in place now), the 8 KB scratch as a flop array, the int8 8x8x8 tile, the mix optimiser over the layer-1 band, and the re-fold of lane 3's SRAM rows at W = 4 and 8 in both conventions. -### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 19:30 UK; everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured) +### 10.3 The SRAM die and the dataset floor (floor lane 3, `docs/analysis/class-v6/floor/sram-and-floor.md` on class-v6-floor-sram at 708d01b4, first reading 12:0x UTC; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30); everything chip-side modelled on lane B's wire figure, 1.3 pJ per bit plus 0.1 nJ per macro access and 0.05 nJ of controller; the GPU side measured) **The read width is the only wire lever on the SRAM die, and the floor is a ticket, not a joule.** Lane B's 1.0 nJ was the 64-byte row; the die reads what the hash asks for, and the wire scales with the bits moved while the macro access does not: @@ -455,14 +455,14 @@ W = 16 is the one width that moves the die more than 2x at zero shadow and it co What it means: against the 16 GB Blackwell card at its knee (2.1 measured on the 5080; 1.7 to 2.1 modelled on the 5070 Ti and 5070) the GDDR7 chip is under 2x at k 1 and the SRAM die 2.7x; against the M5 Max 1.3x and 1.8x; against the LPDDR6 Apple part five years out the GDDR7 chip is at parity at k 1 and the SRAM die 1.5x. The only software lever is the operating point (the lock on Blackwell and Ada, the cap on Ampere, the ADLX offsets on AMD, nothing on Apple or Intel); the occupancy is worth 1 to 2 percent (10.1), the memory clock untouched, the block shape and the cache hint closed. The lane's per-tier scoring rule: size the shadow by the Apple tier's 5 percent point as the 2.0 rule already does and keep N at 100,000 (sizing to the Apple ceiling of 130,000 costs every honest miner 8 to 15 percent of electricity for about 0.15x of edge: the 5090's GDDR7 edge at k 1 goes from 2.13x to 1.95x); do not score the acceptance floor per tier (one network-wide number). The re-measure rule after a class flip: a stored set is stale under a new class, the stored point is the provisional start of the re-measure, max is stock and never stale; the measured v4 to v5 flip moved the 5090's knee by nothing. Owed and labelled: every Ada and Ampere knee is modelled because no rented host allows -lgc (14 pods today, every one refused); the rented stock rows are measured; the sweep adds class v5 watts measured on 16 card classes and a memory-clock try on each. -### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 19:30 UK) +### 10.5 Invention beyond the four layers (floor lane 5, `docs/analysis/class-v6/floor/invention.md` on class-v6-floor-invention, first KEEP/KILL table 13:4x UK; full by 15:30 UK (the earlier 19:30 clock pulled by the coordinator at 13:30)) **One lever found, conditional on one measurement: the second DRAM sector.** The identity's only `k = 1` work is the memory's own, and the record forced one 32-byte sector per read while the dataset item is 64 bytes. Reading the whole item (W = 16 words; the read-width branch's class w64, fingerprint 836e56e7d496e980, the verifier 0.630 against 0.604 ms) forces a second sector in the same open row: on the chip model's own inputs (909 pJ activate plus 1,150 pJ of movement per sector) the GDDR7 chip's read goes from 2.0 to 3.2 nJ and E_mem from 0.466 to 0.62 microjoules (+33 percent at an unchanged 166 MH/s), one HBM3 stack 0.321 to 0.40 (+24 percent), and on floor lane 3's wire figure the N2 SRAM die 0.036 to 0.126 microjoules (66x to 19x). So the band {1, 4} words is inside one sector and "the chip pays nothing" is right to 32 bytes and wrong at 64; this is the measured reason W = 8 is free for the chip too (section 10.3) and W = 16 is the only width that costs it. The card is the whole question: the 5090 ran w64 at 71.9 MH/s (-47 percent; the exported kernel at full occupancy, four uint4 loads per item, no request hint) while its own 64-byte probe reached 15.7 G reads a second at its best lane count (-13 percent); 589 GB/s is a third of the stream, so the bind is the request path, not the pins. The 9070 XT (-3 percent) and the M5 Max (0) are free at 64 bytes (measured). | Candidate | Verdict | Number (5090 knee, GDDR7 chip) | Label | |---|---|---|---| | The second sector, W = 16 words (w64) | KEEP rank 1, conditional | outcome A (the exported form, -47 percent of rate): 5.3x, dead; B (the probe's -13 percent): 3.4x; C (a one-request form within 5 percent of the 4-byte rate): 3.0x to 3.2x at zero shadow, 2.7x at k 0.5, 2.0x at k 1; the chip +33 percent whatever the card does; the M5 Max and the 9070 XT pay nothing | card measured at two occupancies (unlocked, no watts); chip modelled | -| The job that decides it | one PC 1 run, under an hour; asked of the hash lane at 13:5x UK, after kit d | the w64 pack plus the w4 control, the race's occupancy variants, a variant with `ld.global.L2::64B` on the first load (a variantSource anchor rewrite), both clocks, nvidia-smi at 1 Hz; the pass line 95 percent of the 4-byte rate, the kill line under it | owed; default if not run by 19:30 UK: outcome A stands, W = 16 stays `never` | +| The job that decides it | one PC 1 run, under an hour; asked of the hash lane at 13:5x UK, after kit d | the w64 pack plus the w4 control, the race's occupancy variants, a variant with `ld.global.L2::64B` on the first load (a variantSource anchor rewrite), both clocks, nvidia-smi at 1 Hz; the pass line 95 percent of the 4-byte rate, the kill line under it | owed; default if not run by 15:30 UK: outcome A stands, W = 16 stays `never` | | Tensor tiles | KILL as content; the k band corrected | merchant silicon at a nominal 0.30 to 0.56 pJ per INT8 MAC (MTIA v2 0.51, AI 100 Ultra 0.34, B200 0.44, MI355X 0.56, M4 ANE 0.30 measured; the H800 loop 0.34 measured); against the class's tile (1.5 pJ at the knee) k 0.2 to 0.4, against the dense tile (0.83) 0.4 to 0.7: never above the ALU band; "3x to 30x" needs an INT4 layer at 0.46 V (the VSQ chip reads 0.052 at its nominal 0.67 V); Apple -35 percent at 1,024 tiles kills it | claimed, measured | | The RT core | KILL | traversal implementation-defined (Vulkan, DXR, the NVIDIA forum 14 Oct 2024, Blender's Cycles); the BVH opaque; dedicated units 4 to 33 nJ per ray against 290 to 750 measured board-level on an RTX 2080: k 0.01 to 0.2 | claimed, measured | | Channels per dollar, controller plus PHY | KILL as a lever; KEEP as a model correction (rank 2) | no 28 nm GDDR7 PHY exists (Cadence N3 2024, Innosilicon FinFET; the oldest GDDR6 PHY 12 nm): the record's USD 5 M project and 17 M cap row is unbuildable; the floor is 12 nm GDDR6 at USD 20 to 30 M (cap 70 to 100 M) or N5 GDDR7 at USD 50 to 75 M (cap 170 to 250 M); SK hynix ISSCC 2024 PAM3 I/O 1.10 TX plus 0.76 RX pJ per bit measured, E_mem moves under 0.06 | claimed, modelled |