Chip model 5.11 and new-pow 5.1: the per-MAC figures corrected by the per-lane tile count (32 MACs per tile, not 1,024; the 4090's 0.056 pJ is 1.8 pJ, the 5090 reads 2.9 and 1.5 at the shadow's premium), the tensor-tile column withdrawn (a chip's k on tile work 0.03 to 0.3, below the ALU shadow's); Counter ASIC 3.0 status: the microbench's correction, the night's closing sentence on measured rows, the CA4 file's commits

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-08 06:08:07 +00:00
parent aa71d750c6
commit b7ea50ef69
3 changed files with 5 additions and 3 deletions

View file

@ -363,6 +363,8 @@ Per tier: a miner on class v4 pays the premium and gets the 2.1x to 3.9x chip ce
The k column. k is the chip core's energy per op over the GPU's at the same operating point, and the model's 3.4x row takes k about 0.33 for an ALU-shaped core. On the public figures the int8 tensor tile is the one GPU block whose energy per op a chip at the same node cannot undercut with certainty: the 4090 measures 0.056 pJ per MAC; NVIDIA's 5 nm INT4 test chip reads 0.021 pJ per MAC at 0.46 V and about 0.1 at nominal (JSSC 2023, via Dally's NASEM slides; claimed), so INT8 at 2x to 4x that gives k 0.7 to 3 with the centre near 1; every ALU-shaped block reads k 0.3 to 0.8 on the same sources. A shadow built of tensor tiles at the ALU shadow's premium (about 11,400 u8 tiles per hash) therefore gives 2.1x at k = 1 and 1.6x at k = 1.5 and removes the k 0.3 column from the table; it needs a SIMD byte-dot verifier (the scalar one at 12.4 ms fails the 10 ms gate). This is a design candidate, not the shipped stream: the shipped shadow is ALU-shaped and its row stays 2.1x at k = 1 and 3.4x at k about 0.33. The k column. k is the chip core's energy per op over the GPU's at the same operating point, and the model's 3.4x row takes k about 0.33 for an ALU-shaped core. On the public figures the int8 tensor tile is the one GPU block whose energy per op a chip at the same node cannot undercut with certainty: the 4090 measures 0.056 pJ per MAC; NVIDIA's 5 nm INT4 test chip reads 0.021 pJ per MAC at 0.46 V and about 0.1 at nominal (JSSC 2023, via Dally's NASEM slides; claimed), so INT8 at 2x to 4x that gives k 0.7 to 3 with the centre near 1; every ALU-shaped block reads k 0.3 to 0.8 on the same sources. A shadow built of tensor tiles at the ALU shadow's premium (about 11,400 u8 tiles per hash) therefore gives 2.1x at k = 1 and 1.6x at k = 1.5 and removes the k 0.3 column from the table; it needs a SIMD byte-dot verifier (the scalar one at 12.4 ms fails the 10 ms gate). This is a design candidate, not the shipped stream: the shipped shadow is ALU-shaped and its row stays 2.1x at k = 1 and 3.4x at k about 0.33.
Correction, 8 October 2026 (the research lane's microbench on the 5090, counter-asic-4-research.md 15.1a and the corrected 20.3 and 20.4 at 71fd465b): the per-MAC figures above are wrong by a factor of 32. A `mma.m8n8k16` tile is 1,024 multiply-adds per warp, 32 per lane, so a hash does 32 MACs per tile, not 1,024; the 4090's "0.056 pJ per MAC" is 1.8 pJ, and the 5090 at the ALU shadow's premium reads 2.9 pJ per MAC unlocked and 1.5 pJ at the 1,300 MHz lock (the packs job, 366,080 MACs per hash; the microbench's dependent u8 tile 4.1 and 2.2, the wide s8 m16n8k32 tile 1.36 and 0.83). Against the same 5 nm MAC array figures (0.04 to 0.4 pJ per INT8-class MAC, claimed) a chip's k on tile work is therefore 0.03 to 0.3, below the ALU shadow's 0.3 to 0.8, not near 1: at the same premium a tensor-shaped shadow leaves the chip 3.5x to 6.7x where the ALU shadow leaves it 2.1x to 3.5x. The tensor-tile column (2.1x at k = 1, 1.6x at k = 1.5) is withdrawn as a candidate; its premise, that a chip's MAC is no cheaper than the GPU's, is false by 4x to 30x on the public figures. The shipped row is unchanged: 2.1x at k = 1 and 3.4x at k about 0.33, the ALU shadow at the operating point's knee, measured four times at 82 to 90 W.
The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job. The capex column. The `f = 1` GDDR7 chip of 5.5 is USD 2.8 per MH/s of silicon and memory, which is USD 0.00016 per MH/s-hour of capex over two years against USD 0.000023 of electricity: capex-dominated 7x, as the 5090 is (USD 14.7 per MH/s at MSRP, 10x). A 64 MiB hot table adds about USD 15 of N5 die, the shadow core USD 25 to 40, an interposer USD 200, so the chip's capex reaches at most about USD 4.3 per MH/s: the per-unit capex wall is unreachable by 3x to 7x, and the break-even market cap moves only through the project cost (the mission lane's model: about USD 100 M with the N5 shadow core, about 200 M if the shadow runs per load and forces one die or an interposer; the per-load form behind that figure, the 16 x 27 placement, was closed on 7 October 2026 at night when it failed the value-level acceptance test across drawn eras, so the 200 M row rests on no construction shown to exist until a sound per-load class, one pass of a 432-instruction sub-block per load, is drawn, accepted and measured). Every figure here is modelled on cited or claimed parts; the research lane's microbench (20 probes, the mma_u8 and l2 rows the ones this model would take) is on PC 1's queue after the hot-table job.
## 6. The per-day derivation (item 2) ## 6. The per-day derivation (item 2)

View file

@ -189,7 +189,7 @@ Sustained rates over the power window 63.07 to 63.03 MH/s at every R; wall and e
What the rows say: What the rows say:
1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5). 1. The tensor block is free in hash rate to R = 512 on the 4090: 4,096 tile instructions per hash leave the rate at 63.08 MH/s to the third decimal. The kernel is latency-bound on its 128 dependent loads and the tensor work fills stalls that were already there, as the ALU shadow did on the 5090 to 150,800 ops (`latency-shadow-2026-10-06.md` 5).
2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised), 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that. 2. The block costs the honest card almost nothing in energy: 2.9 to 14.7 W, 0.70 pJ per multiply-add at R = 8 falling to 0.056 pJ at R = 512 (the tensor path's fixed cost amortised) [corrected 8 October 2026: these per-MAC figures count 1,024 multiply-adds per tile per lane where a tile is 1,024 per warp and 32 per lane, so they are low by 32x: 22 pJ at R = 8 falling to 1.8 pJ at R = 512; the watts and microjoules per hash stand; counter-asic-4-research.md 15.1a at 71fd465b], 0.05 to 0.23 microjoules per hash on a 3.19 microjoule hash (+1.6 to +7.2 percent). The ALU shadow at N = 100,000 costs the 5090 0.6 microjoules per hash (11 pJ per counted op, `latency-shadow-2026-10-06.md` 5, item 4); the tensor block at its free-band ceiling costs a third of that.
3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force. 3. That is the finding, and it is negative for the scheme's purpose (section 5.3): a shadow lever moves the chip's edge only by the joules it makes the HONEST card spend on work the chip cannot do more cheaply. The tensor path is so efficient on the GPU that the block adds 0.23 microjoules at most, so at `k = 1` the chip's edge falls from 6.9x to 4.9x on GDDR7 against this 4090, where the ALU shadow took the 5090 from 5.6x to 2.1x, and the tensor block costs the verifier 26x more per unit of chip-forcing energy (4.39 ms scalar per 0.23 microjoules against 0.17 ms per 0.6 microjoules). The property the design hoped for (`k_mma` near 1 because the GPU's tensor engine is near the floor) is real and is exactly why the lever is weak: there are no joules in it to force.
4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5. 4. The correctness chain holds at every rung: the PTX fragment read and the plain-integer reference agree on all 2^24 lanes at every R, and the CPU interpreter matches the GPU on 1,024 lanes at every R; the probe's fragment layout (`family-probe.cu` mm8 `warp_ref`) was used as written and needed no correction. Registers 29 to 36, occupancy unchanged. This is the first class-shaped evidence that an `mm8` family is cheap and exact for the honest NVIDIA card, which is what the reserve entry R8 needs; it is not evidence for a class v5.
@ -273,7 +273,7 @@ Reading: B moves the f = 1 chip's edge by 1.4x to 2x at `k = 1` and by 1.1x at `
| Scheme | Verdict | Why, in one line | | Scheme | Verdict | Why, in one line |
|---|---|---| |---|---|---|
| A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source | | A, mining is proving | NEVER (A1, A2); A0 folds into C | one proof per segment is not a distribution of puzzles; the bytes (2.9 MB of openings per block) or the verify (32 to 40 ms) kill every form that is not "hold the trace", and holding the trace is C with a worse data source |
| B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add, 15 W for 4,096 tiles per hash) is the reason. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules | | B, tensor-shaped shadow | NEVER as class v5 content for the anti-chip purpose; the measurement (0.056 to 0.70 pJ per multiply-add as first counted, 1.8 to 22 pJ with the per-lane count corrected on 8 October 2026, 15 W for 4,096 tiles per hash) is the reason, and the correction strengthens it: a chip's MAC at 0.04 to 0.4 pJ (claimed) against the GPU's 1.8 pJ gives a chip k of 0.03 to 0.3 on tile work, below the ALU shadow's. KEEP the `mm8` family as reserve R8 with the two-output correction, for datapath diversity, not for joules |
| C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold | | C, stored state | SHIP AS CLASS v5 CANDIDATE (through the spec items of 4.3 and the Devnet 2 gate): hash rate and watts unchanged by construction and measured equal, build +1.4 ms, verifier +0.11 to 0.21 ms per unit, bit-exact on 1,024 items and 128 lanes; a new property per block (a random sample of state) and a new requirement per mining operation (hold the state); the open decision is what a header verifier is asked to hold |
**A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C. **A, in full.** The mandate asked for something never done, and "mining is proving" is the thing everybody has wanted and nobody has shipped; this lane's contribution is the reason, stated as a bound rather than a feeling: the useful fraction of a proving-as-lottery scheme is (proving work per segment) / (network hashes per segment), 8 percent at 1 GH/s and 0.08 percent at 100 GH/s on this chain's measured figures, because gas sets one and the security budget sets the other, and a puzzle whose verifier either recomputes the piece or verifies a 32 to 40 ms proof cannot sit under a 10 ms gate. The 80/20 split stays. Ledger F13's answer stands and gains this bound. What survives (A0) is scheme C.

File diff suppressed because one or more lines are too long