From ec20c5e19c4db6cee375fbe96ba18c402a9fa39d Mon Sep 17 00:00:00 2001 From: igneum-labs <337424239+igneum-labs@users.noreply.github.com> Date: Tue, 6 Oct 2026 01:33:37 +0000 Subject: [PATCH] Counter ASIC 2.0 status 01:34: the PC 2 chain closed (job 6's curve, the ledger suites, the job-runner finding) Co-Authored-By: Claude Fable 5.1 --- docs/plans/counter-asic-2-status.md | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/docs/plans/counter-asic-2-status.md b/docs/plans/counter-asic-2-status.md index 15e17b101..d48122b9f 100644 --- a/docs/plans/counter-asic-2-status.md +++ b/docs/plans/counter-asic-2-status.md @@ -768,3 +768,13 @@ The crossing watcher (pid 49195) at 00:55:29Z: DAA 144,209, epoch 40, class v2, run-repro-pc2-20261006 (00:44:29Z fetch, the run 00:49:30 to 01:09:38Z, 1,509 s, exit 0, "all cards restored; done"). The package's bit-exact checks PASS on every backend (fingerprint 25f96e7dce90bd4e at 1 GiB on CUDA, NVIDIA OpenCL and gfx1036 OpenCL; program bcc1248b10cc90f2, 128 loads). The rates: RTX 5090 CUDA 62.412 MH/s (120 s, batch 2^24), 5090 OpenCL 62.283 (30 s), gfx1036 3.312 (the integrated RDNA 2, 1 CU); the memprobes 17.5 G dependent reads per second at 1 GiB on CUDA (chase 1,766 ns), 339 GB/s stream, 13.4 T ALU op/s. The 5090 rows are HALF this card's rate in every other run tonight (M16's honest 132.20 at 00:40Z; the app 105 to 116 wall) and the package's proving step took 33.9 s compressed and 20.2 s core on block-338-shard1 against 16.8 s in sweep 3, so the run most likely shared the card with the app's live CUDA worker: its card-off step reads nothing from /api/state (the "{}" defect) and waits 30 s without confirming a stopped worker. Reading passed to the repro agent to confirm from the job's own worker-count lines; if so, the PC 2 5090 rows are "beside the live miner" rows and the morning re-runs the 5090 block alone (gfx1036 and the memprobes stand). The script's close failed twice (ConvertTo-Json out of memory on the sample set; Set-Content with an empty path for the markdown), so no result file was written on PC 2; every row is in the intake. Package fixes, the class: the card-off step confirms by nvidia-smi's compute-apps list, the writer drops the samples from the JSON. "go PC 2" to agg-cost-pc2-5 at 01:10 (about 12 min), then the ledger-fixes-0311 suites. 01:13. Confirmed by the repro agent: the run's 5090 rows (62.4 CUDA, 62.3 OpenCL, 20.2 s core, 33.9 s compressed, 60.0 at 2^21 against 63.2 at 2^24) are beside-the-live-miner rows (the card-off POST was accepted but nothing confirmed a stopped worker; nvidia-smi showed 10,176 MiB used and 347 W after the run); the morning re-runs the 5090 block alone with the card-off confirmed by nvidia-smi's compute-apps list. And a withdrawal: the script read the prover's state from /api/state, got "{}", and did NOT switch the prover back on, so PC 2 has not proved since 00:49Z; the restore job run-prover-on-pc2-20261006 (POST api/prove on, about 10 s, no card) is published and runs on the app's queue; the time without the prover is the repro run plus the queue, about 25 to 35 minutes of PC 2's shard payouts. The close failures were a PowerShell case-insensitivity collision ($Cpu the result table over $cpu its own name string, so the table contained itself and ConvertTo-Json ran out of memory; $md the lines over $Md the path, so Set-Content got an empty path): both renamed; the class goes on the hygiene item for every playbook owner (no two variables that differ by case). agg-cost-pc2-5 failed at parse (01:10:44Z, 1 s, `$RestoreIdentities:` in a string at line 113); the fixed agg-cost-pc2-6 has the go (01:13), then the ledger suites. + +## 01:34 the PC 2 chain is closed: job 6's curve, the ledger suites on PC 2, the job-runner finding; the coordinator stops here + +agg-cost-pc2-6 (01:12:09 to 01:24:14Z, 725 s, exit 0; the parse fix held; the job's own miner alone on the 5090, the finally block put the card and the prover back at 01:24:09Z). The curve, the same four live blocks 96556 to 96559, one empty shard each: batch-log2 22: shards 7.8 to 8.1 s, chained aggregation 10.0 to 10.4 s, 18.1 s a block, 103.9 MH/s; 20: the same (18.0 s, 103.7); 18: 6.7 to 7.0 and 8.8 to 8.9 s, 15.6 s a block, 99.3 MH/s (minus 4.4%); 16: 4.9 to 5.1 and 6.1 to 6.2 s, 11.1 s a block, 83.8 MH/s (minus 19%), reproduced (11.1 s, 84.0). Against the card alone (4.1 s a block, 2.1 s an aggregation): the shortest kernel buys the prover 1.6x for a fifth of the hash rate and the slowdown stays 2.7x, so the 3-second aggregation on a mining card is not reachable by the kernel length; the defaults stay; the 2^16 trade and the batch fold (a new pinned guest, about 0.7 s a block alone by the step costs, estimated) are the project lead's decision in docs/plans/proving-v1.md "Aggregation cost (5 October, night)". Branch agg-cost ea38ece (worktree igneum-wt-agg-cost, on proving-v1's docs tip 9be5817, not pushed). + +The ledger-fixes-0311 suites on PC 2 (build-20261006-012543, published by me from the ledger-rebase worktree at 01:25:43Z, 01:26:15 to 01:32:33Z, 378 s): the Linux node build 137 s, the Windows node build 163 s, the seven suites exit 0 in 46 s, 7 files uploaded. The 46 s says the suites ran WITHOUT the igneum-pow feature (the G6 caveat of 00:29 applies: the lottery-hash consensus tests were not compiled; the Mac run with the feature is the suite evidence, 99/1/4 on kaspa-consensus with the M20 era test the one failure, 108/0/2, 52/0, 36/0, 14/0, 23/0, 20/0 on the rest). Fork tip fbb0082a, docs tip 9872299 (ledger-rebase). + +The job-runner finding (the aggregation-cost agent, 01:30): the app ran three jobs on PC 2 at once from 01:24Z (run-prover-on-pc2 at 01:24:21Z and the suites build at 01:26:15Z landed while nothing else of the queue was expected to run), whereas at 23:40Z PC 2's update-now queued behind the hung sweep ("1 new for this machine, 1 queued"). So the app serialises the jobs it fetches in one poll and runs jobs from different polls side by side; "one job on PC 2 at a time" held tonight by coordination only. Next-cut item (with the update-now item): the runner takes one job at a time per machine whatever the poll, and a `collect` or `update-now` may pre-empt. + +PC 2's prover: off from 00:49Z (the repro job) to 01:24:09Z (job 6's finally block), then confirmed on by run-prover-on-pc2 at 01:24:21Z; about 35 minutes of PC 2's shard payouts lost. Every PC 2 job of the night is closed; nothing is queued. The crossing watcher (pid 49195) runs to 08:00Z; the morning reads /tmp/igneum-devnet/crossing-154800.out first. The coordinator's final report to main follows this entry; the status file ends here unless the morning adds to it.