Counter ASIC 2.0 status 01:16: the repro reading confirmed, the prover off on PC 2 since 00:49Z, job 5's parse failure
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
parent
bdf4237b23
commit
ea6f946587
1 changed files with 2 additions and 0 deletions
|
|
@ -766,3 +766,5 @@ The crossing watcher (pid 49195) at 00:55:29Z: DAA 144,209, epoch 40, class v2,
|
|||
## 01:10 the repro run closed; its 5090 rows read as beside the live miner; agg-cost-pc2-5 has its go
|
||||
|
||||
run-repro-pc2-20261006 (00:44:29Z fetch, the run 00:49:30 to 01:09:38Z, 1,509 s, exit 0, "all cards restored; done"). The package's bit-exact checks PASS on every backend (fingerprint 25f96e7dce90bd4e at 1 GiB on CUDA, NVIDIA OpenCL and gfx1036 OpenCL; program bcc1248b10cc90f2, 128 loads). The rates: RTX 5090 CUDA 62.412 MH/s (120 s, batch 2^24), 5090 OpenCL 62.283 (30 s), gfx1036 3.312 (the integrated RDNA 2, 1 CU); the memprobes 17.5 G dependent reads per second at 1 GiB on CUDA (chase 1,766 ns), 339 GB/s stream, 13.4 T ALU op/s. The 5090 rows are HALF this card's rate in every other run tonight (M16's honest 132.20 at 00:40Z; the app 105 to 116 wall) and the package's proving step took 33.9 s compressed and 20.2 s core on block-338-shard1 against 16.8 s in sweep 3, so the run most likely shared the card with the app's live CUDA worker: its card-off step reads nothing from /api/state (the "{}" defect) and waits 30 s without confirming a stopped worker. Reading passed to the repro agent to confirm from the job's own worker-count lines; if so, the PC 2 5090 rows are "beside the live miner" rows and the morning re-runs the 5090 block alone (gfx1036 and the memprobes stand). The script's close failed twice (ConvertTo-Json out of memory on the sample set; Set-Content with an empty path for the markdown), so no result file was written on PC 2; every row is in the intake. Package fixes, the class: the card-off step confirms by nvidia-smi's compute-apps list, the writer drops the samples from the JSON. "go PC 2" to agg-cost-pc2-5 at 01:10 (about 12 min), then the ledger-fixes-0311 suites.
|
||||
|
||||
01:16. Confirmed by the repro agent: the run's 5090 rows (62.4 CUDA, 62.3 OpenCL, 20.2 s core, 33.9 s compressed, 60.0 at 2^21 against 63.2 at 2^24) are beside-the-live-miner rows (the card-off POST was accepted but nothing confirmed a stopped worker; nvidia-smi showed 10,176 MiB used and 347 W after the run); the morning re-runs the 5090 block alone with the card-off confirmed by nvidia-smi's compute-apps list. And a withdrawal: the script read the prover's state from /api/state, got "{}", and did NOT switch the prover back on, so PC 2 has not proved since 00:49Z; the restore job run-prover-on-pc2-20261006 (POST api/prove on, about 10 s, no card) is published and runs on the app's queue; the time without the prover is the repro run plus the queue, about 25 to 35 minutes of PC 2's shard payouts. The close failures were a PowerShell case-insensitivity collision ($Cpu the result table over $cpu its own name string, so the table contained itself and ConvertTo-Json ran out of memory; $md the lines over $Md the path, so Set-Content got an empty path): both renamed; the class goes on the hygiene item for every playbook owner (no two variables that differ by case). agg-cost-pc2-5 failed at parse (01:10:44Z, 1 s, `$RestoreIdentities:` in a string at line 113); the fixed agg-cost-pc2-6 has the go (01:13), then the ledger suites.
|
||||
|
|
|
|||
Loading…
Reference in a new issue