Bench log: the observer gap was the Mac hibernating

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
igneum-labs 2026-10-04 14:41:05 +00:00
parent 53cb492de3
commit 441b773033

View file

@ -942,3 +942,12 @@ Hash-rate steps (hop.sh; 4 threads on a 2-vCPU VM is x1.6 to x2.7 per node, meas
Reading. Latency: the real inter-region path is 0.3 to 0.7 s per block for 99% of blocks, under the 5-s bound behind GHOSTDAG k and well under spec 03 C1's 2-s inter-region assumption. Partition: the floor rule did what the design says (DECIDED 4 Oct 2026, O-3.15): the 15.8% side locked nothing, the 84.2% side locked every 30-s interval, no index was certified twice, and both sides were on one chain within 15 s of the heal with the minority's 420 to 500 blocks reorged out. The margin is thinner than the weights suggest: certificates carried 8 to 10 of 12 votes with everything connected (signed 66.7% to 89% of total at 14:00 UTC), so during the cut the majority locked at 66.8% twice; the missing votes are the open question (finality owner). Controller: with 12 small steady miners the dual-lane rule overshoots every step (x1.6 to x1.7 on x1.3 to x1.4, x0.53 on x0.75), swings ±20% for 20 min after a step up and then damps, and holds the rate 10 to 20% under target for 10 to 15 min after a step down; none of the settle times is near the simulator's 62 s for x50, and the /1.42 step had not settled in 15 min. The live devnet's two bursty GPU miners oscillate without damping; this network damps. Both regimes are now measured on real timestamps.
What failed: the first partition ran before the window had filled (no lock possible) and a second was added after the first lock, about USD 0.30; the script's sink-count heal criterion reported 539 and 768 s and is tip churn (replaced with the chain-removal heal time); its first-lock poll reported 807 s because it started after that wait (replaced with the journals); the Mac hibernated on a flat battery from 11:56 to 13:16 UTC during the hop schedule, so the 4-thread phase ran 94 min instead of 15 (the scripts now hold the machine awake with caffeinate); the schedule's 4 threads gave x1.42, not x2.5. Another agent's difficulty-v2 rollout (`infra/cloud-devnet/rollout-v2.sh`, activation DAA 16,170) restarted every node one at a time between 14:14:40 and 14:18:05 UTC, the last 56 s of the second cut and its heal window: the cut-phase locks, the reorg depths and both minority heal times are unaffected (the minority reorgs at 14:15:47 and 14:15:51 came before those nodes' rollout restarts); the majority's first lock after the heal (448, 13 s) was on igneum-01 69 s after its own restart and carries that caveat. Finding for the execution owner: both minority nodes logged "[igneum-exec] reorg deeper than the snapshot ring; replaying from genesis" at both heals. Cost: USD 2.82 net by 14:34 UTC, USD 0.57 per hour while the network stays up (it was left running). Caveats: CPU hash rate only, one afternoon, 12 nodes not 1,000, clocks by chrony.
## 4 October 2026, correction: the 78-minute observer gap was the Mac hibernating
The "observer fell 45 minutes behind" entry left the cause open. The cloud-devnet agent's logs show the Mac hibernated
on a 1% battery from 11:56 to 13:16 UTC: node 1, the observer node, `observer.mjs`, the Mac miner and every agent on
the machine stopped; the two PCs, the seed and the cloud network carried the chain (no gap in the chain itself: every
block the observer later stored carried its original header time). The lag-proofing stays (it is right on its own), the
restart loop stays, and the public-face watch now runs from the same machine, so it also sleeps when the Mac does; the
fix is the charger and keep-awake, not software. The cloud scripts now re-exec under `caffeinate -i`.