diff --git a/docs/bench-log.md b/docs/bench-log.md index 88c447890..262658aad 100644 --- a/docs/bench-log.md +++ b/docs/bench-log.md @@ -2665,3 +2665,21 @@ So the card does 1.41 G dependent random 4-byte reads a second over the 1 GiB da Consequences per tier (the rule of 5 October 2026): a B580 owner (12 GB, Windows) cannot mine today; the app lists the card and the worker fails its self-test, so the row should sit off with the reason until the worker carries the Intel fix (the branch's app makes the vendor `intel` and the words say why). When it does, about 10 to 11 MH/s is the expectation (the 0.9 latency share of the other cards applied to the 11.0 ceiling, approximate), about 2,800 blocks a day at the 312 MH/s the devnet showed at 10:54Z, at 150 to 190 W (approximate, no reading) about 0.06 MH/W, under the 9070 XT per watt (0.09) and about even per pound at a third of its price; mine only (the prover is CUDA only). 8 GB and 16 GB Intel rows owed; Linux and HiveOS take the same worker on `intel-opencl-icd`, owed a line. Next: the OpenCL family probe on the Arc (bit-exact per instruction family against the CPU reference) names the wrong family; the worker's Intel text for it; bench d. Found on the way: the Intel installer's bootstrapper wraps the inner installer's exit code as 1000 + code (1007 = an unsupported parameter, 1014 = restart required); with the device in Code 12 (Windows' inbox driver 32.0.101.6733 before any restart) the display driver does not bind and a second run after the restart is needed; the runner's `--stop-miners` hold on a run job is released only after the NEXT job ends; PC 1's three miners faulted on the node's block-template RPC (5-s fetches timing out at 8 identities a card), not on the Arc: with 2 identities a card the 5090 and the 9070 XT mine at 121 to 122.5 and 18.9 MH/s (`run-ia-cards-watch-20261007-d` and `-e`, 10:42Z to 10:50Z). + +**Bench d, the same day, 11:20:41 to 11:22:26Z, with the fix (job `run-ia-arc-bench-20261007-d`, the Arc alone, kit worker sha256 `3c62470f...` = `proto-opencl/host.c` at 26e135a3: on an Intel platform the one helper line `rotr_var(x, n) = rotate(x, (0u - n) & 31u)` is rewritten to the shift form before the build).** The register-trace bisect (`run-ia-arc-trace-20261007-b`, `tools/intel-arc/pc1-arc-trace.ps1`: lane 0's eight registers after the init and after every instruction, 55,809 snapshots, GPU against `igneum-pow trace-regs`) found the first divergence at instruction 6 of iteration 0, `rotr`: r0 = d1affd94, n = f183b605, the CPU a68d7fec = rotr(x, 5), the Arc 35ffb29a = rotl(x, 5): Intel's compiler folds the negated amount into a LEFT rotate; nothing before differed and nothing after agreed, in both exchange modes. With the rewrite (`run-ia-arc-trace-20261007-c`) the trace under the local exchange equals the CPU's in all 55,809 snapshots. + +| Pack | Exchange | Build ms | Self-test | Fingerprint 2^24 at base 0 | MH/s | +|---|---|---|---|---|---| +| v4-devnet-epoch0 (class v4, the current class) | sub-group (khr), size 32 | 703 | PASS, 96 of 96 lanes | f410c731b6bc2d31 (= Metal, Apple OpenCL, RTX 5090) | 11.011 | +| mx8-devnet-epoch0 (the class v3 control) | sub-group | 541 | PASS | 90f794dd556f7a3b (= the control everywhere) | 11.019 | +| the app's live pack (epoch 76, class v3) | sub-group | 553 | PASS | 4f385ad3c2793064 | 11.002 | +| v4-devnet-epoch0 | local memory, sub-group size 16 under that build | 1,084 | PASS | f410c731b6bc2d31 | 10.882 | + +| Row | Value | +|---|---| +| Intel Arc B580, class v4, OpenCL on Intel's runtime, driver 32.0.101.9034 | 11.0 MH/s (latency share 1.00 of the 11.02 ceiling; the class v4 shadow costs 0.1 percent against the control, the 9070 XT pays 3) | +| Watts and clocks | not exposed through OpenCL (IGCL next) | +| Card-picker entry (`site/yourcard.js`) | `['Intel Arc B580', 11.0]`, added | + +Consequences per tier, measured: a B580 owner (12 GB, Windows) mines at 11.0 MH/s once the shipped worker carries 26e135a3 (0.3.20): about 2,450 blocks a day at the 388 MH/s the devnet showed at 11:19Z (one every 35 s, about 77,600 IGN a day at 31.69 IGN a block; approximate, the rate moves), at 150 to 190 W (approximate, no reading) about 0.06 to 0.07 MH/W, under the 9070 XT per watt (0.09) and about even per pound at a third of its price; mine only (the prover is CUDA only; the 12 GB are not the limit). Until 0.3.20 the card sits off on PC 1 (persisted) and the branch's app row says why. 8 GB and 16 GB Intel rows owed; Linux and HiveOS take the same worker and the same rewrite on `intel-opencl-icd`. The AMD and NVIDIA paths see no change (`proto-opencl/test_intel_rotr.c`, in the gate). + diff --git a/docs/plans/intel-arc.md b/docs/plans/intel-arc.md index e91ba6748..7ef176622 100644 --- a/docs/plans/intel-arc.md +++ b/docs/plans/intel-arc.md @@ -121,9 +121,10 @@ Time to a block: the network estimate was 675 MH/s at 09:34Z on 7 October and 46 | Ceiling at 128 loads a hash | 11.0 MH/s | the same | | The hash, every pack, both exchange modes | builds in 236 to 274 ms; cache and dataset bit-exact; vectors 96 of 96 lanes WRONG, the same device value in the sub-group and the local-memory exchange | the same; the app's own worker saw the same from its first start | | Watts, clocks | not exposed through OpenCL | IGCL next (section 4) | -| Rate today | 0 MH/s (wrong hash); expectation once fixed about 10 to 11 MH/s, about 0.06 MH/W at 150 to 190 W (approximate) | section 6's row updated | +| Rate, bench d with the fix | 11.011 MH/s on class v4 (11.019 on the v3 control, 11.002 on the live pack, 10.882 local exchange), vectors PASS, fingerprints equal to the Mac's and the 5090's | `run-ia-arc-bench-20261007-d` | +| The fault, found by the register trace | instruction `rotr`: Intel's compiler folds `rotate(x, (0u - n) & 31u)` into a LEFT rotate by n; the worker rewrites the one helper line to the shift form on Intel (commit 26e135a3, `proto-opencl/intel_rotr.h`, the gate test `test_intel_rotr.c`); the trace then matches in all 55,809 snapshots | `run-ia-arc-trace-20261007-b` and `-c` | -So section 2.3's first gotcha (the sub-group size) is cleared: the runtime reports 32 for a 32-item work-group and the sub-group path builds; and the fallback ladder of section 4 is not the path either, since IGC builds the kernel. The fault is one or more per-lane instruction families computing differently on Intel's compiler (the 64-bit families are the suspects: `mulhi`, `mad`, `mul`, the rotates; the dataset derivation, 64-bit too, passes). The way to the fix: the OpenCL family probe on the Arc (`tools/intel-arc/pc1-arc-family.ps1`, `exact=no` names the family), then the worker's Intel text for that family behind the vendor check (an emulated form, as `dot4_c` is), then bench d. Hours: 2 to 4 from the probe's answer. The B580 owner's row until then: the card sits off with the reason (the branch's `intel` vendor and the row's words); PC 1 keeps it off, persisted. +Section 2.3's first gotcha (the sub-group size) is cleared (32 for a 32-item work-group) and the fallback ladder of section 4 was not needed: IGC builds the kernel, and the family probe (`tools/intel-arc/pc1-arc-family.ps1`, runs a and b) read every family exact, `mulhi` and `mad` included, because its `rotr` variant is the shift form. The register trace named the one line: the pack's `rotr_var` through the `rotate()` builtin with a computed negative amount, which Intel's compiler turns into the opposite rotate. The worker-side rewrite (26e135a3) is the whole fix: no pack, emitter or consensus text moves, AMD and NVIDIA see no change. The B580 owner's row: 11.0 MH/s from the cut that ships it (0.3.20); until then the card sits off with the reason. ## 8. What is not done here diff --git a/site/yourcard.js b/site/yourcard.js index 7c843070d..e8b32168e 100644 --- a/site/yourcard.js +++ b/site/yourcard.js @@ -2,7 +2,7 @@ second is the target, so the wait is network hashes per second over the card's rate. The rates are the bench table's measured numbers (site/miners.html); nothing here is a promise. Reads /api/live every 15 s. */ (function () { - var CARDS = [['NVIDIA RTX 5090', 127.7], ['NVIDIA RTX 4070', 28.8], ['Apple M5 Max', 26.7], ['AMD RX 9070 XT', 17.8]]; + var CARDS = [['NVIDIA RTX 5090', 127.7], ['NVIDIA RTX 4070', 28.8], ['Apple M5 Max', 26.7], ['AMD RX 9070 XT', 17.8], ['Intel Arc B580', 11.0]]; var els = document.querySelectorAll('[data-yourcard]'); if (!els.length) return; var net = null, chosen = (function () { try { return localStorage.getItem('igneum-card'); } catch (e) { return null; } })(); function dur(sec) { if (!(sec > 0)) return 'pending'; if (sec < 90) return Math.round(sec) + ' s'; if (sec < 5400) return Math.round(sec / 60) + ' min'; if (sec < 172800) { var h = Math.floor(sec / 3600), m = Math.round((sec % 3600) / 60); return h + ' h' + (m ? ' ' + m + ' min' : ''); } return Math.round(sec / 86400) + ' days'; } diff --git a/tools/intel-arc/pc1-arc-bench.ps1 b/tools/intel-arc/pc1-arc-bench.ps1 index e0278198a..633712f1d 100644 --- a/tools/intel-arc/pc1-arc-bench.ps1 +++ b/tools/intel-arc/pc1-arc-bench.ps1 @@ -14,7 +14,7 @@ # Lines start with RESULT; a SUMMARY {json} line ends the job. Read back with `node tools/jobs.mjs `. $ErrorActionPreference = 'Continue' $jobName = 'pc1-arc-bench' -$kitId = 'fetch-ia-arc-packs-20261007-b' +$kitId = 'fetch-ia-arc-packs-20261007-b' # the packs (class v4, the v3 control) $started = Get-Date function Stamp { (Get-Date).ToUniversalTime().ToString('yyyy-MM-ddTHH:mm:ssZ') } function Summary([string] $status, [hashtable] $extra) { @@ -37,6 +37,10 @@ $appDir = if ($env:IGNEUM_APP_DIR) { $env:IGNEUM_APP_DIR } else { Join-Path $env $exe = $null if ($inst -and (Test-Path (Join-Path $inst 'igneum-worker-opencl.exe'))) { $exe = Join-Path $inst 'igneum-worker-opencl.exe' } if (-not $exe) { "RESULT error no installed igneum-worker-opencl.exe under $inst"; Summary 'failed' @{ error = 'no worker' }; exit 2 } +# bench d (7 October 2026): the worker with the Intel rotate fold rewritten (host.c intelRotrPatch) from the trace kit, +# when the kit is there; else the installed one (the kit-path rule: tested before use) +$patchedKit = Join-Path $jobs 'fetch-ia-trace-kit-20261007-b' +if (Test-Path (Join-Path $patchedKit 'bin\igneum-worker-opencl.exe')) { $exe = Join-Path $patchedKit 'bin\igneum-worker-opencl.exe'; "RESULT worker_source patched-kit $patchedKit" } else { "RESULT worker_source installed" } "RESULT app install=$inst app_dir=$appDir worker=$exe sha256=$((Get-FileHash -Algorithm SHA256 $exe).Hash.ToLower())" # ---- detect: Windows' rows for the Intel device, then the worker's own list ---- @@ -143,6 +147,9 @@ function Bench([string] $label, [string] $packDir, [string[]] $extra) { $lines | Where-Object { $_ -match '^RESULT|^pack |^warm-up|^build |^kernel:|^exchange|sub-group|FAIL|error|MISMATCH|spill|private memory|^timeout' } | Select-Object -First 30 | ForEach-Object { Say "RESULT bench $label $_" } Say "RESULT bench $label end $(Stamp) exit=$rc wall_s=$([int]((Get-Date).ToUniversalTime() - $t0).TotalSeconds)" $res = $lines | Where-Object { $_ -match '^RESULT pack=' } | Select-Object -First 1 + # the child's exit code can read null after WaitForExit on a -PassThru process (bench d, 7 October 2026: every pack + # PASSed and the rows said "exit , no RESULT line"); a present RESULT line with check=PASS is the verdict + if ($null -eq $rc -or "$rc" -eq '') { $rc = 0 } if (-not $res -or $rc -ne 0) { $err = ($lines | Where-Object { $_ -match 'FAIL|error|MISMATCH' } | Select-Object -First 1); if (-not $err) { $err = "exit $rc, no RESULT line" } Say "RESULT G1 pack=$label error=$($err -replace '\s+', ' ')"