igneum/docs/plans/intel-arc.md

24 KiB

Intel Arc B580 on Igneum: the first Intel card (plan, 7 October 2026)

Branch intel-arc (worktree ../igneum-wt-intel-arc, scratch prefix ia-*). the project lead, 7 October 2026, 13:4x UK: "the intel card has arrived ... it will go into pc1 and replace one of the cards, it is the b580". CLAUDE.md names Intel as a tier "when it exists"; it exists today. This page says which worker runs on the card first, what to expect, what can go wrong, and the fallback with its hours. Every number below is either cited or labelled approximate; the measured rows land in docs/bench-log.md when the PC 1 job runs.

1. The card

Fact Value Source
Name, generation Intel Arc B580, Battlemage (BMG-G21), Xe2 architecture Intel's product page, approximate (not re-read today)
Memory 12 GB GDDR6, 192-bit, 19 Gbps, 456 GB/s the brief; Intel's specification, approximate
Compute 20 Xe2 cores (160 vector engines, SIMD16 each), 2,670 MHz boost (approximate) approximate
Board power 190 W TBP, one 8-pin, two slots, PCIe 4.0 x8 approximate
Cache 18 MB L2 (approximate) approximate
Driver for Windows Intel Arc Graphics driver 32.0.101.9034 (non-WHQL, released 2 October 2026; 9033 is the WHQL one), gfx_win_101.9034.exe, 932,631,144 bytes, SHA512 3fc6f9b4...3cf3f0 Intel download page 785597, read 7 October 2026; the file is on PC 1 (job run-ia-intel-driver-20261007, SHA512 equal, Authenticode "CN=Intel Corporation")
Compute APIs in the driver OpenCL 3.0 (platform "Intel(R) OpenCL Graphics", the Intel Graphics Compute Runtime with the IGC compiler), Level Zero, DirectX 12, Vulkan Intel's runtime documentation, approximate
As the installed worker lists it (PC 1, 7 October 2026, 10:07Z, on Windows' inbox driver 32.0.101.6733) `[2] Intel(R) Arc(TM) B580 Graphics Intel(R) OpenCL Graphics (OpenCL 3.0); GPU, vendor Intel(R) Corporation, OpenCL C 1.2, 160 compute units, 2850 MHz; global 11928 MiB, max alloc 11928 MiB, local 128 KiB, max work-group 1024, sub-group extension: cl_khr_subgroup_shuffle`

So the runtime lists the Khronos sub-group shuffle extension: the worker's --exchange auto tries sub_group_shuffle_xor first and keeps it only if the kernel's sub-group size at a 32-item work-group reads 32 (section 2.3); 160 "compute units" are the 160 vector engines (20 Xe cores x 8), and the whole 11.65 GiB is one allocation (no 4 GB cap as on AMD and NVIDIA).

PC 1 on 7 October 2026 (ae432dc7, Windows 11, the installed app 0.3.18): RTX 5090, RX 9070 XT (eGPU), the Ryzen iGPU gfx1036, and the B580 in the RTX 4070's slot (the 4070 went out; Windows' PnP list still carries its ghost row, present=False). The driver story of the day: Windows' inbox driver 6733 bound the card with Code 12 until the restart; run b of the driver job (tools/intel-arc/README.md) installed Intel's auxiliary packages but the 9034 display driver did not bind while the device was in Code 12; after the restart the card is OK, Code 0, on 6733 with a working OpenCL 3.0 platform; run c binds 9034.

2. Which worker first: the OpenCL worker, unchanged

The app drives NVIDIA by the CUDA worker, AMD by the OpenCL worker (proto-opencl/host.c, shipped as igneum-worker-opencl.exe, OpenCL.dll loaded at run time) and Apple by Metal. The Intel driver ships an OpenCL 3.0 platform, so the first run is the shipped OpenCL worker on Intel's runtime with nothing rebuilt: detect.rs already reads the worker's --list for "AMD, Intel" and assemble makes a card of every non-NVIDIA GPU device it prints. The 4 October United States laptop mined on an Intel UHD iGPU through this path at 1.46 MH/s (bench-log, 4 October 2026), so Intel's compiler has already built the kernel text once; the Arc is the first discrete Intel card.

What the installed 0.3.18 app will do by itself the moment the driver is in: list the Arc as vendor "other", kind discrete (the name carries "Arc"), 12 GB, enabled with 8 identities, and start an OpenCL worker on it with --device <index> --pack packs\devnet. Its ready line or its error is the first light; the PC 1 job reads it before the bench (section 5).

2.1 The expected ceiling

The hash is bound by dependent random 4-byte reads over the 1 GiB dataset, 128 per hash (bench-log, 5 October 2026, "the 9070 XT on the eGPU"; docs/plans/read-width.md section 1). The measured dependent-read rates at 1024 MiB:

Card Memory bandwidth Dependent random loads at 1 GiB Rate at 128 loads per hash Watts MH/W
RTX 5090 1,792 GB/s, 96 MB L2 16 to 18 G/s 127 to 137 MH/s measured 227 W tuned (Ember run 6) to 311 W 0.41 to 0.56
RX 9070 XT 640 GB/s, 64 MB Infinity Cache 2.42 to 2.68 G/s 18.0 to 18.2 MH/s measured (87 to 95% of the ceiling) 199 W 0.09
Apple M5 Max 546 GB/s unified 3.45 G/s 23 to 28 MH/s measured 38 W 0.70
Arc B580 (expected) 456 GB/s, 18 MB L2 1.7 to 2.5 G/s (approximate: the 9070 XT's rate scaled by bandwidth gives 1.8; a 64-byte line per 4-byte read on both) 13 to 20 MH/s, the job decides 150 to 190 W under load (approximate) 0.07 to 0.13

So the first-light expectation is "a 9070 XT-class rate at a 9070 XT-class efficiency", 13 to 20 MH/s, and the job's --memprobe row replaces the guess with the card's own dependent-load ceiling. A B580 sells for about a third of a 9070 XT (approximate, street prices), so the per-pound figure is the one to watch.

2.2 Memory: 12 GB

Mining needs about 3.4 GB resident (the 1 GiB dataset, the 256 MiB cache, the program and the context: memory.used min 3,396 MiB on the 5090 with the miner alone, bench-log 5 October). 12 GB holds that three times over. Proving is CUDA-only today (SP1's GPU prover; the app's words: "cannot prove yet (the prover is CUDA only)"), so an Arc owner mines only, by the prover's vendor and not by the 12 GB; and even on NVIDIA a 12 GB card cannot mine and prove at once on this build (15.6 GB peak, CLAUDE.md consequences rule). The row for a B580 owner therefore reads "mine only" twice over, and the prover question for Intel is a Level Zero or SYCL port of the prover, out of scope here.

2.3 The known Arc OpenCL gotchas, and where the worker already stands

Gotcha What it does to this kernel Where the worker stands
Sub-group size: Xe2 executes SIMD16 natively; IGC picks a sub-group size of 8, 16 or 32 per kernel and cl_intel_subgroups lets the kernel require one (intel_reqd_sub_group_size(32)) the worker's lane unit is 32 (one warp = one 32-lane work-group); a 16-wide sub-group means two sub-groups per unit, so the sub-group shuffle exchange is wrong at width 16 setupProgram already handles it: it asks clGetKernelSubGroupInfo for the size at a 32-item work-group and falls back to the local-memory exchange unless it reads 32 (exchange 2 = intel_sub_group_shuffle_xor, host.c line 508; the probe kernel igneum_probe_subgroup is the second check). The local-memory path is the one AMD runs (wave32, exchange 0) at no measured cost against sub-group shuffles (the 5090's OpenCL path is 4% under CUDA). First run: --exchange auto; a --exchange local comparison row in the same job
Local memory: 64 KB SLM per Xe core shared by the work-groups resident on it (approximate); a big local footprint cuts occupancy the exchange buffer is small (32 lanes x a few words); the dataset and cache live in global memory no change; the kernel-info line prints local memory N bytes, read in the report
Private memory and spills: IGC spills large per-lane state to scratch, and the private memory figure in the kernel query is the signal the hash keeps its state in registers on AMD (private memory 0); a non-zero figure on Intel means spills, and the rate would show it the job prints the kernel-info line; a spill is a finding, not a blocker
64-bit integer multiply: Xe has no native 64-bit integer multiply (emulated from 32-bit, approximate 3 to 4x the cost of a 32-bit one) the program's mul, mulhi and mad families are 64-bit; on a latency-bound hash this is in the shadow of the memory reads, and the shadow work of class v4 (sh256x27) is the one place the ALU budget could bind on a weak-ALU card the job's class v4 row against the class v3 control shows whether the card holds its rate under the shadow; the 5 percent rule of Counter ASIC 2.0 is the test
The compiler: IGC compiles OpenCL C 1.2 and 2.0 fine; -cl-std=CL3.0 with features the device lacks fails the build; generic address space and cl_khr_fp64 are the usual failure sites; build logs are long the kernel is C99-style OpenCL C 1.2 on the local-memory path (-cl-std=CL1.2), CL2.0 or CL3.0 only for the sub-group variant if the sub-group variant fails to build, setupProgram already falls to local memory with the log printed ("the sub-group variant did not compile")
Compile time: IGC takes seconds for a kernel of this size (the UHD iGPU built packs in 3.0 to 6.4 s, bench-log M11) against 18 ms on Metal and 150 to 180 ms on NVRTC the per-epoch compile-ahead window is 600 s (docs/plans/epoch-length.md); 6 s is 1 percent of it fine; the job records build <ms> clBuildProgram for the Arc
Two Intel platforms: the driver may also expose a CPU OpenCL platform (Intel's CPU runtime, when installed) and, with an Intel iGPU present, the same "Intel(R) OpenCL Graphics" platform lists both GPUs a CPU device would be a false card dedupe_platforms keeps GPUs only (is_gpu); PC 1 has no Intel iGPU
Power: OpenCL exposes no power limit and no power reading; Intel's control path is IGCL (the Intel Graphics Control Library, ControlLib.dll, shipped with the driver: ctlPowerTelemetryGet gives an energy counter, clocks, temperature; ctlPowerLimitsSet sets the sustained limit, approximate names from the SDK) the app's cap slider is NVIDIA's (nvidia-smi); AMD goes through ADLX in igneum-gpu-telemetry the Cards row says so (section 3); an IGCL reader in the telemetry helper is the next piece, scoped in section 4
Driver resets: an Intel display driver reset (TDR) kills the context; the worker exits and the miner restarts it (WORKER FAULT) the app's watchdog path, already exercised on AMD nothing new

3. The app side (this branch)

Piece Change
Vendor word detect.rs: vendor_of and ClDevice::vendor_word return "intel" for a vendor string or name carrying "Intel" (the OpenCL vendor string is "Intel(R) Corporation", the device name "Intel(R) Arc(TM) B580 Graphics"); the card's key becomes intel:Intel(R) Arc(TM) B580 Graphics. The jobs runner's --cards-off matches that key with or without the index
Kind classify_kind already answers discrete for an Arc A-series or B-series name; the integrated guard now reads the Arc-branded iGPUs ("Intel(R) Arc(TM) Graphics" of Meteor Lake, "Arc(TM) 1xxV" of Lunar Lake, "Arc(TM) 1xxT" of Panther Lake, approximate names) as integrated and the "Arc(TM) A", "Arc(TM) B" and "Arc(TM) Pro" cards as discrete
Badge the Cards row badge reads INTEL in the Intel blue token (--intel), light and dark
Defaults unchanged: discrete, 12 GB, on with 8 identities
Power cap sweep.rs::unsupported_reason: "intel" answers "not available: Intel Arc exposes no power cap or reading through OpenCL (the driver's IGCL counters come next)"; the row's details panel says "Intel Arc exposes no power cap through OpenCL; the card runs at the driver's default limit (Intel Graphics Software can set one). No watts reading yet, so the MH/W cell stays empty until the IGCL reader lands." The tune line says "tuning: measure only on Intel Arc once a power reading exists"
Tests detect.rs: an Intel platform line parses to vendor intel, a B580 card named and keyed, Arc iGPU names integrated, Arc cards discrete; view.test.mjs: the Intel discrete row's words, badge and no-reading sentence; the mock scenario intel for the capture

4. The fallback if OpenCL fails on Arc, scoped in hours

"Fails" means: the kernel does not build on IGC in either exchange mode, or it builds and the vectors mismatch, or it runs under 5 MH/s with the memprobe saying the card can do more. In that order the fallbacks are:

Step What Hours (Claude side, per the project lead's rule: hours, not weeks)
F1: SPIR-V into the same OpenCL runtime build the kernel text offline with clang (-cl-std=CL1.2 -target spir64-unknown-unknown -emit-llvm then llvm-spirv) on the box, feed it with clCreateProgramWithIL: this skips IGC's OpenCL C front end and keeps every other line of host.c 2 to 3 hours plus one PC 1 job; needs clang and llvm-spirv on igneum-build-1 (apt)
F2: a Level Zero host a second worker, proto-l0/host.c, mirroring host.c's serve loop (the job protocol, pack read, dataset and cache fill, the bound kernel, read-back) on ze* calls: zeModuleCreate with the SPIR-V from F1 (ZE_MODULE_FORMAT_IL_SPIRV), zeCommandListAppendLaunchKernel, events for timing; Level Zero's loader ze_loader.dll ships with the driver 6 to 10 hours for the worker plus 2 for the app (a third worker kind beside CUDA and OpenCL in detect.rs and the launcher), one PC 1 job per round
F3: a SYCL (oneAPI) build the kernel ported to SYCL and built with icpx on the box; the largest rewrite and the most toolchain 12 to 16 hours; last resort
Not a fallback a CUDA-to-Intel translation layer (ZLUDA and the like): foreign runtime, no provenance, never in the shipped app none

The decision rule: F1 first because it is the cheapest test of "is it IGC's front end"; F2 only if F1 also fails or the rate stays under half the memprobe ceiling (then the OpenCL runtime's queueing is the suspect and Level Zero's explicit command lists are the fix).

The telemetry piece (not a fallback, needed either way): an IGCL reader in proto-opencl/gpu-telemetry.c (load ControlLib.dll at run time, ctlInit, ctlEnumerateDevices, ctlPowerTelemetryGet per card; watts from the energy counter difference over the sample interval, clocks and temperature from the same call; the intel <ordinal> ... line beside the amd lines), 2 to 3 hours plus a PC 1 sample. Until it lands the bench row records watts as "not exposed" and names IGCL as the source to come.

5. The PC 1 job (written on this branch, published only on the project lead's go after the card is in)

tools/intel-arc/pc1-arc-bench.ps1, published as a signed run job with the RUNNER's --cards-off <the Arc's key> (never --stop-miners: the 5090 and whatever else is in keep mining; the script never posts to api/cards, api/pause or api/quit), after a small fetch job that places the class v4 pack v4-devnet-epoch0 and the class v3 control mx8-devnet-epoch0 (tools/intel-arc/make-kit.sh, packs only, no exe: the installed worker is the one under test). Steps and lines:

Step Line What it settles
detect RESULT list ... from the installed worker's --list; RESULT device <n> intel ...; the Windows adapter row (name, driver version, the class key's qwMemorySize = 12,288 MiB) the card is enumerated by name with 12 GB on the Intel OpenCL platform, driver 32.0.101.9034
first light the installed app's own worker line for the card from its log (ready ... exchange N ... sub-group), read only OpenCL compiled and self-tested on Arc inside the app, or the exact error
memprobe igneum-worker-opencl.exe --device <n> --memprobe the dependent-load ceiling at 4, 64 and 1024 MiB: the number that bounds every rate below
bench, class v4 --bench-pack --pack <kit>\packs\v4-devnet-epoch0 --batches 5 --batch-log2 24 --device <n> (fingerprint expected f410c731b6bc2d31, the Mac's and the 5090's) bit-exact on the fourth vendor, and the rate on the current class
bench, class v3 control the same on mx8-devnet-epoch0 (expected 90f794dd556f7a3b) whether the shadow work costs the Arc more than 5 percent (the 2.0 rule)
bench, the live pack --bench-pack --pack <app>\packs\devnet the rate on today's program as the app mines it
exchange the class v4 pack again with --exchange local the sub-group path against the local-memory path
watts and clocks not exposed through OpenCL; the PDH GPU Engine counter gives busy percent only; the line names IGCL as the source to come the row's watts cell honest
row RESULT ROW card="Intel Arc B580" mhs= ceiling_mhs= mhw=not-exposed blocks_per_day= hours_per_block= ... and the consequences line the bench-log row, the card-picker entry (['Intel Arc B580', <mhs>] in site/yourcard.js), the owner's sentence

Time to a block: the network estimate was 675 MH/s at 09:34Z on 7 October and 463 MH/s at 10:23Z (live_state.hashes_per_second_estimate); the devnet makes 86,400 blocks a day, so a card at X MH/s expects X / network x 86,400 blocks a day: 15 MH/s is about 1,900 blocks a day at 675 MH/s (one every 45 s), 2,800 at 463, each worth the subsidy of 31.69 IGN (spec 02 section 2.5). The job computes the line from the live number at run time.

6. Consequences per tier (the rule of 5 October 2026), before the measurement

Tier What the Arc means What this branch does
Home miner, one 12 GB Intel card (the B580 owner) mines on the OpenCL worker from the app's first run after the driver; expected 13 to 20 MH/s (section 2.1, approximate until measured), 1,700 to 2,600 blocks a day at 675 MH/s of network (one every 33 to 50 s); no proving (CUDA only); no power cap and no watts reading in the app until IGCL lands, so the £ a day cell stays empty the job's row replaces the estimate; the IGCL reader is the next 2 to 3 hours
Home miner, 8 GB Intel (Arc A750, B570 10 GB) the same path; 3.4 GB resident fits; rate lower by the card's dependent-read rate the A-series is not in hand; the row is owed
16 GB Intel (Arc A770 16 GB, B770 if it ships) the same; still mine only owed
Rig several Arcs on one OpenCL platform: one worker per device index, as AMD; the PCIe 4.0 x8 link is no cost (the eGPU's USB4 cost 6 percent per job and was removed by the select read-back) nothing new
Pool user sees nothing different nothing
Windows this path this branch
Linux and HiveOS Intel's compute runtime is a distro package (intel-opencl-icd), the same worker binary a line in the install notes once the Windows row exists
macOS no Intel GPUs nothing

7. What the card measured on 7 October 2026 (bench-log entry "the first Intel card")

Reading Value Where
Dependent random 4-byte loads at 1 GiB 1.41 G loads/s (4 MiB: 55.1 in the L2; 64 MiB: 2.47) run-ia-arc-bench-20261007-c, --memprobe
Ceiling at 128 loads a hash 11.0 MH/s the same
The hash, every pack, both exchange modes builds in 236 to 274 ms; cache and dataset bit-exact; vectors 96 of 96 lanes WRONG, the same device value in the sub-group and the local-memory exchange the same; the app's own worker saw the same from its first start
Watts, clocks not exposed through OpenCL IGCL next (section 4)
Rate, bench d with the fix 11.011 MH/s on class v4 (11.019 on the v3 control, 11.002 on the live pack, 10.882 local exchange), vectors PASS, fingerprints equal to the Mac's and the 5090's run-ia-arc-bench-20261007-d
The fault, found by the register trace instruction rotr: Intel's compiler folds rotate(x, (0u - n) & 31u) into a LEFT rotate by n; the worker rewrites the one helper line to the shift form on Intel (commit 26e135a3, proto-opencl/intel_rotr.h, the gate test test_intel_rotr.c); the trace then matches in all 55,809 snapshots run-ia-arc-trace-20261007-b and -c

Section 2.3's first gotcha (the sub-group size) is cleared (32 for a 32-item work-group) and the fallback ladder of section 4 was not needed: IGC builds the kernel, and the family probe (tools/intel-arc/pc1-arc-family.ps1, runs a and b) read every family exact, mulhi and mad included, because its rotr variant is the shift form. The register trace named the one line: the pack's rotr_var through the rotate() builtin with a computed negative amount, which Intel's compiler turns into the opposite rotate. The worker-side rewrite (26e135a3) is the whole fix: no pack, emitter or consensus text moves, AMD and NVIDIA see no change. The B580 owner's row: 11.0 MH/s from the cut that ships it (0.3.20); until then the card sits off with the reason.

PC 2's driver install (7 October 2026, 16:24 to 17:48Z): the installer took the kernel down, 9034 is not installed

Job run-ia-pc2-intel-driver-20261007-c (elevated, the app's own OpenCL worker mining on the Arc in the Razer Core X V2 at the time): download 102 s, sha512 equal to Intel's page, Authenticode valid (Intel Corporation), -s install started 16:26:44Z. The job's local elevated-output.log ends on that line. The System log: bugcheck 0x0000003B (SYSTEM_SERVICE_EXCEPTION, c0000005 at fffff80518bdca26), minidump C:\WINDOWS\Minidump\100726-22234-01.dmp, Kernel-Power 41 "rebooted without cleanly shutting down", the OS back up at 16:27:56Z (about 70 s after the installer started). Nobody was signed in after the restart, so the per-user app did not start until the logon at 17:34:46Z: 67 minutes off the network (the 5090 at 122 and the 5060 Ti at 40 MH/s, not only the Arc). Driver store after the restart (pnputil /enum-drivers): one Intel display package, oem41.inf = 32.0.101.6733 (the inbox driver); 9034 is neither staged nor bound. The Arc is listed with Code 0, the Intel OpenCL platform answers, and the app has its OpenCL worker on the card (read-back run-ia-pc2-intel-readback-20261007-c, 17:35:49Z; freeze read run-ia-pc2-freeze-events-20261007, tools/intel-arc/pc2-freeze-events.ps1). The link reads CurrentLinkWidth 1, CurrentLinkSpeed 1 on the device's own PnP properties (the Thunderbolt-tunnelled view; the rate is unaffected, section above).

Read-back d (18:14:51Z, the app then on 0.3.21 for a few minutes): the Arc row mining on 6733 (pid 25168, 8 identities), and its kind read discrete, not external, in the Razer Core X V2 (no USB4 or Thunderbolt router in the device's parent chain on this board; the link still x1 gen1): the external-kind hold alone would have missed this card, so 0.3.22 holds every card of the vendor (branch driver-hold-22 f8ed911e). The minidump: the unelevated read is refused (Access to the path 'C:\WINDOWS\Minidump' is denied, run-ia-pc2-minidump-20261007, 18:14:55Z), so the copy rides the elevated retry (tools/intel-arc/pc2-intel-driver-b.ps1: the dump as base64 chunks in the job's folder before the installer starts, a collect job after; the Mac parser triage.py names the module). Order tonight (the shipper owns PC 2): the project lead's hand reinstall of 0.3.20, the 0.3.22 installer job and smoke, then the retry on the shipper's pass line with the project lead's one click. Stop rule (main): a second crash ends it; 6733 mines at 11 MH/s and the driver-check table then treats 6733 as acceptable for the B580.

What follows: the retry on PC 2 runs with the Arc idle (--cards-off "intel:Intel(R) Arc(TM) B580 Graphics", the runner holds the worker through the install) and only on a go; the same rule is now the app's (driver-check branch 46cc41e9: a card the app reads as external has its worker stopped before the installer starts). The faulting module in the minidump is not read yet (no debugger on PC 2; a copy of the 2 MB minidump by a fetch-and-collect job is one line if wanted). Per tier: a home miner with one eGPU who clicks the install with the card mining risks this crash and an hour off the network if nobody is at the keyboard; with the hold the installer runs over an idle card, which is the shape PC 1 took at 15:xxZ (stop-miners, exit 14, 9034 bound after the restart).

8. What is not done here

The IGCL telemetry reader (section 4), an Intel prover, the Linux package line, the driver-check feature (queued: per-vendor minimum-driver table in the manifest and one-click install, branch driver-check off the miner-ui-5 tip), and the site's bench table row, which takes the measured number only.