Copy, from origin/master 0a63474 (site audit). The home and miner pages named "PC 1" in the figure captions, the
page-by-page lead, two alt texts, the Ember Tune table title, the Ember row's source line and the home pill; they now
say "a Windows rig with an RTX 5090, RTX 4070 and RX 9070 XT" at the first mention on each page and "the Windows rig"
after. The generated pages carried the same names from the bench-log (80 on /bench, 7 on /evidence, the inlined
journey on the home page): site/scrub.mjs now maps PC 1 to "the three-card Windows rig (RTX 5090, RTX 4070, RX 9070
XT)" and PC 2 to "the RTX 5090 Windows rig", build.mjs re-scrubs the stored journey entries, two /bench anchors on the
miner page follow the renamed headings, and \bPC [12]\b joins site/forbidden-strings.txt so the build fails if a number
returns. The scrub also covers the audit's other page-leak shapes (a pid, 0.0.0.0:port, ~/.config paths, --rpclisten=),
which the bench page carried and which now fail the build if they return.
CI: tools/ci/forbidden-strings.txt's appended audit block sat on one physical line with literal \n text, so none of its
patterns was active; \bPC [12]\b is now a real line there and the identity check's export scrub maps the two machines
the same way (igneum-public/tools/sync.sh must carry the same two rules). The other four audit patterns moved to
site/forbidden-strings.txt, since the public export carries simulator schedule logs where a pid is a pid. The identity
check gains a second pass over every served file under site/ (html, json, txt, xml, webmanifest, css, js; not api/ or
the build scripts), unscrubbed, with both pattern lists; dl\.igneum is narrowed to the tokened path so the public
download buttons pass. Shown to fire on a page naming PC 1 (exit 1) and to pass on the tree (0 hits over 232 export
files and 32 served site files).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
80 lines
5.8 KiB
Bash
Executable file
80 lines
5.8 KiB
Bash
Executable file
#!/usr/bin/env bash
|
|
# Identity grep of the public export list, gh-free, for CI (tools/ci/forbidden-strings.txt).
|
|
#
|
|
# Copies the paths that igneum-public/tools/sync.sh publishes into a temporary directory, applies the same generic
|
|
# scrub that sync.sh applies (model names for machines, <lan-ip> for LAN addresses, ~ for home paths, UTC stamps),
|
|
# then greps the result with the committed pattern list. A hit means a change would reach the public mirror with
|
|
# a machine name, a LAN address, a home path or the log-intake key pattern that the generic scrub does not catch.
|
|
# The private rules of the mirror (sync.local.sed, identity.local) are not here; they run at export time.
|
|
#
|
|
# tools/ci/identity-check.sh # exit 1 on any hit, with file:line
|
|
set -euo pipefail
|
|
HERE="$(cd "$(dirname "$0")" && pwd)"
|
|
REPO="$(cd "$HERE/../.." && pwd)"
|
|
PATTERNS="$HERE/forbidden-strings.txt"
|
|
|
|
# The export list of igneum-public/tools/sync.sh (keep in step with it), plus the two files published with the repository
|
|
# at the public testnet (decision of 5 October 2026: the criticism ledger and its fixes file). They are not in sync.sh,
|
|
# because the spec mirror is public today and the repository is not until the public testnet.
|
|
DIRS=(docs/spec docs/analysis sim igneum-pow igneum-census tools/harness proto-cuda/packs docs/benchmarks tools/finality-attacks tools/exec-attacks)
|
|
FILES=(site/ledger.html docs/provenance.md docs/bench-log.md docs/evidence.md docs/fud-ledger.md docs/fud-fixes.md proto-cuda/README.md proto-cuda/CHECKLIST.md proto-cuda/host.cu proto-cuda/build.sh
|
|
proto-cuda/build.bat proto-cuda/.gitignore proto-cuda/emu/emu.sh proto-cuda/emu/shim.cpp proto-cuda/emu/cuda_runtime.h proto-metal/README.md
|
|
proto-metal/MEMHARD.md proto-metal/TESTS.md proto-metal/main.swift proto-opencl/README.md proto-opencl/WAVEFRONT.md proto-opencl/host.c
|
|
proto-opencl/build.sh proto-opencl/build.bat proto-opencl/.gitignore proto-opencl/emu/emu.sh proto-opencl/emu/emu_main.cpp proto-opencl/emu/emu_opencl.h)
|
|
|
|
TMP="$(mktemp -d)"; trap 'rm -rf "$TMP"' EXIT
|
|
for d in "${DIRS[@]}"; do [ -d "$REPO/$d" ] && { mkdir -p "$TMP/$(dirname "$d")"; cp -R "$REPO/$d" "$TMP/$d"; }; done
|
|
for f in "${FILES[@]}"; do [ -f "$REPO/$f" ] && { mkdir -p "$TMP/$(dirname "$f")"; cp "$REPO/$f" "$TMP/$f"; }; done
|
|
# prune what sync.sh prunes
|
|
find "$TMP" \( -name target -o -name __pycache__ -o -name out -o -name 'build-*' -o -name node_modules -o -name results -o -name runs \) -prune -exec rm -rf {} + 2>/dev/null || true
|
|
find "$TMP" \( -name .DS_Store -o -name '*.pyc' \) -type f -delete
|
|
|
|
# the generic scrub of sync.sh step 3 (its public half; the rules are public text, not secrets)
|
|
TEXT_FILES="$(find "$TMP" -type f \( -name '*.md' -o -name '*.rs' -o -name '*.py' -o -name '*.mjs' -o -name '*.sh' -o -name '*.bat' \
|
|
-o -name '*.c' -o -name '*.cu' -o -name '*.cl' -o -name '*.h' -o -name '*.cpp' -o -name '*.swift' -o -name '*.metal' -o -name '*.json' \
|
|
-o -name '*.csv' -o -name '*.toml' -o -name '*.txt' -o -name '*.log' -o -name '*.html' \) -print)"
|
|
while IFS= read -r f; do
|
|
[ -n "$f" ] || continue
|
|
perl -pi -e '
|
|
s/the PC node at 192\.168\.[0-9.]+/the RTX 5090 node on the LAN/g;
|
|
s/\bPC 1\x27s\b/the three-card Windows rig\x27s/g; s/\bPC 1\b/the three-card Windows rig (RTX 5090, RTX 4070, RX 9070 XT)/g;
|
|
s/\bPC 2\x27s\b/the RTX 5090 Windows rig\x27s/g; s/\bPC 2\b/the RTX 5090 Windows rig/g;
|
|
s/\bthe PC node\b/the RTX 5090 node/g;
|
|
s/\bWindows PC\b/an RTX 5090 on Windows/g;
|
|
s/\bthe PC\x27s\b/the RTX 5090 machine\x27s/g;
|
|
s/\bthe PC\b/the RTX 5090 machine/g;
|
|
s/\bPC (joins|start|period)\b/RTX 5090 $1/g;
|
|
s/192\.168\.[0-9]+\.[0-9]+/<lan-ip>/g;
|
|
s/DESKTOP-[A-Z0-9]{7}/<pc-hostname>/g;
|
|
s/~\/Desktop\//`/g; s/``/`/g;
|
|
s/~\/\.cargo\/bin\/cargo/cargo/g;
|
|
s/\/Users\/[A-Za-z0-9_.-]+/~/g;
|
|
s/C:\\Users\\[A-Za-z0-9_.-]+/%USERPROFILE%/g;
|
|
s/(\d{1,2}:\d{2}(?::\d{2})? UTC) = \d{1,2}:\d{2} B[S]T/$1/g;
|
|
s/(\d{1,2}):(\d{2})(:\d{2})? to (\d{1,2}):(\d{2})(:\d{2})? B[S]T/sprintf("%02d:%s%s to %02d:%s%s UTC",($1+23)%24,$2,$3\/\/"",($4+23)%24,$5,$6\/\/"")/ge;
|
|
s/(\d{1,2}):(\d{2})(:\d{2})? B[S]T/sprintf("%02d:%s%s UTC",($1+23)%24,$2,$3\/\/"")/ge;
|
|
' "$f"
|
|
done <<< "$TEXT_FILES"
|
|
perl -pi -e 's/\(Mac side only;/(Apple M5 Max side only;/g; s/\bthe Mac\x27s\b/the Apple M5 Max\x27s/g; s/\bthe Mac\b/the Apple M5 Max/g;' "$TMP/docs/bench-log.md" 2>/dev/null || true
|
|
|
|
PAT="$(grep -vE '^\s*(#|$)' "$PATTERNS")"
|
|
HITS="$(grep -rEn -f <(printf '%s\n' "$PAT") "$TMP" || true)"
|
|
if [ -n "$HITS" ]; then
|
|
echo "identity grep: HITS in the public export list (after the generic scrub):"
|
|
printf '%s\n' "$HITS" | sed "s#^$TMP/##" | cut -c1-200
|
|
exit 1
|
|
fi
|
|
# The served site (6 October 2026, site audit): every file Vercel serves from site/ (the committed pages, the generated
|
|
# pages, the JSON, the manifest, the sitemap), grepped as it is, no scrub, with the export patterns above and the site's
|
|
# own list (site/forbidden-strings.txt, which the build also enforces on the generated pages). The API sources and the
|
|
# build scripts are not served and are not in this pass.
|
|
SITE_FILES="$(find "$REPO/site" -type f \( -name '*.html' -o -name '*.json' -o -name '*.txt' -o -name '*.xml' -o -name '*.webmanifest' -o -name '*.css' -o -name '*.js' \) \
|
|
-not -path '*/node_modules/*' -not -path '*/.vercel/*' -not -path '*/api/*' -not -name 'forbidden-strings.txt' -print)"
|
|
SITE_PAT="$(printf '%s\n%s\n' "$PAT" "$(grep -vE '^\s*(#|$)' "$REPO/site/forbidden-strings.txt")")"
|
|
SITE_HITS="$(printf '%s\n' "$SITE_FILES" | xargs grep -En -f <(printf '%s\n' "$SITE_PAT") 2>/dev/null || true)"
|
|
if [ -n "$SITE_HITS" ]; then
|
|
echo "identity grep: HITS in the served site files (site/, unscrubbed):"
|
|
printf '%s\n' "$SITE_HITS" | sed "s#^$REPO/##" | cut -c1-200
|
|
exit 1
|
|
fi
|
|
echo "identity grep: 0 hits over $(printf '%s\n' "$TEXT_FILES" | grep -c .) export files and $(printf '%s\n' "$SITE_FILES" | grep -c .) served site files"
|