perf(ci): serve node_modules snapshots from O(1) protected sticky disks#109752
Conversation
Every Blacksmith mount of the dependency sticky disk currently 429s: the v2 key minted one backing disk per PR and per manifest hash, which saturated Blacksmith's installation-wide 1000-sticky-disk budget (run 29559333389, job checks-node-compact-small-15: "Sticky disk limit exceeded"). Same-repo shards then fall back to a cold, storeless pnpm install (~40s) on every run - worse than the ~22s actions/cache path the sticky rollout replaced. Re-key the snapshot to one stable disk per node-version and move the install inputs into a runtime fingerprint marker: - Key is now `<repo>-node-deps-bind-v3-<node-version>`; dependency changes refresh the disk in place instead of allocating a new one. - The fingerprint (manifest hashFiles set + node version + lockfile mode) is evaluated on the bind step, before the mount lands, so the '**/package.json' glob cannot sweep snapshot-internal manifests. - Consumers mount read-only (commit: false); on fingerprint match they restore importer archives and skip pnpm install entirely, on mismatch they install against the clone's on-disk store (warm store, no actions/cache download) without capturing. - Writers are trusted non-PR jobs only: build-artifacts on canonical pushes plus the scheduled vitest-cache-warm run as a deadline writer (main pushes cancel each other under merge traffic and would starve the snapshot). commit: on-change keeps warm no-op runs from committing. - Roll the consumer flag out to every Linux Blacksmith lane that installs dependencies (build-artifacts, check-shard, check-additional-shard, check-docs, checks-ui, control-ui-i18n, native-i18n, qa-smoke-ci-profile, checks-fast-core, both contract shards) with the same fork/dispatch gates as the nondist shard. Guards now enumerate the consumer set, pin the O(1) key shape, enforce single-writer commit expressions, and cover the fingerprint marker in the importer capture/restore helper. Verified locally: importer archive capture 0.96s/228K and restore 0.07s against the real 2.0GB hoisted tree; a snapshot untarred into a fresh workspace resolves modules, runs bin shims, `pnpm exec`, and a Vitest suite (one-time pnpm reconcile only when the absolute path changes, which Blacksmith's fixed /home/runner/_work layout avoids).
…uristic stickydisk's on-change mode compares allocated disk bytes with a 4KB threshold, so a fingerprint refresh whose reinstall keeps usage stable (metadata-only manifest edits, same-sized dependency swaps) could be silently discarded, stranding every consumer on a stale marker and a permanent reinstall path. Writers now commit explicitly, mirroring the Vitest transform disk's rationale; warm no-op writer runs re-commit an identical snapshot, which is cheap and safe. Also document that the non-PR commit gate binds cooperating code only - the enforced trust boundary stays the fork/dispatch runner gate, matching the protected node-compile disk's posture.
…splits Reviewer follow-up: the snapshot key is partitioned by node-version and both writers (build-artifacts, vitest-cache-warm) rely on the action default, so a consumer pinning any other version would split onto a key nobody seeds and silently regress to permanently cold installs. Pin the action default to 24.x and assert every sticky consumer resolves to that same key segment.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dd12ce9823
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| # the '**/package.json' glob only sees checkout manifests. Recomputing | ||
| # this hash any later would sweep snapshot-internal package.json files | ||
| # into the fingerprint and permanently miss the warm path. | ||
| DEPS_FINGERPRINT: node-${{ inputs.node-version }}-frozen-${{ inputs.frozen-lockfile }}-${{ hashFiles('**/package.json', 'pnpm-lock.yaml', 'pnpm-workspace.yaml', '.npmrc', '.pnpmfile.cjs', 'pnpmfile.cjs', '.github/actions/setup-node-env/sticky-importers.sh', 'scripts/postinstall-bundled-plugins.mjs', 'scripts/preinstall-package-manager-warning.mjs', 'scripts/prepare-git-hooks.mjs') }} |
There was a problem hiding this comment.
Hash pnpm patch files before skipping install
When only a pnpm patch file changes, this fingerprint stays unchanged because it hashes pnpm-workspace.yaml/pnpm-lock.yaml but not patches/**; the repo currently has patchedDependencies pointing at patches/@[email protected]. In a warm same-repo Blacksmith job, the marker will still match and the action skips pnpm install, so pnpm never reapplies or validates the changed patch file and CI can test stale patched dependency contents instead of the checkout.
Useful? React with 👍 / 👎.
The filter's runtime narrowing did not carry into the map callback, so step.with indexing failed check-test-types (TS18048). Collect a narrowed stepWith instead.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: dd92e5a234
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| # the '**/package.json' glob only sees checkout manifests. Recomputing | ||
| # this hash any later would sweep snapshot-internal package.json files | ||
| # into the fingerprint and permanently miss the warm path. | ||
| DEPS_FINGERPRINT: node-${{ inputs.node-version }}-frozen-${{ inputs.frozen-lockfile }}-${{ hashFiles('**/package.json', 'pnpm-lock.yaml', 'pnpm-workspace.yaml', '.npmrc', '.pnpmfile.cjs', 'pnpmfile.cjs', '.github/actions/setup-node-env/sticky-importers.sh', 'scripts/postinstall-bundled-plugins.mjs', 'scripts/preinstall-package-manager-warning.mjs', 'scripts/prepare-git-hooks.mjs') }} |
There was a problem hiding this comment.
Hash the setup action in the sticky fingerprint
Now that this value is the only freshness gate for the shared v3 dependency disk, changing .github/actions/setup-node-env/action.yml itself without touching a package manifest or lockfile can still hit an old marker and skip pnpm install. That means a same-repo PR that changes the install command, pnpm flags, or sticky setup logic can test against the previous snapshot, and the main writer can keep committing it because the fingerprint never changes; include the action file (and any setup helper whose install behavior affects the snapshot) in this hash.
Useful? React with 👍 / 👎.
check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in #109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in #109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in #109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in #109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in #109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
…109929) The installation hit Blacksmith's backing-disk cap because sticky keys embedded PR numbers and content hashes, minting a new disk per PR and per input change until every mount 429-failed fleet-wide. node-deps (#109752) and ext-boundary (#109804) were already re-keyed; this converts the last two minters and audits the third suspect. Keys changed: - Vitest fs transform cache: `vitest-fs-v2-<pr-N|protected>-...` -> the one existing `vitest-fs-v2-protected-<os>-<arch>-node-<ver>` disk. Entries are content-hash keyed (vitest sha1 over id+content+NODE_ENV+version+env config), so cross-PR sharing is safe by construction: PRs now mount the protected snapshot directly (read-only, commit false) and the separate PR-only seed mount is deleted. Writers (non-PR shard writer + scheduled warm run) keep explicit commit true. Disk count: 1 + O(open PRs) -> 1. - Gradle: `gradle-v1-<task>-<pr-N|protected>-<hashFiles(deps)>` -> `gradle-v2-<task>`. Task scope stays (light ktlint must not seed heavy build lanes); PR number and dependency hash leave the key. The dependency hash moved to an in-job fingerprint marker that makes the non-PR writer rebuild its snapshot cold when inputs change, bounding disk growth; PR mounts are read-only and safely reuse stale snapshots because Gradle caches are content-addressed. Disk count: O(tasks x PRs x dependency bumps) -> O(tasks) (7 today). - Docker builder (useblacksmith/setup-docker-builder): audited, out of scope. The action owns its key internally and always uses the repo name only (src/setup_builder.ts getStickyDisk), so it is already one disk per installation repo and cannot mint per-PR disks. PR warm-path behavior changes: - Vitest: PRs lose only PR-local warm entries for files the PR itself changed (previously persisted on the pr-N disk between pushes); changed files re-transform once per push, unchanged files still hit the protected snapshot. PRs touching transform inputs (lockfile/tsconfig/package.json) run cold per push since the generation wipe is now local to the discarded clone. - Gradle: dependency-bump PRs get warmer (stale-but-valid protected snapshot instead of a cold fresh key); other PRs are unchanged. Deleted `.github/workflows/pr-cache-cleanup.yml`: it existed for the per-PR cache layer (#109425) and only deleted GitHub actions/cache archives, which GitHub's own LRU/TTL eviction already handles for the remaining fork/Windows per-PR archive paths. Guards: pinned the O(1) vitest key + non-PR commit gate, added a Gradle per-task key/writer/fingerprint guard, added a repo-wide scan asserting no useblacksmith/stickydisk key ever contains github.event.pull_request.number or hashFiles(), and replaced the cleanup-workflow assertions with a stays-deleted check.
…ks (openclaw#109752) * perf(ci): serve node_modules snapshots from O(1) protected sticky disks Every Blacksmith mount of the dependency sticky disk currently 429s: the v2 key minted one backing disk per PR and per manifest hash, which saturated Blacksmith's installation-wide 1000-sticky-disk budget (run 29559333389, job checks-node-compact-small-15: "Sticky disk limit exceeded"). Same-repo shards then fall back to a cold, storeless pnpm install (~40s) on every run - worse than the ~22s actions/cache path the sticky rollout replaced. Re-key the snapshot to one stable disk per node-version and move the install inputs into a runtime fingerprint marker: - Key is now `<repo>-node-deps-bind-v3-<node-version>`; dependency changes refresh the disk in place instead of allocating a new one. - The fingerprint (manifest hashFiles set + node version + lockfile mode) is evaluated on the bind step, before the mount lands, so the '**/package.json' glob cannot sweep snapshot-internal manifests. - Consumers mount read-only (commit: false); on fingerprint match they restore importer archives and skip pnpm install entirely, on mismatch they install against the clone's on-disk store (warm store, no actions/cache download) without capturing. - Writers are trusted non-PR jobs only: build-artifacts on canonical pushes plus the scheduled vitest-cache-warm run as a deadline writer (main pushes cancel each other under merge traffic and would starve the snapshot). commit: on-change keeps warm no-op runs from committing. - Roll the consumer flag out to every Linux Blacksmith lane that installs dependencies (build-artifacts, check-shard, check-additional-shard, check-docs, checks-ui, control-ui-i18n, native-i18n, qa-smoke-ci-profile, checks-fast-core, both contract shards) with the same fork/dispatch gates as the nondist shard. Guards now enumerate the consumer set, pin the O(1) key shape, enforce single-writer commit expressions, and cover the fingerprint marker in the importer capture/restore helper. Verified locally: importer archive capture 0.96s/228K and restore 0.07s against the real 2.0GB hoisted tree; a snapshot untarred into a fresh workspace resolves modules, runs bin shims, `pnpm exec`, and a Vitest suite (one-time pnpm reconcile only when the absolute path changes, which Blacksmith's fixed /home/runner/_work layout avoids). * fix(ci): commit dependency snapshots explicitly, not via on-change heuristic stickydisk's on-change mode compares allocated disk bytes with a 4KB threshold, so a fingerprint refresh whose reinstall keeps usage stable (metadata-only manifest edits, same-sized dependency swaps) could be silently discarded, stranding every consumer on a stale marker and a permanent reinstall path. Writers now commit explicitly, mirroring the Vitest transform disk's rationale; warm no-op writer runs re-commit an identical snapshot, which is cheap and safe. Also document that the non-PR commit gate binds cooperating code only - the enforced trust boundary stays the fork/dispatch runner gate, matching the protected node-compile disk's posture. * test(ci): guard sticky consumers against writerless node-version key splits Reviewer follow-up: the snapshot key is partitioned by node-version and both writers (build-artifacts, vitest-cache-warm) rely on the action default, so a consumer pinning any other version would split onto a key nobody seeds and silently regress to permanently cold installs. Pin the action default to 24.x and assert every sticky consumer resolves to that same key segment. * test: narrow sticky-consumer step.with before indexing The filter's runtime narrowing did not carry into the map callback, so step.with indexing failed check-test-types (TS18048). Collect a narrowed stepWith instead.
…law#109804) check-additional-extension-package-boundary is the light-run critical-path pole. Measured on runs 29564442266/29564411446/29564259862 (8 vCPU, sticky mounts 429-failing): Setup Node ~40s cold install, shard step 142-161s = ~101s cold plugin-sdk dts prep (nine parallel tsgo builds) + ~58s compiling 122 plugins at concurrency 6 + ~1s canary; warm actions/cache path on a main push (29551077288) runs the shard in 67s. Levers chosen from that data: - Re-key the ext-boundary sticky disk from per-PR/per-config keys (v1) to one O(1) protected key (v2), following the node_modules snapshot fix in openclaw#109752. Blacksmith's installation-wide 1000-disk cap is saturated and every mount 429s; per-PR keys are a direct contributor. The config/scripts/ lockfile hash moves from the key into the in-job .source-trees marker next to the existing tree-OID + node/pnpm gate. Single semantic writer: the boundary lane commits explicitly (true, not if-missing) on protected pushes only; PR mounts and the check-lint consumer stay read-only, and the seed step is now writer-only (saves ~11s per PR cold run). - Bump the lane 8 -> 32 vCPU and raise OPENCLAW_EXTENSION_BOUNDARY_CONCURRENCY 6 -> 16: both the parallel dts prep and the 122 plugin compiles scale with cores at similar billed core-minutes (same precedent as check-lint/deps). Expected pole time: warm PR runs (protected snapshot tracks every main push) drop to roughly checkout + install-skip + restore + incremental shard, well under 2 minutes; cold runs that touch boundary trees shrink via the core bump. Guarded in ci-workflow-guards so the O(1) key, single-writer commit, and fingerprint composition cannot silently regress.
…penclaw#109929) The installation hit Blacksmith's backing-disk cap because sticky keys embedded PR numbers and content hashes, minting a new disk per PR and per input change until every mount 429-failed fleet-wide. node-deps (openclaw#109752) and ext-boundary (openclaw#109804) were already re-keyed; this converts the last two minters and audits the third suspect. Keys changed: - Vitest fs transform cache: `vitest-fs-v2-<pr-N|protected>-...` -> the one existing `vitest-fs-v2-protected-<os>-<arch>-node-<ver>` disk. Entries are content-hash keyed (vitest sha1 over id+content+NODE_ENV+version+env config), so cross-PR sharing is safe by construction: PRs now mount the protected snapshot directly (read-only, commit false) and the separate PR-only seed mount is deleted. Writers (non-PR shard writer + scheduled warm run) keep explicit commit true. Disk count: 1 + O(open PRs) -> 1. - Gradle: `gradle-v1-<task>-<pr-N|protected>-<hashFiles(deps)>` -> `gradle-v2-<task>`. Task scope stays (light ktlint must not seed heavy build lanes); PR number and dependency hash leave the key. The dependency hash moved to an in-job fingerprint marker that makes the non-PR writer rebuild its snapshot cold when inputs change, bounding disk growth; PR mounts are read-only and safely reuse stale snapshots because Gradle caches are content-addressed. Disk count: O(tasks x PRs x dependency bumps) -> O(tasks) (7 today). - Docker builder (useblacksmith/setup-docker-builder): audited, out of scope. The action owns its key internally and always uses the repo name only (src/setup_builder.ts getStickyDisk), so it is already one disk per installation repo and cannot mint per-PR disks. PR warm-path behavior changes: - Vitest: PRs lose only PR-local warm entries for files the PR itself changed (previously persisted on the pr-N disk between pushes); changed files re-transform once per push, unchanged files still hit the protected snapshot. PRs touching transform inputs (lockfile/tsconfig/package.json) run cold per push since the generation wipe is now local to the discarded clone. - Gradle: dependency-bump PRs get warmer (stale-but-valid protected snapshot instead of a cold fresh key); other PRs are unchanged. Deleted `.github/workflows/pr-cache-cleanup.yml`: it existed for the per-PR cache layer (openclaw#109425) and only deleted GitHub actions/cache archives, which GitHub's own LRU/TTL eviction already handles for the remaining fork/Windows per-PR archive paths. Guards: pinned the O(1) vitest key + non-PR commit gate, added a Gradle per-task key/writer/fingerprint guard, added a repo-wide scan asserting no useblacksmith/stickydisk key ever contains github.event.pull_request.number or hashFiles(), and replaced the cleanup-workflow assertions with a stays-deleted check.
What Problem This Solves
Every same-repo CI job pays ~40s of "Setup Node environment" doing a cold, storeless pnpm install — because the landed sticky bind-mount design (#108386) mints one Blacksmith backing disk per PR x per manifest-hash, and the installation is now at Blacksmith's 1000-disk cap: every sticky mount 429s ("Sticky disk limit exceeded"), warm hits never happen, and the sticky path also disabled the actions/cache fallback those jobs used to have (~22s). Live evidence: run 29559333389 job checks-node-compact-small-15.
Why This Change Was Made
Re-key the dependency snapshot onto O(1) protected disks and make freshness a runtime check instead of a key dimension:
<repo>-node-deps-bind-v3-<node-version>only. All install inputs (pnpm-lock.yaml,**/package.json, workspace yaml, npmrc/pnpmfile, postinstall scripts, frozen mode; pnpm version rides inpackageManager) move into a.openclaw-deps-fingerprintmarker compared at install time. Per-PR/per-hash keys are what saturated the fleet.hashFiles('**/package.json')after the mount would sweep snapshot-internal manifests and permanently miss the warm path (guarded by test).commit: "false"; writers are build-artifacts (per non-docs main push) plus the scheduled vitest-cache-warm run as a deadline writer (main-push concurrency cancellation would otherwise starve the snapshot under merge traffic). Writer commits are explicittrue— stickydisk's on-change allocated-byte heuristic can silently drop a fingerprint refresh.12 consumers flipped (build-artifacts as writer; check-shard, check-additional-shard, check-docs, checks-ui, control-ui-i18n, native-i18n, qa-smoke-ci-profile, checks-fast-core, both contract shards, nondist shard); pnpm-store-warmup, dispatch-only compat jobs, and macOS/iOS jobs intentionally not flipped.
User Impact
None at runtime — CI-only. Once stale v2/per-PR disks are purged (Blacksmith dashboard/support; no API in the pinned action), warm same-repo jobs skip the install entirely; until then the fallback is no worse than today's cold path.
Evidence
node scripts/run-vitest.mjs test/scripts/ci-workflow-guards.test.ts— 76 passed, 1 skipped (new guards: consumer enumeration, O(1) key shape, single writer, fingerprint-before-mount, node-version key splits), rerun green after rebase onto current main.git diff --checkclean..binshims andpnpm exec, passes a Vitest suite.Follow-ups