What
ACTIONS_RUNNER_HOOK_JOB_COMPLETED=/usr/local/bin/cleanup.sh runs after every job. It shows as the "Complete runner" step at 21–39 s (often with Disk free: 111G -> 111G, i.e. it freed nothing). The cost comes from tree walks that scale with cache size, not with what the job did:
lru_cap "$HOME/.gradle" and lru_cap "$HOME/.bun" each run du -sb over the whole tree (up to 20 GiB, hundreds of thousands of files), even when under the cap;
find "$HOME/.gradle" -name '*.lock' -delete walks the entire Gradle home;
find /tmp -mindepth 1 -depth walks all of /tmp.
The hook runs before the runner accepts its next job, so it is on the critical path of every multi-job workflow: checks → build, matrix legs, and release flavours running one after another.
Evidence (one week, 2026-09-16 → 09-23)
- The self-hosted fleet runs roughly 650 jobs/week.
- Measured hook time: 21–39 s on the large pool and 12–30 s on the light pool. Across several PR workflows, about 50 s of a 2.6 min median wall time was this hook.
- Estimate: 650 × ~25 s ≈ 4.5 runner-hours/week, plus ~30–50 s of developer wait per multi-job PR run.
Proposal
- Gate the expensive part on need: only run
lru_cap when df --output=avail for the volume is below a threshold (e.g. 25 %), or at most once per hour, using a timestamp file in the volume. The cap is a safety net; it does not need to run every job.
- Bound the lock sweep:
find "$HOME/.gradle/caches" -maxdepth 3 -name '*.lock' (the locks live in known shallow paths), or drop it and rely on pre-job.sh's poison sweep.
- Keep
/tmp cleaning, but skip walking the preserved workdir subtree with -prune instead of visiting and sparing each path.
- Print the elapsed seconds per section, so regressions are visible.
This also covers #7's concern: extending the LRU cap to cargo/npm/pnpm is cheap once it runs only when disk is actually short. Done the current way, #7 would add three more full-tree du walks per job.
Expected result
The hook drops to ~1–3 s in the common case. That saves ≈ 4 runner-hours/week, and ~30–50 s off every multi-job PR run (≈ 15–20 % of median PR wall time on the Next.js consumers).
Filed by an automated CI/CD scout. Figures come from Actions job step timings for the stated window.
What
ACTIONS_RUNNER_HOOK_JOB_COMPLETED=/usr/local/bin/cleanup.shruns after every job. It shows as the "Complete runner" step at 21–39 s (often withDisk free: 111G -> 111G, i.e. it freed nothing). The cost comes from tree walks that scale with cache size, not with what the job did:lru_cap "$HOME/.gradle"andlru_cap "$HOME/.bun"each rundu -sbover the whole tree (up to 20 GiB, hundreds of thousands of files), even when under the cap;find "$HOME/.gradle" -name '*.lock' -deletewalks the entire Gradle home;find /tmp -mindepth 1 -depthwalks all of/tmp.The hook runs before the runner accepts its next job, so it is on the critical path of every multi-job workflow:
checks → build, matrix legs, and release flavours running one after another.Evidence (one week, 2026-09-16 → 09-23)
Proposal
lru_capwhendf --output=availfor the volume is below a threshold (e.g. 25 %), or at most once per hour, using a timestamp file in the volume. The cap is a safety net; it does not need to run every job.find "$HOME/.gradle/caches" -maxdepth 3 -name '*.lock'(the locks live in known shallow paths), or drop it and rely onpre-job.sh's poison sweep./tmpcleaning, but skip walking the preserved workdir subtree with-pruneinstead of visiting and sparing each path.This also covers #7's concern: extending the LRU cap to cargo/npm/pnpm is cheap once it runs only when disk is actually short. Done the current way, #7 would add three more full-tree
duwalks per job.Expected result
The hook drops to ~1–3 s in the common case. That saves ≈ 4 runner-hours/week, and ~30–50 s off every multi-job PR run (≈ 15–20 % of median PR wall time on the Next.js consumers).
Filed by an automated CI/CD scout. Figures come from Actions job step timings for the stated window.