Skip to content

cleanup.sh adds 20–39 s to every self-hosted job: du -sb + full find over ~20 GB caches on each job-completed hook #16

Description

@azlekov

What

ACTIONS_RUNNER_HOOK_JOB_COMPLETED=/usr/local/bin/cleanup.sh runs after every job. It shows as the "Complete runner" step at 21–39 s (often with Disk free: 111G -> 111G, i.e. it freed nothing). The cost comes from tree walks that scale with cache size, not with what the job did:

  • lru_cap "$HOME/.gradle" and lru_cap "$HOME/.bun" each run du -sb over the whole tree (up to 20 GiB, hundreds of thousands of files), even when under the cap;
  • find "$HOME/.gradle" -name '*.lock' -delete walks the entire Gradle home;
  • find /tmp -mindepth 1 -depth walks all of /tmp.

The hook runs before the runner accepts its next job, so it is on the critical path of every multi-job workflow: checks → build, matrix legs, and release flavours running one after another.

Evidence (one week, 2026-09-16 → 09-23)

  • The self-hosted fleet runs roughly 650 jobs/week.
  • Measured hook time: 21–39 s on the large pool and 12–30 s on the light pool. Across several PR workflows, about 50 s of a 2.6 min median wall time was this hook.
  • Estimate: 650 × ~25 s ≈ 4.5 runner-hours/week, plus ~30–50 s of developer wait per multi-job PR run.

Proposal

  1. Gate the expensive part on need: only run lru_cap when df --output=avail for the volume is below a threshold (e.g. 25 %), or at most once per hour, using a timestamp file in the volume. The cap is a safety net; it does not need to run every job.
  2. Bound the lock sweep: find "$HOME/.gradle/caches" -maxdepth 3 -name '*.lock' (the locks live in known shallow paths), or drop it and rely on pre-job.sh's poison sweep.
  3. Keep /tmp cleaning, but skip walking the preserved workdir subtree with -prune instead of visiting and sparing each path.
  4. Print the elapsed seconds per section, so regressions are visible.

This also covers #7's concern: extending the LRU cap to cargo/npm/pnpm is cheap once it runs only when disk is actually short. Done the current way, #7 would add three more full-tree du walks per job.

Expected result

The hook drops to ~1–3 s in the common case. That saves ≈ 4 runner-hours/week, and ~30–50 s off every multi-job PR run (≈ 15–20 % of median PR wall time on the Next.js consumers).


Filed by an automated CI/CD scout. Figures come from Actions job step timings for the stated window.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions