Skip to content

ci: add manual production deployment with a shared deploy driver - #2

Open
wan9chi wants to merge 3 commits into
mainfrom
claude/ci-deployment-main-push-fbc93a
Open

wan9chi wants to merge 3 commits into
mainfrom
claude/ci-deployment-main-push-fbc93a

Conversation

@wan9chi

@wan9chi wan9chi commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds a production deployment of the remote cache that a maintainer starts by hand, and moves the deploy logic that staging and production share into one script.

  • Staging (remote-cache-deploy.yml) still deploys on every PR and every push to main. Its only workflow change is pnpm ci:deploy staging.
  • Production (remote-cache-production.yml, new) runs only on workflow_dispatch. It runs pnpm check, pnpm smoke, then pnpm ci:deploy production in a production environment.

What changed

  • scripts/ci/deploy.ts (new) is the shared driver. runContext validates the GitHub run identity. deployTarget:

    • runs operator setup for every bound namespace, with the deploy deferred and DEPLOYMENT_ID injected;
    • runs an optional prepare hook;
    • deploys the Worker once;
    • asserts R2 is private;
    • polls each namespace endpoint until it serves the new deployment ID.

    reportDeployment writes the url output and a step summary.

  • .github/production-repositories.jsonc (new) is the production repository list, a map of cache namespace to public owner/repo. It starts with rolldown → rolldown/rolldown. Adding or removing a repository is a one-line edit.

  • scripts/ci/production.ts (new) names the production Worker, D1 and R2 (voidzero-remote-cache), and reads and validates the repository list when it deploys. On each deploy it enables every listed namespace and disables every other one. It refuses:

    • events other than workflow_dispatch;
    • refs other than refs/heads/main;
    • repositories other than this one, so repos created by Deploy to Cloudflare can't run it;
    • names that look like staging.
  • scripts/ci.ts keeps the staging rules (-ci prefix, pinned subdomain, writes only on push to the default branch). It now calls runContext and deployTarget, and staging's extra steps (scope-ownership guard, other/manual scopes, policy recovery) are its prepare hook. The commands are now deploy staging, deploy production, and test.

  • scripts/deploy.ts: extracted waitForDeployment and namespaceEnabled for reuse. The button deploy's behavior is unchanged.

  • Docs: a "Production deployment" section in docs/e2e-plan.md, linked from the README.

Reviewer notes

  • Two small staging changes:
    • After deploying, staging now also checks the e2e endpoint for the new deployment ID.
    • NAMESPACES is now built from the scopes stored in D1. The ownership guard limits those to the same e2e, other and manual.
    • The run summary gets a "Remote cache deployment" section.
  • Shared token: staging and production use the same CLOUDFLARE_* repository secrets, and internal PR staging runs receive them. A production-only token can be added later as a production environment secret without workflow changes. Consider adding required reviewers to that environment.
  • The repository list decides each namespace's switches in production. Every deploy turns listed namespaces on and all others off, so a manual pnpm operator policy change lasts only until the next deploy. Only rows whose switches change are updated, because the changed_policy trigger bumps policy_version and stops in-flight uploads. The cache-wide pnpm operator deployment switches stay manual and act as the emergency stop. To roll back, re-run an earlier production run.
  • Not in this PR: the Deploy to Cloudflare button has no live end-to-end test yet. The acceptance exercise in docs/e2e-plan.md has no recorded results. Follow-up options: a CI job that runs pnpm deploy against resources provisioned by ID, plus the manual exercise.

After merge

  1. Turn on staging. The variables have never been set, so its deploy job has been skipped:
    gh variable set REMOTE_CACHE_WORKERS_SUBDOMAIN --body <account-subdomain>
    gh variable set REMOTE_CACHE_DEPLOY_ENABLED --body true
  2. Confirm the API token has Workers Admin (the first production run creates the Worker), D1 Edit, and Workers R2 Storage Edit.
  3. Once the main staging run passes, start production: gh workflow run remote-cache-production.yml --ref main.

Testing

  • pnpm check passes.
  • pnpm smoke: 37/37 tests pass and the bundle dry run succeeds.
  • New test/ci.test.ts:
    • runContext validation;
    • two bindings produce one Worker deploy with prepare run first;
    • a withdrawn namespace stays disabled on a fresh-runner retry, and its endpoint must return 404;
    • repository reassignment and a subdomain mismatch fail before deploy;
    • production guards.
  • I ran pnpm ci:deploy production locally with push, copied-repo and feature-branch contexts. Each is rejected before contacting Cloudflare, and a valid context stops at the missing credentials.
  • Not yet run against real Cloudflare. The staging run on this PR will exercise the shared driver once the staging variables are set.

🤖 Generated with Claude Code

Staging and production now share scripts/ci/deploy.ts, which validates the
GitHub run identity, runs operator setup for every bound namespace, deploys the
Worker once, asserts private R2, and checks each endpoint for the new
deployment ID.

The production target (scripts/ci/production.ts) binds rolldown/rolldown to
the voidzero-remote-cache Worker. It deploys only from manual runs of this
repository's main branch, so repositories created by Deploy to Cloudflare
cannot run it. Staging keeps its own rules and fixtures, and the button deploy
reuses the extracted endpoint polling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Remote cache staging

Commit: 904e6ea339985655e3981db98a8df2e28fe145cf

Cloudflare deployment is disabled. Configure the repository secrets and variables described in the e2e plan, then set REMOTE_CACHE_DEPLOY_ENABLED=true and rerun this workflow. No live deployment passed verification.

Manual checks and complete e2e plan.

wan9chi and others added 2 commits October 7, 2026 15:50
Production repositories now live in .github/production-repositories.jsonc,
a namespace-to-repository map, so adding or removing one is a one-line edit.
The production target reads and validates the list only when it deploys, and
the test checks the committed file without depending on its entries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each production deployment now enables every namespace listed in
.github/production-repositories.jsonc and disables every other namespace, so
removing an entry withdraws its repository. Only rows whose switches change are
updated, because the changed_policy trigger bumps policy_version on any switch
update and that stops in-flight uploads from publishing. The cache-wide
deployment switches stay manual.

The shared driver passes its operator IO to prepare hooks, so the staging and
production hooks run against the same IO as the rest of the deployment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant