You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Diagnose and permanently fix the intermittent Supabase max_concurrency stall in edge-worker:e2e:portable-runtimes. It repeatedly blocks Version Packages PRs, while the same workload passes under Node and Bun. Add bounded progress/database/worker diagnostics first, reproduce the actual cause, and protect the fix with a deterministic regression test. Do not mask the failure with automatic reruns or a longer timeout.
User report or idea
The user originally requested investigation of failing CI on Version Packages PR #687 (0.17.1). The same symptom reappeared on Version Packages PR #693 after the separate telemetry endpoint fix (#691, implemented by #692) merged.
The user asked whether this was the same issue and how to fix it permanently for future runs, then requested this self-contained issue as a handoff for a new worktree. Scope: debugging instrumentation, diagnosis, regression coverage, and the smallest root-cause fix. Issue creation does not create a worktree, change code, rerun CI, merge a PR, or release packages.
Test location reported by Deno: pkgs/edge-worker/tests/e2e-portable-runtimes/portable-runtimes.test.ts:507:8.
Exact error text (ANSI color codes omitted):
Error: Timeout after 60000ms waiting for test_seq to increment by 200
The failing test is portable runtimes - max_concurrency works in supabase. Deno reports 13 passed and 1 failed in about 1m52s. The Supabase test consumes its 60-second deadline; the Node and Bun variants each pass in about 6 seconds. Other Supabase examples (conn_max_pg_default, conn_max_pg_override, conn_env_var) pass, as do both process-based step-queue tests.
All other executed checks pass, including edge-worker-e2e, edge-worker-integration, build/test, CLI/client E2E, and core pgTAP. Website/demo deployment jobs skip because of the failed prerequisite.
Failure-time function-server log now exists
PR #690 added printing of the function-server log when the monitored suite fails. This log header appears in #693:
Monitored command exited with status 1. Function server log:
The dumped runtime log contains these exact lines (outer GitHub job timestamps/prefixes omitted):
2026-10-02T06:18:43.304957652Z Using supabase-edge-runtime-1.69.25 (compatible with Deno v2.1.4)
2026-10-02T06:18:46.522222649Z serving the request with supabase/functions/max_concurrency
2026-10-02T06:18:46.888569093Z serving the request with supabase/functions/max_concurrency
2026-10-02T06:18:47.890172012Z serving the request with supabase/functions/max_concurrency
Similar max_concurrency request lines repeat throughout the timeout and later tests. The dumped log has no CPU-limit shutdown, forced SQL close, or connection errors for this run. It does show startup warnings about the missing vendored src/flows.ts, and the expected auth_test readiness probe warning Auth validation failed: Missing Authorization header. Those warnings also appeared during locally passing tests; neither is established as causal.
Repeated request lines prove HTTP invocations, not worker execution progress, healthy heartbeats, successful cron responses, or message completion.
Download the complete current failed-job evidence:
gh run view 36972575630 --repo pgflow-dev/pgflow --log-failed > /tmp/pgflow-693-failed.log
This is preferable to reconstructing the full log from excerpts. Preserve the complete output from every new reproduction.
PR #687, head dad163005929a97af95c021debbd37c69c42b01d: attempt 1 failed the same portable test; rerunning only failed jobs produced a successful attempt 2 on the same commit.
Full portable suite passed 14/14 on the PR base 6cc225de5275164a8f7032206162b821758bb74a; Supabase concurrency took about 3 seconds.
Temporarily applying the exact Version Packages #687 release diff also passed 14/14; Supabase concurrency took about 4 seconds. The diff was then reverted.
A combined Nx E2E invocation hit the agent's 300-second command deadline after portable tests printed success. The environment was recovered via the supported fresh target before running standard E2E separately. Do not count that interrupted invocation as a clean final gate.
The standalone standard E2E performance test processed 2000 messages but paused near completion: initial throughput was roughly 52–63/s, total time about 218.7 seconds. Its local server log included CPU-limit/shutdown errors. Exact diagnostic strings:
Supabase shutdown exceeded the 5000ms deadline; forcing sql close
Error: write CONNECTION_DESTROYED pooler:6543
Error: write CONNECTION_ENDED pooler:6543
Supabase shutdown timed out after 5000ms
These are evidence from a different, higher-volume local run, not proof of the #693 cause. The earlier CPU-shutdown hypothesis must not be upgraded into a finding; #693's dumped log does not support it.
Local setup initially failed because pkgs/cli/node_modules/.bin/supabase still pointed at a stale Supabase 2.34.3 installation with no binary. Root pnpm rebuild supabase, package-filtered rebuild, and frozen-lockfile install alone did not fix that stale shim. Moving that untracked shim aside let the root Supabase 2.63.1 binary resolve, after which fresh setup and tests succeeded. This was a local reproduction obstacle, not the CI root cause; do not bake that ad hoc repair into the worker fix.
Investigation and findings
Baseline and test flow
Current main after the telemetry fix: fb5b8670e (fix(telemetry): use the stable pgflow collector endpoint (#692)). Start the new worktree from current origin/main, not the older checkout used during investigation. Worker runtime, portable tests, lifecycle helpers, and pkgs/edge-worker/project.json showed no differences between that main and the earlier inspected worker baseline.
The portable fixture lives in pkgs/edge-worker/supabase/functions/_shared/portable_examples.js. Its max_concurrency settings are:
Supabase path calls resetExample(..., { preserveWorkerMetadata: true }), purges/creates the queue, restarts the sequence, starts the HTTP worker, waits for worker_functions.start_mode = 'http', captures sequence state, sends the batch, then waits for sequence progress.
startWorker() in tests/e2e/_helpers.ts invokes the function and checks for a recent worker heartbeat; that check does not establish future processing health.
waitForExampleAssertion() polls every 500ms, uses a 60000ms deadline for max_concurrency, and reports only the target on timeout—not the observed final sequence value.
Node/Bun paths use unique queue/worker names but the shared sequence; they run after the Supabase case. They explicitly check a fresh worker, process exit, and stopped_at. Keep observations separate by runtime/queue.
Suite files share long-lived workers and database state. A single-file E2E command is not the supported isolation strategy.
Runtime paths already read during investigation:
src/core/Worker.ts: heartbeat and batch-processing main loop, retry backoff, stop/start/deprecation handling.
src/core/BatchProcessor.ts: slot checks, waitForSlot(), polling, and asynchronous execution scheduling.
Existing @henrygd/queue dependency handles scheduling and queue-size accounting. Read its installed implementation if evidence points to completion/slot timing; do not blame it without a reproduction.
Lifecycle and CI ownership
Nx target: pnpm nx run edge-worker:e2e:portable-runtimes (uncached, depends on worker build).
Runner: scripts/run-e2e.sh tests/e2e-portable-runtimes owns migration setup, vendored dependency sync, function serve, readiness, suite, and exact-resource cleanup under resource locks.
scripts/functions-server.sh monitors the CLI server and retains/dumps its output on failure. Server process liveness is not worker-loop liveness.
CI job in .github/workflows/ci.yml uses ubuntu-latest, repository setup, Bun setup, affected-project check, scripts/ci-prestart-supabase.sh edge-worker, then the Nx portable target.
tests/e2e/_helpers.ts already provides fetchWorkers() and startWorkersMonitor() with worker-function, cron, and pg_net inspection. Reuse useful existing queries rather than introduce a second diagnostics framework. Its debug helper activates under DEBUG=1 or VERBOSE=1.
What is and is not proved
Proved: the exact symptom recurs; Node/Bun and other Supabase tests pass; a same-commit rerun passed previously; ordinary server logs alone cannot explain the stall.
Not proved: how many of the 200 messages complete; whether the worker heartbeat stalls; whether messages remain visible/invisible or get retried; whether a SQL connection blocks; whether startup/purge/cron races or scheduler notifications explain the failure.
Current strongest lead: an HTTP-responsive function with a processing loop that stops progressing. This is only a hypothesis. Database/pool contention and test setup races remain alternatives.
The merged telemetry fix #692 changes sender URLs, tests, docs, and a SQL migration, not worker execution source. The timeout predates that change; timing after #692 merge is not evidence that the URL change caused it.
Proposed solution or design
1. Establish the reproduction environment and evidence budget
Create a fresh worktree from current main using the repository-supported worktree workflow; suggested branch fix/portable-supabase-stall.
Read root/package instructions and load developing-pgflow; it governs setup, full-suite cadence, and bounded retries. Load schema/pgTAP/migration skills only if the eventual fix touches SQL.
Establish one fresh pgflow-owned environment with pnpm nx test-env:fresh edge-worker, then reuse it. Only one database-mutating worktree at a time until existing shared-resource limitations are addressed. Do not stop unrelated Supabase projects.
Inspect resolved Nx targets, then use the full portable suite. Preserve complete command output and exit status. Recover after interrupted tests rather than reuse uncertain infrastructure evidence.
2. Add bounded diagnostic evidence at the actual stall
Prefer one failure-time snapshot in the portable test, plus narrowly enabled worker logging. The snapshot should identify the runtime/example/queue and include:
Evidence
Purpose
Sequence baseline, expected threshold, last observed value, elapsed time, and recent progress timestamps
Distinguish zero processing, partial completion, and a threshold/setup issue. Record the sequence value promptly: later Node/Bun tests reset the shared sequence.
Queue totals for visible/invisible/archived messages and bounded read-count/visibility summaries
Worker IDs, startup/last-heartbeat age, stopped_at, deprecated_at, registered function mode, and invocation metadata
Identify the actual worker and whether lifecycle progress stops despite HTTP requests.
Bounded pg_stat_activity rows with wait events, query age, and blocking PID relationships
Distinguish SQL/pool waits, blocked polling/handlers, and an application stall. Do not publish secrets, connection strings, payloads, or arbitrary user queries.
Relevant bounded cron/pg_net status if invocation/replacement evidence points there
Check actual HTTP status/errors instead of inferring success from request lines.
Instrument at the failure boundary, not merely at suite end. Keep snapshots read-only and bounded; diagnostics must not hang on the same exhausted worker pool. Use an independent diagnostic connection/timeout if needed, preserve the original timeout as the primary error, and treat diagnostic query failures as secondary evidence.
Enable debug output for the max_concurrency fixture during diagnosis to show heartbeat/polling/batch/slot/execution progress. Keep this confined to the test fixture or diagnostic mode; avoid enabling noisy logs or recording payloads in production. Retain #690's server-log dump. If logs remain too large, report bounded progress/last-known states rather than add an observability framework.
Add a small runnable check for snapshot failure handling/format if new nontrivial diagnostic logic warrants it.
3. Use evidence to produce a deterministic regression
Run the instrumented portable suite in the relevant CI environment and inspect the first failure; do not repeatedly rerun broad suites until green.
If polling/SQL stalls, reproduce the demonstrated pool/lock condition. If slot scheduling stalls, reproduce the completion/interleaving condition. If startup/reset/cron state races, reproduce that lifecycle sequence.
Select the smallest repository-supported unit/integration regression once the mechanism is known. The regression must fail without the root-cause fix, not merely assert that an occasional full-suite run passes.
Fix the shared runtime path where all affected callers route, or fix test setup if evidence proves that is the source. Preserve both Supabase and Node/Bun behavior.
4. Validate and deliver separately from telemetry
Run focused affected checks, the batch gate, and the applicable full portable and standard E2E suites. Include integration/types/lint/build or lifecycle checks according to affected Nx inputs and repository rules.
Preserve the 200-message workload, 10-worker-slot/four-connection fixture settings, and current 60-second deadline during diagnosis. Do not make retries, fewer messages, longer timeouts, arbitrary sleeps, disabled assertions, or cache bypasses the permanent solution.
A test-only instrumentation/setup change need not imply a package release. If the proven fix changes public worker behavior, follow the package changeset/release policy.
Open a separate correctness/reliability PR. Include evidence linking the observed queue/worker/SQL state to the fix, regression results, and CI outcome.
Acceptance criteria
A failed Supabase concurrency test emits the actual sequence progress and bounded queue, worker, and database-wait evidence at failure time without masking its original error.
The diagnostic mode exposes relevant worker progress, preserves cleanup/server monitoring, and does not leak payloads, credentials, or private environment values.
The cause is demonstrated with a deterministic failing regression; hypotheses and alternative explanations are explicitly resolved or retained as unknown.
The smallest runtime/setup fix passes the regression plus applicable full portable and standard E2E checks on CI, preserving Node/Bun support and test coverage.
The recurring release-blocking condition is addressed without weakening the workload, deadline, assertions, or adding retry-until-green behavior. Any remaining inability to reproduce is reported as an open gate, not a claimed permanent fix.
Related work
Version Packages #687 — earlier Version Packages PR with identical portable failure; passed on a same-commit failed-job rerun and later merged.
test: simplify lifecycle and defer worktree isolation #676 — open lifecycle simplification/worktree-isolation issue. Its body and all available comments were read (no comments). It shares runner/resource-ownership concerns but does not describe or fix this specific processing stall. Full worktree isolation remains separate scope, not a prerequisite for this investigation.
Open/closed searches for max_concurrency, portable-runtime, test_seq, and CI timeout terms found no probable duplicate. This is an ordinary bug issue; no Delivery, roadmap, or dependency relationships are changed.
Open questions and risks
Which message/worker/SQL state actually accompanies the CI stall? Current evidence lacks those snapshots.
Can the interleaving be forced locally, or does it need instrumented GitHub runners? Same-commit success proves intermittency, not a healthy implementation.
Debug logs and database snapshots can change timing; keep them bounded and obtain a regression independent of incidental timing.
Supabase worker/cron activity and shared sequence resets can obscure evidence after failure. Capture before subsequent test cases mutate state.
The warning about vendored src/flows.ts and the separate local CPU/shutdown errors must not be treated as causal without supporting evidence.
The investigation uses only public CI evidence and repository code; local identity/path details and credential-bearing information are omitted/redacted from this issue.
Summary
Diagnose and permanently fix the intermittent Supabase
max_concurrencystall inedge-worker:e2e:portable-runtimes. It repeatedly blocks Version Packages PRs, while the same workload passes under Node and Bun. Add bounded progress/database/worker diagnostics first, reproduce the actual cause, and protect the fix with a deterministic regression test. Do not mask the failure with automatic reruns or a longer timeout.User report or idea
The user originally requested investigation of failing CI on Version Packages PR #687 (0.17.1). The same symptom reappeared on Version Packages PR #693 after the separate telemetry endpoint fix (#691, implemented by #692) merged.
The user asked whether this was the same issue and how to fix it permanently for future runs, then requested this self-contained issue as a handoff for a new worktree. Scope: debugging instrumentation, diagnosis, regression coverage, and the smallest root-cause fix. Issue creation does not create a worktree, change code, rerun CI, merge a PR, or release packages.
Evidence supplied
Current failing CI: PR #693
0ad72226296a8638c3a3e57bd292b00f8071cd25, branchchangeset-release/main.pull_request, failed).pkgs/edge-worker/tests/e2e-portable-runtimes/portable-runtimes.test.ts:507:8.Exact error text (ANSI color codes omitted):
The failing test is
portable runtimes - max_concurrency works in supabase. Deno reports 13 passed and 1 failed in about 1m52s. The Supabase test consumes its 60-second deadline; the Node and Bun variants each pass in about 6 seconds. Other Supabase examples (conn_max_pg_default,conn_max_pg_override,conn_env_var) pass, as do both process-based step-queue tests.All other executed checks pass, including
edge-worker-e2e,edge-worker-integration, build/test, CLI/client E2E, and core pgTAP. Website/demo deployment jobs skip because of the failed prerequisite.Failure-time function-server log now exists
PR #690 added printing of the function-server log when the monitored suite fails. This log header appears in #693:
The dumped runtime log contains these exact lines (outer GitHub job timestamps/prefixes omitted):
Similar
max_concurrencyrequest lines repeat throughout the timeout and later tests. The dumped log has no CPU-limit shutdown, forced SQL close, or connection errors for this run. It does show startup warnings about the missing vendoredsrc/flows.ts, and the expectedauth_testreadiness probe warningAuth validation failed: Missing Authorization header. Those warnings also appeared during locally passing tests; neither is established as causal.Repeated request lines prove HTTP invocations, not worker execution progress, healthy heartbeats, successful cron responses, or message completion.
Download the complete current failed-job evidence:
gh run view 36972575630 --repo pgflow-dev/pgflow --log-failed > /tmp/pgflow-693-failed.logThis is preferable to reconstructing the full log from excerpts. Preserve the complete output from every new reproduction.
Prior failures and checks
dad163005929a97af95c021debbd37c69c42b01d: attempt 1 failed the same portable test; rerunning only failed jobs produced a successful attempt 2 on the same commit.Local checks during #687 investigation:
6cc225de5275164a8f7032206162b821758bb74a; Supabase concurrency took about 3 seconds.These are evidence from a different, higher-volume local run, not proof of the #693 cause. The earlier CPU-shutdown hypothesis must not be upgraded into a finding; #693's dumped log does not support it.
Local setup initially failed because
pkgs/cli/node_modules/.bin/supabasestill pointed at a stale Supabase 2.34.3 installation with no binary. Rootpnpm rebuild supabase, package-filtered rebuild, and frozen-lockfile install alone did not fix that stale shim. Moving that untracked shim aside let the root Supabase 2.63.1 binary resolve, after which fresh setup and tests succeeded. This was a local reproduction obstacle, not the CI root cause; do not bake that ad hoc repair into the worker fix.Investigation and findings
Baseline and test flow
Current main after the telemetry fix:
fb5b8670e(fix(telemetry): use the stable pgflow collector endpoint (#692)). Start the new worktree from currentorigin/main, not the older checkout used during investigation. Worker runtime, portable tests, lifecycle helpers, andpkgs/edge-worker/project.jsonshowed no differences between that main and the earlier inspected worker baseline.The portable fixture lives in
pkgs/edge-worker/supabase/functions/_shared/portable_examples.js. Itsmax_concurrencysettings are:In
portable-runtimes.test.ts:resetExample(..., { preserveWorkerMetadata: true }), purges/creates the queue, restarts the sequence, starts the HTTP worker, waits forworker_functions.start_mode = 'http', captures sequence state, sends the batch, then waits for sequence progress.startWorker()intests/e2e/_helpers.tsinvokes the function and checks for a recent worker heartbeat; that check does not establish future processing health.waitForExampleAssertion()polls every 500ms, uses a 60000ms deadline formax_concurrency, and reports only the target on timeout—not the observed final sequence value.stopped_at. Keep observations separate by runtime/queue.Runtime paths already read during investigation:
src/core/Worker.ts: heartbeat and batch-processing main loop, retry backoff, stop/start/deprecation handling.src/core/BatchProcessor.ts: slot checks,waitForSlot(), polling, and asynchronous execution scheduling.src/core/ExecutionController.ts: promise-queue accounting, slot waiters, completion notifications.src/queue/ReadWithPollPoller.ts,src/queue/Queue.ts,src/queue/MessageExecutor.ts: queue claims, visibility, completion/retry operations.src/queue/createQueueWorker.ts: defaults includemaxPollSeconds: 5,pollIntervalMs: 200,visibilityTimeout: 10,batchSize: 10,maxPgConnections: 4.src/core/WorkerLifecycle.ts,src/core/Queries.ts: worker registration and heartbeat persistence.src/platform/SupabasePlatformAdapter.ts: authenticated HTTP startup, in-isolate worker replacement, lifetime extension, 5-second shutdown deadline, worker marking and SQL cleanup.@henrygd/queuedependency handles scheduling and queue-size accounting. Read its installed implementation if evidence points to completion/slot timing; do not blame it without a reproduction.Lifecycle and CI ownership
pnpm nx run edge-worker:e2e:portable-runtimes(uncached, depends on worker build).scripts/run-e2e.sh tests/e2e-portable-runtimesowns migration setup, vendored dependency sync, function serve, readiness, suite, and exact-resource cleanup under resource locks.scripts/functions-server.shmonitors the CLI server and retains/dumps its output on failure. Server process liveness is not worker-loop liveness..github/workflows/ci.ymlusesubuntu-latest, repository setup, Bun setup, affected-project check,scripts/ci-prestart-supabase.sh edge-worker, then the Nx portable target.tests/e2e/_helpers.tsalready providesfetchWorkers()andstartWorkersMonitor()with worker-function, cron, and pg_net inspection. Reuse useful existing queries rather than introduce a second diagnostics framework. Its debug helper activates underDEBUG=1orVERBOSE=1.What is and is not proved
Proved: the exact symptom recurs; Node/Bun and other Supabase tests pass; a same-commit rerun passed previously; ordinary server logs alone cannot explain the stall.
Not proved: how many of the 200 messages complete; whether the worker heartbeat stalls; whether messages remain visible/invisible or get retried; whether a SQL connection blocks; whether startup/purge/cron races or scheduler notifications explain the failure.
Current strongest lead: an HTTP-responsive function with a processing loop that stops progressing. This is only a hypothesis. Database/pool contention and test setup races remain alternatives.
The merged telemetry fix #692 changes sender URLs, tests, docs, and a SQL migration, not worker execution source. The timeout predates that change; timing after #692 merge is not evidence that the URL change caused it.
Proposed solution or design
1. Establish the reproduction environment and evidence budget
fix/portable-supabase-stall.developing-pgflow; it governs setup, full-suite cadence, and bounded retries. Load schema/pgTAP/migration skills only if the eventual fix touches SQL.pnpm nx test-env:fresh edge-worker, then reuse it. Only one database-mutating worktree at a time until existing shared-resource limitations are addressed. Do not stop unrelated Supabase projects.2. Add bounded diagnostic evidence at the actual stall
Prefer one failure-time snapshot in the portable test, plus narrowly enabled worker logging. The snapshot should identify the runtime/example/queue and include:
stopped_at,deprecated_at, registered function mode, and invocation metadatapg_stat_activityrows with wait events, query age, and blocking PID relationshipsInstrument at the failure boundary, not merely at suite end. Keep snapshots read-only and bounded; diagnostics must not hang on the same exhausted worker pool. Use an independent diagnostic connection/timeout if needed, preserve the original timeout as the primary error, and treat diagnostic query failures as secondary evidence.
Enable debug output for the
max_concurrencyfixture during diagnosis to show heartbeat/polling/batch/slot/execution progress. Keep this confined to the test fixture or diagnostic mode; avoid enabling noisy logs or recording payloads in production. Retain #690's server-log dump. If logs remain too large, report bounded progress/last-known states rather than add an observability framework.Add a small runnable check for snapshot failure handling/format if new nontrivial diagnostic logic warrants it.
3. Use evidence to produce a deterministic regression
4. Validate and deliver separately from telemetry
Acceptance criteria
Related work
Open/closed searches for
max_concurrency,portable-runtime,test_seq, and CI timeout terms found no probable duplicate. This is an ordinary bug issue; no Delivery, roadmap, or dependency relationships are changed.Open questions and risks
src/flows.tsand the separate local CPU/shutdown errors must not be treated as causal without supporting evidence.