From 1886938e2b8c60b313ab84ee3ec499215c3168b3 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Sun, 20 Sep 2026 21:49:52 +0000 Subject: [PATCH 01/23] fix(lifecycle): stop rejected subtask replay loops --- .../architecture/task-lifecycle-gap-report.md | 108 ++++++++------ docs/architecture/task-lifecycle-model.md | 23 +-- .../task-lifecycle-remediation-blocks.md | 99 ++++++++----- scripts/check-task-lifecycle.ts | 135 ++++++++++++++++-- .../ClineProvider.delegation.spec.ts | 123 ++++++++++++++++ .../__tests__/taskLifecycle.spec.ts | 83 +++++++++++ src/core/task-persistence/index.ts | 1 + src/core/task-persistence/taskLifecycle.ts | 19 ++- src/core/webview/ClineProvider.ts | 30 +++- 9 files changed, 513 insertions(+), 108 deletions(-) diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index 66db3eb608..800389f40b 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -6,6 +6,10 @@ This report inventories Zoo Code task lifecycle state, mutation, persistence, sc The audit covers tracked TypeScript, JSON, YAML, and Markdown under `packages/`, `src/`, `apps/cli`, `apps/vscode-e2e`, `scripts/`, `.github/workflows`, and `docs/architecture`. It traces production symbols to bounded models, focused tests, extension-host E2E, and CI entry points. +Inventory completeness is not composed verification. A closed symbol and ownership inventory lists every boundary with its local evidence. It does not prove that persistence, restart, replay, UI approval, scheduling, and external effects compose into one correct protocol. Local evidence classes support an end-to-end claim only through the joint-checker or trace-validation work tracked by `LIFE-GAP-015`. + +Issue [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714) is recorded evidence for this distinction. Every involved symbol (`setPendingTaskAction`, `resumeTaskFromHistory`, `resumePendingTaskAction`, `delegateParentAndOpenChild`, `delegateTaskToChild`) was inventoried with mapped ownership, yet the composed loop across restart, pending-action replay, auto-approval, and an authoritative transition rejection was found by incident report, not by inventory. The #1714 changes add the `settleRejectedCreateSubtaskAction` reducer, `stage` and `settle-rejected` model actions, three settlement witnesses, and matching focused tests. Their direct model coverage is specified in [the model suite](./task-lifecycle-model.md). They address the reported loop for the authoritative-rejection path; LIFE-GAP-039 and LIFE-GAP-040 record the remaining crash-window, startup-repair, and replay-bound obligations. + “Exhaustive” means exhaustive over the repository paths, symbol families, and search terms listed here at the audited commit. It does not include ignored/generated output, deployment branch-protection settings, runtime telemetry, dynamically constructed names that evade text search, or behavior in dependencies. GitHub issue links are historical provenance only; stable `LIFE-GAP-*` IDs own the active burn-down. ## Methodology and audit criteria @@ -163,63 +167,68 @@ CI runs lint, typecheck, and `pnpm lifecycle:model-check` in `.github/workflows/ - Parser replay fixes two scopes, one raw index, and local action order. Transport transforms and arbitrary malformed histories are excluded. - Focused tests and E2E are representative histories, not exhaustive interleavings. - No public-runtime telemetry or production traces were available for trace validation. +- Scheduler liveness is an intentional exclusion. No checker defines fairness for eventual admission, queue drain, cleanup, or retry progress. Proving those properties requires a separate temporal-model audit with explicit fairness assumptions; this report does not claim liveness coverage. +- External tool-effect idempotency is an intentional exclusion. Lifecycle records track call and action identity (`LIFE-GAP-038`), but no model, test, or claim in this audit covers whether a replayed or re-executed tool call repeats effects on the editor, filesystem, or external services. That boundary needs a separate future audit. ## Ranked GAP register Severity reflects plausible data loss, ownership corruption, permission/context errors, or stuck work. Confidence reflects direct source evidence, deterministic witness, or inference. -| ID | Severity | Confidence | Gap and production impact | Witness/reproducer | Dependencies | Objective closure criteria | -| ------------ | ----------- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-GAP-001 | Critical | High | Stale cross-host completion can clear a newer handoff and orphan its child. | Exact shortest witness in shared-store checker. | Disk-authoritative ownership guard or generation. | Deterministic two-store test passes; authoritative child ID is revalidated under disk lock; witness becomes universal invariant. | -| LIFE-GAP-002 | High | High | Stale message save can restore abandoned lineage. | Exact shortest stale-save/abandon witness. | Lifecycle-field ownership, tombstone, or generation. | Stale save cannot alter detached lineage in production test or bounded model; witness promoted. | -| LIFE-GAP-003 | High | High | Legacy dequeue-before-submit can lose queued user feedback on submission failure. | `Task.processQueuedMessages` removes before async submit. | Standardize claim/persist/ack path. | Every queue consumer claims first, removes after durable acceptance, releases on failure; failure test retains message. | -| LIFE-GAP-004 | High | High | Pair writes are not cross-host/crash atomic; partial lifecycle state is observable. | Modeled second-write failure landmark. | WAL/repair intent or explicitly idempotent recovery per pair operation. | Crash injection at each write boundary converges to a documented legal state. | -| LIFE-GAP-005 | High | High | Delegation creates child state and commits parent ownership in separate failure domains. | Provider rollback tests cover selected failures. | Idempotent operation ID or repair intent. | Failure injection before/after every persistence/publication step restores parent or completes delegation without orphan state. | -| LIFE-GAP-006 | High | Medium-high | Parent completion messages can persist before lifecycle completion commit fails. | Source ordering in `reopenParentFromDelegation`. | Transaction/intent or reorder with recovery. | Inject pair-write failure and prove no stale result context, or prove replay safely completes the lifecycle. | -| LIFE-GAP-007 | Medium | High | Three confirmed task-local mode readers are fixed and production-tested, but the repository-wide reader inventory and enforceable task-local boundary remain incomplete. | Merged focused tests cover environment details, built-in validation, and custom-tool execution; the pure checker proves only divergence and selector storage. | Reader inventory and task-local API boundary. | Every mode-sensitive reader is classified; required task-local consumers have divergent production tests in both permission directions; checker/docs distinguish executed readers from proxy evidence. | -| LIFE-GAP-008 | High | High | Responses API argument-only deltas may be dropped; absent indices can alias calls. | Transform requires ID/name and defaults index to zero. | Provider transform correlation design. | Multi-delta and concurrent index-less production tests reconstruct isolated calls; parser model includes mapped transform events. | -| LIFE-GAP-009 | Medium | High | Async lifecycle event listeners have no awaited settlement contract. | Async listeners registered on Node EventEmitter. | Classify notification versus barrier events. | Barrier side effects move to awaited methods; notification listeners have contained rejection tests and documented ordering. | -| LIFE-GAP-010 | Medium | Medium-high | Detached usage drain can write after abort, replacement, delegation, or a newer request generation. | Background iterator mutates/persists without generation guard. | Request-generation ownership. | Late drain may account valid usage but cannot mutate stale UI/message/lifecycle state; controlled delayed-stream test passes. | -| LIFE-GAP-011 | Medium | High | Completion model excludes terminal status persistence and downstream public consumers. | Model ends at emission readiness. | Consumer inventory and contract. | Consequential consumers are enumerated; required status/metadata ordering has production refinement tests. | -| LIFE-GAP-012 | Medium | High | No persisted attempt/generation distinguishes delayed pre-interruption completion from valid post-resume completion of the same child. | Documented model exclusion. | Persisted generation token. | Reducer/store/API/model reject stale generation while accepting resumed generation; restart E2E covers it. | -| LIFE-GAP-013 | Medium | High | Task lifecycle status vocabulary is copied across schema, task metadata, Task, CLI, and history reader. | Literal union inventory. | Shared exported schema-derived type. | Consumers import one owner; CI/static check rejects incompatible local copies. | -| LIFE-GAP-014 | Medium | High | The serial production contract relies on singular reducer ownership and a default one-permit provider scheduler; optional fan-out must not be mistaken for baseline coverage. | The production-backed lifecycle model rejects multiple active awaited children, while the optional fan-out model has no production imports or E2E. | Serial baseline ratchet; separately ticketed fan-out decision. | Baseline: close cross-host violations, assert provider scheduler capacity and serial ordering, and keep fan-out outside baseline CI. Optional fan-out: implement adapters/E2E before reclassification. | -| LIFE-GAP-015 | Medium | High | Independent checks do not establish end-to-end refinement. | Six baseline state spaces plus one optional fan-out state space remain disjoint. | Boundary mappings and tractable joint bounds. | Add joint checker/trace validation for each cross-model claim, or keep every claim explicitly local. | -| LIFE-GAP-016 | Medium | High | Traceability is documentary and can drift from scripts, symbols, tests, and CI. | Shared-store scenario count has drifted in documentation. | Machine-readable manifest/checker summaries. | CI validates stable IDs, model membership, bounds, symbol/test paths, and workflow invocation. | -| LIFE-GAP-017 | Medium | High | Store cache records are exposed without cloning; external mutation may bypass locking. | `get`/`getAll` return cached objects. | Immutability/read API decision. | Freeze/clone records or prove callers cannot mutate; mutation regression test. | -| LIFE-GAP-018 | Medium | High | Ordinary history-file reads cast JSON instead of applying the shared schema. | Store reconciliation/read path. | Validation/quarantine policy. | Malformed records are rejected or quarantined deterministically with tests and recovery documentation. | -| LIFE-GAP-019 | Medium | High | Generic task IDs lack the importer’s explicit path-safety validation. | Importer validates IDs; generic paths interpolate IDs. | Shared safe-ID boundary. | Every filesystem task ID passes one validator; traversal and separator tests cover all entry points. | -| LIFE-GAP-020 | Medium | Medium-high | Watch/reconcile convergence is eventual and failure-tolerant, not coherent. | Debounced watcher plus periodic scan. | Version/notification or documented eventual contract. | Define stale-read window and convergence property; multi-host test covers missed watcher event and concurrent update. | -| LIFE-GAP-021 | Medium | High | Store disposal does not await queued writes. | Synchronous `dispose` stops watcher/timer only. | Async drain/close contract. | Disposal awaits or explicitly cancels writes; no post-dispose writes in deterministic test. | -| LIFE-GAP-022 | Medium | High | Public clear and webview clear use different delegated-child semantics. | API uses eviction; webview removes directly. | One clear contract. | All ingress paths converge on the same lifecycle transition and tests assert identical persisted results. | -| LIFE-GAP-023 | Medium | High | History deletion can be resurrected after swallowed unlink failure. | Cache removal precedes best-effort unlink. | Tombstone or surfaced failure/retry. | Inject unlink failure and prove item stays deleted or operation reports failure without false success. | -| LIFE-GAP-024 | Medium | High | Parser cleanup depends on normal finalization or garbage collection. | Weak scope maps lack universal request `finally`. | Request-level cleanup owner. | Abort/error/success all clear active parser state in production integration tests. | -| LIFE-GAP-025 | Medium | High | Abort listeners may accumulate during successful stream chunks. | Per-chunk listener removed only by abort. | Settle-time listener cleanup. | Long stream keeps bounded listener count and removes listeners on both race outcomes. | -| LIFE-GAP-026 | Medium | Medium-high | Usage-drain timeout cannot interrupt a permanently pending `iterator.next()`. | Elapsed time checked before await. | Deadline race/abort. | Hung iterator settles drain within wall-clock bound in fake-timer test. | -| LIFE-GAP-027 | Medium | High | Task-level delegation listeners are untyped/dead while provider listeners own the same public events. | `src/extension/api.ts` registers both paths. | Single typed event owner. | Remove duplicate/dead listeners or define one source; public API test proves exactly-once emission. | -| LIFE-GAP-028 | Low-medium | High | `TaskSpawned` payload semantics differ across task/provider/public surfaces. | Child-only, ambiguous task ID, and parent+child forms. | Event contract normalization. | Payloads use explicit names and adapters are type-checked with compatibility tests. | -| LIFE-GAP-029 | Low-medium | High | `Task.taskStatus` and `TaskRegistry.getRunning` are projections, not scheduler/persistence truth. | Ask markers and abort flags only. | Naming/contract clarification. | Rename or document exact predicates; callers stop using them as stronger lifecycle evidence. | -| LIFE-GAP-030 | Low-medium | Medium | `Task.run()` may resolve immediately if another path already started the task. | `_started` short-circuit versus scheduler callback. | Single start owner. | Scheduler-facing start returns the actual run promise; duplicate-start test proves settlement identity. | -| LIFE-GAP-031 | Low-medium | High | Queue state is memory-only and cleared on disposal. | `MessageQueueService.dispose`. | Product durability decision. | Document intentional loss or persist claims/messages with restart tests. | -| LIFE-GAP-032 | Low-medium | Medium | Webview abandonment handler exists without a confirmed production UI sender. | Protocol/handler search only. | Reachability decision. | Add supported sender/E2E or remove/deprecate unreachable command. | -| LIFE-GAP-033 | Low-medium | High | Resume ingress differs in awaiting and error propagation. | Webview/API/IPC adapters diverge. | Shared resume operation contract. | Contract tests compare result/error/publication semantics for each surface. | -| LIFE-GAP-034 | Low | High | Bounds and model metadata are handwritten and not mechanically synchronized. | Constants, prose, and console summaries duplicate values. | Machine-readable checker metadata. | CI compares emitted metadata with docs/manifest and rejects undocumented bound/action/landmark changes. | -| LIFE-GAP-035 | Medium | High | Tool-originated child initialization lacks a complete durable ownership contract; initial todo state is the confirmed witness and can disappear after rehydration. | Create a child with explicit initial todos, switch or restart before `update_todo_list`, then reopen it; restoration finds no todo message and yields an empty list. | Canonical durable task-ID-scoped child-initialization owner and publication contract. | Inventory every `new_task`-originated child field; persist required initial state before visibility/run; restore deep-equal independent state across switching, interrupted resume, checkpoint restore, and fresh-host restart; preserve later-update and explicit-empty precedence; add constructor deep-copy, persistence, webview scoping, and E2E witnesses. | -| LIFE-GAP-036 | High | High | Interactive todo approval edit state is process-global and uncorrelated; one task's delayed edit can be consumed by another task's pending approval. | Start approvals for tasks A and B, send A's edited list through `setPendingTodoList`, then resolve B; B reads the shared `approvedTodoList`. | Task/action/tool-call-correlated approval state and webview protocol. | Carry task ID and action/tool-call ID through proposal, webview edit, approval, cancellation, and settlement; reject stale/mismatched edits; deep-clone inputs; test two interleaved approvals, denial, cancellation, task switch, and delayed edits. | -| LIFE-GAP-037 | Medium-high | High | Singleton tool handlers share partial presentation state across calls/tasks, so interleaved paths can cause false or missed stabilization. | Interleave A:`x`, B:`y`, A:`x` or A:`x`, B:`x` through one handler's `lastSeenPartialPath`. | Per-call handler state keyed by task and tool-call identity. | Isolate partial state by `(taskId, toolCallId)` or handler instance; prove independent stabilization and cleanup after success, malformed finalization, rejection, cancellation, abandonment, and incomplete streams. | -| LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. | +| ID | Severity | Confidence | Gap and production impact | Witness/reproducer | Dependencies | Objective closure criteria | +| ------------ | ----------- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-GAP-001 | Critical | High | Stale cross-host completion can clear a newer handoff and orphan its child. | Exact shortest witness in shared-store checker. | Disk-authoritative ownership guard or generation. | Deterministic two-store test passes; authoritative child ID is revalidated under disk lock; witness becomes universal invariant. | +| LIFE-GAP-002 | High | High | Stale message save can restore abandoned lineage. | Exact shortest stale-save/abandon witness. | Lifecycle-field ownership, tombstone, or generation. | Stale save cannot alter detached lineage in production test or bounded model; witness promoted. | +| LIFE-GAP-003 | High | High | Legacy dequeue-before-submit can lose queued user feedback on submission failure. | `Task.processQueuedMessages` removes before async submit. | Standardize claim/persist/ack path. | Every queue consumer claims first, removes after durable acceptance, releases on failure; failure test retains message. | +| LIFE-GAP-004 | High | High | Pair writes are not cross-host/crash atomic; partial lifecycle state is observable. | Modeled second-write failure landmark. | WAL/repair intent or explicitly idempotent recovery per pair operation. | Crash injection at each write boundary converges to a documented legal state. | +| LIFE-GAP-005 | High | High | Delegation creates child state and commits parent ownership in separate failure domains. | Provider rollback tests cover selected failures. | Idempotent operation ID or repair intent. | Failure injection before/after every persistence/publication step restores parent or completes delegation without orphan state. | +| LIFE-GAP-006 | High | Medium-high | Parent completion messages can persist before lifecycle completion commit fails. | Source ordering in `reopenParentFromDelegation`. | Transaction/intent or reorder with recovery. | Inject pair-write failure and prove no stale result context, or prove replay safely completes the lifecycle. The composed obligation spans the message write, both pair files, and child pending-action state; per-layer passes do not compose. Closure must enumerate one phase table across those writes and state the recovery interaction with resumed pending-action replay (012). | +| LIFE-GAP-007 | Medium | High | Three confirmed task-local mode readers are fixed and production-tested, but the repository-wide reader inventory and enforceable task-local boundary remain incomplete. | Merged focused tests cover environment details, built-in validation, and custom-tool execution; the pure checker proves only divergence and selector storage. | Reader inventory and task-local API boundary. | Every mode-sensitive reader is classified; required task-local consumers have divergent production tests in both permission directions; checker/docs distinguish executed readers from proxy evidence. | +| LIFE-GAP-008 | High | High | Responses API argument-only deltas may be dropped; absent indices can alias calls. | Transform requires ID/name and defaults index to zero. | Provider transform correlation design. | Multi-delta and concurrent index-less production tests reconstruct isolated calls; parser model includes mapped transform events. | +| LIFE-GAP-009 | Medium | High | Async lifecycle event listeners have no awaited settlement contract. | Async listeners registered on Node EventEmitter. | Classify notification versus barrier events. | Barrier side effects move to awaited methods; notification listeners have contained rejection tests and documented ordering. | +| LIFE-GAP-010 | Medium | Medium-high | Detached usage drain can write after abort, replacement, delegation, or a newer request generation. | Background iterator mutates/persists without generation guard. | Request-generation ownership. | Late drain may account valid usage but cannot mutate stale UI/message/lifecycle state; controlled delayed-stream test passes. | +| LIFE-GAP-011 | Medium | High | Completion model excludes terminal status persistence and downstream public consumers. | Model ends at emission readiness. | Consumer inventory and contract. | Consequential consumers are enumerated; required status/metadata ordering has production refinement tests. | +| LIFE-GAP-012 | Medium | High | No persisted attempt/generation distinguishes delayed pre-interruption completion from valid post-resume completion of the same child. | Documented model exclusion. | Persisted generation token. | Reducer/store/API/model reject stale generation while accepting resumed generation; restart E2E covers it. The token must also cover staged pending-action replay, because `resumePendingTaskAction` re-executes a staged action after restart with no attempt counter or bound, as #1714 demonstrated. Settlement of a stale action must record the generation so a later legitimate action is accepted. | +| LIFE-GAP-013 | Medium | High | Task lifecycle status vocabulary is copied across schema, task metadata, Task, CLI, and history reader. | Literal union inventory. | Shared exported schema-derived type. | Consumers import one owner; CI/static check rejects incompatible local copies. | +| LIFE-GAP-014 | Medium | High | The serial production contract relies on singular reducer ownership and a default one-permit provider scheduler; optional fan-out must not be mistaken for baseline coverage. | The production-backed lifecycle model rejects multiple active awaited children, while the optional fan-out model has no production imports or E2E. | Serial baseline ratchet; separately ticketed fan-out decision. | Baseline: close cross-host violations, assert provider scheduler capacity and serial ordering, and keep fan-out outside baseline CI. Optional fan-out: implement adapters/E2E before reclassification. | +| LIFE-GAP-015 | Medium | High | Independent checks do not establish end-to-end refinement. | Six baseline state spaces plus one optional fan-out state space remain disjoint. | Boundary mappings and tractable joint bounds. | Add joint checker/trace validation for each cross-model claim, or keep every claim explicitly local. | +| LIFE-GAP-016 | Medium | High | Traceability is documentary and can drift from scripts, symbols, tests, and CI. | Shared-store scenario count has drifted in documentation. | Machine-readable manifest/checker summaries. | CI validates stable IDs, model membership, bounds, symbol/test paths, and workflow invocation. | +| LIFE-GAP-017 | Medium | High | Store cache records are exposed without cloning; external mutation may bypass locking. | `get`/`getAll` return cached objects. | Immutability/read API decision. | Freeze/clone records or prove callers cannot mutate; mutation regression test. | +| LIFE-GAP-018 | Medium | High | Ordinary history-file reads cast JSON instead of applying the shared schema. | Store reconciliation/read path. | Validation/quarantine policy. | Malformed records are rejected or quarantined deterministically with tests and recovery documentation. | +| LIFE-GAP-019 | Medium | High | Generic task IDs lack the importer’s explicit path-safety validation. | Importer validates IDs; generic paths interpolate IDs. | Shared safe-ID boundary. | Every filesystem task ID passes one validator; traversal and separator tests cover all entry points. | +| LIFE-GAP-020 | Medium | Medium-high | Watch/reconcile convergence is eventual and failure-tolerant, not coherent. | Debounced watcher plus periodic scan. | Version/notification or documented eventual contract. | Define stale-read window and convergence property; multi-host test covers missed watcher event and concurrent update. Convergence must define a pending-action rule: reconciliation and startup repair change status and lineage without any pending-action decision, so a converged record can keep an action whose request no longer exists (the #1714 precondition survived startup reconciliation). The stale-read window must also cover the second settlement write that follows an authoritative rejection (039). | +| LIFE-GAP-021 | Medium | High | Store disposal does not await queued writes. | Synchronous `dispose` stops watcher/timer only. | Async drain/close contract. | Disposal awaits or explicitly cancels writes; no post-dispose writes in deterministic test. | +| LIFE-GAP-022 | Medium | High | Public clear and webview clear use different delegated-child semantics. | API uses eviction; webview removes directly. | One clear contract. | All ingress paths converge on the same lifecycle transition and tests assert identical persisted results. | +| LIFE-GAP-023 | Medium | High | History deletion can be resurrected after swallowed unlink failure. | Cache removal precedes best-effort unlink. | Tombstone or surfaced failure/retry. | Inject unlink failure and prove item stays deleted or operation reports failure without false success. | +| LIFE-GAP-024 | Medium | High | Parser cleanup depends on normal finalization or garbage collection. | Weak scope maps lack universal request `finally`. | Request-level cleanup owner. | Abort/error/success all clear active parser state in production integration tests. | +| LIFE-GAP-025 | Medium | High | Abort listeners may accumulate during successful stream chunks. | Per-chunk listener removed only by abort. | Settle-time listener cleanup. | Long stream keeps bounded listener count and removes listeners on both race outcomes. | +| LIFE-GAP-026 | Medium | Medium-high | Usage-drain timeout cannot interrupt a permanently pending `iterator.next()`. | Elapsed time checked before await. | Deadline race/abort. | Hung iterator settles drain within wall-clock bound in fake-timer test. | +| LIFE-GAP-027 | Medium | High | Task-level delegation listeners are untyped/dead while provider listeners own the same public events. | `src/extension/api.ts` registers both paths. | Single typed event owner. | Remove duplicate/dead listeners or define one source; public API test proves exactly-once emission. | +| LIFE-GAP-028 | Low-medium | High | `TaskSpawned` payload semantics differ across task/provider/public surfaces. | Child-only, ambiguous task ID, and parent+child forms. | Event contract normalization. | Payloads use explicit names and adapters are type-checked with compatibility tests. | +| LIFE-GAP-029 | Low-medium | High | `Task.taskStatus` and `TaskRegistry.getRunning` are projections, not scheduler/persistence truth. | Ask markers and abort flags only. | Naming/contract clarification. | Rename or document exact predicates; callers stop using them as stronger lifecycle evidence. | +| LIFE-GAP-030 | Low-medium | Medium | `Task.run()` may resolve immediately if another path already started the task. | `_started` short-circuit versus scheduler callback. | Single start owner. | Scheduler-facing start returns the actual run promise; duplicate-start test proves settlement identity. | +| LIFE-GAP-031 | Low-medium | High | Queue state is memory-only and cleared on disposal. | `MessageQueueService.dispose`. | Product durability decision. | Document intentional loss or persist claims/messages with restart tests. The durability decision must state the restart composition: which queue states can coexist with a staged pending action after restart, and whether action replay can consume feedback that was never acknowledged. | +| LIFE-GAP-032 | Low-medium | Medium | Webview abandonment handler exists without a confirmed production UI sender. | Protocol/handler search only. | Reachability decision. | Add supported sender/E2E or remove/deprecate unreachable command. | +| LIFE-GAP-033 | Low-medium | High | Resume ingress differs in awaiting and error propagation. | Webview/API/IPC adapters diverge. | Shared resume operation contract. | Contract tests compare result/error/publication semantics for each surface. | +| LIFE-GAP-034 | Low | High | Bounds and model metadata are handwritten and not mechanically synchronized. | Constants, prose, and console summaries duplicate values. | Machine-readable checker metadata. | CI compares emitted metadata with docs/manifest and rejects undocumented bound/action/landmark changes. | +| LIFE-GAP-035 | Medium | High | Tool-originated child initialization lacks a complete durable ownership contract; initial todo state is the confirmed witness and can disappear after rehydration. | Create a child with explicit initial todos, switch or restart before `update_todo_list`, then reopen it; restoration finds no todo message and yields an empty list. | Canonical durable task-ID-scoped child-initialization owner and publication contract. | Inventory every `new_task`-originated child field; persist required initial state before visibility/run; restore deep-equal independent state across switching, interrupted resume, checkpoint restore, and fresh-host restart; preserve later-update and explicit-empty precedence; add constructor deep-copy, persistence, webview scoping, and E2E witnesses. | +| LIFE-GAP-036 | High | High | Interactive todo approval edit state is process-global and uncorrelated; one task's delayed edit can be consumed by another task's pending approval. | Start approvals for tasks A and B, send A's edited list through `setPendingTodoList`, then resolve B; B reads the shared `approvedTodoList`. | Task/action/tool-call-correlated approval state and webview protocol. | Carry task ID and action/tool-call ID through proposal, webview edit, approval, cancellation, and settlement; reject stale/mismatched edits; deep-clone inputs; test two interleaved approvals, denial, cancellation, task switch, and delayed edits. Correlation must extend to the staged pending action: `NewTaskTool` persists the action before approval, so the protocol must bind `pendingAction.actionId` to the approving task and tool call and reject cross-task edits before staging; denied, cancelled, and errored approvals need durable settlement rules equal to the rejected-delegation settlement; deep-copy extends to `pendingAction.todos` at staging. | +| LIFE-GAP-037 | Medium-high | High | Singleton tool handlers share partial presentation state across calls/tasks, so interleaved paths can cause false or missed stabilization. | Interleave A:`x`, B:`y`, A:`x` or A:`x`, B:`x` through one handler's `lastSeenPartialPath`. | Per-call handler state keyed by task and tool-call identity. | Isolate partial state by `(taskId, toolCallId)` or handler instance; prove independent stabilization and cleanup after success, malformed finalization, rejection, cancellation, abandonment, and incomplete streams. | +| LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. The bijection must include the staged pending action: `pendingAction.actionId` derives from `sanitizeToolUseId(toolCallId)`, so non-injective sanitization can make a settlement targeted at one raw call clear another call's staged action; the model's stale-settlement witness depends on collision-free action IDs. | +| LIFE-GAP-039 | High | High | Settlement after an authoritative pending-action rejection is a second, best-effort store write. A crash or settlement-write failure between the rejected transition and the settlement write leaves the rejected action staged, and startup replay can re-execute it without an attempt bound. | Source: the `delegateParentAndOpenChild` settlement block logs and swallows settlement errors; `pendingAction` carries no attempt count; the #1714 bounded-retry fix suggestion is unimplemented. | Durable operation intent (P2-004); attempt/generation identity (012). | Fault injection between rejection and settlement converges to a documented legal state; settlement-write failure is retried or surfaced without false success; replay of a rejected action is bounded by a persisted attempt/backoff record; model or focused witnesses cover the crash window. | +| LIFE-GAP-040 | Medium | Medium-high | Startup topology repair reconciles status and lineage but has no pending-action rule. `reconcileDelegationState`, `repairActiveDelegation`, and repair-intent replay can leave a staged action on a repaired record, and the two interrupt paths disagree on whether `pendingAction` survives. | #1714 reproduced the loop precondition through startup reconciliation of a persisted active child; `DelegationRepairIntent` guards and targets carry no pending-action field; `applyDelegationRepairIntent` spreads records without a pending-action decision. | Durable operation intent (P2-004); correlated action identity (036, 038). | Every repair outcome (orphaned delegation, orphaned active child, completed child, intent replay, quarantine) defines preserve, settle, or drop for pending actions; deterministic repair tests and one fresh-host restart test prove the #1714 precondition cannot survive repair; both interrupt paths persist or settle the action consistently. | +| LIFE-GAP-041 | Medium | Medium | Cascade deletion traverses persisted `childIds` with no cycle or duplicate guard. `deleteTaskWithId`'s `collectChildIds` recurses through unvalidated records, so a cyclic or duplicated persisted graph recurses without bound, while cost aggregation already carries a visited set. | Structural: `collectChildIds` has no visited set or depth bound; `aggregateTaskCostsRecursive` does; persisted reads are unvalidated (018). | Schema-validated history reads (018); shared safe-ID boundary (019). | Traversal terminates on cyclic, self-referencing, and duplicate persisted graphs with deterministic fixture tests; cascade deletion deletes the acyclic closure exactly once or fails closed with a surfaced error; the guard is shared with other persisted-graph traversals or listed with exclusions. | ## Portfolio remediation plan The [1-SP remediation block register](./task-lifecycle-remediation-blocks.md) decomposes this portfolio into small modeling/documentation increments. It assigns every GAP exactly one primary block, preserves dependencies across workstreams, and keeps optional fan-out separate from baseline ownership. -The 38 IDs are not 38 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. +The 41 IDs are not 41 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. | Cluster | Gap IDs | Root fix and likely ownership | Complexity | Engineering risk | Objective portfolio evidence | | ------------------------------------------ | -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | P1. Persisted ownership and generation | 001, 002, 012, 017, 020 | Disk-authoritative lifecycle ownership/generation and immutable store reads across history types, lifecycle reducers, `TaskHistoryStore`, provider delegation, and reconciliation. | XL | High: persisted compatibility and cross-host races | Two-host stale-write tests, promoted invariants, restart/reconciliation evidence, backward-compatible optional data. | -| P2. Durable operation and crash recovery | 004, 005, 006, 021, 023 | Operation intent/replay or explicit idempotent recovery for pair writes, delegation, completion messages, shutdown, and deletion. | XL | Very high: failure ordering can create new corruption | Fault injection at every durable boundary, crash/restart convergence, no false success, recovery reachability. | -| P3. Schema, path, and lifecycle vocabulary | 013, 018, 019 | Schema-derived status ownership, validated ordinary history reads, and one safe task-ID boundary across types, persistence, metadata, CLI, and import paths. | M | Medium: malformed legacy data and downgrade behavior | Migration/quarantine fixtures, traversal tests, type/static ratchets. | +| P2. Durable operation and crash recovery | 004, 005, 006, 021, 023, 039, 040 | Operation intent/replay or explicit idempotent recovery for pair writes, delegation, completion messages, shutdown, and deletion. | XL | Very high: failure ordering can create new corruption | Fault injection at every durable boundary, crash/restart convergence, no false success, recovery reachability. | +| P3. Schema, path, and lifecycle vocabulary | 013, 018, 019, 041 | Schema-derived status ownership, validated ordinary history reads, and one safe task-ID boundary across types, persistence, metadata, CLI, and import paths. | M | Medium: malformed legacy data and downgrade behavior | Migration/quarantine fixtures, traversal tests, type/static ratchets. | | P4. Request, stream, and tool identity | 008, 010, 024, 025, 026, 030, 037, 038 | Request generation plus canonical call identity, then task/generation/call-scoped parser and partial-handler state. Owners include provider transforms, parser, `Task`, `BaseTool`, editing handlers, and tool-ID utilities. | XL | High: provider compatibility and duplicate execution | Adversarial IDs/index-less streams, delayed/cancelled generation tests, cleanup/deadline checks, production-backed call-state model. | | P5. Tool-owned task state and queueing | 003, 007, 031, 035, 036 | Task-local context, durable child initialization, correlated approval identity, and claim/persist/ack queueing across tools, `Task`, provider/webview, message queue, and history schema. | XL | High: cross-task contamination and persistence precedence | Omitted/explicit child controls, switch/restart E2E, two-approval schedules, queue failure retention, mode-permission tests. | | P6. Event and ingress contracts | 009, 011, 022, 027, 028, 029, 032, 033 | Classify barriers versus notifications; normalize lifecycle payloads and clear/resume semantics across Task, provider, public API, IPC, and webview. | L | Medium-high: public compatibility and ordering | Exactly-once event tests, consumer inventory, cross-surface contract matrix, compatibility adapters where required. | @@ -229,10 +238,11 @@ The 38 IDs are not 38 independent projects. They group into eight programs with ### Root fixes that close multiple gaps - One persisted generation and disk-authoritative ownership design should close 001, 002, and 012; immutable reads and explicit reconciliation semantics address 017/020 around that owner. -- One durable operation-intent/replay framework can support 004, 005, 006, 021, and 023, but each operation still needs its own legal recovery states and fault-injection matrix. +- One durable operation-intent/replay framework can support 004, 005, 006, 021, 023, 039, and 040, but each operation still needs its own legal recovery states and fault-injection matrix. The rejection-settlement crash window and startup-repair pending-action rules join the same framework. - One request-generation/canonical-call identity established before parser indexing can support 008, 010, 024–026, 030, 037, and 038. - One correlated `(taskId, actionId, toolCallId)` approval protocol can close 036 and support 007/035; it does not itself make child state durable. - One typed lifecycle operation layer can normalize P6, but public compatibility requires separate adapters rather than a flag-day payload rewrite. +- One validated-read and cycle-safe traversal guard closes 041 beside 018 and 019. ### Independent work that should not be collapsed @@ -291,9 +301,9 @@ Four workstreams can proceed concurrently after foundation decisions: ## Burn-down dependencies | Dependency | Enables | -| ---------------------------------------------- | ------------------------------------------------------------- | +| ---------------------------------------------- | ------------------------------------------------------------- | --- | | Disk-authoritative ownership/generation design | LIFE-GAP-001, 002, 012, 020 | -| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023 | +| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023, 039, 040 | | Task-local execution-context owner | LIFE-GAP-007 and optional future fan-out | | Request generation and terminal cleanup owner | LIFE-GAP-010, 024, 025, 026 | | Event notification/barrier contract | LIFE-GAP-009, 011, 027, 028 | @@ -304,6 +314,7 @@ Four workstreams can proceed concurrently after foundation decisions: | Correlated approval ownership | LIFE-GAP-036 | | Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | | Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | +| Validated-read and traversal-guard design | LIFE-GAP-018, 019, 041 | | ## Mechanically useful follow-up checklist @@ -324,6 +335,9 @@ Four workstreams can proceed concurrently after foundation decisions: - [ ] For LIFE-GAP-036, correlate every approval edit and settlement with task ID plus action/tool-call ID; reject stale cross-task edits. - [ ] For LIFE-GAP-037, interleave equal and unequal partial paths across two calls and two tasks, then verify terminal cleanup. - [ ] For LIFE-GAP-038, use adversarial raw IDs to verify one-to-one durable call, approval, execution, result, pending-action, and replay identity. +- [ ] For LIFE-GAP-039, inject failure between the authoritative rejection and the settlement write and prove bounded, convergent recovery without unbounded replay. +- [ ] For LIFE-GAP-040, run startup repair and reconciliation against records that carry staged pending actions and prove the #1714 loop precondition cannot survive repair. +- [ ] For LIFE-GAP-041, delete through cyclic, self-referencing, and duplicate persisted `childIds` graphs and prove termination with exact-once deletion or a surfaced failure. ## Completeness statement @@ -331,4 +345,4 @@ At the audited commit, this report covers every tracked definition and directly ## Historical provenance -Relevant historical reports include [#1469](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1469), [#1021](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1021), [#1623](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1623), [#369](https://github.com/Zoo-Code-Org/Zoo-Code/issues/369), [#372](https://github.com/Zoo-Code-Org/Zoo-Code/issues/372), [#612](https://github.com/Zoo-Code-Org/Zoo-Code/issues/612), [#1453](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1453), [#1279](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1279), [#921](https://github.com/Zoo-Code-Org/Zoo-Code/issues/921), [#920](https://github.com/Zoo-Code-Org/Zoo-Code/issues/920), and [#1468](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1468). These links provide provenance only; closure is governed by the objective criteria above. +Relevant historical reports include [#1469](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1469), [#1021](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1021), [#1623](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1623), [#369](https://github.com/Zoo-Code-Org/Zoo-Code/issues/369), [#372](https://github.com/Zoo-Code-Org/Zoo-Code/issues/372), [#612](https://github.com/Zoo-Code-Org/Zoo-Code/issues/612), [#1453](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1453), [#1279](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1279), [#921](https://github.com/Zoo-Code-Org/Zoo-Code/issues/921), [#920](https://github.com/Zoo-Code-Org/Zoo-Code/issues/920), and [#1468](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1468), with rejected-delegation settlement, its crash window, and its startup-repair aftermath recorded in [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714). These links provide provenance only; closure is governed by the objective criteria above. diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index 355b08aa7e..22b7ded257 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -48,11 +48,14 @@ TLA+/PlusCal or Quint with TLC becomes a better fit when the lifecycle needs tem | `interrupt(child)` | cancellation or eviction through `markDelegatedChildInterrupted` | | `complete(child)` | `ClineProvider.reopenParentFromDelegation` | | `abandon(child)` | `ClineProvider.abandonSubtask` | +| Pending-action settlement | `settleRejectedCreateSubtaskAction` inside the delegation update path (#1714) | | Atomic event step | `atomicReadAndUpdate`, `atomicUpdatePair`, and per-parent delegation transition lock | | Event interleaving | Competing completion, cancellation, abandonment, and new delegation calls | The model has three fixed task slots, enough to cover competing siblings and a nested parent-child-grandchild chain. It explores every reachable interleaving through depth 12, deduplicating canonical states. Representative checks also exercise rejected operations that do not create a new state: a second concurrent delegation while the first child is active, stale completion after re-delegation, late completion after abandonment, completion after interruption, and nested completion. Named semantic landmarks require the graph to retain interrupted-child re-delegation and nested delegation even when the raw state total changes. +Each task slot can also hold one of two pending `create_subtask` actions. A `stage` action mirrors `setPendingTaskAction` overwrite semantics, delegation and completion clear the action their request carried, and a `settle-rejected` action models the settlement that follows an authoritative delegation rejection (#1714). Production settles through the typed `LifecycleTransitionError` from the shared guards: the provider writes the settled record with one extra `atomicReadAndUpdate` call, then propagates the original rejection. Three named witnesses must remain reachable: settlement from an interrupted record after rejection, settlement through a successful active delegation, and stale-action protection where a settlement targeting one action ID leaves a replacement action intact. A mismatched pending-action request keeps its production behavior: the atomic update throws before any transition, and no settlement runs. + Production completion also accepts a recovery-compatible `active` parent that still awaits the returning child, then clears the stale pointers. Normal model transitions never create that intermediate state, so it is covered by a focused reducer test rather than admitted as a generally valid reachable state. ## Shared-store concurrency model @@ -134,6 +137,7 @@ The task delegation checker currently enforces: 5. Parent-child lineage is acyclic. 6. Completed task records cannot be changed by later lifecycle events. 7. Active-child re-delegation, stale completion after ownership moves to another child, duplicate/late completion, and abandonment of a live child are rejected by the shared production guards. +8. A rejected delegation settles only the exact matching pending `create_subtask` action. Settlement preserves status, lineage, and accounting, never clears a replacement or different-kind action, and never mutates a completed record. The completion persistence checker additionally enforces: @@ -147,15 +151,15 @@ These are safety claims within the documented bounds. The checks do not claim li ## Coverage audit -| Protocol area | Coverage status | Production/model relationship | Explicit limits and open points | -| ------------------------------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Delegation lifecycle | Production-backed bounded universal | The explorer calls the four production reducers for three task slots through depth 12. | Excludes provider instances, persistence failures, scheduler state, most live `Task` behavior, and generation identity for delayed pre-interruption completion. Recovery-compatible active-parent completion is test-only. | -| Shared-store concurrency | Production-backed bounded scenarios plus known-unsafe witnesses | The explorer imports production delta/merge functions and reducers; a real-filesystem test is a smoke check. | Does not prove crash safety, filesystem/lock semantics, arbitrary processes, or loss-free same-field merging. #1469 and #1021 remain unsafe. | -| Provider handoff and scheduler | Mixed: production-backed reducers/selector plus abstract bounded protocol | Commits use production reducers; provider ownership, publication, transition locks, and permits are model abstractions through depth 15. | Selector correctness does not refine all downstream readers. Scheduler tests cover concrete permit behavior separately. | -| Optional fan-out scope | Planned-only abstract bounded protocol outside baseline CI | The model has two sibling slots and two abstract permits and imports no production fan-out transition. | Excluded from baseline closure; production fan-out remains separately scoped future functionality. | -| Cleanup | Abstract bounded universal plus adapter tests | Abort, disposal, settlement, rejection, and provider shutdown are modeled as protocol/environment actions. | No direct execution of all production cleanup methods, filesystem/editor promises, timing liveness, fairness, or arbitrary task counts. | -| Parser request scope | Production-backed bounded schedule replay | The checker executes production parser APIs across 924 order-preserving schedules for two scopes. | Assumes callers stop invoking a finalized scope; transport behavior, arbitrary request counts, indices, and malformed histories are outside the claim. | -| Completion persistence | Abstract bounded universal plus production tests and one fresh-host E2E path | The model abstracts persistence as a durable phase with at most two write starts; production guards and retry paths are tested separately. | Production permits more retries; no power-loss/filesystem proof, fairness, arbitrary retry count, complete delegated fallback, provider status metadata, or downstream event-consumer model. | +| Protocol area | Coverage status | Production/model relationship | Explicit limits and open points | +| ------------------------------ | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Delegation lifecycle | Production-backed bounded universal | The explorer calls the production lifecycle and settlement reducers for three task slots through depth 12. | Excludes provider instances, persistence failures, scheduler state, most live `Task` behavior, and generation identity for delayed pre-interruption completion. Recovery-compatible active-parent completion is test-only. Settlement models the matching-action request only; a mismatched action aborts the update unmodeled. | +| Shared-store concurrency | Production-backed bounded scenarios plus known-unsafe witnesses | The explorer imports production delta/merge functions and reducers; a real-filesystem test is a smoke check. | Does not prove crash safety, filesystem/lock semantics, arbitrary processes, or loss-free same-field merging. #1469 and #1021 remain unsafe. | +| Provider handoff and scheduler | Mixed: production-backed reducers/selector plus abstract bounded protocol | Commits use production reducers; provider ownership, publication, transition locks, and permits are model abstractions through depth 15. | Selector correctness does not refine all downstream readers. Scheduler tests cover concrete permit behavior separately. | +| Optional fan-out scope | Planned-only abstract bounded protocol outside baseline CI | The model has two sibling slots and two abstract permits and imports no production fan-out transition. | Excluded from baseline closure; production fan-out remains separately scoped future functionality. | +| Cleanup | Abstract bounded universal plus adapter tests | Abort, disposal, settlement, rejection, and provider shutdown are modeled as protocol/environment actions. | No direct execution of all production cleanup methods, filesystem/editor promises, timing liveness, fairness, or arbitrary task counts. | +| Parser request scope | Production-backed bounded schedule replay | The checker executes production parser APIs across 924 order-preserving schedules for two scopes. | Assumes callers stop invoking a finalized scope; transport behavior, arbitrary request counts, indices, and malformed histories are outside the claim. | +| Completion persistence | Abstract bounded universal plus production tests and one fresh-host E2E path | The model abstracts persistence as a durable phase with at most two write starts; production guards and retry paths are tested separately. | Production permits more retries; no power-loss/filesystem proof, fairness, arbitrary retry count, complete delegated fallback, provider status metadata, or downstream event-consumer model. | The production mapping above names primary lifecycle transitions, not every mutation or consumer. Generic store upserts, reconciliation, repair replay, migrations, tool entry points, webview/public API abandonment, provider status updates, and public `TaskCompleted` re-emission remain outside the persisted reducer graph unless explicitly named by a submodel or focused test. For task-local mode, the consumers named under #921/#1623 are a confirmed set, not an exhaustive repository-wide inventory; other mode-sensitive tools must be audited before claiming universal reader isolation. @@ -197,6 +201,7 @@ The [Task lifecycle verification GAP report](./task-lifecycle-gap-report.md) is - Shared status vocabulary: [#612](https://github.com/Zoo-Code-Org/Zoo-Code/issues/612). - Completion visibility history: [#1453](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1453) and [#1279](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1279). - Cross-instance history preservation: [#920](https://github.com/Zoo-Code-Org/Zoo-Code/issues/920). +- Rejected-delegation pending-action settlement: [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714). - Parser request scoping: [#1468](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1468). ## Extending the model diff --git a/docs/architecture/task-lifecycle-remediation-blocks.md b/docs/architecture/task-lifecycle-remediation-blocks.md index ffbcbb334b..110fc7a054 100644 --- a/docs/architecture/task-lifecycle-remediation-blocks.md +++ b/docs/architecture/task-lifecycle-remediation-blocks.md @@ -13,7 +13,7 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## Ownership rules -- Every `LIFE-GAP-001` through `LIFE-GAP-038` has exactly one primary block below. +- Every `LIFE-GAP-001` through `LIFE-GAP-041` has exactly one primary block below. - A block owns exactly one GAP ID. Dependencies may reference other blocks but do not duplicate ownership. - Block IDs are stable: `LIFE-BLK-P-`. - Baseline blocks describe current serial production behavior. Optional fan-out is isolated under `FANOUT-BLK-*` and does not own a baseline `LIFE-GAP`. @@ -21,54 +21,57 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## P1: Persisted ownership and generation -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | --------------------------- | -------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P1-001 | 001 | Encode authoritative awaited-child revalidation as a model boundary and retain the shortest stale-completion witness. | `TaskHistoryStore.atomicUpdatePair`, `ClineProvider.reopenParentFromDelegation`; shared-store checker; cross-instance tests. | None | Model names lock-time ownership check, witness, bounds, and production test required for promotion. | -| LIFE-BLK-P1-002 | 002 | Specify lifecycle-owned lineage fields versus metadata writes and the stale-save witness. | `Task.saveClineMessages`, `taskMetadata`, `mergeHistoryDelta`; shared-store checker. | P1-001 ownership vocabulary | Field ownership table and monotonic-detachment invariant are explicit; no claim of current safety. | -| LIFE-BLK-P1-012 | 012 | Add attempt-generation state and stale-versus-resumed completion scenarios to the specification. | `PendingTaskAction.actionId`, interruption/resume/completion reducers; lifecycle checker exclusion. | P1-001 | Two generations and acceptance/rejection landmarks are specified with a bounded future checker shape. | -| LIFE-BLK-P1-017 | 017 | Inventory mutable cache read consumers and define immutable read semantics. | `TaskHistoryStore.get/getAll`; store tests. | None | Every direct caller is classified; clone/freeze test criteria and compatibility exclusions are recorded. | -| LIFE-BLK-P1-020 | 020 | Define observable stale-cache and convergence histories. | watcher, `invalidate`, `reconcile`; shared-store landmarks and cross-instance tests. | P1-001 | Missed-watch and explicit-refresh histories have bounded properties and objective convergence evidence. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P1-001 | 001 | Encode authoritative awaited-child revalidation as a model boundary and retain the shortest stale-completion witness. | `TaskHistoryStore.atomicUpdatePair`, `ClineProvider.reopenParentFromDelegation`; shared-store checker; cross-instance tests. | None | Model names lock-time ownership check, witness, bounds, and production test required for promotion. | +| LIFE-BLK-P1-002 | 002 | Specify lifecycle-owned lineage fields versus metadata writes and the stale-save witness. | `Task.saveClineMessages`, `taskMetadata`, `mergeHistoryDelta`; shared-store checker. | P1-001 ownership vocabulary | Field ownership table and monotonic-detachment invariant are explicit; no claim of current safety. | +| LIFE-BLK-P1-012 | 012 | Add attempt-generation state and stale-versus-resumed completion scenarios to the specification. | `PendingTaskAction.actionId`, interruption/resume/completion reducers; lifecycle checker exclusion. | P1-001 | Two generations and acceptance/rejection landmarks are specified with a bounded future checker shape. The contract covers staged pending-action replay attempts, bounded retry identity, and settlement of stale actions. | +| LIFE-BLK-P1-017 | 017 | Inventory mutable cache read consumers and define immutable read semantics. | `TaskHistoryStore.get/getAll`; store tests. | None | Every direct caller is classified; clone/freeze test criteria and compatibility exclusions are recorded. | +| LIFE-BLK-P1-020 | 020 | Define observable stale-cache and convergence histories. | watcher, `invalidate`, `reconcile`; shared-store landmarks and cross-instance tests. | P1-001 | Missed-watch and explicit-refresh histories have bounded properties and objective convergence evidence. Convergence states a pending-action rule for reconciled records with P2-040. | ## P2: Durable operation and crash recovery -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | -------------------------------------------------------------------------- | ----------------------------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------------------------- | -| LIFE-BLK-P2-004 | 004 | Enumerate pair-write interruption points and legal recovered states. | `atomicUpdatePair`; pair-failure landmark/tests. | P1-001 | Every pre/post-write cut has one legal outcome and required fault-injection assertion. | -| LIFE-BLK-P2-005 | 005 | Map delegation create/persist/publish/start cuts and rollback obligations. | `delegateParentAndOpenChild`; provider handoff model/tests. | P1-001, P2-004 | Transition table covers every cut without claiming child/parent atomicity. | -| LIFE-BLK-P2-006 | 006 | Specify completion message/lifecycle commit phases and replay outcomes. | `reopenParentFromDelegation`; completion and shared-store models. | P1-012, P2-004 | Result visibility and lifecycle state are mapped for each injected failure point. | -| LIFE-BLK-P2-021 | 021 | Define store close/drain semantics and post-dispose write exclusion. | `TaskHistoryStore.dispose`, write lock; store tests. | P2-004 recovery vocabulary | A bounded close-state machine and deterministic pending-write test criteria are documented. | -| LIFE-BLK-P2-023 | 023 | Specify deletion unlink failure and reconciliation histories. | `delete/deleteMany`, task directory/checkpoint cleanup; deletion tests. | P2-004 | False-success and resurrection outcomes are explicit with tombstone/retry closure choices. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P2-004 | 004 | Enumerate pair-write interruption points and legal recovered states. | `atomicUpdatePair`; pair-failure landmark/tests. | P1-001 | Every pre/post-write cut has one legal outcome and required fault-injection assertion. | +| LIFE-BLK-P2-005 | 005 | Map delegation create/persist/publish/start cuts and rollback obligations. | `delegateParentAndOpenChild`; provider handoff model/tests. | P1-001, P2-004 | Transition table covers every cut without claiming child/parent atomicity. | +| LIFE-BLK-P2-006 | 006 | Specify completion message/lifecycle commit phases and replay outcomes. | `reopenParentFromDelegation`; completion and shared-store models. | P1-012, P2-004 | Result visibility and lifecycle state are mapped for each injected failure point. The phase table includes child pending-action state at every cut and its interaction with resumed replay (012). | +| LIFE-BLK-P2-021 | 021 | Define store close/drain semantics and post-dispose write exclusion. | `TaskHistoryStore.dispose`, write lock; store tests. | P2-004 recovery vocabulary | A bounded close-state machine and deterministic pending-write test criteria are documented. | +| LIFE-BLK-P2-023 | 023 | Specify deletion unlink failure and reconciliation histories. | `delete/deleteMany`, task directory/checkpoint cleanup; deletion tests. | P2-004 | False-success and resurrection outcomes are explicit with tombstone/retry closure choices. | +| LIFE-BLK-P2-039 | 039 | Specify the crash window between authoritative rejection and settlement, plus bounded replay of a rejected action. | `ClineProvider.delegateParentAndOpenChild` settlement block, `settleRejectedCreateSubtaskAction`, `Task.resumePendingTaskAction`; lifecycle settlement witnesses; provider failure-injection tests. | P1-012, P2-004 | Every cut between rejection and settlement has one legal outcome; settlement-write failure is retried or surfaced without false success; replay of a rejected action is bounded and test-covered. | +| LIFE-BLK-P2-040 | 040 | Define pending-action preserve/settle/drop rules for startup repair and reconciliation outcomes. | `reconcileDelegationState`, `repairActiveDelegation`, `applyDelegationRepairIntent`, `DelegationRepairIntent`; reconciliation tests; fresh-host restart E2E. | P2-004 | Every repair outcome has an explicit pending-action rule; the #1714 loop precondition cannot survive repair; both interrupt paths persist or settle the action identically. | ## P3: Schema, path, and vocabulary -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ---------- | ---------------------------------------------------------------------------------------- | -| LIFE-BLK-P3-013 | 013 | Publish the canonical persisted-status owner and copied-union inventory. | `historyItemSchema`, task metadata, Task, CLI/history reader; typecheck/tests. | None | Every copy is listed with replacement/static-ratchet criteria. | -| LIFE-BLK-P3-018 | 018 | Define normal-read validation and quarantine outcomes for malformed history. | `readTaskFile`, reconciliation, shared Zod schema; fixtures. | None | Missing/invalid/legacy records have distinct expected outcomes and test fixtures. | -| LIFE-BLK-P3-019 | 019 | Inventory every task-ID-to-path entry and one shared safe-ID contract. | store paths, imports, deletion, checkpoints; traversal tests. | P3-018 | All path constructors are mapped and separator/traversal acceptance tests are specified. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P3-013 | 013 | Publish the canonical persisted-status owner and copied-union inventory. | `historyItemSchema`, task metadata, Task, CLI/history reader; typecheck/tests. | None | Every copy is listed with replacement/static-ratchet criteria. | +| LIFE-BLK-P3-018 | 018 | Define normal-read validation and quarantine outcomes for malformed history. | `readTaskFile`, reconciliation, shared Zod schema; fixtures. | None | Missing/invalid/legacy records have distinct expected outcomes and test fixtures. | +| LIFE-BLK-P3-019 | 019 | Inventory every task-ID-to-path entry and one shared safe-ID contract. | store paths, imports, deletion, checkpoints; traversal tests. | P3-018 | All path constructors are mapped and separator/traversal acceptance tests are specified. | +| LIFE-BLK-P3-041 | 041 | Specify cycle-safe, duplicate-safe traversal for cascade deletion over unvalidated persisted graphs. | `ClineProvider.deleteTaskWithId` `collectChildIds`, `TaskHistoryStore.deleteMany`; package-local traversal fixtures with cyclic, self-referencing, and duplicate graphs. | P3-018 | Traversal terminates on malformed graphs; deletion removes the acyclic closure exactly once or fails closed with a surfaced error; fixture tests prove both outcomes. | ## P4: Request, stream, and tool identity -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -------------------------- | ------------------------------------------------------------------------------------------- | -| LIFE-BLK-P4-008 | 008 | Add transform-to-parser cases for argument-only deltas and absent indices. | Responses transform, parser APIs/tests. | None | Two-call bounded schedules and expected isolated reconstruction are specified. | -| LIFE-BLK-P4-010 | 010 | Define request-generation ownership for detached usage writes. | Task request/drain paths; delayed-stream tests. | P4-038 identity vocabulary | Old/new generation mutations and allowed accounting-only updates are explicit. | -| LIFE-BLK-P4-024 | 024 | Map parser cleanup on success, abort, provider error, and replacement. | parser scope plus Task request terminal paths. | P4-010 | Every terminal path owns cleanup; late-event exclusions are stated. | -| LIFE-BLK-P4-025 | 025 | Specify listener lifetime for one chunk race and long streams. | `nextChunkWithAbort`; listener-count tests. | None | Both race outcomes remove listeners and a bounded stream cannot accumulate them. | -| LIFE-BLK-P4-026 | 026 | Model a true wall-clock deadline around pending iterator reads. | detached usage drain; fake-timer tests. | P4-010 | Permanently pending `next()` has a terminal deadline transition and no stale mutations. | -| LIFE-BLK-P4-030 | 030 | Define duplicate-start/run-promise identity. | `Task.start/run`, scheduler callback; Task tests. | P4-010 | Repeated starts share the actual settlement and cannot bypass scheduler ownership. | -| LIFE-BLK-P4-037 | 037 | Specify call-scoped partial path state and two-call interleavings. | `BaseTool.lastSeenPartialPath`, editing tool singletons; focused tests. | P4-038, P4-010 | Equal/different path interleavings and sibling-safe cleanup are bounded and reachable. | -| LIFE-BLK-P4-038 | 038 | Define canonical raw-to-durable call identity and collision witnesses. | tool-ID utility, parser, Task history, results, pending actions; duplicate-ID tests. | None | Adversarial IDs preserve or explicitly reject one-to-one call/result/replay correspondence. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P4-008 | 008 | Add transform-to-parser cases for argument-only deltas and absent indices. | Responses transform, parser APIs/tests. | None | Two-call bounded schedules and expected isolated reconstruction are specified. | +| LIFE-BLK-P4-010 | 010 | Define request-generation ownership for detached usage writes. | Task request/drain paths; delayed-stream tests. | P4-038 identity vocabulary | Old/new generation mutations and allowed accounting-only updates are explicit. | +| LIFE-BLK-P4-024 | 024 | Map parser cleanup on success, abort, provider error, and replacement. | parser scope plus Task request terminal paths. | P4-010 | Every terminal path owns cleanup; late-event exclusions are stated. | +| LIFE-BLK-P4-025 | 025 | Specify listener lifetime for one chunk race and long streams. | `nextChunkWithAbort`; listener-count tests. | None | Both race outcomes remove listeners and a bounded stream cannot accumulate them. | +| LIFE-BLK-P4-026 | 026 | Model a true wall-clock deadline around pending iterator reads. | detached usage drain; fake-timer tests. | P4-010 | Permanently pending `next()` has a terminal deadline transition and no stale mutations. | +| LIFE-BLK-P4-030 | 030 | Define duplicate-start/run-promise identity. | `Task.start/run`, scheduler callback; Task tests. | P4-010 | Repeated starts share the actual settlement and cannot bypass scheduler ownership. | +| LIFE-BLK-P4-037 | 037 | Specify call-scoped partial path state and two-call interleavings. | `BaseTool.lastSeenPartialPath`, editing tool singletons; focused tests. | P4-038, P4-010 | Equal/different path interleavings and sibling-safe cleanup are bounded and reachable. | +| LIFE-BLK-P4-038 | 038 | Define canonical raw-to-durable call identity and collision witnesses. | tool-ID utility, parser, Task history, results, pending actions; duplicate-ID tests. | None | Adversarial IDs preserve or explicitly reject one-to-one call/result/replay correspondence. The bijection includes the staged pending action, whose `actionId` derives from sanitized call IDs, across restart between approval and settlement. | ## P5: Tool-owned task state and queueing -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | -| LIFE-BLK-P5-003 | 003 | Map every queue consumer to claim/persist/ack or dequeue-before-submit. | `MessageQueueService`, Task queue paths; failure tests. | None | Every consumer is classified and message-retention failure evidence is specified. | -| LIFE-BLK-P5-007 | 007 | Inventory remaining mode-sensitive readers and authoritative task/provider source after the merged three-reader fix. | handoff selector, named production readers `getEnvironmentDetails`, `validateToolUse` call sites in `presentAssistantMessage`, custom tool execution, merged environment/validation/custom-tool tests, delegated reader checker. | None | Confirmed readers are marked production-tested; unclassified readers remain listed; the pure checker is not described as executing downstream readers. | -| LIFE-BLK-P5-031 | 031 | Define intentional versus accidental queue loss across task disposal/restart. | queue service disposal and task lifecycle; E2E boundary. | P5-003 | Product contract, excluded durability, and restart witness are explicit. | -| LIFE-BLK-P5-035 | 035 | Specify durable child initialization precedence using initial todos as witness. | `NewTaskTool`, Task constructor, history/messages, rehydration, UI state. | P2-005 | Omitted, explicit-empty, initial, updated, switched, and restarted cases are mapped. | -| LIFE-BLK-P5-036 | 036 | Model two approval identities and stale/cross-task todo edits. | `approvedTodoList`, webview handler, approval callbacks/tests. | P4-038 | Two-task schedules require task/action/call correlation; current unsafe witness is explicit. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P5-003 | 003 | Map every queue consumer to claim/persist/ack or dequeue-before-submit. | `MessageQueueService`, Task queue paths; failure tests. | None | Every consumer is classified and message-retention failure evidence is specified. | +| LIFE-BLK-P5-007 | 007 | Inventory remaining mode-sensitive readers and authoritative task/provider source after the merged three-reader fix. | handoff selector, named production readers `getEnvironmentDetails`, `validateToolUse` call sites in `presentAssistantMessage`, custom tool execution, merged environment/validation/custom-tool tests, delegated reader checker. | None | Confirmed readers are marked production-tested; unclassified readers remain listed; the pure checker is not described as executing downstream readers. | +| LIFE-BLK-P5-031 | 031 | Define intentional versus accidental queue loss across task disposal/restart. | queue service disposal and task lifecycle; E2E boundary. | P5-003 | Product contract, excluded durability, and restart witness are explicit. The contract states the restart composition between queue loss and pending-action replay ordering. | +| LIFE-BLK-P5-035 | 035 | Specify durable child initialization precedence using initial todos as witness. | `NewTaskTool`, Task constructor, history/messages, rehydration, UI state. | P2-005 | Omitted, explicit-empty, initial, updated, switched, and restarted cases are mapped. | +| LIFE-BLK-P5-036 | 036 | Model two approval identities and stale/cross-task todo edits. | `approvedTodoList`, webview handler, approval callbacks/tests. | P4-038 | Two-task schedules require task/action/call correlation; current unsafe witness is explicit. Correlation binds the staged `pendingAction.actionId` to task and call; denied and cancelled approvals have durable settlement rules. | ## P6: Event and ingress contracts @@ -108,7 +111,7 @@ Optional future fan-out blocks do not own `LIFE-GAP-014` and do not participate ## Mechanical coverage check -The primary tables above map the closed integer range `001..038` exactly once. Reviewers should verify this mechanically before changing the register: +The primary tables above map the closed integer range `001..041` exactly once. Reviewers should verify this mechanically before changing the register: ```sh rg -o '^\| LIFE-BLK-P[0-9]-[0-9]{3} \|' docs/architecture/task-lifecycle-remediation-blocks.md \ @@ -121,9 +124,27 @@ The command must print nothing. It matches only primary table rows, so dependenc Separately compare block suffixes with the GAP column to detect omissions or mismatches: ```sh -node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:38},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==38||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' +node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:41},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==41||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' ``` +## Issue-update plan for #1688 and child issues + +The next documentation subtask applies this mapping to [#1688](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1688) and its children. GitHub issue numbers are tracking links only; `LIFE-GAP-*` IDs own the burn-down. + +| Finding | GAP IDs | Issue to update | Update | +| -------------------------------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------- | +| Inventory-versus-composed scope statement | register-wide | #1688 | Replace the 001..038 range with 001..041 in the umbrella body and link the scope statement. | +| #1714 settlement model evidence | context for 039 | #1688, #1690 | Record the settlement reducer, witnesses, and focused tests as current evidence. | +| Crash-safe rejection settlement | 039 | #1690 | Add LIFE-GAP-039 to the P2 child checklist. | +| Startup repair pending-action reconciliation | 040 | #1690 | Add LIFE-GAP-040 to the P2 child checklist. | +| Cycle-safe cascade deletion traversal | 041 | #1691 | Add LIFE-GAP-041 to the P3 child checklist. | +| Strengthened completion commit obligations | 006 | #1690 | Append the composed phase-table and pending-action obligations to the child's 006 item. | +| Strengthened generation obligations | 012 | #1689 | Append pending-action replay and attempt identity to the child's 012 item. | +| Strengthened convergence obligations | 020 | #1689 | Append the pending-action convergence rule to the child's 020 item. | +| Strengthened queue durability contract | 031 | #1693 | Append the restart-composition obligation to the child's 031 item. | +| Strengthened approval correlation | 036 | #1693 | Append staged-action correlation and durable denial settlement to the child's 036 item. | +| Strengthened call identity | 038 | #1692 | Append the staged-action bijection obligation to the child's 038 item. | + ## Block completion template - [ ] Stable block and parent GAP IDs are in the PR description. diff --git a/scripts/check-task-lifecycle.ts b/scripts/check-task-lifecycle.ts index 73e9078366..42546c12d4 100644 --- a/scripts/check-task-lifecycle.ts +++ b/scripts/check-task-lifecycle.ts @@ -1,12 +1,13 @@ import assert from "node:assert/strict" -import type { HistoryItem } from "../packages/types/src/history" +import type { HistoryItem, PendingTaskAction } from "../packages/types/src/history" import { abandonDelegatedChild, completeDelegatedChild, delegateTaskToChild, interruptDelegatedChild, + settleRejectedCreateSubtaskAction, } from "../src/core/task-persistence/taskLifecycle" const taskIds = ["parent", "child-a", "child-b"] as const @@ -16,6 +17,8 @@ type ModelState = Record interface Transition { name: string next: ModelState + delegation?: { parentId: TaskId } + settlement?: { taskId: TaskId; actionId: string } } interface TraceStep { @@ -23,9 +26,16 @@ interface TraceStep { state: ModelState } +interface WitnessContext { + prev: ModelState + next: ModelState + transition: Transition +} + const MAX_DEPTH = 12 const MAX_STATES = 10_000 -const expectedActions = ["delegate", "interrupt", "complete", "abandon"] as const +const actionIds = ["action-1", "action-2"] as const +const expectedActions = ["delegate", "interrupt", "complete", "abandon", "stage", "settle-rejected"] as const const semanticLandmarks = { "interrupted-child-redelegation": (state: ModelState) => state.parent?.status === "delegated" && @@ -37,6 +47,27 @@ const semanticLandmarks = { state["child-a"]?.status === "delegated" && state["child-a"].awaitingChildId === "child-b", } satisfies Record boolean> +const semanticWitnesses = { + "interrupted-pending-delegation-settled": ({ prev, next, transition }: WitnessContext) => + transition.settlement !== undefined && + prev[transition.settlement.taskId]?.status === "interrupted" && + prev[transition.settlement.taskId]?.pendingAction?.actionId === transition.settlement.actionId && + next[transition.settlement.taskId]?.pendingAction === undefined, + "successful-active-delegation-settlement": ({ prev, next, transition }: WitnessContext) => + transition.delegation !== undefined && + prev[transition.delegation.parentId]?.status === "active" && + prev[transition.delegation.parentId]?.pendingAction?.kind === "create_subtask" && + next[transition.delegation.parentId]?.pendingAction === undefined, + "stale-settlement-rejected": ({ prev, next, transition }: WitnessContext) => { + if (transition.settlement === undefined) return false + const beforeAction = prev[transition.settlement.taskId]?.pendingAction + return ( + beforeAction?.kind === "create_subtask" && + beforeAction.actionId !== transition.settlement.actionId && + next[transition.settlement.taskId]?.pendingAction?.actionId === beforeAction.actionId + ) + }, +} satisfies Record boolean> function task(id: TaskId, parentTaskId?: TaskId): HistoryItem { return { @@ -54,6 +85,17 @@ function task(id: TaskId, parentTaskId?: TaskId): HistoryItem { } } +function createSubtaskAction(actionId: (typeof actionIds)[number]): PendingTaskAction { + return { + kind: "create_subtask", + actionId, + approvalText: "{}", + mode: "code", + message: `message for ${actionId}`, + todos: [], + } +} + function initialState(): ModelState { return { parent: task("parent"), "child-a": undefined, "child-b": undefined } } @@ -70,18 +112,38 @@ function transitions(state: ModelState): Transition[] { const parent = state[parentId] if (!parent) continue + const awaitedStatus = parent.awaitingChildId ? state[parent.awaitingChildId as TaskId]?.status : undefined + const delegationValid = + parent.status === "active" || (parent.status === "delegated" && awaitedStatus === "interrupted") + for (const childId of taskIds) { if (childId === parentId || state[childId]) continue - const awaitedStatus = parent.awaitingChildId ? state[parent.awaitingChildId as TaskId]?.status : undefined - if (parent.status !== "active" && !(parent.status === "delegated" && awaitedStatus === "interrupted")) { - continue - } - const delegated = delegateTaskToChild(parent, childId, awaitedStatus) + if (!delegationValid) continue + const delegated = { ...delegateTaskToChild(parent, childId, awaitedStatus), pendingAction: undefined } result.push({ name: `delegate(${parentId}, ${childId})`, next: replace(state, delegated, task(childId, parentId)), + delegation: { parentId }, }) } + + for (const actionId of actionIds) { + if (parent.status === "active" || parent.status === "interrupted") { + result.push({ + name: `stage(${parentId}, ${actionId})`, + next: replace(state, { ...parent, pendingAction: createSubtaskAction(actionId) }), + }) + } + + const pending = parent.pendingAction + if (!delegationValid && parent.status !== "completed" && pending?.kind === "create_subtask") { + result.push({ + name: `settle-rejected(${parentId}, ${actionId})`, + next: replace(state, settleRejectedCreateSubtaskAction(parent, actionId)), + settlement: { taskId: parentId, actionId }, + }) + } + } } for (const childId of taskIds) { @@ -103,7 +165,7 @@ function transitions(state: ModelState): Transition[] { const completed = completeDelegatedChild(parent, child, `${childId} result`) result.push({ name: `complete(${childId})`, - next: replace(state, completed.parent, completed.child), + next: replace(state, completed.parent, { ...completed.child, pendingAction: undefined }), }) } @@ -189,10 +251,36 @@ function checkTransitionInvariants(previous: ModelState, transition: Transition) violations.push(`${id}: completed task changed after ${transition.name}`) } } + + const settlement = transition.settlement + if (settlement) { + const before = previous[settlement.taskId] + const after = transition.next[settlement.taskId] + const beforeAction = before?.pendingAction + const afterAction = after?.pendingAction + + if (afterAction && canonicalTask(afterAction) !== canonicalTask(beforeAction)) { + violations.push(`${settlement.taskId}: settlement after ${transition.name} modified a replacement action`) + } + if ( + !afterAction && + beforeAction && + !(beforeAction.kind === "create_subtask" && beforeAction.actionId === settlement.actionId) + ) { + violations.push(`${settlement.taskId}: settlement after ${transition.name} cleared a non-matching action`) + } + const beforeRest = { ...before, pendingAction: undefined } + const afterRest = { ...after, pendingAction: undefined } + if (canonicalTask(beforeRest) !== canonicalTask(afterRest)) { + violations.push( + `${settlement.taskId}: settlement after ${transition.name} changed status, lineage, or accounting`, + ) + } + } return violations } -function canonicalTask(value: HistoryItem | undefined): string { +function canonicalTask(value: unknown): string { return JSON.stringify(value ?? null) } @@ -204,6 +292,7 @@ function runModelCheck(): number { const visited = new Set([canonical(start)]) const reachedActions = new Set() const reachedLandmarks = new Set() + const reachedWitnesses = new Set() const frontier: ModelState[] = [] for (let index = 0; index < queue.length; index++) { @@ -220,6 +309,9 @@ function runModelCheck(): number { for (const transition of transitions(node.state)) { reachedActions.add(transition.name.slice(0, transition.name.indexOf("("))) + for (const [name, matches] of Object.entries(semanticWitnesses)) { + if (matches({ prev: node.state, next: transition.next, transition })) reachedWitnesses.add(name) + } const transitionViolations = checkTransitionInvariants(node.state, transition) const trace = [...node.trace, { action: transition.name, state: transition.next }] if (transitionViolations.length) { @@ -244,6 +336,10 @@ function runModelCheck(): number { if (missingLandmarks.length) { throw new Error(`Task lifecycle model has unreachable semantic landmarks: ${missingLandmarks.join(", ")}`) } + const missingWitnesses = Object.keys(semanticWitnesses).filter((name) => !reachedWitnesses.has(name)) + if (missingWitnesses.length) { + throw new Error(`Task lifecycle model has unreachable semantic witnesses: ${missingWitnesses.join(", ")}`) + } const unexploredSuccessor = frontier .flatMap((state) => transitions(state)) .find((transition) => !visited.has(canonical(transition.next))) @@ -278,10 +374,29 @@ function runRepresentativeScenarios(): void { const interruptedCompletion = completeDelegatedChild(delegated, interruptedA, "resumed result") assert.equal(interruptedCompletion.child.status, "completed") assert.equal(interruptedCompletion.parent.status, "active") + + const rejectedParent: HistoryItem = { + ...parent, + status: "interrupted", + pendingAction: createSubtaskAction("action-1"), + } + assert.throws(() => delegateTaskToChild(rejectedParent, childA.id), /Invalid task status transition/) + const settled = settleRejectedCreateSubtaskAction(rejectedParent, "action-1") + assert.equal(settled.status, "interrupted") + assert.equal(settled.pendingAction, undefined) + assert.equal(settled.childIds, rejectedParent.childIds) + + const replacement = { ...rejectedParent, pendingAction: createSubtaskAction("action-2") } + assert.equal(settleRejectedCreateSubtaskAction(replacement, "action-1"), replacement) + assert.equal( + settleRejectedCreateSubtaskAction({ ...rejectedParent, status: "completed" }, "action-1").pendingAction + ?.actionId, + "action-1", + ) } runRepresentativeScenarios() const checkedStates = runModelCheck() console.log( - `Task lifecycle model check passed: ${checkedStates} reachable states, ${expectedActions.length}/${expectedActions.length} actions reachable, ${Object.keys(semanticLandmarks).length}/${Object.keys(semanticLandmarks).length} landmarks reached, depth <= ${MAX_DEPTH}, ${taskIds.length} task slots`, + `Task lifecycle model check passed: ${checkedStates} reachable states, ${expectedActions.length}/${expectedActions.length} actions reachable, ${Object.keys(semanticLandmarks).length}/${Object.keys(semanticLandmarks).length} landmarks reached, ${Object.keys(semanticWitnesses).length}/${Object.keys(semanticWitnesses).length} settlement witnesses reached, depth <= ${MAX_DEPTH}, ${taskIds.length} task slots`, ) diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index 422c264e2c..68f8e2b1fc 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -753,4 +753,127 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { expect(deleteTaskWithId).toHaveBeenCalledWith("child-1", false) expect(createTaskWithHistoryItem).toHaveBeenCalledWith(parentHistoryItem) }) + + it("settles an interrupted parent's pending action when delegation is rejected", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + let current: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const taskHistoryStore = { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => current), + atomicReadAndUpdate: vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + current = updater(current) + return [current] + }), + } + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: current }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId: vi.fn().mockResolvedValue(undefined), + createTaskWithHistoryItem: vi.fn().mockResolvedValue(undefined), + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore, + } as unknown as ClineProvider + + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow("Invalid task status transition: interrupted → delegated") + + expect(current.pendingAction).toBeUndefined() + expect(current.status).toBe("interrupted") + expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + }) + + it("does not restore a rejected pending action when its settlement write fails", async () => { + const settlementError = new Error("pending-action settlement failed") + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + const interruptedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const atomicReadAndUpdate = vi + .fn() + .mockImplementationOnce(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + updater(interruptedParent) + return [] + }) + .mockRejectedValueOnce(settlementError) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: interruptedParent }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId: vi.fn().mockResolvedValue(undefined), + createTaskWithHistoryItem, + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => interruptedParent), + atomicReadAndUpdate, + }, + } as unknown as ClineProvider + + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow("Invalid task status transition: interrupted → delegated") + + expect(atomicReadAndUpdate).toHaveBeenCalledTimes(2) + expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(createTaskWithHistoryItem).not.toHaveBeenCalled() + }) }) diff --git a/src/core/task-persistence/__tests__/taskLifecycle.spec.ts b/src/core/task-persistence/__tests__/taskLifecycle.spec.ts index fe415f09f8..5d3129e023 100644 --- a/src/core/task-persistence/__tests__/taskLifecycle.spec.ts +++ b/src/core/task-persistence/__tests__/taskLifecycle.spec.ts @@ -5,6 +5,8 @@ import { completeDelegatedChild, delegateTaskToChild, interruptDelegatedChild, + LifecycleTransitionError, + settleRejectedCreateSubtaskAction, } from "../taskLifecycle" function item(id: string, overrides: Partial = {}): HistoryItem { @@ -102,3 +104,84 @@ describe("task lifecycle transitions", () => { expect(abandoned.child).toMatchObject({ parentTaskId: undefined, rootTaskId: undefined }) }) }) + +describe("settleRejectedCreateSubtaskAction", () => { + const createSubtaskAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + + it("clears only the matching pending create_subtask action", () => { + const parent = item("parent", { + status: "interrupted", + parentTaskId: "root", + rootTaskId: "root", + awaitingChildId: undefined, + tokensIn: 12, + totalCost: 0.5, + pendingAction: createSubtaskAction, + }) + + const settled = settleRejectedCreateSubtaskAction(parent, "create-action") + + expect(settled).toEqual({ + ...parent, + pendingAction: undefined, + }) + expect(settled).toMatchObject({ + status: "interrupted", + parentTaskId: "root", + rootTaskId: "root", + tokensIn: 12, + totalCost: 0.5, + }) + }) + + it("never clears a replacement action with a different ID", () => { + const parent = item("parent", { + status: "interrupted", + pendingAction: { ...createSubtaskAction, actionId: "replacement-action" }, + }) + + expect(settleRejectedCreateSubtaskAction(parent, "stale-action")).toBe(parent) + }) + + it("never clears a pending action of a different kind", () => { + const parent = item("parent", { + pendingAction: { + kind: "finish_subtask", + actionId: "create-action", + approvalText: "{}", + parentTaskId: "root", + result: "done", + }, + }) + + expect(settleRejectedCreateSubtaskAction(parent, "create-action")).toBe(parent) + }) + + it("leaves a record without a pending action unchanged", () => { + const parent = item("parent", { status: "interrupted" }) + + expect(settleRejectedCreateSubtaskAction(parent, "create-action")).toBe(parent) + }) + + it("never mutates a completed record", () => { + const parent = item("parent", { status: "completed", pendingAction: createSubtaskAction }) + + expect(settleRejectedCreateSubtaskAction(parent, "create-action")).toBe(parent) + }) + + it("rejects an interrupted parent's delegation with a typed transition error", () => { + const parent = item("parent", { status: "interrupted", pendingAction: createSubtaskAction }) + + expect(() => delegateTaskToChild(parent, "child")).toThrow(LifecycleTransitionError) + expect(() => delegateTaskToChild(parent, "child")).toThrow( + "Invalid task status transition: interrupted → delegated", + ) + }) +}) diff --git a/src/core/task-persistence/index.ts b/src/core/task-persistence/index.ts index 14adeedc68..0baf9ce57e 100644 --- a/src/core/task-persistence/index.ts +++ b/src/core/task-persistence/index.ts @@ -21,6 +21,7 @@ export { delegateTaskToChild, interruptDelegatedChild, LifecycleTransitionError, + settleRejectedCreateSubtaskAction, type HistoryItemStatus, VALID_TASK_STATUS_TRANSITIONS, } from "./taskLifecycle" diff --git a/src/core/task-persistence/taskLifecycle.ts b/src/core/task-persistence/taskLifecycle.ts index efd2e1148f..b8a5e571f0 100644 --- a/src/core/task-persistence/taskLifecycle.ts +++ b/src/core/task-persistence/taskLifecycle.ts @@ -20,10 +20,27 @@ export class LifecycleTransitionError extends Error { export function assertValidTransition(from: HistoryItemStatus | undefined, to: HistoryItemStatus): void { const fromStatus: HistoryItemStatus = from ?? "active" if (!VALID_TASK_STATUS_TRANSITIONS[fromStatus].includes(to)) { - throw new Error(`Invalid task status transition: ${fromStatus} → ${to}`) + throw new LifecycleTransitionError(`Invalid task status transition: ${fromStatus} → ${to}`) } } +/** + * Settles the pending create_subtask action whose delegation the authoritative + * parent record rejected (#1714). Only the exact matching action ID is cleared; + * status, lineage, accounting, and unrelated fields are preserved. A replaced + * or different-kind pending action is never cleared. + */ +export function settleRejectedCreateSubtaskAction(parent: HistoryItem, pendingActionId: string): HistoryItem { + const pending = parent.pendingAction + if (parent.status === "completed") { + return parent + } + if (pending?.kind !== "create_subtask" || pending.actionId !== pendingActionId) { + return parent + } + return { ...parent, pendingAction: undefined } +} + export function delegateTaskToChild( parent: HistoryItem, childId: string, diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 86ce5d8e67..ba333acfe4 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -126,6 +126,8 @@ import { completeDelegatedChild, delegateTaskToChild, interruptDelegatedChild, + LifecycleTransitionError, + settleRejectedCreateSubtaskAction, } from "../task-persistence" import { readTaskMessages } from "../task-persistence/taskMessages" import { getNonce } from "./getNonce" @@ -4044,6 +4046,25 @@ export class ClineProvider (err as Error)?.message ?? String(err) }`, ) + // The authoritative parent record rejected this delegation (#1714). + // Settle the matching pending create_subtask action durably so a retry + // cannot replay a rejected action, then propagate the original error. + let settlementFailed = false + if (pendingActionId && err instanceof LifecycleTransitionError) { + try { + await this.taskHistoryStore.atomicReadAndUpdate(parentTaskId, (historyItem) => + settleRejectedCreateSubtaskAction(historyItem, pendingActionId), + ) + this.recentTasksCache = undefined + } catch (settlementError) { + settlementFailed = true + this.log( + `[delegateParentAndOpenChild] Failed to settle pending action ${pendingActionId} for parent ${parentTaskId}: ${ + (settlementError as Error)?.message ?? String(settlementError) + }`, + ) + } + } try { // Only pop the stack if the child we just created is still on top. // A concurrent delegation could have pushed another child since we created ours. @@ -4067,8 +4088,13 @@ export class ClineProvider ) } try { - const { historyItem: parentHistory } = await this.getTaskWithId(parentTaskId) - await this.createTaskWithHistoryItem(parentHistory) + // A failed settlement write leaves the rejected pending action in + // durable storage. Restoring the stored parent would replay the + // rejected action, so leave the parent unrestored instead. + if (!settlementFailed) { + const { historyItem: parentHistory } = await this.getTaskWithId(parentTaskId) + await this.createTaskWithHistoryItem(parentHistory) + } } catch (rollbackError) { this.log( `[delegateParentAndOpenChild] Failed to restore parent ${parentTaskId} during rollback: ${ From a7f94b4c8ff5d55b8ad6c5c91e06ab437c33ef03 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Sun, 20 Sep 2026 23:39:32 +0000 Subject: [PATCH 02/23] fix(lifecycle): preserve pending actions in model checks --- .../architecture/task-lifecycle-gap-report.md | 6 +- docs/architecture/task-lifecycle-model.md | 2 +- .../task-lifecycle-remediation-blocks.md | 35 +++------- scripts/check-task-lifecycle.ts | 13 +++- .../ClineProvider.delegation.spec.ts | 67 ++++++++++++++++++- 5 files changed, 89 insertions(+), 34 deletions(-) diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index 800389f40b..266a3ebf00 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -8,7 +8,7 @@ The audit covers tracked TypeScript, JSON, YAML, and Markdown under `packages/`, Inventory completeness is not composed verification. A closed symbol and ownership inventory lists every boundary with its local evidence. It does not prove that persistence, restart, replay, UI approval, scheduling, and external effects compose into one correct protocol. Local evidence classes support an end-to-end claim only through the joint-checker or trace-validation work tracked by `LIFE-GAP-015`. -Issue [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714) is recorded evidence for this distinction. Every involved symbol (`setPendingTaskAction`, `resumeTaskFromHistory`, `resumePendingTaskAction`, `delegateParentAndOpenChild`, `delegateTaskToChild`) was inventoried with mapped ownership, yet the composed loop across restart, pending-action replay, auto-approval, and an authoritative transition rejection was found by incident report, not by inventory. The #1714 changes add the `settleRejectedCreateSubtaskAction` reducer, `stage` and `settle-rejected` model actions, three settlement witnesses, and matching focused tests. Their direct model coverage is specified in [the model suite](./task-lifecycle-model.md). They address the reported loop for the authoritative-rejection path; LIFE-GAP-039 and LIFE-GAP-040 record the remaining crash-window, startup-repair, and replay-bound obligations. +Issue [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714) is recorded evidence for this distinction. Every involved symbol (`setPendingTaskAction`, `resumeTaskFromHistory`, `resumePendingTaskAction`, `delegateParentAndOpenChild`, `delegateTaskToChild`) was inventoried with mapped ownership, yet the composed loop across restart, pending-action replay, auto-approval, and an authoritative transition rejection was found by incident report, not by inventory. The #1714 changes add the `settleRejectedCreateSubtaskAction` reducer, `stage` and `settle-rejected` model actions, four pending-action witnesses, and matching focused tests. Their direct model coverage is specified in [the model suite](./task-lifecycle-model.md). They address the reported loop for the authoritative-rejection path; LIFE-GAP-039 and LIFE-GAP-040 record the remaining crash-window, startup-repair, and replay-bound obligations. “Exhaustive” means exhaustive over the repository paths, symbol families, and search terms listed here at the audited commit. It does not include ignored/generated output, deployment branch-protection settings, runtime telemetry, dynamically constructed names that evade text search, or behavior in dependencies. GitHub issue links are historical provenance only; stable `LIFE-GAP-*` IDs own the active burn-down. @@ -216,13 +216,12 @@ Severity reflects plausible data loss, ownership corruption, permission/context | LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. The bijection must include the staged pending action: `pendingAction.actionId` derives from `sanitizeToolUseId(toolCallId)`, so non-injective sanitization can make a settlement targeted at one raw call clear another call's staged action; the model's stale-settlement witness depends on collision-free action IDs. | | LIFE-GAP-039 | High | High | Settlement after an authoritative pending-action rejection is a second, best-effort store write. A crash or settlement-write failure between the rejected transition and the settlement write leaves the rejected action staged, and startup replay can re-execute it without an attempt bound. | Source: the `delegateParentAndOpenChild` settlement block logs and swallows settlement errors; `pendingAction` carries no attempt count; the #1714 bounded-retry fix suggestion is unimplemented. | Durable operation intent (P2-004); attempt/generation identity (012). | Fault injection between rejection and settlement converges to a documented legal state; settlement-write failure is retried or surfaced without false success; replay of a rejected action is bounded by a persisted attempt/backoff record; model or focused witnesses cover the crash window. | | LIFE-GAP-040 | Medium | Medium-high | Startup topology repair reconciles status and lineage but has no pending-action rule. `reconcileDelegationState`, `repairActiveDelegation`, and repair-intent replay can leave a staged action on a repaired record, and the two interrupt paths disagree on whether `pendingAction` survives. | #1714 reproduced the loop precondition through startup reconciliation of a persisted active child; `DelegationRepairIntent` guards and targets carry no pending-action field; `applyDelegationRepairIntent` spreads records without a pending-action decision. | Durable operation intent (P2-004); correlated action identity (036, 038). | Every repair outcome (orphaned delegation, orphaned active child, completed child, intent replay, quarantine) defines preserve, settle, or drop for pending actions; deterministic repair tests and one fresh-host restart test prove the #1714 precondition cannot survive repair; both interrupt paths persist or settle the action consistently. | -| LIFE-GAP-041 | Medium | Medium | Cascade deletion traverses persisted `childIds` with no cycle or duplicate guard. `deleteTaskWithId`'s `collectChildIds` recurses through unvalidated records, so a cyclic or duplicated persisted graph recurses without bound, while cost aggregation already carries a visited set. | Structural: `collectChildIds` has no visited set or depth bound; `aggregateTaskCostsRecursive` does; persisted reads are unvalidated (018). | Schema-validated history reads (018); shared safe-ID boundary (019). | Traversal terminates on cyclic, self-referencing, and duplicate persisted graphs with deterministic fixture tests; cascade deletion deletes the acyclic closure exactly once or fails closed with a surfaced error; the guard is shared with other persisted-graph traversals or listed with exclusions. | ## Portfolio remediation plan The [1-SP remediation block register](./task-lifecycle-remediation-blocks.md) decomposes this portfolio into small modeling/documentation increments. It assigns every GAP exactly one primary block, preserves dependencies across workstreams, and keeps optional fan-out separate from baseline ownership. -The 41 IDs are not 41 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. +The 40 IDs are not 40 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. | Cluster | Gap IDs | Root fix and likely ownership | Complexity | Engineering risk | Objective portfolio evidence | | ------------------------------------------ | -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | @@ -337,7 +336,6 @@ Four workstreams can proceed concurrently after foundation decisions: - [ ] For LIFE-GAP-038, use adversarial raw IDs to verify one-to-one durable call, approval, execution, result, pending-action, and replay identity. - [ ] For LIFE-GAP-039, inject failure between the authoritative rejection and the settlement write and prove bounded, convergent recovery without unbounded replay. - [ ] For LIFE-GAP-040, run startup repair and reconciliation against records that carry staged pending actions and prove the #1714 loop precondition cannot survive repair. -- [ ] For LIFE-GAP-041, delete through cyclic, self-referencing, and duplicate persisted `childIds` graphs and prove termination with exact-once deletion or a surfaced failure. ## Completeness statement diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index 22b7ded257..f8c412b334 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -54,7 +54,7 @@ TLA+/PlusCal or Quint with TLC becomes a better fit when the lifecycle needs tem The model has three fixed task slots, enough to cover competing siblings and a nested parent-child-grandchild chain. It explores every reachable interleaving through depth 12, deduplicating canonical states. Representative checks also exercise rejected operations that do not create a new state: a second concurrent delegation while the first child is active, stale completion after re-delegation, late completion after abandonment, completion after interruption, and nested completion. Named semantic landmarks require the graph to retain interrupted-child re-delegation and nested delegation even when the raw state total changes. -Each task slot can also hold one of two pending `create_subtask` actions. A `stage` action mirrors `setPendingTaskAction` overwrite semantics, delegation and completion clear the action their request carried, and a `settle-rejected` action models the settlement that follows an authoritative delegation rejection (#1714). Production settles through the typed `LifecycleTransitionError` from the shared guards: the provider writes the settled record with one extra `atomicReadAndUpdate` call, then propagates the original rejection. Three named witnesses must remain reachable: settlement from an interrupted record after rejection, settlement through a successful active delegation, and stale-action protection where a settlement targeting one action ID leaves a replacement action intact. A mismatched pending-action request keeps its production behavior: the atomic update throws before any transition, and no settlement runs. +Each task slot can also hold one of two pending `create_subtask` actions. A `stage` action mirrors `setPendingTaskAction` overwrite semantics, delegation clears the action its request carried, completion preserves unrelated actions, and a `settle-rejected` action models the settlement that follows an authoritative delegation rejection (#1714). Production settles through the typed `LifecycleTransitionError` from the shared guards: the provider writes the settled record with one extra `atomicReadAndUpdate` call, then propagates the original rejection. Four named witnesses must remain reachable: settlement from an interrupted record after rejection, settlement through a successful active delegation, unrelated-action preservation during completion, and stale-action protection where a settlement targeting one action ID leaves a replacement action intact. A mismatched pending-action request keeps its production behavior: the atomic update throws before any transition, and no settlement runs. Production completion also accepts a recovery-compatible `active` parent that still awaits the returning child, then clears the stale pointers. Normal model transitions never create that intermediate state, so it is covered by a focused reducer test rather than admitted as a generally valid reachable state. diff --git a/docs/architecture/task-lifecycle-remediation-blocks.md b/docs/architecture/task-lifecycle-remediation-blocks.md index 110fc7a054..d30cccc960 100644 --- a/docs/architecture/task-lifecycle-remediation-blocks.md +++ b/docs/architecture/task-lifecycle-remediation-blocks.md @@ -13,7 +13,7 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## Ownership rules -- Every `LIFE-GAP-001` through `LIFE-GAP-041` has exactly one primary block below. +- Every `LIFE-GAP-001` through `LIFE-GAP-040` has exactly one primary block below. - A block owns exactly one GAP ID. Dependencies may reference other blocks but do not duplicate ownership. - Block IDs are stable: `LIFE-BLK-P-`. - Baseline blocks describe current serial production behavior. Optional fan-out is isolated under `FANOUT-BLK-*` and does not own a baseline `LIFE-GAP`. @@ -43,12 +43,11 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## P3: Schema, path, and vocabulary -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | ---------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P3-013 | 013 | Publish the canonical persisted-status owner and copied-union inventory. | `historyItemSchema`, task metadata, Task, CLI/history reader; typecheck/tests. | None | Every copy is listed with replacement/static-ratchet criteria. | -| LIFE-BLK-P3-018 | 018 | Define normal-read validation and quarantine outcomes for malformed history. | `readTaskFile`, reconciliation, shared Zod schema; fixtures. | None | Missing/invalid/legacy records have distinct expected outcomes and test fixtures. | -| LIFE-BLK-P3-019 | 019 | Inventory every task-ID-to-path entry and one shared safe-ID contract. | store paths, imports, deletion, checkpoints; traversal tests. | P3-018 | All path constructors are mapped and separator/traversal acceptance tests are specified. | -| LIFE-BLK-P3-041 | 041 | Specify cycle-safe, duplicate-safe traversal for cascade deletion over unvalidated persisted graphs. | `ClineProvider.deleteTaskWithId` `collectChildIds`, `TaskHistoryStore.deleteMany`; package-local traversal fixtures with cyclic, self-referencing, and duplicate graphs. | P3-018 | Traversal terminates on malformed graphs; deletion removes the acyclic closure exactly once or fails closed with a surfaced error; fixture tests prove both outcomes. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------ | ---------- | ---------------------------------------------------------------------------------------- | +| LIFE-BLK-P3-013 | 013 | Publish the canonical persisted-status owner and copied-union inventory. | `historyItemSchema`, task metadata, Task, CLI/history reader; typecheck/tests. | None | Every copy is listed with replacement/static-ratchet criteria. | +| LIFE-BLK-P3-018 | 018 | Define normal-read validation and quarantine outcomes for malformed history. | `readTaskFile`, reconciliation, shared Zod schema; fixtures. | None | Missing/invalid/legacy records have distinct expected outcomes and test fixtures. | +| LIFE-BLK-P3-019 | 019 | Inventory every task-ID-to-path entry and one shared safe-ID contract. | store paths, imports, deletion, checkpoints; traversal tests. | P3-018 | All path constructors are mapped and separator/traversal acceptance tests are specified. | ## P4: Request, stream, and tool identity @@ -111,7 +110,7 @@ Optional future fan-out blocks do not own `LIFE-GAP-014` and do not participate ## Mechanical coverage check -The primary tables above map the closed integer range `001..041` exactly once. Reviewers should verify this mechanically before changing the register: +The primary tables above map the closed integer range `001..040` exactly once. Reviewers should verify this mechanically before changing the register: ```sh rg -o '^\| LIFE-BLK-P[0-9]-[0-9]{3} \|' docs/architecture/task-lifecycle-remediation-blocks.md \ @@ -124,27 +123,9 @@ The command must print nothing. It matches only primary table rows, so dependenc Separately compare block suffixes with the GAP column to detect omissions or mismatches: ```sh -node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:41},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==41||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' +node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:40},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==40||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' ``` -## Issue-update plan for #1688 and child issues - -The next documentation subtask applies this mapping to [#1688](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1688) and its children. GitHub issue numbers are tracking links only; `LIFE-GAP-*` IDs own the burn-down. - -| Finding | GAP IDs | Issue to update | Update | -| -------------------------------------------- | --------------- | --------------- | ------------------------------------------------------------------------------------------- | -| Inventory-versus-composed scope statement | register-wide | #1688 | Replace the 001..038 range with 001..041 in the umbrella body and link the scope statement. | -| #1714 settlement model evidence | context for 039 | #1688, #1690 | Record the settlement reducer, witnesses, and focused tests as current evidence. | -| Crash-safe rejection settlement | 039 | #1690 | Add LIFE-GAP-039 to the P2 child checklist. | -| Startup repair pending-action reconciliation | 040 | #1690 | Add LIFE-GAP-040 to the P2 child checklist. | -| Cycle-safe cascade deletion traversal | 041 | #1691 | Add LIFE-GAP-041 to the P3 child checklist. | -| Strengthened completion commit obligations | 006 | #1690 | Append the composed phase-table and pending-action obligations to the child's 006 item. | -| Strengthened generation obligations | 012 | #1689 | Append pending-action replay and attempt identity to the child's 012 item. | -| Strengthened convergence obligations | 020 | #1689 | Append the pending-action convergence rule to the child's 020 item. | -| Strengthened queue durability contract | 031 | #1693 | Append the restart-composition obligation to the child's 031 item. | -| Strengthened approval correlation | 036 | #1693 | Append staged-action correlation and durable denial settlement to the child's 036 item. | -| Strengthened call identity | 038 | #1692 | Append the staged-action bijection obligation to the child's 038 item. | - ## Block completion template - [ ] Stable block and parent GAP IDs are in the PR description. diff --git a/scripts/check-task-lifecycle.ts b/scripts/check-task-lifecycle.ts index 42546c12d4..8dfb1c202a 100644 --- a/scripts/check-task-lifecycle.ts +++ b/scripts/check-task-lifecycle.ts @@ -18,6 +18,7 @@ interface Transition { name: string next: ModelState delegation?: { parentId: TaskId } + completion?: { childId: TaskId } settlement?: { taskId: TaskId; actionId: string } } @@ -58,6 +59,15 @@ const semanticWitnesses = { prev[transition.delegation.parentId]?.status === "active" && prev[transition.delegation.parentId]?.pendingAction?.kind === "create_subtask" && next[transition.delegation.parentId]?.pendingAction === undefined, + "completion-preserves-unrelated-pending-action": ({ prev, next, transition }: WitnessContext) => { + if (transition.completion === undefined) return false + const beforeAction = prev[transition.completion.childId]?.pendingAction + return ( + beforeAction !== undefined && + next[transition.completion.childId]?.status === "completed" && + canonicalTask(next[transition.completion.childId]?.pendingAction) === canonicalTask(beforeAction) + ) + }, "stale-settlement-rejected": ({ prev, next, transition }: WitnessContext) => { if (transition.settlement === undefined) return false const beforeAction = prev[transition.settlement.taskId]?.pendingAction @@ -165,7 +175,8 @@ function transitions(state: ModelState): Transition[] { const completed = completeDelegatedChild(parent, child, `${childId} result`) result.push({ name: `complete(${childId})`, - next: replace(state, completed.parent, { ...completed.child, pendingAction: undefined }), + next: replace(state, completed.parent, completed.child), + completion: { childId }, }) } diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index 68f8e2b1fc..e0633964fa 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -789,7 +789,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { getCurrentTask, removeClineFromStack: vi.fn().mockResolvedValue(undefined), createTask, - getTaskWithId: vi.fn().mockResolvedValue({ historyItem: current }), + getTaskWithId: vi.fn().mockImplementation(async () => ({ historyItem: current })), handleModeSwitch: vi.fn().mockResolvedValue(undefined), deleteTaskWithId: vi.fn().mockResolvedValue(undefined), createTaskWithHistoryItem: vi.fn().mockResolvedValue(undefined), @@ -811,6 +811,71 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { expect(current.pendingAction).toBeUndefined() expect(current.status).toBe("interrupted") expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(provider.createTaskWithHistoryItem).toHaveBeenCalledWith( + expect.objectContaining({ + status: "interrupted", + pendingAction: undefined, + }), + ) + }) + + it("does not settle a pending action after an unrelated persistence failure", async () => { + const persistenceError = new Error("parent persistence failed") + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + const interruptedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const atomicReadAndUpdate = vi.fn().mockRejectedValue(persistenceError) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: interruptedParent }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId: vi.fn().mockResolvedValue(undefined), + createTaskWithHistoryItem, + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => interruptedParent), + atomicReadAndUpdate, + }, + } as unknown as ClineProvider + + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow(persistenceError) + + expect(atomicReadAndUpdate).toHaveBeenCalledTimes(1) + expect(interruptedParent.pendingAction).toBe(pendingAction) + expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(createTaskWithHistoryItem).toHaveBeenCalledWith(interruptedParent) }) it("does not restore a rejected pending action when its settlement write fails", async () => { From 72141a775ab239005b5eeb29aa5ec4ee62f6b15e Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 21 Sep 2026 00:40:11 +0000 Subject: [PATCH 03/23] fix(lifecycle): preserve replacement pending actions --- .../architecture/task-lifecycle-gap-report.md | 6 +- docs/architecture/task-lifecycle-model.md | 28 ++-- scripts/check-task-lifecycle.ts | 50 ++++++- .../ClineProvider.delegation.spec.ts | 93 +++++++++++- src/core/task-persistence/TaskHistoryStore.ts | 41 ++++++ .../TaskHistoryStore.realConcurrency.spec.ts | 139 ++++++++++++++++++ src/core/webview/ClineProvider.ts | 11 +- 7 files changed, 336 insertions(+), 32 deletions(-) diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index 266a3ebf00..9f8d34e5d3 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -227,7 +227,7 @@ The 40 IDs are not 40 independent projects. They group into eight programs with | ------------------------------------------ | -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | P1. Persisted ownership and generation | 001, 002, 012, 017, 020 | Disk-authoritative lifecycle ownership/generation and immutable store reads across history types, lifecycle reducers, `TaskHistoryStore`, provider delegation, and reconciliation. | XL | High: persisted compatibility and cross-host races | Two-host stale-write tests, promoted invariants, restart/reconciliation evidence, backward-compatible optional data. | | P2. Durable operation and crash recovery | 004, 005, 006, 021, 023, 039, 040 | Operation intent/replay or explicit idempotent recovery for pair writes, delegation, completion messages, shutdown, and deletion. | XL | Very high: failure ordering can create new corruption | Fault injection at every durable boundary, crash/restart convergence, no false success, recovery reachability. | -| P3. Schema, path, and lifecycle vocabulary | 013, 018, 019, 041 | Schema-derived status ownership, validated ordinary history reads, and one safe task-ID boundary across types, persistence, metadata, CLI, and import paths. | M | Medium: malformed legacy data and downgrade behavior | Migration/quarantine fixtures, traversal tests, type/static ratchets. | +| P3. Schema, path, and lifecycle vocabulary | 013, 018, 019 | Schema-derived status ownership, validated ordinary history reads, and one safe task-ID boundary across types, persistence, metadata, CLI, and import paths. | M | Medium: malformed legacy data and downgrade behavior | Migration/quarantine fixtures, traversal tests, type/static ratchets. | | P4. Request, stream, and tool identity | 008, 010, 024, 025, 026, 030, 037, 038 | Request generation plus canonical call identity, then task/generation/call-scoped parser and partial-handler state. Owners include provider transforms, parser, `Task`, `BaseTool`, editing handlers, and tool-ID utilities. | XL | High: provider compatibility and duplicate execution | Adversarial IDs/index-less streams, delayed/cancelled generation tests, cleanup/deadline checks, production-backed call-state model. | | P5. Tool-owned task state and queueing | 003, 007, 031, 035, 036 | Task-local context, durable child initialization, correlated approval identity, and claim/persist/ack queueing across tools, `Task`, provider/webview, message queue, and history schema. | XL | High: cross-task contamination and persistence precedence | Omitted/explicit child controls, switch/restart E2E, two-approval schedules, queue failure retention, mode-permission tests. | | P6. Event and ingress contracts | 009, 011, 022, 027, 028, 029, 032, 033 | Classify barriers versus notifications; normalize lifecycle payloads and clear/resume semantics across Task, provider, public API, IPC, and webview. | L | Medium-high: public compatibility and ordering | Exactly-once event tests, consumer inventory, cross-surface contract matrix, compatibility adapters where required. | @@ -241,7 +241,7 @@ The 40 IDs are not 40 independent projects. They group into eight programs with - One request-generation/canonical-call identity established before parser indexing can support 008, 010, 024–026, 030, 037, and 038. - One correlated `(taskId, actionId, toolCallId)` approval protocol can close 036 and support 007/035; it does not itself make child state durable. - One typed lifecycle operation layer can normalize P6, but public compatibility requires separate adapters rather than a flag-day payload rewrite. -- One validated-read and cycle-safe traversal guard closes 041 beside 018 and 019. +- One validated-read and cycle-safe traversal guard closes 018 and 019. ### Independent work that should not be collapsed @@ -313,7 +313,7 @@ Four workstreams can proceed concurrently after foundation decisions: | Correlated approval ownership | LIFE-GAP-036 | | Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | | Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | -| Validated-read and traversal-guard design | LIFE-GAP-018, 019, 041 | | +| Validated-read and traversal-guard design | LIFE-GAP-018, 019 | | ## Mechanically useful follow-up checklist diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index f8c412b334..c987effad2 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -41,20 +41,20 @@ TLA+/PlusCal or Quint with TLC becomes a better fit when the lifecycle needs tem ## Production mapping -| Model concept | Production concept | -| ------------------------- | ------------------------------------------------------------------------------------ | -| Task record and status | `HistoryItem` persisted by `TaskHistoryStore` | -| `delegate(parent, child)` | `ClineProvider.delegateParentAndOpenChild` | -| `interrupt(child)` | cancellation or eviction through `markDelegatedChildInterrupted` | -| `complete(child)` | `ClineProvider.reopenParentFromDelegation` | -| `abandon(child)` | `ClineProvider.abandonSubtask` | -| Pending-action settlement | `settleRejectedCreateSubtaskAction` inside the delegation update path (#1714) | -| Atomic event step | `atomicReadAndUpdate`, `atomicUpdatePair`, and per-parent delegation transition lock | -| Event interleaving | Competing completion, cancellation, abandonment, and new delegation calls | +| Model concept | Production concept | +| ------------------------- | -------------------------------------------------------------------------------------------------------------------- | +| Task record and status | `HistoryItem` persisted by `TaskHistoryStore` | +| `delegate(parent, child)` | `ClineProvider.delegateParentAndOpenChild` | +| `interrupt(child)` | cancellation or eviction through `markDelegatedChildInterrupted` | +| `complete(child)` | `ClineProvider.reopenParentFromDelegation` | +| `abandon(child)` | `ClineProvider.abandonSubtask` | +| Pending-action settlement | `TaskHistoryStore.clearPendingActionIfMatching` compare-and-clear in the rejected-delegation settlement path (#1714) | +| Atomic event step | `atomicReadAndUpdate`, `atomicUpdatePair`, and per-parent delegation transition lock | +| Event interleaving | Competing completion, cancellation, abandonment, and new delegation calls | The model has three fixed task slots, enough to cover competing siblings and a nested parent-child-grandchild chain. It explores every reachable interleaving through depth 12, deduplicating canonical states. Representative checks also exercise rejected operations that do not create a new state: a second concurrent delegation while the first child is active, stale completion after re-delegation, late completion after abandonment, completion after interruption, and nested completion. Named semantic landmarks require the graph to retain interrupted-child re-delegation and nested delegation even when the raw state total changes. -Each task slot can also hold one of two pending `create_subtask` actions. A `stage` action mirrors `setPendingTaskAction` overwrite semantics, delegation clears the action its request carried, completion preserves unrelated actions, and a `settle-rejected` action models the settlement that follows an authoritative delegation rejection (#1714). Production settles through the typed `LifecycleTransitionError` from the shared guards: the provider writes the settled record with one extra `atomicReadAndUpdate` call, then propagates the original rejection. Four named witnesses must remain reachable: settlement from an interrupted record after rejection, settlement through a successful active delegation, unrelated-action preservation during completion, and stale-action protection where a settlement targeting one action ID leaves a replacement action intact. A mismatched pending-action request keeps its production behavior: the atomic update throws before any transition, and no settlement runs. +Each task slot can also hold one of two pending `create_subtask` actions. A `stage` action mirrors `setPendingTaskAction` overwrite semantics, delegation clears the action its request carried, completion clears the child's action only when its event carries the matching action ID, and a `settle-rejected` action models the settlement that follows an authoritative delegation rejection (#1714). Production settles through the typed `LifecycleTransitionError` from the shared guards: the provider calls the disk-authoritative `TaskHistoryStore.clearPendingActionIfMatching` compare-and-clear under the per-file lock, then propagates the original rejection. Six named witnesses must remain reachable: settlement from an interrupted record after rejection, settlement through a successful active delegation, unrelated-action preservation during completion, stale-action protection where a settlement targeting one action ID leaves a replacement action intact, matching-ID completion clearing, and replacement-ID completion preservation. A mismatched pending-action request keeps its production behavior: the atomic update throws before any transition, and no settlement runs. Production completion also accepts a recovery-compatible `active` parent that still awaits the returning child, then clears the stale pointers. Normal model transitions never create that intermediate state, so it is covered by a focused reducer test rather than admitted as a generally valid reachable state. @@ -66,6 +66,7 @@ The same `pnpm lifecycle:model-check` command also runs a second bounded explore - store read/update operations hold the host mutex, while live-task snapshots used by completion and message saves may outlive it; - a write delta is computed relative to that host's cache; - revalidation under the per-file disk lock checks only status-transition legality; +- the rejected-delegation settlement compare-and-clear decides inside the disk merge callback, so a replacement action persisted by another host is never cleared; - fields absent from the delta preserve the current disk value, `childIds` are unioned, and other same-field conflicts are last-writer-wins; - `atomicUpdatePair` commits its files in order, with another host able to act between file commits; - successful pair-operation cache entries publish together after both file writes; if the second write fails, the cache publishes only the first committed record; @@ -82,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with one synchronized integration smoke check through the real `proper-lockfile` and filesystem rename path; broader VS Code E2E remains reserved for restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including the stale-settlement compare-and-clear regression; broader VS Code E2E remains reserved for restart and extension-host behavior. ## Task cleanup protocol model @@ -137,7 +138,8 @@ The task delegation checker currently enforces: 5. Parent-child lineage is acyclic. 6. Completed task records cannot be changed by later lifecycle events. 7. Active-child re-delegation, stale completion after ownership moves to another child, duplicate/late completion, and abandonment of a live child are rejected by the shared production guards. -8. A rejected delegation settles only the exact matching pending `create_subtask` action. Settlement preserves status, lineage, and accounting, never clears a replacement or different-kind action, and never mutates a completed record. +8. A rejected delegation settles only the exact matching pending `create_subtask` action. Settlement preserves status, lineage, and accounting, never clears a replacement or different-kind action, and never mutates a completed record. The settlement compare-and-clear reads the persisted record under the per-file lock, so it never clears from a stale host cache. +9. A completion clears the completing child's pending action only when the completion event carries the exact matching action ID. A completion with no action ID or a different ID preserves the pending action. The completion persistence checker additionally enforces: diff --git a/scripts/check-task-lifecycle.ts b/scripts/check-task-lifecycle.ts index 8dfb1c202a..bf05ad68e9 100644 --- a/scripts/check-task-lifecycle.ts +++ b/scripts/check-task-lifecycle.ts @@ -18,7 +18,7 @@ interface Transition { name: string next: ModelState delegation?: { parentId: TaskId } - completion?: { childId: TaskId } + completion?: { childId: TaskId; pendingActionId?: string } settlement?: { taskId: TaskId; actionId: string } } @@ -77,6 +77,28 @@ const semanticWitnesses = { next[transition.settlement.taskId]?.pendingAction?.actionId === beforeAction.actionId ) }, + "matching-completion-clears-pending-action": ({ prev, next, transition }: WitnessContext) => { + if (transition.completion?.pendingActionId === undefined) return false + const beforeAction = prev[transition.completion.childId]?.pendingAction + const after = next[transition.completion.childId] + return ( + beforeAction?.kind === "create_subtask" && + beforeAction.actionId === transition.completion.pendingActionId && + after?.status === "completed" && + after.pendingAction === undefined + ) + }, + "replacement-completion-preserves-pending-action": ({ prev, next, transition }: WitnessContext) => { + if (transition.completion?.pendingActionId === undefined) return false + const beforeAction = prev[transition.completion.childId]?.pendingAction + const after = next[transition.completion.childId] + return ( + beforeAction?.kind === "create_subtask" && + beforeAction.actionId !== transition.completion.pendingActionId && + after?.status === "completed" && + canonicalTask(after.pendingAction) === canonicalTask(beforeAction) + ) + }, } satisfies Record boolean> function task(id: TaskId, parentTaskId?: TaskId): HistoryItem { @@ -178,6 +200,17 @@ function transitions(state: ModelState): Transition[] { next: replace(state, completed.parent, completed.child), completion: { childId }, }) + for (const actionId of actionIds) { + const completedChild: HistoryItem = + child.pendingAction?.actionId === actionId + ? { ...completed.child, pendingAction: undefined } + : completed.child + result.push({ + name: `complete(${childId}, ${actionId})`, + next: replace(state, completed.parent, completedChild), + completion: { childId, pendingActionId: actionId }, + }) + } } if (parent.status === "delegated" && parent.awaitingChildId === child.id && child.status === "interrupted") { @@ -288,6 +321,19 @@ function checkTransitionInvariants(previous: ModelState, transition: Transition) ) } } + + const completion = transition.completion + if (completion) { + const beforeAction = previous[completion.childId]?.pendingAction + const afterAction = transition.next[completion.childId]?.pendingAction + + if (afterAction && canonicalTask(afterAction) !== canonicalTask(beforeAction)) { + violations.push(`${completion.childId}: completion after ${transition.name} replaced a pending action`) + } + if (!afterAction && beforeAction && beforeAction.actionId !== completion.pendingActionId) { + violations.push(`${completion.childId}: completion after ${transition.name} cleared a non-matching action`) + } + } return violations } @@ -409,5 +455,5 @@ function runRepresentativeScenarios(): void { runRepresentativeScenarios() const checkedStates = runModelCheck() console.log( - `Task lifecycle model check passed: ${checkedStates} reachable states, ${expectedActions.length}/${expectedActions.length} actions reachable, ${Object.keys(semanticLandmarks).length}/${Object.keys(semanticLandmarks).length} landmarks reached, ${Object.keys(semanticWitnesses).length}/${Object.keys(semanticWitnesses).length} settlement witnesses reached, depth <= ${MAX_DEPTH}, ${taskIds.length} task slots`, + `Task lifecycle model check passed: ${checkedStates} reachable states, ${expectedActions.length}/${expectedActions.length} actions reachable, ${Object.keys(semanticLandmarks).length}/${Object.keys(semanticLandmarks).length} landmarks reached, ${Object.keys(semanticWitnesses).length}/${Object.keys(semanticWitnesses).length} semantic witnesses reached, depth <= ${MAX_DEPTH}, ${taskIds.length} task slots`, ) diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index e0633964fa..5eb49eb164 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -782,6 +782,12 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { current = updater(current) return [current] }), + clearPendingActionIfMatching: vi.fn(async (_taskId: string, actionId: string) => { + if (current.pendingAction?.kind === "create_subtask" && current.pendingAction.actionId === actionId) { + current = { ...current, pendingAction: undefined } + } + return current + }), } const provider = { taskScheduler: new TaskScheduler(), @@ -842,6 +848,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { return child }) const atomicReadAndUpdate = vi.fn().mockRejectedValue(persistenceError) + const clearPendingActionIfMatching = vi.fn().mockResolvedValue(interruptedParent) const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) const provider = { taskScheduler: new TaskScheduler(), @@ -859,6 +866,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { invalidate: vi.fn().mockResolvedValue(undefined), get: vi.fn(() => interruptedParent), atomicReadAndUpdate, + clearPendingActionIfMatching, }, } as unknown as ClineProvider @@ -873,6 +881,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { ).rejects.toThrow(persistenceError) expect(atomicReadAndUpdate).toHaveBeenCalledTimes(1) + expect(clearPendingActionIfMatching).not.toHaveBeenCalled() expect(interruptedParent.pendingAction).toBe(pendingAction) expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) expect(createTaskWithHistoryItem).toHaveBeenCalledWith(interruptedParent) @@ -900,13 +909,11 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { getCurrentTask.mockReturnValue(child) return child }) - const atomicReadAndUpdate = vi - .fn() - .mockImplementationOnce(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { - updater(interruptedParent) - return [] - }) - .mockRejectedValueOnce(settlementError) + const atomicReadAndUpdate = vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + updater(interruptedParent) + return [] + }) + const clearPendingActionIfMatching = vi.fn().mockRejectedValue(settlementError) const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) const provider = { taskScheduler: new TaskScheduler(), @@ -924,6 +931,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { invalidate: vi.fn().mockResolvedValue(undefined), get: vi.fn(() => interruptedParent), atomicReadAndUpdate, + clearPendingActionIfMatching, }, } as unknown as ClineProvider @@ -937,8 +945,77 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { }), ).rejects.toThrow("Invalid task status transition: interrupted → delegated") - expect(atomicReadAndUpdate).toHaveBeenCalledTimes(2) + expect(atomicReadAndUpdate).toHaveBeenCalledTimes(1) + expect(clearPendingActionIfMatching).toHaveBeenCalledTimes(1) expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) expect(createTaskWithHistoryItem).not.toHaveBeenCalled() }) + + it("restores the authoritative parent record when settlement preserves a replacement action", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + const replacementAction = { ...pendingAction, actionId: "replacement-action" } + const cachedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const replacedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction: replacementAction, + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const clearPendingActionIfMatching = vi.fn(async (_taskId: string, _actionId: string) => replacedParent) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: replacedParent }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId: vi.fn().mockResolvedValue(undefined), + createTaskWithHistoryItem, + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => cachedParent), + atomicReadAndUpdate: vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + updater(cachedParent) + return [] + }), + clearPendingActionIfMatching, + }, + } as unknown as ClineProvider + + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow("Invalid task status transition: interrupted → delegated") + + expect(clearPendingActionIfMatching).toHaveBeenCalledTimes(1) + expect(replacedParent.pendingAction).toEqual(replacementAction) + expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(createTaskWithHistoryItem).toHaveBeenCalledWith(replacedParent) + }) }) diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 3d4cc47604..dab08a1d2f 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -1060,6 +1060,47 @@ export class TaskHistoryStore { }) } + /** + * Disk-authoritative compare-and-clear for a rejected `create_subtask` + * pending action (#1714). The comparison runs inside the per-file + * advisory lock's merge callback, so the decision reads the persisted + * record rather than this store's possibly stale cache. A missing, + * different-kind, or replacement pending action is preserved unchanged. + * The authoritative record is written back, the store cache is refreshed + * with it, and it is returned to the caller. + * + * @throws If the task ID is not present in the cache. + */ + public async clearPendingActionIfMatching(taskId: string, expectedActionId: string): Promise { + return this.withLock(async () => { + const cached = this.cache.get(taskId) + if (!cached) { + throw new Error(`[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`) + } + const filePath = await this.getTaskFilePath(taskId) + let authoritative: HistoryItem = cached + await safeWriteJson(filePath, cached, { + merge: (existing, incoming) => { + const disk = + existing && typeof existing === "object" && "id" in existing + ? (existing as HistoryItem) + : (incoming as HistoryItem) + const pendingAction = disk.pendingAction + authoritative = + pendingAction?.kind === "create_subtask" && pendingAction.actionId === expectedActionId + ? { ...disk, pendingAction: undefined } + : disk + return authoritative + }, + }) + this.cache.set(taskId, authoritative) + if (this.onWrite) { + await this.onWrite(this.getAll()) + } + return authoritative + }) + } + // ────────────────────────────── Private: Write lock ────────────────────────────── /** diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index d94ca8f782..be1c0c444b 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -72,7 +72,146 @@ function item(id: string): HistoryItem { } } +function createAction(actionId: string, message: string) { + return { + kind: "create_subtask" as const, + actionId, + approvalText: "{}", + mode: "code", + message, + todos: [], + } +} + describe("TaskHistoryStore real cross-host locking", () => { + it("preserves a replacement pending action when a stale store settles the prior action", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-stale-settlement-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const actionA = { + kind: "create_subtask" as const, + actionId: "action-a", + approvalText: "{}", + mode: "code", + message: "action A", + todos: [], + } + const actionB = { ...actionA, actionId: "action-b", message: "action B" } + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: actionA }) + await storeB.initialize() + + await storeB.atomicReadAndUpdate("shared-task", (current) => ({ ...current, pendingAction: actionB })) + expect(storeA.get("shared-task")?.pendingAction).toEqual(actionA) + + const authoritative = await storeA.clearPendingActionIfMatching("shared-task", actionA.actionId) + expect(authoritative.pendingAction).toEqual(actionB) + expect(storeA.get("shared-task")?.pendingAction).toEqual(actionB) + await storeB.invalidate("shared-task") + + expect(storeB.get("shared-task")?.pendingAction).toEqual(actionB) + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("clears a matching create_subtask action from disk and refreshes the cache", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-compare-clear-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const actionA = createAction("action-a", "action A") + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: actionA }) + await storeB.initialize() + + const authoritative = await storeA.clearPendingActionIfMatching("shared-task", actionA.actionId) + + expect(authoritative.pendingAction).toBeUndefined() + expect(storeA.get("shared-task")?.pendingAction).toBeUndefined() + await storeB.invalidate("shared-task") + expect(storeB.get("shared-task")?.pendingAction).toBeUndefined() + expect(storeB.get("shared-task")).toMatchObject({ id: "shared-task", status: "active" }) + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("clears a matching action from disk even when the calling store cache is stale", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-disk-compare-clear-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const actionA = createAction("action-a", "action A") + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + + await storeB.atomicReadAndUpdate("shared-task", (current) => ({ ...current, pendingAction: actionA })) + + const authoritative = await storeA.clearPendingActionIfMatching("shared-task", actionA.actionId) + + expect(authoritative.pendingAction).toBeUndefined() + await storeB.invalidate("shared-task") + expect(storeB.get("shared-task")?.pendingAction).toBeUndefined() + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("preserves a different-kind pending action with the same action ID", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-different-kind-")) + const storeA = new TaskHistoryStore(storagePath) + const finishAction = { + kind: "finish_subtask" as const, + actionId: "action-a", + approvalText: "{}", + parentTaskId: "parent-1", + result: "done", + } + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: finishAction }) + + const authoritative = await storeA.clearPendingActionIfMatching("shared-task", finishAction.actionId) + + expect(authoritative.pendingAction).toEqual(finishAction) + expect(storeA.get("shared-task")?.pendingAction).toEqual(finishAction) + } finally { + storeA.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("preserves the record when no pending action is persisted", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-no-action-")) + const storeA = new TaskHistoryStore(storagePath) + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + + const authoritative = await storeA.clearPendingActionIfMatching("shared-task", "action-a") + + expect(authoritative.pendingAction).toBeUndefined() + expect(authoritative).toMatchObject({ id: "shared-task", status: "active" }) + } finally { + storeA.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("preserves independent stale-cache deltas through the real per-file lock", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-real-lock-")) const storeA = new TaskHistoryStore(storagePath) diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index ba333acfe4..aa75414203 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -127,7 +127,6 @@ import { delegateTaskToChild, interruptDelegatedChild, LifecycleTransitionError, - settleRejectedCreateSubtaskAction, } from "../task-persistence" import { readTaskMessages } from "../task-persistence/taskMessages" import { getNonce } from "./getNonce" @@ -4047,14 +4046,14 @@ export class ClineProvider }`, ) // The authoritative parent record rejected this delegation (#1714). - // Settle the matching pending create_subtask action durably so a retry - // cannot replay a rejected action, then propagate the original error. + // Settle the matching pending create_subtask action durably through + // the disk-authoritative compare-and-clear so a retry cannot replay + // a rejected action and a replacement action from another host is + // never cleared, then propagate the original error. let settlementFailed = false if (pendingActionId && err instanceof LifecycleTransitionError) { try { - await this.taskHistoryStore.atomicReadAndUpdate(parentTaskId, (historyItem) => - settleRejectedCreateSubtaskAction(historyItem, pendingActionId), - ) + await this.taskHistoryStore.clearPendingActionIfMatching(parentTaskId, pendingActionId) this.recentTasksCache = undefined } catch (settlementError) { settlementFailed = true From 10473437916e5eae9da1cf91b417f742f0ceb62d Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 21 Sep 2026 03:23:50 +0000 Subject: [PATCH 04/23] fix(lifecycle): honor cross-host task deletion --- .../architecture/task-lifecycle-gap-report.md | 32 ++++----- .../ClineProvider.delegation.spec.ts | 65 +++++++++++++++++++ src/core/task-persistence/TaskHistoryStore.ts | 22 +++++-- .../TaskHistoryStore.realConcurrency.spec.ts | 27 ++++++++ 4 files changed, 124 insertions(+), 22 deletions(-) diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index 9f8d34e5d3..2d73f48534 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -241,7 +241,7 @@ The 40 IDs are not 40 independent projects. They group into eight programs with - One request-generation/canonical-call identity established before parser indexing can support 008, 010, 024–026, 030, 037, and 038. - One correlated `(taskId, actionId, toolCallId)` approval protocol can close 036 and support 007/035; it does not itself make child state durable. - One typed lifecycle operation layer can normalize P6, but public compatibility requires separate adapters rather than a flag-day payload rewrite. -- One validated-read and cycle-safe traversal guard closes 018 and 019. +- Validated ordinary reads close 018. One shared filesystem task-ID validator covering traversal and separator cases across store paths, imports, deletion, and checkpoints closes 019. ### Independent work that should not be collapsed @@ -299,21 +299,21 @@ Four workstreams can proceed concurrently after foundation decisions: ## Burn-down dependencies -| Dependency | Enables | -| ---------------------------------------------- | ------------------------------------------------------------- | --- | -| Disk-authoritative ownership/generation design | LIFE-GAP-001, 002, 012, 020 | -| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023, 039, 040 | -| Task-local execution-context owner | LIFE-GAP-007 and optional future fan-out | -| Request generation and terminal cleanup owner | LIFE-GAP-010, 024, 025, 026 | -| Event notification/barrier contract | LIFE-GAP-009, 011, 027, 028 | -| Machine-readable lifecycle manifest | LIFE-GAP-013, 016, 034 | -| Serial scheduler baseline | LIFE-GAP-014 | -| Optional fan-out product program | Historical #369/#372 scope, outside baseline | -| Durable task-scoped child initialization | LIFE-GAP-035 and future tool/lifecycle composition | -| Correlated approval ownership | LIFE-GAP-036 | -| Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | -| Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | -| Validated-read and traversal-guard design | LIFE-GAP-018, 019 | | +| Dependency | Enables | +| -------------------------------------------------------------------- | ------------------------------------------------------------- | --- | +| Disk-authoritative ownership/generation design | LIFE-GAP-001, 002, 012, 020 | +| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023, 039, 040 | +| Task-local execution-context owner | LIFE-GAP-007 and optional future fan-out | +| Request generation and terminal cleanup owner | LIFE-GAP-010, 024, 025, 026 | +| Event notification/barrier contract | LIFE-GAP-009, 011, 027, 028 | +| Machine-readable lifecycle manifest | LIFE-GAP-013, 016, 034 | +| Serial scheduler baseline | LIFE-GAP-014 | +| Optional fan-out product program | Historical #369/#372 scope, outside baseline | +| Durable task-scoped child initialization | LIFE-GAP-035 and future tool/lifecycle composition | +| Correlated approval ownership | LIFE-GAP-036 | +| Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | +| Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | +| Validated ordinary reads and one shared filesystem task-ID validator | LIFE-GAP-018, 019 | | ## Mechanically useful follow-up checklist diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index 5eb49eb164..a741f1815d 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -5,6 +5,7 @@ import type { HistoryItem } from "@roo-code/types" import { providerIdentifiers, RooCodeEventName } from "@roo-code/types" import { ClineProvider } from "../core/webview/ClineProvider" import { TaskScheduler } from "../core/task/TaskScheduler" +import { LifecycleTransitionError } from "../core/task-persistence" const parentHistoryItem: HistoryItem = { id: "parent-1", @@ -825,6 +826,70 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { ) }) + it("rolls back a typed lifecycle rejection without a pending action ID", async () => { + const interruptedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + let transitionError: LifecycleTransitionError | undefined + const atomicReadAndUpdate = vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + try { + updater(interruptedParent) + } catch (error) { + if (error instanceof LifecycleTransitionError) transitionError = error + throw error + } + return [] + }) + const clearPendingActionIfMatching = vi.fn() + const deleteTaskWithId = vi.fn().mockResolvedValue(undefined) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: interruptedParent }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId, + createTaskWithHistoryItem, + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => interruptedParent), + atomicReadAndUpdate, + clearPendingActionIfMatching, + }, + } as unknown as ClineProvider + + let rejection: unknown + try { + await ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: "Do something", + initialTodos: [], + mode: "code", + }) + } catch (error) { + rejection = error + } + + expect(rejection).toBe(transitionError) + expect(transitionError).toBeInstanceOf(LifecycleTransitionError) + expect(clearPendingActionIfMatching).not.toHaveBeenCalled() + expect(deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(createTaskWithHistoryItem).toHaveBeenCalledWith(interruptedParent) + }) + it("does not settle a pending action after an unrelated persistence failure", async () => { const persistenceError = new Error("parent persistence failed") const pendingAction = { diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index dab08a1d2f..9317431b8c 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -1069,7 +1069,11 @@ export class TaskHistoryStore { * The authoritative record is written back, the store cache is refreshed * with it, and it is returned to the caller. * - * @throws If the task ID is not present in the cache. + * Deletion by another host is authoritative (#1726): when no persisted + * record exists, the merge callback removes the stale cache entry and + * throws instead of writing the cached record back to disk. + * + * @throws If the task ID is not present in the cache or no persisted record remains on disk. */ public async clearPendingActionIfMatching(taskId: string, expectedActionId: string): Promise { return this.withLock(async () => { @@ -1080,11 +1084,17 @@ export class TaskHistoryStore { const filePath = await this.getTaskFilePath(taskId) let authoritative: HistoryItem = cached await safeWriteJson(filePath, cached, { - merge: (existing, incoming) => { - const disk = - existing && typeof existing === "object" && "id" in existing - ? (existing as HistoryItem) - : (incoming as HistoryItem) + merge: (existing) => { + if (!existing || typeof existing !== "object" || !("id" in existing)) { + // Writing the cached record back would recreate a task + // another host deleted, so drop the stale entry first. + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, + ) + } + const disk = existing as HistoryItem const pendingAction = disk.pendingAction authoritative = pendingAction?.kind === "create_subtask" && pendingAction.actionId === expectedActionId diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index be1c0c444b..b853bbf72c 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -212,6 +212,33 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("does not recreate a task deleted by another host before settlement", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-deleted-settlement-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const actionA = createAction("action-a", "action A") + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: actionA }) + await storeB.initialize() + + await storeB.delete("shared-task") + expect(storeA.get("shared-task")?.pendingAction).toEqual(actionA) + + await expect(storeA.clearPendingActionIfMatching("shared-task", actionA.actionId)).rejects.toThrow( + "task shared-task not found", + ) + expect(storeA.get("shared-task")).toBeUndefined() + await storeB.invalidate("shared-task") + expect(storeB.get("shared-task")).toBeUndefined() + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("preserves independent stale-cache deltas through the real per-file lock", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-real-lock-")) const storeA = new TaskHistoryStore(storagePath) From d6304f36adab6aadd629d4a7daca72a242f9dff7 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Tue, 22 Sep 2026 00:29:31 +0000 Subject: [PATCH 05/23] fix(lifecycle): preserve completed task records --- src/core/task-persistence/TaskHistoryStore.ts | 18 ++++++------ .../TaskHistoryStore.realConcurrency.spec.ts | 28 +++++++++++++++++++ 2 files changed, 36 insertions(+), 10 deletions(-) diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 9317431b8c..d6c42e3280 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -9,7 +9,7 @@ import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" -import { assertValidTransition, type HistoryItemStatus } from "./taskLifecycle" +import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" export { assertValidTransition, type HistoryItemStatus } from "./taskLifecycle" @@ -1064,10 +1064,12 @@ export class TaskHistoryStore { * Disk-authoritative compare-and-clear for a rejected `create_subtask` * pending action (#1714). The comparison runs inside the per-file * advisory lock's merge callback, so the decision reads the persisted - * record rather than this store's possibly stale cache. A missing, - * different-kind, or replacement pending action is preserved unchanged. - * The authoritative record is written back, the store cache is refreshed - * with it, and it is returned to the caller. + * record rather than this store's possibly stale cache. Settlement goes + * through the shared `settleRejectedCreateSubtaskAction` reducer, so a + * completed record is never mutated and a missing, different-kind, or + * replacement pending action is preserved unchanged. The authoritative + * record is written back, the store cache is refreshed with it, and it + * is returned to the caller. * * Deletion by another host is authoritative (#1726): when no persisted * record exists, the merge callback removes the stale cache entry and @@ -1095,11 +1097,7 @@ export class TaskHistoryStore { ) } const disk = existing as HistoryItem - const pendingAction = disk.pendingAction - authoritative = - pendingAction?.kind === "create_subtask" && pendingAction.actionId === expectedActionId - ? { ...disk, pendingAction: undefined } - : disk + authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) return authoritative }, }) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index b853bbf72c..bf338b96e0 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -144,6 +144,34 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("preserves a matching create_subtask action on a completed record", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-completed-settlement-")) + const store = new TaskHistoryStore(storagePath) + const pendingAction = createAction("action-a", "action A") + const completed: HistoryItem = { ...item("shared-task"), status: "completed", pendingAction } + const filePath = path.join(storagePath, "tasks", completed.id, "history_item.json") + + try { + await store.initialize() + await store.upsert(completed) + + const beforeDisk = JSON.parse(await fs.readFile(filePath, "utf8")) as HistoryItem + expect({ cache: store.get(completed.id), disk: beforeDisk }).toEqual({ cache: completed, disk: completed }) + + const returned = await store.clearPendingActionIfMatching(completed.id, pendingAction.actionId) + const afterDisk = JSON.parse(await fs.readFile(filePath, "utf8")) as HistoryItem + + expect({ returned, disk: afterDisk, cache: store.get(completed.id) }).toEqual({ + returned: completed, + disk: completed, + cache: completed, + }) + } finally { + store.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("clears a matching action from disk even when the calling store cache is stale", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-disk-compare-clear-")) const storeA = new TaskHistoryStore(storagePath) From 0e04c1e3cee538bf77927b34b58cb27064b36f24 Mon Sep 17 00:00:00 2001 From: "coderabbitai[bot]" <136622811+coderabbitai[bot]@users.noreply.github.com> Date: Tue, 22 Sep 2026 01:20:09 +0000 Subject: [PATCH 06/23] fix(lifecycle): prevent rejected subtask replay on restart and serialize task deletion with writes --- .../architecture/task-lifecycle-gap-report.md | 2 +- docs/architecture/task-lifecycle-model.md | 2 +- src/core/task-persistence/TaskHistoryStore.ts | 59 ++++++++++---- .../TaskHistoryStore.realConcurrency.spec.ts | 50 ++++++++++++ src/core/task/Task.ts | 34 ++++++++- .../task/__tests__/Task.persistence.spec.ts | 76 +++++++++++++++++++ src/core/webview/ClineProvider.ts | 5 +- src/utils/fileLock.ts | 28 +++++++ src/utils/safeWriteJson.ts | 24 +----- 9 files changed, 239 insertions(+), 41 deletions(-) create mode 100644 src/utils/fileLock.ts diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index 2d73f48534..b22288318b 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -214,7 +214,7 @@ Severity reflects plausible data loss, ownership corruption, permission/context | LIFE-GAP-036 | High | High | Interactive todo approval edit state is process-global and uncorrelated; one task's delayed edit can be consumed by another task's pending approval. | Start approvals for tasks A and B, send A's edited list through `setPendingTodoList`, then resolve B; B reads the shared `approvedTodoList`. | Task/action/tool-call-correlated approval state and webview protocol. | Carry task ID and action/tool-call ID through proposal, webview edit, approval, cancellation, and settlement; reject stale/mismatched edits; deep-clone inputs; test two interleaved approvals, denial, cancellation, task switch, and delayed edits. Correlation must extend to the staged pending action: `NewTaskTool` persists the action before approval, so the protocol must bind `pendingAction.actionId` to the approving task and tool call and reject cross-task edits before staging; denied, cancelled, and errored approvals need durable settlement rules equal to the rejected-delegation settlement; deep-copy extends to `pendingAction.todos` at staging. | | LIFE-GAP-037 | Medium-high | High | Singleton tool handlers share partial presentation state across calls/tasks, so interleaved paths can cause false or missed stabilization. | Interleave A:`x`, B:`y`, A:`x` or A:`x`, B:`x` through one handler's `lastSeenPartialPath`. | Per-call handler state keyed by task and tool-call identity. | Isolate partial state by `(taskId, toolCallId)` or handler instance; prove independent stabilization and cleanup after success, malformed finalization, rejection, cancellation, abandonment, and incomplete streams. | | LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. The bijection must include the staged pending action: `pendingAction.actionId` derives from `sanitizeToolUseId(toolCallId)`, so non-injective sanitization can make a settlement targeted at one raw call clear another call's staged action; the model's stale-settlement witness depends on collision-free action IDs. | -| LIFE-GAP-039 | High | High | Settlement after an authoritative pending-action rejection is a second, best-effort store write. A crash or settlement-write failure between the rejected transition and the settlement write leaves the rejected action staged, and startup replay can re-execute it without an attempt bound. | Source: the `delegateParentAndOpenChild` settlement block logs and swallows settlement errors; `pendingAction` carries no attempt count; the #1714 bounded-retry fix suggestion is unimplemented. | Durable operation intent (P2-004); attempt/generation identity (012). | Fault injection between rejection and settlement converges to a documented legal state; settlement-write failure is retried or surfaced without false success; replay of a rejected action is bounded by a persisted attempt/backoff record; model or focused witnesses cover the crash window. | +| LIFE-GAP-039 | Medium | High | Settlement after an authoritative pending-action rejection remains a second store write, but an interrupted task is now a non-replayable durable marker: restart must settle its exact pending `create_subtask` action before replay, and a failed recovery write stops without creating another child. A crash can still leave the rejected action staged and require recovery on every restart. | Source: the `delegateParentAndOpenChild` settlement block logs settlement errors and skips in-process restoration; `Task.resumeTaskFromHistory` retries exact-action settlement for interrupted create-subtask actions and surfaces failure before replay. `pendingAction` still carries no persisted attempt/backoff count. | Durable operation intent (P2-004); attempt/generation identity (012). | Persist settlement intent or attempt/backoff state so recovery converges without depending on storage becoming available; fault injection across every cut preserves the no-replay invariant and proves bounded recovery work. | | LIFE-GAP-040 | Medium | Medium-high | Startup topology repair reconciles status and lineage but has no pending-action rule. `reconcileDelegationState`, `repairActiveDelegation`, and repair-intent replay can leave a staged action on a repaired record, and the two interrupt paths disagree on whether `pendingAction` survives. | #1714 reproduced the loop precondition through startup reconciliation of a persisted active child; `DelegationRepairIntent` guards and targets carry no pending-action field; `applyDelegationRepairIntent` spreads records without a pending-action decision. | Durable operation intent (P2-004); correlated action identity (036, 038). | Every repair outcome (orphaned delegation, orphaned active child, completed child, intent replay, quarantine) defines preserve, settle, or drop for pending actions; deterministic repair tests and one fresh-host restart test prove the #1714 precondition cannot survive repair; both interrupt paths persist or settle the action consistently. | ## Portfolio remediation plan diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index c987effad2..1fad1e22d5 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -83,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including the stale-settlement compare-and-clear regression; broader VS Code E2E remains reserved for restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear and concurrent settlement/deletion regressions. Deletion uses the same per-file advisory lock as settlement, so it cannot interleave with the settlement rename window. Restart recovery also treats an interrupted task's pending `create_subtask` action as rejected: it must settle that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. ## Task cleanup protocol model diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index d6c42e3280..17ef670029 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -7,7 +7,8 @@ import deepEqual from "fast-deep-equal" import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" +import { acquireFileLock, LOCK_STALE_MS } from "../../utils/fileLock" +import { safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" @@ -266,17 +267,10 @@ export class TaskHistoryStore { */ async delete(taskId: string): Promise { return this.withLock(async () => { + await this.deleteTaskFile(taskId) this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) - // Remove per-task file (best-effort) - try { - const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) - } catch { - // File may already be deleted - } - // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { await this.onWrite(this.getAll()) @@ -290,15 +284,9 @@ export class TaskHistoryStore { async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { for (const taskId of taskIds) { + await this.deleteTaskFile(taskId) this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) - - try { - const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) - } catch { - // File may already be deleted - } } // Call onWrite callback inside the lock for serialized write-through @@ -883,6 +871,45 @@ export class TaskHistoryStore { } } + /** Delete through the same cross-process lock used by safeWriteJson. */ + private async deleteTaskFile(taskId: string): Promise { + const filePath = await this.getTaskFilePath(taskId) + const taskDir = path.dirname(filePath) + try { + await fs.access(taskDir) + } catch (error: unknown) { + const code = + error && typeof error === "object" && "code" in error ? (error as { code?: string }).code : undefined + if (code !== "ENOENT") { + throw error + } + // safeWriteJson creates the directory before taking this file lock. + // If it is absent, no writer has reached the shared protocol yet and + // deletion can linearize here without creating an empty task directory. + return + } + + let releaseLock: (() => Promise) | undefined + try { + releaseLock = await acquireFileLock(filePath) + await fs.unlink(filePath) + } catch (error: unknown) { + const code = + error && typeof error === "object" && "code" in error ? (error as { code?: string }).code : undefined + if (code !== "ENOENT") { + throw error + } + } finally { + if (releaseLock) { + try { + await releaseLock() + } catch (error) { + console.error(`Failed to release lock for ${filePath}:`, error) + } + } + } + } + // ────────────────────────────── Private: fs.watch ────────────────────────────── /** diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index bf338b96e0..48e930113f 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -267,6 +267,56 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("rejects settlement for a task absent from the cache without creating it", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-cache-miss-settlement-")) + const store = new TaskHistoryStore(storagePath) + const filePath = path.join(storagePath, "tasks", "missing-task", "history_item.json") + + try { + await store.initialize() + + await expect(store.clearPendingActionIfMatching("missing-task", "action-a")).rejects.toThrow( + "task missing-task not found in cache", + ) + expect(store.get("missing-task")).toBeUndefined() + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + } finally { + store.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("serializes deletion with an in-flight settlement so the task stays deleted", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-during-settlement-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const actionA = createAction("action-a", "action A") + const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") + const lockPath = `${filePath}.lock` + const largeTask = "x".repeat(16 * 1024 * 1024) + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), task: largeTask, pendingAction: actionA }) + await storeB.initialize() + + const settlement = storeA.clearPendingActionIfMatching("shared-task", actionA.actionId) + await vi.waitFor(() => expect(fs.stat(lockPath)).resolves.toBeDefined(), { interval: 1, timeout: 2_000 }) + const deletion = storeB.delete("shared-task") + + await expect(settlement).resolves.toMatchObject({ id: "shared-task", pendingAction: undefined }) + await expect(deletion).resolves.toBeUndefined() + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + expect(storeB.get("shared-task")).toBeUndefined() + await storeA.invalidate("shared-task") + expect(storeA.get("shared-task")).toBeUndefined() + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("preserves independent stale-cache deltas through the real per-file lock", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-real-lock-")) const storeA = new TaskHistoryStore(storagePath) diff --git a/src/core/task/Task.ts b/src/core/task/Task.ts index 55798437c3..7731944c08 100644 --- a/src/core/task/Task.ts +++ b/src/core/task/Task.ts @@ -609,7 +609,7 @@ export class Task extends EventEmitter implements TaskLike { this.parentTask = parentTask this.taskNumber = taskNumber - this.initialStatus = initialStatus + this.initialStatus = initialStatus ?? historyItem?.status this.pendingAction = historyItem?.pendingAction // Store the task's mode and API config name when it's created. @@ -994,6 +994,37 @@ export class Task extends EventEmitter implements TaskLike { } } + /** + * An interrupted task cannot legally delegate, so its staged create-subtask + * action is a durable rejection marker rather than replayable work. Reconcile + * it before restart replay; if persistence is still unavailable, propagate the + * error and leave the task stopped instead of creating another doomed child. + */ + private async settleInterruptedCreateSubtaskBeforeReplay(): Promise { + const action = this.pendingAction + if (this.initialStatus !== "interrupted" || action?.kind !== "create_subtask") { + return + } + + const provider = this.providerRef.deref() + if (!provider) { + throw new Error( + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Provider unavailable for task ${this.taskId}`, + ) + } + + const authoritative = await provider.taskHistoryStore.clearPendingActionIfMatching(this.taskId, action.actionId) + if (this.pendingAction?.actionId === action.actionId) { + this.pendingAction = authoritative.pendingAction + } + + if (this.pendingAction?.kind === "create_subtask") { + throw new Error( + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Task ${this.taskId} still has a rejected create-subtask action`, + ) + } + } + private handleQueuedAskResponse(message: QueuedMessage, resolution: QueuedAskResolution): string | undefined { this.handleWebviewAskResponse(resolution.response, message.text, message.images) if (resolution.requiresDurableAck) { @@ -2372,6 +2403,7 @@ export class Task extends EventEmitter implements TaskLike { // This is important in case the user deletes messages without resuming // the task first. this.hydrateApiConversationHistory(savedApiConversationHistory) + await this.settleInterruptedCreateSubtaskBeforeReplay() if ( this.pendingAction && this.apiConversationHistory.some( diff --git a/src/core/task/__tests__/Task.persistence.spec.ts b/src/core/task/__tests__/Task.persistence.spec.ts index 8d3314a9a6..03aab21ec5 100644 --- a/src/core/task/__tests__/Task.persistence.spec.ts +++ b/src/core/task/__tests__/Task.persistence.spec.ts @@ -1340,6 +1340,14 @@ describe("Task persistence", () => { parentTaskId: "parent-1", result: "Done", } + const createSubtaskAction: PendingTaskAction = { + kind: "create_subtask", + actionId: "create-action", + approvalText: JSON.stringify({ tool: "newTask" }), + mode: "code", + message: "Child task", + todos: [], + } it("replays an unresolved pending action instead of a generic resume ask", async () => { const messages: ClineMessage[] = [ @@ -1383,6 +1391,74 @@ describe("Task persistence", () => { expect(mockSaveTaskMessages).not.toHaveBeenCalled() }) + it("settles an interrupted create-subtask action before restart replay", async () => { + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + const clearRejectedAction = vi.fn().mockResolvedValue({ + id: "parent-1", + status: "interrupted", + pendingAction: undefined, + }) + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + vi.spyOn(task, "ask").mockResolvedValue({ response: "noButtonClicked" }) + vi.spyOn(getTaskPersistenceAccess(task), "initiateTaskLoop").mockResolvedValue(undefined) + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + expect(clearRejectedAction).toHaveBeenCalledWith("parent-1", "create-action") + expect(replay).not.toHaveBeenCalled() + expect(task.ask).toHaveBeenCalledWith("resume_task") + }) + + it("does not replay an interrupted create-subtask action when restart settlement fails", async () => { + const settlementError = new Error("settlement unavailable") + mockReadTaskMessages.mockResolvedValue([]) + mockReadApiMessages.mockResolvedValue([]) + mockProvider.taskHistoryStore.clearPendingActionIfMatching = vi.fn().mockRejectedValue(settlementError) + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const ask = vi.spyOn(task, "ask") + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + + await expect(getTaskPersistenceAccess(task).resumeTaskFromHistory()).rejects.toThrow(settlementError) + + expect(replay).not.toHaveBeenCalled() + expect(ask).not.toHaveBeenCalled() + }) + it("reconciles an already-persisted tool result before generic resume", async () => { mockReadTaskMessages.mockResolvedValue([{ ts: 1, type: "say", say: "text", text: "Child" }]) mockReadApiMessages.mockResolvedValue([ diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index aa75414203..ee8b3f19e5 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -4088,8 +4088,9 @@ export class ClineProvider } try { // A failed settlement write leaves the rejected pending action in - // durable storage. Restoring the stored parent would replay the - // rejected action, so leave the parent unrestored instead. + // durable storage. Restoring the stored parent would replay it in + // this process, so leave the parent unrestored. Restart recovery also + // settles interrupted create-subtask actions before allowing replay. if (!settlementFailed) { const { historyItem: parentHistory } = await this.getTaskWithId(parentTaskId) await this.createTaskWithHistoryItem(parentHistory) diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts new file mode 100644 index 0000000000..9563c666d0 --- /dev/null +++ b/src/utils/fileLock.ts @@ -0,0 +1,28 @@ +import * as path from "path" +import * as lockfile from "proper-lockfile" + +export const LOCK_STALE_MS = 31_000 + +/** + * Acquire the advisory lock shared by JSON writes and destructive mutations. + * The target may not exist yet, so callers can use the same protocol for + * creation, replacement, and deletion. + */ +export function acquireFileLock(filePath: string): Promise<() => Promise> { + const absoluteFilePath = path.resolve(filePath) + return lockfile.lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10_000, + realpath: false, + retries: { + retries: 5, + factor: 2, + minTimeout: 100, + maxTimeout: 1_000, + }, + onCompromised: (error) => { + console.error(`Lock at ${absoluteFilePath} was compromised:`, error) + throw error + }, + }) +} diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 957a0bb20f..b20aaba405 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -1,9 +1,10 @@ import * as fs from "fs/promises" import * as fsSync from "fs" import * as path from "path" -import * as lockfile from "proper-lockfile" import { JsonStreamStringify } from "json-stream-stringify" +import { acquireFileLock, LOCK_STALE_MS } from "./fileLock" + /** * Options for safeWriteJson function */ @@ -63,22 +64,7 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Acquire the lock before any file operations try { - releaseLock = await lockfile.lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long - realpath: false, // the file may not exist yet, which is acceptable - retries: { - // Configuration for retrying lock acquisition - retries: 5, // Number of retries after the initial attempt - factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) - minTimeout: 100, // Minimum time to wait before the first retry (in ms) - maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) - }, - onCompromised: (err) => { - console.error(`Lock at ${absoluteFilePath} was compromised:`, err) - throw err - }, - }) + releaseLock = await acquireFileLock(absoluteFilePath) } catch (lockError) { // If lock acquisition fails, we throw immediately. // The releaseLock remains a no-op, so the finally block in the main file operations @@ -247,6 +233,4 @@ async function _streamDataToFile(targetPath: string, data: any, prettyPrint = fa }) } -export const LOCK_STALE_MS = 31_000 - -export { safeWriteJson } +export { LOCK_STALE_MS, safeWriteJson } From 29ae8ec2ab3670fb4c72d3cd7e78fec8146e5212 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Wed, 23 Sep 2026 01:28:04 +0000 Subject: [PATCH 07/23] fix(persistence): serialize task deletion with history writes --- docs/architecture/task-lifecycle-model.md | 2 +- src/__tests__/delegation-concurrent.spec.ts | 11 + src/core/task-persistence/TaskHistoryStore.ts | 184 ++++++++++---- .../TaskHistoryStore.crossInstance.spec.ts | 18 +- .../TaskHistoryStore.guardCompromise.spec.ts | 149 ++++++++++++ .../TaskHistoryStore.realConcurrency.spec.ts | 101 ++++++++ src/core/webview/ClineProvider.ts | 19 +- .../ClineProvider.sticky-profile.spec.ts | 11 + .../ClineProvider.taskHistory.spec.ts | 11 + src/eslint-suppressions.json | 4 +- src/utils/__tests__/safeWriteJson.test.ts | 224 +++++++----------- src/utils/fileLock.ts | 77 ++++-- src/utils/safeWriteJson.ts | 176 ++++++-------- 13 files changed, 653 insertions(+), 334 deletions(-) create mode 100644 src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index 1fad1e22d5..7cdb3fc651 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -83,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear and concurrent settlement/deletion regressions. Deletion uses the same per-file advisory lock as settlement, so it cannot interleave with the settlement rename window. Restart recovery also treats an interrupted task's pending `create_subtask` action as rejected: it must settle that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear and concurrent settlement/deletion regressions. Deletion shares a per-task guard lock outside the task directory with history writes, so it cannot interleave with the settlement rename window or the recursive directory removal. Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. ## Task cleanup protocol model diff --git a/src/__tests__/delegation-concurrent.spec.ts b/src/__tests__/delegation-concurrent.spec.ts index 40d9b49ee5..222fdb6318 100644 --- a/src/__tests__/delegation-concurrent.spec.ts +++ b/src/__tests__/delegation-concurrent.spec.ts @@ -23,6 +23,17 @@ vi.mock("../utils/safeWriteJson", () => ({ safeWriteJson: vi.fn().mockResolvedValue(undefined), })) +// The store's cross-process task guard and advisory file lock need a real +// filesystem, which this spec stubs out. The in-process serialization under +// test does not depend on them. +vi.mock("../utils/fileLock", async () => ({ + ...(await vi.importActual("../utils/fileLock")), + acquireFileLock: vi.fn().mockResolvedValue({ + release: vi.fn().mockResolvedValue(undefined), + isCompromised: vi.fn().mockReturnValue(false), + }), +})) + vi.mock("../utils/storage", () => ({ getStorageBasePath: vi.fn().mockResolvedValue("/tmp/test-storage"), })) diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 17ef670029..7b2dbd8a1f 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -7,7 +7,13 @@ import deepEqual from "fast-deep-equal" import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { acquireFileLock, LOCK_STALE_MS } from "../../utils/fileLock" +import { + acquireFileLock, + assertLockUsable, + DESTRUCTIVE_LOCK_RETRIES, + LOCK_STALE_MS, + type AcquireFileLockOptions, +} from "../../utils/fileLock" import { safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" @@ -804,9 +810,21 @@ export class TaskHistoryStore { await fs.access(filePath) // File already exists, skip (don't overwrite existing per-task files) } catch { - // File doesn't exist, write it - await safeWriteJson(filePath, item) - this.cache.set(item.id, item) + // File doesn't exist, write it under the same task guard as + // every other history write so a concurrent deletion cannot + // interleave with the removal of the task directory. + const guard = await this.acquireTaskIoGuard(item.id) + try { + assertLockUsable(guard, filePath, "write") + await safeWriteJson(filePath, item) + this.cache.set(item.id, item) + } finally { + try { + await guard.release() + } catch (error) { + console.error(`Failed to release task guard for ${item.id}:`, error) + } + } } } @@ -829,6 +847,22 @@ export class TaskHistoryStore { return { id, ...this.computeDelta(cached, incoming) } } + /** + * Cross-process guard for one task. It lives outside the removable task + * directory, so deletion can hold it across the history-file unlink and + * the recursive directory removal while history writers hold it around + * their writes. Neither side can then interleave with the other. + */ + private async acquireTaskIoGuard( + taskId: string, + options?: AcquireFileLockOptions, + ): Promise> { + const tasksDir = await this.getTasksDir() + const guardDir = path.join(tasksDir, ".guards") + await fs.mkdir(guardDir, { recursive: true }) + return acquireFileLock(path.join(guardDir, `${taskId}.guard`), options) + } + /** * Write a HistoryItem to its per-task `history_item.json` file. * @@ -839,20 +873,30 @@ export class TaskHistoryStore { */ private async writeTaskFile(item: HistoryItem, delta?: Partial): Promise { const filePath = await this.getTaskFilePath(item.id) - if (delta) { - let written: HistoryItem = item - const mergeFn = mergeWithDisk(delta) - await safeWriteJson(filePath, item, { - merge: (existing, incoming) => { - const result = mergeFn(existing, incoming) - written = result as HistoryItem - return result - }, - }) - return written - } else { - await safeWriteJson(filePath, item) - return item + const guard = await this.acquireTaskIoGuard(item.id) + try { + assertLockUsable(guard, filePath, "write") + if (delta) { + let written: HistoryItem = item + const mergeFn = mergeWithDisk(delta) + await safeWriteJson(filePath, item, { + merge: (existing, incoming) => { + const result = mergeFn(existing, incoming) + written = result as HistoryItem + return result + }, + }) + return written + } else { + await safeWriteJson(filePath, item) + return item + } + } finally { + try { + await guard.release() + } catch (error) { + console.error(`Failed to release task guard for ${item.id}:`, error) + } } } @@ -871,7 +915,15 @@ export class TaskHistoryStore { } } - /** Delete through the same cross-process lock used by safeWriteJson. */ + /** + * Delete the history file and remove the task directory. + * + * The task guard lives outside the removable directory and spans the + * unlink and the recursive removal, so a concurrent history writer + * cannot interleave with either step. Both acquisitions use a retry + * budget that outlasts LOCK_STALE_MS, so a lock left by a crashed + * process is broken within this call instead of failing the deletion. + */ private async deleteTaskFile(taskId: string): Promise { const filePath = await this.getTaskFilePath(taskId) const taskDir = path.dirname(filePath) @@ -889,23 +941,47 @@ export class TaskHistoryStore { return } - let releaseLock: (() => Promise) | undefined + const guard = await this.acquireTaskIoGuard(taskId, { retries: DESTRUCTIVE_LOCK_RETRIES }) try { - releaseLock = await acquireFileLock(filePath) - await fs.unlink(filePath) - } catch (error: unknown) { - const code = - error && typeof error === "object" && "code" in error ? (error as { code?: string }).code : undefined - if (code !== "ENOENT") { - throw error + let fileLock: Awaited> | undefined + try { + fileLock = await acquireFileLock(filePath, { retries: DESTRUCTIVE_LOCK_RETRIES }) + assertLockUsable(fileLock, filePath, "deletion") + // The guard can be lost while waiting for the file lock above. + assertLockUsable(guard, filePath, "deletion") + await fs.unlink(filePath) + } catch (error: unknown) { + const code = + error && typeof error === "object" && "code" in error + ? (error as { code?: string }).code + : undefined + if (code !== "ENOENT") { + throw error + } + } finally { + if (fileLock) { + try { + await fileLock.release() + } catch (error) { + console.error(`Failed to release lock for ${filePath}:`, error) + } + } + } + + // The guard must still hold exclusion before this second mutation. + assertLockUsable(guard, taskDir, "task directory removal") + try { + await fs.rm(taskDir, { recursive: true, force: true }) + } catch (error) { + // The record is already unlinked. Mirror the provider's historical + // tolerance for a directory that cannot be removed right now. + console.error(`[TaskHistoryStore] Failed to remove task directory ${taskDir}:`, error) } } finally { - if (releaseLock) { - try { - await releaseLock() - } catch (error) { - console.error(`Failed to release lock for ${filePath}:`, error) - } + try { + await guard.release() + } catch (error) { + console.error(`Failed to release task guard for ${taskId}:`, error) } } } @@ -1112,22 +1188,32 @@ export class TaskHistoryStore { } const filePath = await this.getTaskFilePath(taskId) let authoritative: HistoryItem = cached - await safeWriteJson(filePath, cached, { - merge: (existing) => { - if (!existing || typeof existing !== "object" || !("id" in existing)) { - // Writing the cached record back would recreate a task - // another host deleted, so drop the stale entry first. - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - throw new Error( - `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, - ) - } - const disk = existing as HistoryItem - authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) - return authoritative - }, - }) + const guard = await this.acquireTaskIoGuard(taskId) + try { + assertLockUsable(guard, filePath, "write") + await safeWriteJson(filePath, cached, { + merge: (existing) => { + if (!existing || typeof existing !== "object" || !("id" in existing)) { + // Writing the cached record back would recreate a task + // another host deleted, so drop the stale entry first. + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, + ) + } + const disk = existing as HistoryItem + authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) + return authoritative + }, + }) + } finally { + try { + await guard.release() + } catch (error) { + console.error(`Failed to release task guard for ${taskId}:`, error) + } + } this.cache.set(taskId, authoritative) if (this.onWrite) { await this.onWrite(this.getAll()) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts index cef9874e5f..40c555dbaa 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts @@ -142,25 +142,25 @@ describe("TaskHistoryStore cross-instance safety", () => { expect(storeB.get("shared-task")).toBeUndefined() }) - it("delete by instance A is detected even when the task directory remains", async () => { + it("delete by instance A removes the history file, the task directory, and is detected by instance B", async () => { await storeA.initialize() await storeB.initialize() - const item = makeHistoryItem({ id: "file-only-delete" }) + const item = makeHistoryItem({ id: "full-delete" }) await storeA.upsert(item) await storeB.reconcile() - expect(storeB.get("file-only-delete")).toBeDefined() + expect(storeB.get("full-delete")).toBeDefined() - // delete() unlinks history_item.json but leaves the task directory. - await storeA.delete("file-only-delete") + // delete() unlinks history_item.json and removes the task directory + // under the task guard shared with history writes. + await storeA.delete("full-delete") - // Directory still exists (other files like ui_messages.json may remain). - const taskDir = path.join(tmpDir, "tasks", "file-only-delete") - await expect(fs.access(taskDir)).resolves.toBeUndefined() + const taskDir = path.join(tmpDir, "tasks", "full-delete") + await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) await storeB.reconcile() - expect(storeB.get("file-only-delete")).toBeUndefined() + expect(storeB.get("full-delete")).toBeUndefined() }) it("per-task file updates by one instance are visible to another after invalidation", async () => { diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts new file mode 100644 index 0000000000..3dfeec39cf --- /dev/null +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts @@ -0,0 +1,149 @@ +// pnpm --filter roo-cline test core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts + +import * as fs from "fs/promises" +import * as path from "path" +import * as os from "os" + +import type { HistoryItem } from "@roo-code/types" + +import { GlobalFileNames } from "../../../shared/globalFileNames" + +vi.mock("../../../utils/storage", () => ({ + getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => { + return defaultPath + }), +})) + +// Mock safeWriteJson with plain fs writes so only the outer task guard +// behavior is under test. +vi.mock("../../../utils/safeWriteJson", () => ({ + safeWriteJson: vi.fn().mockImplementation(async (filePath: string, data: unknown) => { + await fs.mkdir(path.dirname(filePath), { recursive: true }) + await fs.writeFile(filePath, JSON.stringify(data, null, "\t"), "utf8") + }), +})) + +function makeHistoryItem(overrides: Partial = {}): HistoryItem { + return { + id: `task-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, + number: 1, + ts: Date.now(), + task: "Test task", + tokensIn: 100, + tokensOut: 50, + totalCost: 0.01, + workspace: "/test/workspace", + ...overrides, + } +} + +async function seedTaskOnDisk(tmpDir: string, taskId: string): Promise<{ taskDir: string; historyFile: string }> { + const taskDir = path.join(tmpDir, "tasks", taskId) + const historyFile = path.join(taskDir, GlobalFileNames.historyItem) + await fs.mkdir(taskDir, { recursive: true }) + await fs.writeFile(historyFile, JSON.stringify(makeHistoryItem({ id: taskId }))) + return { taskDir, historyFile } +} + +describe("TaskHistoryStore task guard compromise", () => { + let tmpDir: string + + beforeEach(async () => { + tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-guard-")) + }) + + afterEach(async () => { + vi.doUnmock("proper-lockfile") + vi.resetModules() + await fs.rm(tmpDir, { recursive: true, force: true }).catch(() => {}) + }) + + /** + * Simulate proper-lockfile reporting the guard lock (never the history + * file lock) as compromised the moment it is acquired. Guards live at + * `tasks/.guards/.guard`. + */ + function compromiseGuardsOnAcquire() { + vi.doMock("proper-lockfile", () => ({ + lock: (file: string, options: { onCompromised: (error: Error) => void }) => { + if (file.endsWith(".guard")) { + options.onCompromised(new Error("Guard no longer available")) + } + return Promise.resolve(async () => {}) + }, + })) + } + + async function importStoreClass(): Promise { + const module = await import("../TaskHistoryStore") + return module.TaskHistoryStore + } + + it("aborts a history write when the task guard is compromised", async () => { + compromiseGuardsOnAcquire() + const TaskHistoryStore = await importStoreClass() + const store = new TaskHistoryStore(tmpDir) + await store.initialize() + + await expect(store.upsert(makeHistoryItem({ id: "guard-write" }))).rejects.toThrow("was compromised") + + // The write aborted before reaching the history file. + const historyFile = path.join(tmpDir, "tasks", "guard-write", GlobalFileNames.historyItem) + await expect(fs.access(historyFile)).rejects.toThrow() + expect(store.get("guard-write")).toBeUndefined() + + store.dispose() + }) + + it("aborts a task deletion when the task guard is compromised before the unlink", async () => { + const { taskDir, historyFile } = await seedTaskOnDisk(tmpDir, "guard-delete") + + compromiseGuardsOnAcquire() + const TaskHistoryStore = await importStoreClass() + const store = new TaskHistoryStore(tmpDir) + await store.initialize() + expect(store.get("guard-delete")).toBeDefined() + + await expect(store.delete("guard-delete")).rejects.toThrow("was compromised") + + // Neither the history file nor the task directory was removed. + await expect(fs.access(historyFile)).resolves.toBeUndefined() + await expect(fs.access(taskDir)).resolves.toBeUndefined() + expect(store.get("guard-delete")).toBeDefined() + + store.dispose() + }) + + it("stops before removing the task directory when the guard is compromised mid-deletion", async () => { + const { taskDir, historyFile } = await seedTaskOnDisk(tmpDir, "guard-mid-delete") + + // The history-file lock's release runs after the unlink and before + // the directory removal, so it is the deterministic hook for losing + // the guard between the two mutations. In-flight operations cannot + // be cancelled; the store must re-check before each mutation. + let compromiseGuard: (() => void) | undefined + vi.doMock("proper-lockfile", () => ({ + lock: (file: string, options: { onCompromised: (error: Error) => void }) => { + if (file.endsWith(".guard")) { + compromiseGuard = () => options.onCompromised(new Error("Guard no longer available")) + return Promise.resolve(async () => {}) + } + return Promise.resolve(async () => { + compromiseGuard?.() + }) + }, + })) + + const TaskHistoryStore = await importStoreClass() + const store = new TaskHistoryStore(tmpDir) + await store.initialize() + + await expect(store.delete("guard-mid-delete")).rejects.toThrow("was compromised") + + // The unlink already landed, but the guarded directory removal did not. + await expect(fs.access(historyFile)).rejects.toThrow() + await expect(fs.access(taskDir)).resolves.toBeUndefined() + + store.dispose() + }) +}) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index 48e930113f..b489f571e8 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -4,6 +4,7 @@ import * as path from "path" import type { HistoryItem } from "@roo-code/types" +import { acquireFileLock } from "../../../utils/fileLock" import { TaskHistoryStore } from "../TaskHistoryStore" type WriteTaskFile = (item: HistoryItem, delta?: Partial) => Promise @@ -368,4 +369,104 @@ describe("TaskHistoryStore real cross-host locking", () => { await fs.rm(storagePath, { recursive: true, force: true }) } }) + + it("serializes deletion with an external history-write guard holder and removes the task directory", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-guard-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const tasksDir = path.join(storagePath, "tasks") + const guardPath = path.join(tasksDir, ".guards", "shared-task.guard") + const filePath = path.join(tasksDir, "shared-task", "history_item.json") + const taskDir = path.join(tasksDir, "shared-task") + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + + // Hold the same guard that history writes hold, so the deletion + // must wait instead of unlinking underneath a writer. + const guard = await acquireFileLock(guardPath) + const deletion = storeB.delete("shared-task") + await new Promise((resolve) => setTimeout(resolve, 50)) + await expect(fs.access(filePath)).resolves.toBeUndefined() + + await guard.release() + await expect(deletion).resolves.toBeUndefined() + + // The guard spans the history-file unlink and the recursive + // directory removal. + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) + expect(storeB.get("shared-task")).toBeUndefined() + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("lets a write that starts during deletion run only after the removal window", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-write-during-delete-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + + const deletion = storeB.delete("shared-task") + await vi.waitFor(() => expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }), { + interval: 1, + timeout: 2_000, + }) + + // The write starts while the deletion still owns the task guard, + // so it cannot interleave with the removal of the task directory. + const write = storeA.upsert({ ...item("shared-task"), ts: 2000, task: "rewritten after deletion" }) + + await expect(deletion).resolves.toBeUndefined() + await expect(write).resolves.toBeDefined() + + const persisted = JSON.parse(await fs.readFile(filePath, "utf8")) as HistoryItem + expect(persisted).toMatchObject({ id: "shared-task", task: "rewritten after deletion" }) + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("deletes a task whose file lock is held live beyond the legacy retry window", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-contention-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const tasksDir = path.join(storagePath, "tasks") + const filePath = path.join(tasksDir, "shared-task", "history_item.json") + const taskDir = path.join(tasksDir, "shared-task") + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + + // A live holder keeps the lock mtime fresh, so staleness never + // breaks it. The deletion must wait out the hold instead of failing + // once the legacy retry budget of roughly 2.5 seconds is spent. + const holder = await acquireFileLock(filePath) + const deletion = storeB.delete("shared-task") + await new Promise((resolve) => setTimeout(resolve, 3_000)) + await holder.release() + + await expect(deletion).resolves.toBeUndefined() + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) }) diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index ee8b3f19e5..00802a2c54 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -2375,15 +2375,15 @@ export class ClineProvider } } - // Delete all tasks from state in one batch + // Delete all tasks from state in one batch. The store removes each + // history file and its task directory under the shared task guard, + // so a concurrent history write cannot interleave with the removal. await this.taskHistoryStore.deleteMany(allIdsToDelete) this.recentTasksCache = undefined - // Delete associated shadow repositories or branches and task directories + // Delete associated shadow repositories or branches const globalStorageDir = this.contextProxy.globalStorageUri.fsPath const workspaceDir = this.cwd - const { getTaskDirectoryPath } = await import("../../utils/storage") - const globalStoragePath = this.contextProxy.globalStorageUri.fsPath for (const taskId of allIdsToDelete) { try { @@ -2393,17 +2393,6 @@ export class ClineProvider `[deleteTaskWithId${taskId}] failed to delete associated shadow repository or branch: ${error instanceof Error ? error.message : String(error)}`, ) } - - // Delete the task directory - try { - const dirPath = await getTaskDirectoryPath(globalStoragePath, taskId) - await fs.rm(dirPath, { recursive: true, force: true }) - console.log(`[deleteTaskWithId${taskId}] removed task directory`) - } catch (error) { - console.error( - `[deleteTaskWithId${taskId}] failed to remove task directory: ${error instanceof Error ? error.message : String(error)}`, - ) - } } await this.postStateToWebview() diff --git a/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts b/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts index 7d8493fba3..8f04d36ae6 100644 --- a/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts @@ -88,6 +88,17 @@ vi.mock("../../prompts/sections/custom-instructions") vi.mock("../../../utils/safeWriteJson") +// The store's cross-process task guard and advisory file lock need a real +// filesystem, which this spec stubs out. The store behavior under test does +// not depend on them. +vi.mock("../../../utils/fileLock", async () => ({ + ...(await vi.importActual("../../../utils/fileLock")), + acquireFileLock: vi.fn().mockResolvedValue({ + release: vi.fn().mockResolvedValue(undefined), + isCompromised: vi.fn().mockReturnValue(false), + }), +})) + vi.mock("../../../api", () => ({ buildApiHandler: vi.fn().mockReturnValue({ getModel: vi.fn().mockReturnValue({ diff --git a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts index 2bbf0736c6..f8e1a06d77 100644 --- a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts @@ -59,6 +59,17 @@ vi.mock("../../../utils/safeWriteJson", () => ({ safeWriteJson: vi.fn().mockResolvedValue(undefined), })) +// The store's cross-process task guard and advisory file lock need a real +// filesystem, which this spec stubs out. The store behavior under test does +// not depend on them. +vi.mock("../../../utils/fileLock", async () => ({ + ...(await vi.importActual("../../../utils/fileLock")), + acquireFileLock: vi.fn().mockResolvedValue({ + release: vi.fn().mockResolvedValue(undefined), + isCompromised: vi.fn().mockReturnValue(false), + }), +})) + vi.mock("@modelcontextprotocol/sdk/types.js", () => ({ CallToolResultSchema: {}, ListResourcesResultSchema: {}, diff --git a/src/eslint-suppressions.json b/src/eslint-suppressions.json index 93741e9174..73aa87e3e1 100644 --- a/src/eslint-suppressions.json +++ b/src/eslint-suppressions.json @@ -1686,7 +1686,7 @@ }, "utils/__tests__/safeWriteJson.test.ts": { "@typescript-eslint/no-explicit-any": { - "count": 27 + "count": 26 } }, "utils/__tests__/shell.spec.ts": { @@ -1716,7 +1716,7 @@ }, "utils/safeWriteJson.ts": { "@typescript-eslint/no-explicit-any": { - "count": 4 + "count": 2 } }, "utils/tts.ts": { diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index 79d08678a0..b3f34022b9 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -158,57 +158,44 @@ describe("safeWriteJson", () => { expect(content).toEqual({ initial: "content" }) }) - test("should handle failure when renaming filePath to tempBackupFilePath (filePath exists)", async () => { + test("should leave the original file in place when the commit rename fails", async () => { const initialData = { message: "Initial content, should remain" } const newData = { message: "New content, should not be written" } // Overwrite the pre-created file with specific initial data await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + // The replacement is a single rename, so its failure leaves the + // original file untouched — the target is never missing. vi.mocked(fs.rename).mockImplementationOnce(async () => { - throw new Error("Rename to backup failed") + throw new Error("Rename from temp to final failed") }) - await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename to backup failed") + await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename from temp to final failed") - // Verify the original file still exists with initial content const content = await readFileContent(currentTestFilePath) expect(content).toEqual(initialData) }) - test("should handle failure when renaming tempNewFilePath to filePath (filePath exists, backup succeeded)", async () => { - const initialData = { message: "Initial content, should be restored" } - const newData = { message: "New content" } - - // Overwrite the pre-created file with specific initial data - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + test("should remove leftover temp and backup orphans for the target before writing", async () => { + const orphanNew = path.join(tempDir, ".test-file.json.new_123_abc.tmp") + const orphanBackup = path.join(tempDir, ".test-file.json.bak_456_def.tmp") + const orphanOtherTarget = path.join(tempDir, ".other-file.json.new_789_ghi.tmp") + await fs.writeFile(orphanNew, "{}") + await fs.writeFile(orphanBackup, "{}") + await fs.writeFile(orphanOtherTarget, "{}") - // Track rename calls - let renameCallCount = 0 - - // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn - vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { - renameCallCount++ - if (renameCallCount === 1) { - // First call: filePath -> tempBackupFilePath (should succeed) - return fsPromisesActuals.rename!(oldPath, newPath) - } else if (renameCallCount === 2) { - // Second call: tempNewFilePath -> filePath (should fail) - throw new Error("Rename from temp to final failed") - } else if (renameCallCount === 3) { - // Third call: tempBackupFilePath -> filePath (rollback, should succeed) - return fsPromisesActuals.rename!(oldPath, newPath) - } - // Default: use original implementation - return fsPromisesActuals.rename!(oldPath, newPath) - }) + const data = { message: "written after orphan recovery" } + await safeWriteJson(currentTestFilePath, data) - await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename from temp to final failed") + // Orphans of this target are removed while holding the lock. + await expect(fileExists(orphanNew)).resolves.toBe(false) + await expect(fileExists(orphanBackup)).resolves.toBe(false) + // Temp files of other targets are left alone. + await expect(fileExists(orphanOtherTarget)).resolves.toBe(true) - // Verify the file was restored to initial content const content = await readFileContent(currentTestFilePath) - expect(content).toEqual(initialData) + expect(content).toEqual(data) }) // Tests for directory creation functionality @@ -292,78 +279,6 @@ describe("safeWriteJson", () => { expect(content).toEqual(data) }) - test("should handle failure when deleting tempBackupFilePath (filePath exists, all renames succeed)", async () => { - const initialData = { message: "Initial content" } - const newData = { message: "Successfully written new content" } - - // Overwrite the pre-created file with specific initial data - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - - // fs.unlink is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn - vi.mocked(fs.unlink).mockImplementationOnce(async () => { - throw new Error("Failed to delete backup file") - }) - - // The write should succeed even if backup deletion fails - await safeWriteJson(currentTestFilePath, newData) - - // Verify the new content was written successfully - const content = await readFileContent(currentTestFilePath) - expect(content).toEqual(newData) - }) - - // Test for console error suppression during backup deletion - test("should suppress console.error when backup deletion fails", async () => { - const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) // Suppress console.error - const initialData = { message: "Initial" } - const newData = { message: "New" } - - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - - // fs.unlink is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn - vi.mocked(fs.unlink).mockImplementation(async (filePath: any) => { - if (filePath.toString().includes(".bak_")) { - throw new Error("Backup deletion failed") - } - return fsPromisesActuals.unlink!(filePath) - }) - - await safeWriteJson(currentTestFilePath, newData) - - // Verify console.error was called with the expected message - expect(consoleErrorSpy).toHaveBeenCalledWith(expect.stringContaining("Successfully wrote"), expect.any(Error)) - - consoleErrorSpy.mockRestore() - vi.mocked(fs.unlink).mockRestore() - }) - - // The expected error message might need to change if the mock behaves differently. - test("should handle failure when renaming tempNewFilePath to filePath (filePath initially exists)", async () => { - // currentTestFilePath exists due to beforeEach. - const initialData = { message: "Initial content" } - const newData = { message: "New content" } - - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - - // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn - let renameCallCount = 0 - vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { - renameCallCount++ - if (renameCallCount === 2) { - // Second call: tempNewFilePath -> filePath (should fail) - throw new Error("Rename failed") - } - // For all other calls, use the original implementation - return fsPromisesActuals.rename!(oldPath, newPath) - }) - - await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename failed") - - // The file should be restored to its initial content - const content = await readFileContent(currentTestFilePath) - expect(content).toEqual(initialData) - }) - test("should throw an error if an inter-process lock is already held for the filePath", async () => { vi.resetModules() // Clear module cache to ensure fresh imports for this test @@ -434,41 +349,6 @@ describe("safeWriteJson", () => { expect(vi.mocked(fs.access)).toHaveBeenCalled() }) - // Test for rollback failure scenario - test("should log error and re-throw original if rollback fails", async () => { - const initialData = { message: "Initial, should be lost if rollback fails" } - const newData = { message: "New content" } - - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - - const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) // Suppress console.error - - // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn - let renameCallCount = 0 - vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { - renameCallCount++ - if (renameCallCount === 2) { - // Second call: tempNewFilePath -> filePath (fail) - throw new Error("Primary rename failed") - } else if (renameCallCount === 3) { - // Third call: tempBackupFilePath -> filePath (rollback, also fail) - throw new Error("Rollback rename failed") - } - return fsPromisesActuals.rename!(oldPath, newPath) - }) - - // Should throw the original error, not the rollback error - await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Primary rename failed") - - // Verify console.error was called for the rollback failure - expect(consoleErrorSpy).toHaveBeenCalledWith( - expect.stringContaining("Failed to restore backup"), - expect.objectContaining({ message: "Rollback rename failed" }), - ) - - consoleErrorSpy.mockRestore() - }) - // Merge option tests test("should merge incoming data with existing file content when merge callback is provided", async () => { const initial = { a: 1, b: 2 } @@ -542,4 +422,68 @@ describe("safeWriteJson", () => { const content = await readFileContent(currentTestFilePath) expect(content).toEqual({ c: 3 }) }) + + test("should remove no orphan temp files when the lock is compromised before orphan cleanup", async () => { + vi.resetModules() + + const compromisedPath = path.join(tempDir, "compromised-cleanup-file.json") + await fsPromisesActuals.writeFile!(compromisedPath, JSON.stringify({ initial: "content" })) + + const orphanNew = path.join(tempDir, ".compromised-cleanup-file.json.new_123_abc.tmp") + const orphanBackup = path.join(tempDir, ".compromised-cleanup-file.json.bak_456_def.tmp") + await fs.writeFile(orphanNew, "{}") + await fs.writeFile(orphanBackup, "{}") + + // Simulate proper-lockfile reporting the lost lock right after + // acquisition, before any file operation runs. + vi.doMock("proper-lockfile", () => ({ + lock: (_file: string, options: { onCompromised: (error: Error) => void }) => { + options.onCompromised(new Error("Lock no longer available")) + return Promise.resolve(async () => {}) + }, + })) + + const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") + + await expect(mockedSafeWriteJson(compromisedPath, { replaced: true })).rejects.toThrow("was compromised") + + // The compromised lock must not classify any temp file as an orphan: + // without exclusion it could belong to a live peer writer. + await expect(fileExists(orphanNew)).resolves.toBe(true) + await expect(fileExists(orphanBackup)).resolves.toBe(true) + + // The write aborted, so the target kept its original content. + const content = await readFileContent(compromisedPath) + expect(content).toEqual({ initial: "content" }) + + vi.doUnmock("proper-lockfile") + }) + + test("should abort before replacing the target when the lock is compromised", async () => { + vi.resetModules() + + const compromisedPath = path.join(tempDir, "compromised-lock-file.json") + await fsPromisesActuals.writeFile!(compromisedPath, JSON.stringify({ initial: "content" })) + + // Simulate proper-lockfile detecting the lost lock right after + // acquisition: it reports compromise from its update timer instead of + // rejecting the lock() promise. + vi.doMock("proper-lockfile", () => ({ + lock: (_file: string, options: { onCompromised: (error: Error) => void }) => { + options.onCompromised(new Error("Lock no longer available")) + return Promise.resolve(async () => {}) + }, + })) + + const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") + + await expect(mockedSafeWriteJson(compromisedPath, { replaced: true })).rejects.toThrow("was compromised") + + // The write aborted before the commit rename, so the target kept its + // original content instead of being replaced without exclusion. + const content = await readFileContent(compromisedPath) + expect(content).toEqual({ initial: "content" }) + + vi.unmock("proper-lockfile") + }) }) diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts index 9563c666d0..ded0e701e9 100644 --- a/src/utils/fileLock.ts +++ b/src/utils/fileLock.ts @@ -3,26 +3,71 @@ import * as lockfile from "proper-lockfile" export const LOCK_STALE_MS = 31_000 +/** + * Retry budget for destructive mutations. The backoff outlasts + * LOCK_STALE_MS, so a lock left behind by a crashed process is broken + * within the same acquisition instead of failing the mutation. + */ +export const DESTRUCTIVE_LOCK_RETRIES = { + retries: 36, + factor: 2, + minTimeout: 100, + maxTimeout: 1_000, +} + +export interface AcquiredFileLock { + release: () => Promise + /** + * True once proper-lockfile reported the held lock compromised, for + * example when another host's stale-lock recovery removed the lock + * directory. Callers must check this before mutating the protected + * target and fail the operation instead of writing without exclusion. + */ + isCompromised: () => boolean +} + +/** + * Fail with one deterministic error when the held lock was reported + * compromised, so the caller aborts before the next mutation of the + * protected target instead of writing or deleting without exclusion. + */ +export function assertLockUsable(lock: AcquiredFileLock, targetPath: string, action: string): void { + if (lock.isCompromised()) { + throw new Error(`Lock for ${targetPath} was compromised before ${action}`) + } +} + +export interface AcquireFileLockOptions { + retries?: lockfile.LockOptions["retries"] +} + /** * Acquire the advisory lock shared by JSON writes and destructive mutations. * The target may not exist yet, so callers can use the same protocol for * creation, replacement, and deletion. */ -export function acquireFileLock(filePath: string): Promise<() => Promise> { +export function acquireFileLock(filePath: string, options?: AcquireFileLockOptions): Promise { const absoluteFilePath = path.resolve(filePath) - return lockfile.lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10_000, - realpath: false, - retries: { - retries: 5, - factor: 2, - minTimeout: 100, - maxTimeout: 1_000, - }, - onCompromised: (error) => { - console.error(`Lock at ${absoluteFilePath} was compromised:`, error) - throw error - }, - }) + let compromised = false + return lockfile + .lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10_000, + realpath: false, + retries: options?.retries ?? { + retries: 5, + factor: 2, + minTimeout: 100, + maxTimeout: 1_000, + }, + // proper-lockfile invokes this from a timer callback after + // acquisition, so a throw here never reaches the awaited operation. + // Record the state instead; callers check isCompromised before + // mutating and fail the operation themselves. + onCompromised: (error) => { + compromised = true + console.error(`Lock at ${absoluteFilePath} was compromised:`, error) + }, + }) + .then((release) => ({ release, isCompromised: () => compromised })) } diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index b20aaba405..c5be4940a4 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -3,7 +3,7 @@ import * as fsSync from "fs" import * as path from "path" import { JsonStreamStringify } from "json-stream-stringify" -import { acquireFileLock, LOCK_STALE_MS } from "./fileLock" +import { acquireFileLock, assertLockUsable, LOCK_STALE_MS } from "./fileLock" /** * Options for safeWriteJson function @@ -32,9 +32,13 @@ export interface SafeWriteJsonOptions { * Safely writes JSON data to a file. * - Creates parent directories if they don't exist * - Uses 'proper-lockfile' for inter-process advisory locking to prevent concurrent writes to the same path. - * - Writes to a temporary file first. - * - If the target file exists, it's backed up before being replaced. - * - Attempts to roll back and clean up in case of errors. + * - Removes leftover temp files for this target left by a crashed writer. + * - Writes to a temporary file first, then replaces the target with one + * atomic rename, so the target is never missing between the two states. + * - Cleans up the temporary file in case of errors. + * - Aborts before orphan cleanup and before replacing the target when the + * advisory lock was compromised, so it never writes or deletes without + * exclusion. * - Supports pretty-printing with indentation while maintaining streaming efficiency. * * @param {string} filePath - The absolute path to the target file. @@ -45,43 +49,35 @@ export interface SafeWriteJsonOptions { async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJsonOptions): Promise { const absoluteFilePath = path.resolve(filePath) - let releaseLock = async () => {} // Initialized to a no-op // For directory creation const dirPath = path.dirname(absoluteFilePath) - // Ensure directory structure exists with improved reliability + // Ensure directory structure exists try { - // Create directory with recursive option await fs.mkdir(dirPath, { recursive: true }) - - // Verify directory exists after creation attempt await fs.access(dirPath) - } catch (dirError: any) { + } catch (dirError: unknown) { console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) throw dirError } - // Acquire the lock before any file operations + // Acquire the lock before any file operations. On failure the release + // helper stays unused and the error propagates. + let lock: Awaited> try { - releaseLock = await acquireFileLock(absoluteFilePath) + lock = await acquireFileLock(absoluteFilePath) } catch (lockError) { - // If lock acquisition fails, we throw immediately. - // The releaseLock remains a no-op, so the finally block in the main file operations - // try-catch-finally won't try to release an unacquired lock if this path is taken. console.error(`Failed to acquire lock for ${absoluteFilePath}:`, lockError) - // Propagate the lock acquisition error throw lockError } - // Variables to hold the actual paths of temp files if they are created. + // Path of the temporary file while it exists, so the error path can clean it up. let actualTempNewFilePath: string | null = null - let actualTempBackupFilePath: string | null = null try { // If a merge callback was provided, read the current file under the lock - // and let the caller merge before we write. Must be inside try/finally - // so a throwing merge still releases the lock. + // and let the caller merge before we write. if (options?.merge) { let existing: unknown = null try { @@ -96,108 +92,47 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso data = options.merge(existing, data) } + // A compromised lock no longer excludes a peer writer, so its temp + // files are not orphans. Abort before discovery and again before + // each removal instead of deleting a live writer's temp files. + assertLockUsable(lock, absoluteFilePath, "orphan cleanup") + await removeLeftoverTempFiles(dirPath, path.basename(absoluteFilePath), () => + assertLockUsable(lock, absoluteFilePath, "orphan cleanup"), + ) + // Step 1: Write data to a new temporary file. actualTempNewFilePath = path.join( - path.dirname(absoluteFilePath), + dirPath, `.${path.basename(absoluteFilePath)}.new_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, ) await _streamDataToFile(actualTempNewFilePath, data, options?.prettyPrint) - // Step 2: Check if the target file exists. If so, rename it to a backup path. - try { - // Check for target file existence - await fs.access(absoluteFilePath) - // Target exists, create a backup path and rename. - actualTempBackupFilePath = path.join( - path.dirname(absoluteFilePath), - `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, - ) - await fs.rename(absoluteFilePath, actualTempBackupFilePath) - } catch (accessError: any) { - // Explicitly type accessError - if (accessError.code !== "ENOENT") { - // An error other than "file not found" occurred during access check. - throw accessError - } - // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. - } + // A compromised lock means another host may be mutating the target. + // Fail here instead of replacing the target without exclusion. + assertLockUsable(lock, absoluteFilePath, "commit") - // Step 3: Rename the new temporary file to the target file path. - // This is the main "commit" step. + // Step 2: Replace the target with one atomic rename. The target holds + // either the old content or the new content at every instant, so a + // crash cannot leave it missing. await fs.rename(actualTempNewFilePath, absoluteFilePath) - - // If we reach here, the new file is successfully in place. - // The original actualTempNewFilePath is now the main file, so we shouldn't try to clean it up as "temp". - // Mark as "used" or "committed" actualTempNewFilePath = null - - // Step 4: If a backup was created, attempt to delete it. - if (actualTempBackupFilePath) { - try { - await fs.unlink(actualTempBackupFilePath) - // Mark backup as handled - actualTempBackupFilePath = null - } catch (unlinkBackupError) { - // Log this error, but do not re-throw. The main operation was successful. - // actualTempBackupFilePath remains set, indicating an orphaned backup. - console.error( - `Successfully wrote ${absoluteFilePath}, but failed to clean up backup ${actualTempBackupFilePath}:`, - unlinkBackupError, - ) - } - } } catch (originalError) { - console.error(`Operation failed for ${absoluteFilePath}: [Original Error Caught]`, originalError) - - const newFileToCleanupWithinCatch = actualTempNewFilePath - const backupFileToRollbackOrCleanupWithinCatch = actualTempBackupFilePath - - // Attempt rollback if a backup was made - if (backupFileToRollbackOrCleanupWithinCatch) { - try { - await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) - // Mark as handled, prevent later unlink of this path - actualTempBackupFilePath = null - } catch (rollbackError) { - // actualTempBackupFilePath (outer scope) remains pointing to backupFileToRollbackOrCleanupWithinCatch - console.error( - `[Catch] Failed to restore backup ${backupFileToRollbackOrCleanupWithinCatch} to ${absoluteFilePath}:`, - rollbackError, - ) - } - } + console.error(`Operation failed for ${absoluteFilePath}:`, originalError) - // Cleanup the .new file if it exists - if (newFileToCleanupWithinCatch) { + if (actualTempNewFilePath) { try { - await fs.unlink(newFileToCleanupWithinCatch) + await fs.unlink(actualTempNewFilePath) } catch (cleanupError) { - console.error( - `[Catch] Failed to clean up temporary new file ${newFileToCleanupWithinCatch}:`, - cleanupError, - ) + console.error(`[Catch] Failed to clean up temporary new file ${actualTempNewFilePath}:`, cleanupError) } } - // Cleanup the .bak file if it still needs to be (i.e., wasn't successfully restored) - if (actualTempBackupFilePath) { - try { - await fs.unlink(actualTempBackupFilePath) - } catch (cleanupError) { - console.error( - `[Catch] Failed to clean up temporary backup file ${actualTempBackupFilePath}:`, - cleanupError, - ) - } - } - throw originalError // This MUST be the error that rejects the promise. + throw originalError } finally { // Release the lock in the main finally block. try { - // releaseLock will be the actual unlock function if lock was acquired, - // or the initial no-op if acquisition failed. - await releaseLock() + await lock.release() } catch (unlockError) { // Do not re-throw here, as the originalError from the try/catch (if any) is more important. console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) @@ -205,6 +140,43 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } +/** + * Remove leftover `..new_*.tmp` and `..bak_*.tmp` files. + * Safe while the caller holds the advisory lock for `targetBasename`. + * @param dirPath The directory holding the target file. + * @param targetBasename The target file's base name. + * @param assertLockUsable Called before each removal so the caller can + * abort while the lock is compromised. + */ +async function removeLeftoverTempFiles( + dirPath: string, + targetBasename: string, + assertLockUsable: () => void, +): Promise { + const newPrefix = `.${targetBasename}.new_` + const backupPrefix = `.${targetBasename}.bak_` + let entries: string[] + try { + entries = await fs.readdir(dirPath) + } catch { + return + } + for (const entry of entries) { + if (!entry.endsWith(".tmp")) { + continue + } + if (!entry.startsWith(newPrefix) && !entry.startsWith(backupPrefix)) { + continue + } + assertLockUsable() + try { + await fs.unlink(path.join(dirPath, entry)) + } catch (error) { + console.error(`Failed to clean up leftover temp file ${entry} for ${targetBasename}:`, error) + } + } +} + /** * Helper function to stream JSON data to a file. * @param targetPath The path to write the stream to. From 3c8a82bb61c488c3f9956f5b94b35f2053ac5343 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Wed, 23 Sep 2026 03:51:04 +0000 Subject: [PATCH 08/23] fix(lifecycle): scope branch to exact-action settlement and merge main Drop the broad deletion-serialization work: the shared task guard (taskIoGuard, taskPathSafety, fileLock), safeWriteJson refactors, storage path policy, and writer coordination across messages, tools, and webview handlers. Keep only the rejected create_subtask exact-action settlement: the LifecycleTransitionError reducer, the TaskHistoryStore.clearPendingActionIfMatching compare-and-clear, typed rejection handling in ClineProvider, pre-replay settlement in Task, and their focused tests, model witnesses, and architecture documentation. Merge local main to carry the #1678 repeated-cancel behavior. --- .../architecture/task-lifecycle-gap-report.md | 128 +++++----- docs/architecture/task-lifecycle-model.md | 2 +- .../task-lifecycle-remediation-blocks.md | 70 +++--- src/__tests__/delegation-concurrent.spec.ts | 11 - src/core/task-persistence/TaskHistoryStore.ts | 211 ++++------------ .../TaskHistoryStore.crossInstance.spec.ts | 18 +- .../TaskHistoryStore.guardCompromise.spec.ts | 149 ----------- .../TaskHistoryStore.realConcurrency.spec.ts | 236 ------------------ .../task/__tests__/Task.persistence.spec.ts | 42 ++++ src/core/webview/ClineProvider.ts | 19 +- .../ClineProvider.sticky-profile.spec.ts | 11 - .../ClineProvider.taskHistory.spec.ts | 11 - src/eslint-suppressions.json | 4 +- src/utils/__tests__/safeWriteJson.test.ts | 224 ++++++++++------- src/utils/fileLock.ts | 73 ------ src/utils/safeWriteJson.ts | 196 +++++++++------ 16 files changed, 470 insertions(+), 935 deletions(-) delete mode 100644 src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts delete mode 100644 src/utils/fileLock.ts diff --git a/docs/architecture/task-lifecycle-gap-report.md b/docs/architecture/task-lifecycle-gap-report.md index b22288318b..66db3eb608 100644 --- a/docs/architecture/task-lifecycle-gap-report.md +++ b/docs/architecture/task-lifecycle-gap-report.md @@ -6,10 +6,6 @@ This report inventories Zoo Code task lifecycle state, mutation, persistence, sc The audit covers tracked TypeScript, JSON, YAML, and Markdown under `packages/`, `src/`, `apps/cli`, `apps/vscode-e2e`, `scripts/`, `.github/workflows`, and `docs/architecture`. It traces production symbols to bounded models, focused tests, extension-host E2E, and CI entry points. -Inventory completeness is not composed verification. A closed symbol and ownership inventory lists every boundary with its local evidence. It does not prove that persistence, restart, replay, UI approval, scheduling, and external effects compose into one correct protocol. Local evidence classes support an end-to-end claim only through the joint-checker or trace-validation work tracked by `LIFE-GAP-015`. - -Issue [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714) is recorded evidence for this distinction. Every involved symbol (`setPendingTaskAction`, `resumeTaskFromHistory`, `resumePendingTaskAction`, `delegateParentAndOpenChild`, `delegateTaskToChild`) was inventoried with mapped ownership, yet the composed loop across restart, pending-action replay, auto-approval, and an authoritative transition rejection was found by incident report, not by inventory. The #1714 changes add the `settleRejectedCreateSubtaskAction` reducer, `stage` and `settle-rejected` model actions, four pending-action witnesses, and matching focused tests. Their direct model coverage is specified in [the model suite](./task-lifecycle-model.md). They address the reported loop for the authoritative-rejection path; LIFE-GAP-039 and LIFE-GAP-040 record the remaining crash-window, startup-repair, and replay-bound obligations. - “Exhaustive” means exhaustive over the repository paths, symbol families, and search terms listed here at the audited commit. It does not include ignored/generated output, deployment branch-protection settings, runtime telemetry, dynamically constructed names that evade text search, or behavior in dependencies. GitHub issue links are historical provenance only; stable `LIFE-GAP-*` IDs own the active burn-down. ## Methodology and audit criteria @@ -167,66 +163,62 @@ CI runs lint, typecheck, and `pnpm lifecycle:model-check` in `.github/workflows/ - Parser replay fixes two scopes, one raw index, and local action order. Transport transforms and arbitrary malformed histories are excluded. - Focused tests and E2E are representative histories, not exhaustive interleavings. - No public-runtime telemetry or production traces were available for trace validation. -- Scheduler liveness is an intentional exclusion. No checker defines fairness for eventual admission, queue drain, cleanup, or retry progress. Proving those properties requires a separate temporal-model audit with explicit fairness assumptions; this report does not claim liveness coverage. -- External tool-effect idempotency is an intentional exclusion. Lifecycle records track call and action identity (`LIFE-GAP-038`), but no model, test, or claim in this audit covers whether a replayed or re-executed tool call repeats effects on the editor, filesystem, or external services. That boundary needs a separate future audit. ## Ranked GAP register Severity reflects plausible data loss, ownership corruption, permission/context errors, or stuck work. Confidence reflects direct source evidence, deterministic witness, or inference. -| ID | Severity | Confidence | Gap and production impact | Witness/reproducer | Dependencies | Objective closure criteria | -| ------------ | ----------- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-GAP-001 | Critical | High | Stale cross-host completion can clear a newer handoff and orphan its child. | Exact shortest witness in shared-store checker. | Disk-authoritative ownership guard or generation. | Deterministic two-store test passes; authoritative child ID is revalidated under disk lock; witness becomes universal invariant. | -| LIFE-GAP-002 | High | High | Stale message save can restore abandoned lineage. | Exact shortest stale-save/abandon witness. | Lifecycle-field ownership, tombstone, or generation. | Stale save cannot alter detached lineage in production test or bounded model; witness promoted. | -| LIFE-GAP-003 | High | High | Legacy dequeue-before-submit can lose queued user feedback on submission failure. | `Task.processQueuedMessages` removes before async submit. | Standardize claim/persist/ack path. | Every queue consumer claims first, removes after durable acceptance, releases on failure; failure test retains message. | -| LIFE-GAP-004 | High | High | Pair writes are not cross-host/crash atomic; partial lifecycle state is observable. | Modeled second-write failure landmark. | WAL/repair intent or explicitly idempotent recovery per pair operation. | Crash injection at each write boundary converges to a documented legal state. | -| LIFE-GAP-005 | High | High | Delegation creates child state and commits parent ownership in separate failure domains. | Provider rollback tests cover selected failures. | Idempotent operation ID or repair intent. | Failure injection before/after every persistence/publication step restores parent or completes delegation without orphan state. | -| LIFE-GAP-006 | High | Medium-high | Parent completion messages can persist before lifecycle completion commit fails. | Source ordering in `reopenParentFromDelegation`. | Transaction/intent or reorder with recovery. | Inject pair-write failure and prove no stale result context, or prove replay safely completes the lifecycle. The composed obligation spans the message write, both pair files, and child pending-action state; per-layer passes do not compose. Closure must enumerate one phase table across those writes and state the recovery interaction with resumed pending-action replay (012). | -| LIFE-GAP-007 | Medium | High | Three confirmed task-local mode readers are fixed and production-tested, but the repository-wide reader inventory and enforceable task-local boundary remain incomplete. | Merged focused tests cover environment details, built-in validation, and custom-tool execution; the pure checker proves only divergence and selector storage. | Reader inventory and task-local API boundary. | Every mode-sensitive reader is classified; required task-local consumers have divergent production tests in both permission directions; checker/docs distinguish executed readers from proxy evidence. | -| LIFE-GAP-008 | High | High | Responses API argument-only deltas may be dropped; absent indices can alias calls. | Transform requires ID/name and defaults index to zero. | Provider transform correlation design. | Multi-delta and concurrent index-less production tests reconstruct isolated calls; parser model includes mapped transform events. | -| LIFE-GAP-009 | Medium | High | Async lifecycle event listeners have no awaited settlement contract. | Async listeners registered on Node EventEmitter. | Classify notification versus barrier events. | Barrier side effects move to awaited methods; notification listeners have contained rejection tests and documented ordering. | -| LIFE-GAP-010 | Medium | Medium-high | Detached usage drain can write after abort, replacement, delegation, or a newer request generation. | Background iterator mutates/persists without generation guard. | Request-generation ownership. | Late drain may account valid usage but cannot mutate stale UI/message/lifecycle state; controlled delayed-stream test passes. | -| LIFE-GAP-011 | Medium | High | Completion model excludes terminal status persistence and downstream public consumers. | Model ends at emission readiness. | Consumer inventory and contract. | Consequential consumers are enumerated; required status/metadata ordering has production refinement tests. | -| LIFE-GAP-012 | Medium | High | No persisted attempt/generation distinguishes delayed pre-interruption completion from valid post-resume completion of the same child. | Documented model exclusion. | Persisted generation token. | Reducer/store/API/model reject stale generation while accepting resumed generation; restart E2E covers it. The token must also cover staged pending-action replay, because `resumePendingTaskAction` re-executes a staged action after restart with no attempt counter or bound, as #1714 demonstrated. Settlement of a stale action must record the generation so a later legitimate action is accepted. | -| LIFE-GAP-013 | Medium | High | Task lifecycle status vocabulary is copied across schema, task metadata, Task, CLI, and history reader. | Literal union inventory. | Shared exported schema-derived type. | Consumers import one owner; CI/static check rejects incompatible local copies. | -| LIFE-GAP-014 | Medium | High | The serial production contract relies on singular reducer ownership and a default one-permit provider scheduler; optional fan-out must not be mistaken for baseline coverage. | The production-backed lifecycle model rejects multiple active awaited children, while the optional fan-out model has no production imports or E2E. | Serial baseline ratchet; separately ticketed fan-out decision. | Baseline: close cross-host violations, assert provider scheduler capacity and serial ordering, and keep fan-out outside baseline CI. Optional fan-out: implement adapters/E2E before reclassification. | -| LIFE-GAP-015 | Medium | High | Independent checks do not establish end-to-end refinement. | Six baseline state spaces plus one optional fan-out state space remain disjoint. | Boundary mappings and tractable joint bounds. | Add joint checker/trace validation for each cross-model claim, or keep every claim explicitly local. | -| LIFE-GAP-016 | Medium | High | Traceability is documentary and can drift from scripts, symbols, tests, and CI. | Shared-store scenario count has drifted in documentation. | Machine-readable manifest/checker summaries. | CI validates stable IDs, model membership, bounds, symbol/test paths, and workflow invocation. | -| LIFE-GAP-017 | Medium | High | Store cache records are exposed without cloning; external mutation may bypass locking. | `get`/`getAll` return cached objects. | Immutability/read API decision. | Freeze/clone records or prove callers cannot mutate; mutation regression test. | -| LIFE-GAP-018 | Medium | High | Ordinary history-file reads cast JSON instead of applying the shared schema. | Store reconciliation/read path. | Validation/quarantine policy. | Malformed records are rejected or quarantined deterministically with tests and recovery documentation. | -| LIFE-GAP-019 | Medium | High | Generic task IDs lack the importer’s explicit path-safety validation. | Importer validates IDs; generic paths interpolate IDs. | Shared safe-ID boundary. | Every filesystem task ID passes one validator; traversal and separator tests cover all entry points. | -| LIFE-GAP-020 | Medium | Medium-high | Watch/reconcile convergence is eventual and failure-tolerant, not coherent. | Debounced watcher plus periodic scan. | Version/notification or documented eventual contract. | Define stale-read window and convergence property; multi-host test covers missed watcher event and concurrent update. Convergence must define a pending-action rule: reconciliation and startup repair change status and lineage without any pending-action decision, so a converged record can keep an action whose request no longer exists (the #1714 precondition survived startup reconciliation). The stale-read window must also cover the second settlement write that follows an authoritative rejection (039). | -| LIFE-GAP-021 | Medium | High | Store disposal does not await queued writes. | Synchronous `dispose` stops watcher/timer only. | Async drain/close contract. | Disposal awaits or explicitly cancels writes; no post-dispose writes in deterministic test. | -| LIFE-GAP-022 | Medium | High | Public clear and webview clear use different delegated-child semantics. | API uses eviction; webview removes directly. | One clear contract. | All ingress paths converge on the same lifecycle transition and tests assert identical persisted results. | -| LIFE-GAP-023 | Medium | High | History deletion can be resurrected after swallowed unlink failure. | Cache removal precedes best-effort unlink. | Tombstone or surfaced failure/retry. | Inject unlink failure and prove item stays deleted or operation reports failure without false success. | -| LIFE-GAP-024 | Medium | High | Parser cleanup depends on normal finalization or garbage collection. | Weak scope maps lack universal request `finally`. | Request-level cleanup owner. | Abort/error/success all clear active parser state in production integration tests. | -| LIFE-GAP-025 | Medium | High | Abort listeners may accumulate during successful stream chunks. | Per-chunk listener removed only by abort. | Settle-time listener cleanup. | Long stream keeps bounded listener count and removes listeners on both race outcomes. | -| LIFE-GAP-026 | Medium | Medium-high | Usage-drain timeout cannot interrupt a permanently pending `iterator.next()`. | Elapsed time checked before await. | Deadline race/abort. | Hung iterator settles drain within wall-clock bound in fake-timer test. | -| LIFE-GAP-027 | Medium | High | Task-level delegation listeners are untyped/dead while provider listeners own the same public events. | `src/extension/api.ts` registers both paths. | Single typed event owner. | Remove duplicate/dead listeners or define one source; public API test proves exactly-once emission. | -| LIFE-GAP-028 | Low-medium | High | `TaskSpawned` payload semantics differ across task/provider/public surfaces. | Child-only, ambiguous task ID, and parent+child forms. | Event contract normalization. | Payloads use explicit names and adapters are type-checked with compatibility tests. | -| LIFE-GAP-029 | Low-medium | High | `Task.taskStatus` and `TaskRegistry.getRunning` are projections, not scheduler/persistence truth. | Ask markers and abort flags only. | Naming/contract clarification. | Rename or document exact predicates; callers stop using them as stronger lifecycle evidence. | -| LIFE-GAP-030 | Low-medium | Medium | `Task.run()` may resolve immediately if another path already started the task. | `_started` short-circuit versus scheduler callback. | Single start owner. | Scheduler-facing start returns the actual run promise; duplicate-start test proves settlement identity. | -| LIFE-GAP-031 | Low-medium | High | Queue state is memory-only and cleared on disposal. | `MessageQueueService.dispose`. | Product durability decision. | Document intentional loss or persist claims/messages with restart tests. The durability decision must state the restart composition: which queue states can coexist with a staged pending action after restart, and whether action replay can consume feedback that was never acknowledged. | -| LIFE-GAP-032 | Low-medium | Medium | Webview abandonment handler exists without a confirmed production UI sender. | Protocol/handler search only. | Reachability decision. | Add supported sender/E2E or remove/deprecate unreachable command. | -| LIFE-GAP-033 | Low-medium | High | Resume ingress differs in awaiting and error propagation. | Webview/API/IPC adapters diverge. | Shared resume operation contract. | Contract tests compare result/error/publication semantics for each surface. | -| LIFE-GAP-034 | Low | High | Bounds and model metadata are handwritten and not mechanically synchronized. | Constants, prose, and console summaries duplicate values. | Machine-readable checker metadata. | CI compares emitted metadata with docs/manifest and rejects undocumented bound/action/landmark changes. | -| LIFE-GAP-035 | Medium | High | Tool-originated child initialization lacks a complete durable ownership contract; initial todo state is the confirmed witness and can disappear after rehydration. | Create a child with explicit initial todos, switch or restart before `update_todo_list`, then reopen it; restoration finds no todo message and yields an empty list. | Canonical durable task-ID-scoped child-initialization owner and publication contract. | Inventory every `new_task`-originated child field; persist required initial state before visibility/run; restore deep-equal independent state across switching, interrupted resume, checkpoint restore, and fresh-host restart; preserve later-update and explicit-empty precedence; add constructor deep-copy, persistence, webview scoping, and E2E witnesses. | -| LIFE-GAP-036 | High | High | Interactive todo approval edit state is process-global and uncorrelated; one task's delayed edit can be consumed by another task's pending approval. | Start approvals for tasks A and B, send A's edited list through `setPendingTodoList`, then resolve B; B reads the shared `approvedTodoList`. | Task/action/tool-call-correlated approval state and webview protocol. | Carry task ID and action/tool-call ID through proposal, webview edit, approval, cancellation, and settlement; reject stale/mismatched edits; deep-clone inputs; test two interleaved approvals, denial, cancellation, task switch, and delayed edits. Correlation must extend to the staged pending action: `NewTaskTool` persists the action before approval, so the protocol must bind `pendingAction.actionId` to the approving task and tool call and reject cross-task edits before staging; denied, cancelled, and errored approvals need durable settlement rules equal to the rejected-delegation settlement; deep-copy extends to `pendingAction.todos` at staging. | -| LIFE-GAP-037 | Medium-high | High | Singleton tool handlers share partial presentation state across calls/tasks, so interleaved paths can cause false or missed stabilization. | Interleave A:`x`, B:`y`, A:`x` or A:`x`, B:`x` through one handler's `lastSeenPartialPath`. | Per-call handler state keyed by task and tool-call identity. | Isolate partial state by `(taskId, toolCallId)` or handler instance; prove independent stabilization and cleanup after success, malformed finalization, rejection, cancellation, abandonment, and incomplete streams. | -| LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. The bijection must include the staged pending action: `pendingAction.actionId` derives from `sanitizeToolUseId(toolCallId)`, so non-injective sanitization can make a settlement targeted at one raw call clear another call's staged action; the model's stale-settlement witness depends on collision-free action IDs. | -| LIFE-GAP-039 | Medium | High | Settlement after an authoritative pending-action rejection remains a second store write, but an interrupted task is now a non-replayable durable marker: restart must settle its exact pending `create_subtask` action before replay, and a failed recovery write stops without creating another child. A crash can still leave the rejected action staged and require recovery on every restart. | Source: the `delegateParentAndOpenChild` settlement block logs settlement errors and skips in-process restoration; `Task.resumeTaskFromHistory` retries exact-action settlement for interrupted create-subtask actions and surfaces failure before replay. `pendingAction` still carries no persisted attempt/backoff count. | Durable operation intent (P2-004); attempt/generation identity (012). | Persist settlement intent or attempt/backoff state so recovery converges without depending on storage becoming available; fault injection across every cut preserves the no-replay invariant and proves bounded recovery work. | -| LIFE-GAP-040 | Medium | Medium-high | Startup topology repair reconciles status and lineage but has no pending-action rule. `reconcileDelegationState`, `repairActiveDelegation`, and repair-intent replay can leave a staged action on a repaired record, and the two interrupt paths disagree on whether `pendingAction` survives. | #1714 reproduced the loop precondition through startup reconciliation of a persisted active child; `DelegationRepairIntent` guards and targets carry no pending-action field; `applyDelegationRepairIntent` spreads records without a pending-action decision. | Durable operation intent (P2-004); correlated action identity (036, 038). | Every repair outcome (orphaned delegation, orphaned active child, completed child, intent replay, quarantine) defines preserve, settle, or drop for pending actions; deterministic repair tests and one fresh-host restart test prove the #1714 precondition cannot survive repair; both interrupt paths persist or settle the action consistently. | +| ID | Severity | Confidence | Gap and production impact | Witness/reproducer | Dependencies | Objective closure criteria | +| ------------ | ----------- | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| LIFE-GAP-001 | Critical | High | Stale cross-host completion can clear a newer handoff and orphan its child. | Exact shortest witness in shared-store checker. | Disk-authoritative ownership guard or generation. | Deterministic two-store test passes; authoritative child ID is revalidated under disk lock; witness becomes universal invariant. | +| LIFE-GAP-002 | High | High | Stale message save can restore abandoned lineage. | Exact shortest stale-save/abandon witness. | Lifecycle-field ownership, tombstone, or generation. | Stale save cannot alter detached lineage in production test or bounded model; witness promoted. | +| LIFE-GAP-003 | High | High | Legacy dequeue-before-submit can lose queued user feedback on submission failure. | `Task.processQueuedMessages` removes before async submit. | Standardize claim/persist/ack path. | Every queue consumer claims first, removes after durable acceptance, releases on failure; failure test retains message. | +| LIFE-GAP-004 | High | High | Pair writes are not cross-host/crash atomic; partial lifecycle state is observable. | Modeled second-write failure landmark. | WAL/repair intent or explicitly idempotent recovery per pair operation. | Crash injection at each write boundary converges to a documented legal state. | +| LIFE-GAP-005 | High | High | Delegation creates child state and commits parent ownership in separate failure domains. | Provider rollback tests cover selected failures. | Idempotent operation ID or repair intent. | Failure injection before/after every persistence/publication step restores parent or completes delegation without orphan state. | +| LIFE-GAP-006 | High | Medium-high | Parent completion messages can persist before lifecycle completion commit fails. | Source ordering in `reopenParentFromDelegation`. | Transaction/intent or reorder with recovery. | Inject pair-write failure and prove no stale result context, or prove replay safely completes the lifecycle. | +| LIFE-GAP-007 | Medium | High | Three confirmed task-local mode readers are fixed and production-tested, but the repository-wide reader inventory and enforceable task-local boundary remain incomplete. | Merged focused tests cover environment details, built-in validation, and custom-tool execution; the pure checker proves only divergence and selector storage. | Reader inventory and task-local API boundary. | Every mode-sensitive reader is classified; required task-local consumers have divergent production tests in both permission directions; checker/docs distinguish executed readers from proxy evidence. | +| LIFE-GAP-008 | High | High | Responses API argument-only deltas may be dropped; absent indices can alias calls. | Transform requires ID/name and defaults index to zero. | Provider transform correlation design. | Multi-delta and concurrent index-less production tests reconstruct isolated calls; parser model includes mapped transform events. | +| LIFE-GAP-009 | Medium | High | Async lifecycle event listeners have no awaited settlement contract. | Async listeners registered on Node EventEmitter. | Classify notification versus barrier events. | Barrier side effects move to awaited methods; notification listeners have contained rejection tests and documented ordering. | +| LIFE-GAP-010 | Medium | Medium-high | Detached usage drain can write after abort, replacement, delegation, or a newer request generation. | Background iterator mutates/persists without generation guard. | Request-generation ownership. | Late drain may account valid usage but cannot mutate stale UI/message/lifecycle state; controlled delayed-stream test passes. | +| LIFE-GAP-011 | Medium | High | Completion model excludes terminal status persistence and downstream public consumers. | Model ends at emission readiness. | Consumer inventory and contract. | Consequential consumers are enumerated; required status/metadata ordering has production refinement tests. | +| LIFE-GAP-012 | Medium | High | No persisted attempt/generation distinguishes delayed pre-interruption completion from valid post-resume completion of the same child. | Documented model exclusion. | Persisted generation token. | Reducer/store/API/model reject stale generation while accepting resumed generation; restart E2E covers it. | +| LIFE-GAP-013 | Medium | High | Task lifecycle status vocabulary is copied across schema, task metadata, Task, CLI, and history reader. | Literal union inventory. | Shared exported schema-derived type. | Consumers import one owner; CI/static check rejects incompatible local copies. | +| LIFE-GAP-014 | Medium | High | The serial production contract relies on singular reducer ownership and a default one-permit provider scheduler; optional fan-out must not be mistaken for baseline coverage. | The production-backed lifecycle model rejects multiple active awaited children, while the optional fan-out model has no production imports or E2E. | Serial baseline ratchet; separately ticketed fan-out decision. | Baseline: close cross-host violations, assert provider scheduler capacity and serial ordering, and keep fan-out outside baseline CI. Optional fan-out: implement adapters/E2E before reclassification. | +| LIFE-GAP-015 | Medium | High | Independent checks do not establish end-to-end refinement. | Six baseline state spaces plus one optional fan-out state space remain disjoint. | Boundary mappings and tractable joint bounds. | Add joint checker/trace validation for each cross-model claim, or keep every claim explicitly local. | +| LIFE-GAP-016 | Medium | High | Traceability is documentary and can drift from scripts, symbols, tests, and CI. | Shared-store scenario count has drifted in documentation. | Machine-readable manifest/checker summaries. | CI validates stable IDs, model membership, bounds, symbol/test paths, and workflow invocation. | +| LIFE-GAP-017 | Medium | High | Store cache records are exposed without cloning; external mutation may bypass locking. | `get`/`getAll` return cached objects. | Immutability/read API decision. | Freeze/clone records or prove callers cannot mutate; mutation regression test. | +| LIFE-GAP-018 | Medium | High | Ordinary history-file reads cast JSON instead of applying the shared schema. | Store reconciliation/read path. | Validation/quarantine policy. | Malformed records are rejected or quarantined deterministically with tests and recovery documentation. | +| LIFE-GAP-019 | Medium | High | Generic task IDs lack the importer’s explicit path-safety validation. | Importer validates IDs; generic paths interpolate IDs. | Shared safe-ID boundary. | Every filesystem task ID passes one validator; traversal and separator tests cover all entry points. | +| LIFE-GAP-020 | Medium | Medium-high | Watch/reconcile convergence is eventual and failure-tolerant, not coherent. | Debounced watcher plus periodic scan. | Version/notification or documented eventual contract. | Define stale-read window and convergence property; multi-host test covers missed watcher event and concurrent update. | +| LIFE-GAP-021 | Medium | High | Store disposal does not await queued writes. | Synchronous `dispose` stops watcher/timer only. | Async drain/close contract. | Disposal awaits or explicitly cancels writes; no post-dispose writes in deterministic test. | +| LIFE-GAP-022 | Medium | High | Public clear and webview clear use different delegated-child semantics. | API uses eviction; webview removes directly. | One clear contract. | All ingress paths converge on the same lifecycle transition and tests assert identical persisted results. | +| LIFE-GAP-023 | Medium | High | History deletion can be resurrected after swallowed unlink failure. | Cache removal precedes best-effort unlink. | Tombstone or surfaced failure/retry. | Inject unlink failure and prove item stays deleted or operation reports failure without false success. | +| LIFE-GAP-024 | Medium | High | Parser cleanup depends on normal finalization or garbage collection. | Weak scope maps lack universal request `finally`. | Request-level cleanup owner. | Abort/error/success all clear active parser state in production integration tests. | +| LIFE-GAP-025 | Medium | High | Abort listeners may accumulate during successful stream chunks. | Per-chunk listener removed only by abort. | Settle-time listener cleanup. | Long stream keeps bounded listener count and removes listeners on both race outcomes. | +| LIFE-GAP-026 | Medium | Medium-high | Usage-drain timeout cannot interrupt a permanently pending `iterator.next()`. | Elapsed time checked before await. | Deadline race/abort. | Hung iterator settles drain within wall-clock bound in fake-timer test. | +| LIFE-GAP-027 | Medium | High | Task-level delegation listeners are untyped/dead while provider listeners own the same public events. | `src/extension/api.ts` registers both paths. | Single typed event owner. | Remove duplicate/dead listeners or define one source; public API test proves exactly-once emission. | +| LIFE-GAP-028 | Low-medium | High | `TaskSpawned` payload semantics differ across task/provider/public surfaces. | Child-only, ambiguous task ID, and parent+child forms. | Event contract normalization. | Payloads use explicit names and adapters are type-checked with compatibility tests. | +| LIFE-GAP-029 | Low-medium | High | `Task.taskStatus` and `TaskRegistry.getRunning` are projections, not scheduler/persistence truth. | Ask markers and abort flags only. | Naming/contract clarification. | Rename or document exact predicates; callers stop using them as stronger lifecycle evidence. | +| LIFE-GAP-030 | Low-medium | Medium | `Task.run()` may resolve immediately if another path already started the task. | `_started` short-circuit versus scheduler callback. | Single start owner. | Scheduler-facing start returns the actual run promise; duplicate-start test proves settlement identity. | +| LIFE-GAP-031 | Low-medium | High | Queue state is memory-only and cleared on disposal. | `MessageQueueService.dispose`. | Product durability decision. | Document intentional loss or persist claims/messages with restart tests. | +| LIFE-GAP-032 | Low-medium | Medium | Webview abandonment handler exists without a confirmed production UI sender. | Protocol/handler search only. | Reachability decision. | Add supported sender/E2E or remove/deprecate unreachable command. | +| LIFE-GAP-033 | Low-medium | High | Resume ingress differs in awaiting and error propagation. | Webview/API/IPC adapters diverge. | Shared resume operation contract. | Contract tests compare result/error/publication semantics for each surface. | +| LIFE-GAP-034 | Low | High | Bounds and model metadata are handwritten and not mechanically synchronized. | Constants, prose, and console summaries duplicate values. | Machine-readable checker metadata. | CI compares emitted metadata with docs/manifest and rejects undocumented bound/action/landmark changes. | +| LIFE-GAP-035 | Medium | High | Tool-originated child initialization lacks a complete durable ownership contract; initial todo state is the confirmed witness and can disappear after rehydration. | Create a child with explicit initial todos, switch or restart before `update_todo_list`, then reopen it; restoration finds no todo message and yields an empty list. | Canonical durable task-ID-scoped child-initialization owner and publication contract. | Inventory every `new_task`-originated child field; persist required initial state before visibility/run; restore deep-equal independent state across switching, interrupted resume, checkpoint restore, and fresh-host restart; preserve later-update and explicit-empty precedence; add constructor deep-copy, persistence, webview scoping, and E2E witnesses. | +| LIFE-GAP-036 | High | High | Interactive todo approval edit state is process-global and uncorrelated; one task's delayed edit can be consumed by another task's pending approval. | Start approvals for tasks A and B, send A's edited list through `setPendingTodoList`, then resolve B; B reads the shared `approvedTodoList`. | Task/action/tool-call-correlated approval state and webview protocol. | Carry task ID and action/tool-call ID through proposal, webview edit, approval, cancellation, and settlement; reject stale/mismatched edits; deep-clone inputs; test two interleaved approvals, denial, cancellation, task switch, and delayed edits. | +| LIFE-GAP-037 | Medium-high | High | Singleton tool handlers share partial presentation state across calls/tasks, so interleaved paths can cause false or missed stabilization. | Interleave A:`x`, B:`y`, A:`x` or A:`x`, B:`x` through one handler's `lastSeenPartialPath`. | Per-call handler state keyed by task and tool-call identity. | Isolate partial state by `(taskId, toolCallId)` or handler instance; prove independent stabilization and cleanup after success, malformed finalization, rejection, cancellation, abandonment, and incomplete streams. | +| LIFE-GAP-038 | High | High | Lossy tool-ID canonicalization can deduplicate persisted history without deduplicating execution, results, approvals, or pending-action replay. | Distinct raw IDs such as `call:a` and `call/a` both sanitize to `call_a`; history may retain one call while execution retains both. | One collision-resistant canonical call identity before indexing and persistence. | Reject or disambiguate collisions; prove a bijection among parsed call, durable tool use, approval, execution, result, pending action, and replay; test adversarial native/MCP IDs and restart between approval and settlement. | ## Portfolio remediation plan The [1-SP remediation block register](./task-lifecycle-remediation-blocks.md) decomposes this portfolio into small modeling/documentation increments. It assigns every GAP exactly one primary block, preserves dependencies across workstreams, and keeps optional fan-out separate from baseline ownership. -The 40 IDs are not 40 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. +The 38 IDs are not 38 independent projects. They group into eight programs with shared root causes and implementation surfaces. Complexity classes reflect implementation breadth, coupling, and verification risk rather than schedule or duration. | Cluster | Gap IDs | Root fix and likely ownership | Complexity | Engineering risk | Objective portfolio evidence | | ------------------------------------------ | -------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | | P1. Persisted ownership and generation | 001, 002, 012, 017, 020 | Disk-authoritative lifecycle ownership/generation and immutable store reads across history types, lifecycle reducers, `TaskHistoryStore`, provider delegation, and reconciliation. | XL | High: persisted compatibility and cross-host races | Two-host stale-write tests, promoted invariants, restart/reconciliation evidence, backward-compatible optional data. | -| P2. Durable operation and crash recovery | 004, 005, 006, 021, 023, 039, 040 | Operation intent/replay or explicit idempotent recovery for pair writes, delegation, completion messages, shutdown, and deletion. | XL | Very high: failure ordering can create new corruption | Fault injection at every durable boundary, crash/restart convergence, no false success, recovery reachability. | +| P2. Durable operation and crash recovery | 004, 005, 006, 021, 023 | Operation intent/replay or explicit idempotent recovery for pair writes, delegation, completion messages, shutdown, and deletion. | XL | Very high: failure ordering can create new corruption | Fault injection at every durable boundary, crash/restart convergence, no false success, recovery reachability. | | P3. Schema, path, and lifecycle vocabulary | 013, 018, 019 | Schema-derived status ownership, validated ordinary history reads, and one safe task-ID boundary across types, persistence, metadata, CLI, and import paths. | M | Medium: malformed legacy data and downgrade behavior | Migration/quarantine fixtures, traversal tests, type/static ratchets. | | P4. Request, stream, and tool identity | 008, 010, 024, 025, 026, 030, 037, 038 | Request generation plus canonical call identity, then task/generation/call-scoped parser and partial-handler state. Owners include provider transforms, parser, `Task`, `BaseTool`, editing handlers, and tool-ID utilities. | XL | High: provider compatibility and duplicate execution | Adversarial IDs/index-less streams, delayed/cancelled generation tests, cleanup/deadline checks, production-backed call-state model. | | P5. Tool-owned task state and queueing | 003, 007, 031, 035, 036 | Task-local context, durable child initialization, correlated approval identity, and claim/persist/ack queueing across tools, `Task`, provider/webview, message queue, and history schema. | XL | High: cross-task contamination and persistence precedence | Omitted/explicit child controls, switch/restart E2E, two-approval schedules, queue failure retention, mode-permission tests. | @@ -237,11 +229,10 @@ The 40 IDs are not 40 independent projects. They group into eight programs with ### Root fixes that close multiple gaps - One persisted generation and disk-authoritative ownership design should close 001, 002, and 012; immutable reads and explicit reconciliation semantics address 017/020 around that owner. -- One durable operation-intent/replay framework can support 004, 005, 006, 021, 023, 039, and 040, but each operation still needs its own legal recovery states and fault-injection matrix. The rejection-settlement crash window and startup-repair pending-action rules join the same framework. +- One durable operation-intent/replay framework can support 004, 005, 006, 021, and 023, but each operation still needs its own legal recovery states and fault-injection matrix. - One request-generation/canonical-call identity established before parser indexing can support 008, 010, 024–026, 030, 037, and 038. - One correlated `(taskId, actionId, toolCallId)` approval protocol can close 036 and support 007/035; it does not itself make child state durable. - One typed lifecycle operation layer can normalize P6, but public compatibility requires separate adapters rather than a flag-day payload rewrite. -- Validated ordinary reads close 018. One shared filesystem task-ID validator covering traversal and separator cases across store paths, imports, deletion, and checkpoints closes 019. ### Independent work that should not be collapsed @@ -299,21 +290,20 @@ Four workstreams can proceed concurrently after foundation decisions: ## Burn-down dependencies -| Dependency | Enables | -| -------------------------------------------------------------------- | ------------------------------------------------------------- | --- | -| Disk-authoritative ownership/generation design | LIFE-GAP-001, 002, 012, 020 | -| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023, 039, 040 | -| Task-local execution-context owner | LIFE-GAP-007 and optional future fan-out | -| Request generation and terminal cleanup owner | LIFE-GAP-010, 024, 025, 026 | -| Event notification/barrier contract | LIFE-GAP-009, 011, 027, 028 | -| Machine-readable lifecycle manifest | LIFE-GAP-013, 016, 034 | -| Serial scheduler baseline | LIFE-GAP-014 | -| Optional fan-out product program | Historical #369/#372 scope, outside baseline | -| Durable task-scoped child initialization | LIFE-GAP-035 and future tool/lifecycle composition | -| Correlated approval ownership | LIFE-GAP-036 | -| Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | -| Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | -| Validated ordinary reads and one shared filesystem task-ID validator | LIFE-GAP-018, 019 | | +| Dependency | Enables | +| ---------------------------------------------- | ------------------------------------------------------------- | +| Disk-authoritative ownership/generation design | LIFE-GAP-001, 002, 012, 020 | +| Durable operation intent/recovery design | LIFE-GAP-004, 005, 006, 021, 023 | +| Task-local execution-context owner | LIFE-GAP-007 and optional future fan-out | +| Request generation and terminal cleanup owner | LIFE-GAP-010, 024, 025, 026 | +| Event notification/barrier contract | LIFE-GAP-009, 011, 027, 028 | +| Machine-readable lifecycle manifest | LIFE-GAP-013, 016, 034 | +| Serial scheduler baseline | LIFE-GAP-014 | +| Optional fan-out product program | Historical #369/#372 scope, outside baseline | +| Durable task-scoped child initialization | LIFE-GAP-035 and future tool/lifecycle composition | +| Correlated approval ownership | LIFE-GAP-036 | +| Task/tool-call-scoped partial state | LIFE-GAP-037 with request-generation cleanup gaps 010 and 024 | +| Canonical tool-call identity | LIFE-GAP-038 with generation/replay gap 012 | ## Mechanically useful follow-up checklist @@ -334,8 +324,6 @@ Four workstreams can proceed concurrently after foundation decisions: - [ ] For LIFE-GAP-036, correlate every approval edit and settlement with task ID plus action/tool-call ID; reject stale cross-task edits. - [ ] For LIFE-GAP-037, interleave equal and unequal partial paths across two calls and two tasks, then verify terminal cleanup. - [ ] For LIFE-GAP-038, use adversarial raw IDs to verify one-to-one durable call, approval, execution, result, pending-action, and replay identity. -- [ ] For LIFE-GAP-039, inject failure between the authoritative rejection and the settlement write and prove bounded, convergent recovery without unbounded replay. -- [ ] For LIFE-GAP-040, run startup repair and reconciliation against records that carry staged pending actions and prove the #1714 loop precondition cannot survive repair. ## Completeness statement @@ -343,4 +331,4 @@ At the audited commit, this report covers every tracked definition and directly ## Historical provenance -Relevant historical reports include [#1469](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1469), [#1021](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1021), [#1623](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1623), [#369](https://github.com/Zoo-Code-Org/Zoo-Code/issues/369), [#372](https://github.com/Zoo-Code-Org/Zoo-Code/issues/372), [#612](https://github.com/Zoo-Code-Org/Zoo-Code/issues/612), [#1453](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1453), [#1279](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1279), [#921](https://github.com/Zoo-Code-Org/Zoo-Code/issues/921), [#920](https://github.com/Zoo-Code-Org/Zoo-Code/issues/920), and [#1468](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1468), with rejected-delegation settlement, its crash window, and its startup-repair aftermath recorded in [#1714](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1714). These links provide provenance only; closure is governed by the objective criteria above. +Relevant historical reports include [#1469](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1469), [#1021](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1021), [#1623](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1623), [#369](https://github.com/Zoo-Code-Org/Zoo-Code/issues/369), [#372](https://github.com/Zoo-Code-Org/Zoo-Code/issues/372), [#612](https://github.com/Zoo-Code-Org/Zoo-Code/issues/612), [#1453](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1453), [#1279](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1279), [#921](https://github.com/Zoo-Code-Org/Zoo-Code/issues/921), [#920](https://github.com/Zoo-Code-Org/Zoo-Code/issues/920), and [#1468](https://github.com/Zoo-Code-Org/Zoo-Code/issues/1468). These links provide provenance only; closure is governed by the objective criteria above. diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index 7cdb3fc651..38324af9ba 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -83,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear and concurrent settlement/deletion regressions. Deletion shares a per-task guard lock outside the task directory with history writes, so it cannot interleave with the settlement rename window or the recursive directory removal. Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear. Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. ## Task cleanup protocol model diff --git a/docs/architecture/task-lifecycle-remediation-blocks.md b/docs/architecture/task-lifecycle-remediation-blocks.md index d30cccc960..ffbcbb334b 100644 --- a/docs/architecture/task-lifecycle-remediation-blocks.md +++ b/docs/architecture/task-lifecycle-remediation-blocks.md @@ -13,7 +13,7 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## Ownership rules -- Every `LIFE-GAP-001` through `LIFE-GAP-040` has exactly one primary block below. +- Every `LIFE-GAP-001` through `LIFE-GAP-038` has exactly one primary block below. - A block owns exactly one GAP ID. Dependencies may reference other blocks but do not duplicate ownership. - Block IDs are stable: `LIFE-BLK-P-`. - Baseline blocks describe current serial production behavior. Optional fan-out is isolated under `FANOUT-BLK-*` and does not own a baseline `LIFE-GAP`. @@ -21,25 +21,23 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## P1: Persisted ownership and generation -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P1-001 | 001 | Encode authoritative awaited-child revalidation as a model boundary and retain the shortest stale-completion witness. | `TaskHistoryStore.atomicUpdatePair`, `ClineProvider.reopenParentFromDelegation`; shared-store checker; cross-instance tests. | None | Model names lock-time ownership check, witness, bounds, and production test required for promotion. | -| LIFE-BLK-P1-002 | 002 | Specify lifecycle-owned lineage fields versus metadata writes and the stale-save witness. | `Task.saveClineMessages`, `taskMetadata`, `mergeHistoryDelta`; shared-store checker. | P1-001 ownership vocabulary | Field ownership table and monotonic-detachment invariant are explicit; no claim of current safety. | -| LIFE-BLK-P1-012 | 012 | Add attempt-generation state and stale-versus-resumed completion scenarios to the specification. | `PendingTaskAction.actionId`, interruption/resume/completion reducers; lifecycle checker exclusion. | P1-001 | Two generations and acceptance/rejection landmarks are specified with a bounded future checker shape. The contract covers staged pending-action replay attempts, bounded retry identity, and settlement of stale actions. | -| LIFE-BLK-P1-017 | 017 | Inventory mutable cache read consumers and define immutable read semantics. | `TaskHistoryStore.get/getAll`; store tests. | None | Every direct caller is classified; clone/freeze test criteria and compatibility exclusions are recorded. | -| LIFE-BLK-P1-020 | 020 | Define observable stale-cache and convergence histories. | watcher, `invalidate`, `reconcile`; shared-store landmarks and cross-instance tests. | P1-001 | Missed-watch and explicit-refresh histories have bounded properties and objective convergence evidence. Convergence states a pending-action rule for reconciled records with P2-040. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | --------------------------- | -------------------------------------------------------------------------------------------------------- | +| LIFE-BLK-P1-001 | 001 | Encode authoritative awaited-child revalidation as a model boundary and retain the shortest stale-completion witness. | `TaskHistoryStore.atomicUpdatePair`, `ClineProvider.reopenParentFromDelegation`; shared-store checker; cross-instance tests. | None | Model names lock-time ownership check, witness, bounds, and production test required for promotion. | +| LIFE-BLK-P1-002 | 002 | Specify lifecycle-owned lineage fields versus metadata writes and the stale-save witness. | `Task.saveClineMessages`, `taskMetadata`, `mergeHistoryDelta`; shared-store checker. | P1-001 ownership vocabulary | Field ownership table and monotonic-detachment invariant are explicit; no claim of current safety. | +| LIFE-BLK-P1-012 | 012 | Add attempt-generation state and stale-versus-resumed completion scenarios to the specification. | `PendingTaskAction.actionId`, interruption/resume/completion reducers; lifecycle checker exclusion. | P1-001 | Two generations and acceptance/rejection landmarks are specified with a bounded future checker shape. | +| LIFE-BLK-P1-017 | 017 | Inventory mutable cache read consumers and define immutable read semantics. | `TaskHistoryStore.get/getAll`; store tests. | None | Every direct caller is classified; clone/freeze test criteria and compatibility exclusions are recorded. | +| LIFE-BLK-P1-020 | 020 | Define observable stale-cache and convergence histories. | watcher, `invalidate`, `reconcile`; shared-store landmarks and cross-instance tests. | P1-001 | Missed-watch and explicit-refresh histories have bounded properties and objective convergence evidence. | ## P2: Durable operation and crash recovery -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | ------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P2-004 | 004 | Enumerate pair-write interruption points and legal recovered states. | `atomicUpdatePair`; pair-failure landmark/tests. | P1-001 | Every pre/post-write cut has one legal outcome and required fault-injection assertion. | -| LIFE-BLK-P2-005 | 005 | Map delegation create/persist/publish/start cuts and rollback obligations. | `delegateParentAndOpenChild`; provider handoff model/tests. | P1-001, P2-004 | Transition table covers every cut without claiming child/parent atomicity. | -| LIFE-BLK-P2-006 | 006 | Specify completion message/lifecycle commit phases and replay outcomes. | `reopenParentFromDelegation`; completion and shared-store models. | P1-012, P2-004 | Result visibility and lifecycle state are mapped for each injected failure point. The phase table includes child pending-action state at every cut and its interaction with resumed replay (012). | -| LIFE-BLK-P2-021 | 021 | Define store close/drain semantics and post-dispose write exclusion. | `TaskHistoryStore.dispose`, write lock; store tests. | P2-004 recovery vocabulary | A bounded close-state machine and deterministic pending-write test criteria are documented. | -| LIFE-BLK-P2-023 | 023 | Specify deletion unlink failure and reconciliation histories. | `delete/deleteMany`, task directory/checkpoint cleanup; deletion tests. | P2-004 | False-success and resurrection outcomes are explicit with tombstone/retry closure choices. | -| LIFE-BLK-P2-039 | 039 | Specify the crash window between authoritative rejection and settlement, plus bounded replay of a rejected action. | `ClineProvider.delegateParentAndOpenChild` settlement block, `settleRejectedCreateSubtaskAction`, `Task.resumePendingTaskAction`; lifecycle settlement witnesses; provider failure-injection tests. | P1-012, P2-004 | Every cut between rejection and settlement has one legal outcome; settlement-write failure is retried or surfaced without false success; replay of a rejected action is bounded and test-covered. | -| LIFE-BLK-P2-040 | 040 | Define pending-action preserve/settle/drop rules for startup repair and reconciliation outcomes. | `reconcileDelegationState`, `repairActiveDelegation`, `applyDelegationRepairIntent`, `DelegationRepairIntent`; reconciliation tests; fresh-host restart E2E. | P2-004 | Every repair outcome has an explicit pending-action rule; the #1714 loop precondition cannot survive repair; both interrupt paths persist or settle the action identically. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | -------------------------------------------------------------------------- | ----------------------------------------------------------------------- | -------------------------- | ------------------------------------------------------------------------------------------- | +| LIFE-BLK-P2-004 | 004 | Enumerate pair-write interruption points and legal recovered states. | `atomicUpdatePair`; pair-failure landmark/tests. | P1-001 | Every pre/post-write cut has one legal outcome and required fault-injection assertion. | +| LIFE-BLK-P2-005 | 005 | Map delegation create/persist/publish/start cuts and rollback obligations. | `delegateParentAndOpenChild`; provider handoff model/tests. | P1-001, P2-004 | Transition table covers every cut without claiming child/parent atomicity. | +| LIFE-BLK-P2-006 | 006 | Specify completion message/lifecycle commit phases and replay outcomes. | `reopenParentFromDelegation`; completion and shared-store models. | P1-012, P2-004 | Result visibility and lifecycle state are mapped for each injected failure point. | +| LIFE-BLK-P2-021 | 021 | Define store close/drain semantics and post-dispose write exclusion. | `TaskHistoryStore.dispose`, write lock; store tests. | P2-004 recovery vocabulary | A bounded close-state machine and deterministic pending-write test criteria are documented. | +| LIFE-BLK-P2-023 | 023 | Specify deletion unlink failure and reconciliation histories. | `delete/deleteMany`, task directory/checkpoint cleanup; deletion tests. | P2-004 | False-success and resurrection outcomes are explicit with tombstone/retry closure choices. | ## P3: Schema, path, and vocabulary @@ -51,26 +49,26 @@ Completing one block does not close its `LIFE-GAP` unless the parent GAP closure ## P4: Request, stream, and tool identity -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P4-008 | 008 | Add transform-to-parser cases for argument-only deltas and absent indices. | Responses transform, parser APIs/tests. | None | Two-call bounded schedules and expected isolated reconstruction are specified. | -| LIFE-BLK-P4-010 | 010 | Define request-generation ownership for detached usage writes. | Task request/drain paths; delayed-stream tests. | P4-038 identity vocabulary | Old/new generation mutations and allowed accounting-only updates are explicit. | -| LIFE-BLK-P4-024 | 024 | Map parser cleanup on success, abort, provider error, and replacement. | parser scope plus Task request terminal paths. | P4-010 | Every terminal path owns cleanup; late-event exclusions are stated. | -| LIFE-BLK-P4-025 | 025 | Specify listener lifetime for one chunk race and long streams. | `nextChunkWithAbort`; listener-count tests. | None | Both race outcomes remove listeners and a bounded stream cannot accumulate them. | -| LIFE-BLK-P4-026 | 026 | Model a true wall-clock deadline around pending iterator reads. | detached usage drain; fake-timer tests. | P4-010 | Permanently pending `next()` has a terminal deadline transition and no stale mutations. | -| LIFE-BLK-P4-030 | 030 | Define duplicate-start/run-promise identity. | `Task.start/run`, scheduler callback; Task tests. | P4-010 | Repeated starts share the actual settlement and cannot bypass scheduler ownership. | -| LIFE-BLK-P4-037 | 037 | Specify call-scoped partial path state and two-call interleavings. | `BaseTool.lastSeenPartialPath`, editing tool singletons; focused tests. | P4-038, P4-010 | Equal/different path interleavings and sibling-safe cleanup are bounded and reachable. | -| LIFE-BLK-P4-038 | 038 | Define canonical raw-to-durable call identity and collision witnesses. | tool-ID utility, parser, Task history, results, pending actions; duplicate-ID tests. | None | Adversarial IDs preserve or explicitly reject one-to-one call/result/replay correspondence. The bijection includes the staged pending action, whose `actionId` derives from sanitized call IDs, across restart between approval and settlement. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | -------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | -------------------------- | ------------------------------------------------------------------------------------------- | +| LIFE-BLK-P4-008 | 008 | Add transform-to-parser cases for argument-only deltas and absent indices. | Responses transform, parser APIs/tests. | None | Two-call bounded schedules and expected isolated reconstruction are specified. | +| LIFE-BLK-P4-010 | 010 | Define request-generation ownership for detached usage writes. | Task request/drain paths; delayed-stream tests. | P4-038 identity vocabulary | Old/new generation mutations and allowed accounting-only updates are explicit. | +| LIFE-BLK-P4-024 | 024 | Map parser cleanup on success, abort, provider error, and replacement. | parser scope plus Task request terminal paths. | P4-010 | Every terminal path owns cleanup; late-event exclusions are stated. | +| LIFE-BLK-P4-025 | 025 | Specify listener lifetime for one chunk race and long streams. | `nextChunkWithAbort`; listener-count tests. | None | Both race outcomes remove listeners and a bounded stream cannot accumulate them. | +| LIFE-BLK-P4-026 | 026 | Model a true wall-clock deadline around pending iterator reads. | detached usage drain; fake-timer tests. | P4-010 | Permanently pending `next()` has a terminal deadline transition and no stale mutations. | +| LIFE-BLK-P4-030 | 030 | Define duplicate-start/run-promise identity. | `Task.start/run`, scheduler callback; Task tests. | P4-010 | Repeated starts share the actual settlement and cannot bypass scheduler ownership. | +| LIFE-BLK-P4-037 | 037 | Specify call-scoped partial path state and two-call interleavings. | `BaseTool.lastSeenPartialPath`, editing tool singletons; focused tests. | P4-038, P4-010 | Equal/different path interleavings and sibling-safe cleanup are bounded and reachable. | +| LIFE-BLK-P4-038 | 038 | Define canonical raw-to-durable call identity and collision witnesses. | tool-ID utility, parser, Task history, results, pending actions; duplicate-ID tests. | None | Adversarial IDs preserve or explicitly reject one-to-one call/result/replay correspondence. | ## P5: Tool-owned task state and queueing -| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | -| --------------- | --- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| LIFE-BLK-P5-003 | 003 | Map every queue consumer to claim/persist/ack or dequeue-before-submit. | `MessageQueueService`, Task queue paths; failure tests. | None | Every consumer is classified and message-retention failure evidence is specified. | -| LIFE-BLK-P5-007 | 007 | Inventory remaining mode-sensitive readers and authoritative task/provider source after the merged three-reader fix. | handoff selector, named production readers `getEnvironmentDetails`, `validateToolUse` call sites in `presentAssistantMessage`, custom tool execution, merged environment/validation/custom-tool tests, delegated reader checker. | None | Confirmed readers are marked production-tested; unclassified readers remain listed; the pure checker is not described as executing downstream readers. | -| LIFE-BLK-P5-031 | 031 | Define intentional versus accidental queue loss across task disposal/restart. | queue service disposal and task lifecycle; E2E boundary. | P5-003 | Product contract, excluded durability, and restart witness are explicit. The contract states the restart composition between queue loss and pending-action replay ordering. | -| LIFE-BLK-P5-035 | 035 | Specify durable child initialization precedence using initial todos as witness. | `NewTaskTool`, Task constructor, history/messages, rehydration, UI state. | P2-005 | Omitted, explicit-empty, initial, updated, switched, and restarted cases are mapped. | -| LIFE-BLK-P5-036 | 036 | Model two approval identities and stale/cross-task todo edits. | `approvedTodoList`, webview handler, approval callbacks/tests. | P4-038 | Two-task schedules require task/action/call correlation; current unsafe witness is explicit. Correlation binds the staged `pendingAction.actionId` to task and call; denied and cancelled approvals have durable settlement rules. | +| Block | GAP | 1-SP increment | Production/model/test mapping | Depends on | Acceptance | +| --------------- | --- | -------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | +| LIFE-BLK-P5-003 | 003 | Map every queue consumer to claim/persist/ack or dequeue-before-submit. | `MessageQueueService`, Task queue paths; failure tests. | None | Every consumer is classified and message-retention failure evidence is specified. | +| LIFE-BLK-P5-007 | 007 | Inventory remaining mode-sensitive readers and authoritative task/provider source after the merged three-reader fix. | handoff selector, named production readers `getEnvironmentDetails`, `validateToolUse` call sites in `presentAssistantMessage`, custom tool execution, merged environment/validation/custom-tool tests, delegated reader checker. | None | Confirmed readers are marked production-tested; unclassified readers remain listed; the pure checker is not described as executing downstream readers. | +| LIFE-BLK-P5-031 | 031 | Define intentional versus accidental queue loss across task disposal/restart. | queue service disposal and task lifecycle; E2E boundary. | P5-003 | Product contract, excluded durability, and restart witness are explicit. | +| LIFE-BLK-P5-035 | 035 | Specify durable child initialization precedence using initial todos as witness. | `NewTaskTool`, Task constructor, history/messages, rehydration, UI state. | P2-005 | Omitted, explicit-empty, initial, updated, switched, and restarted cases are mapped. | +| LIFE-BLK-P5-036 | 036 | Model two approval identities and stale/cross-task todo edits. | `approvedTodoList`, webview handler, approval callbacks/tests. | P4-038 | Two-task schedules require task/action/call correlation; current unsafe witness is explicit. | ## P6: Event and ingress contracts @@ -110,7 +108,7 @@ Optional future fan-out blocks do not own `LIFE-GAP-014` and do not participate ## Mechanical coverage check -The primary tables above map the closed integer range `001..040` exactly once. Reviewers should verify this mechanically before changing the register: +The primary tables above map the closed integer range `001..038` exactly once. Reviewers should verify this mechanically before changing the register: ```sh rg -o '^\| LIFE-BLK-P[0-9]-[0-9]{3} \|' docs/architecture/task-lifecycle-remediation-blocks.md \ @@ -123,7 +121,7 @@ The command must print nothing. It matches only primary table rows, so dependenc Separately compare block suffixes with the GAP column to detect omissions or mismatches: ```sh -node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:40},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==40||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' +node -e 'const fs=require("fs");const s=fs.readFileSync("docs/architecture/task-lifecycle-remediation-blocks.md","utf8");const rows=[...s.matchAll(/^\| LIFE-BLK-P\d-(\d{3}) \| (\d{3}) \|/gm)];const gaps=rows.map(r=>r[2]);const want=Array.from({length:38},(_,i)=>String(i+1).padStart(3,"0"));if(rows.length!==38||rows.some(r=>r[1]!==r[2])||want.some(id=>!gaps.includes(id)))process.exit(1)' ``` ## Block completion template diff --git a/src/__tests__/delegation-concurrent.spec.ts b/src/__tests__/delegation-concurrent.spec.ts index 222fdb6318..40d9b49ee5 100644 --- a/src/__tests__/delegation-concurrent.spec.ts +++ b/src/__tests__/delegation-concurrent.spec.ts @@ -23,17 +23,6 @@ vi.mock("../utils/safeWriteJson", () => ({ safeWriteJson: vi.fn().mockResolvedValue(undefined), })) -// The store's cross-process task guard and advisory file lock need a real -// filesystem, which this spec stubs out. The in-process serialization under -// test does not depend on them. -vi.mock("../utils/fileLock", async () => ({ - ...(await vi.importActual("../utils/fileLock")), - acquireFileLock: vi.fn().mockResolvedValue({ - release: vi.fn().mockResolvedValue(undefined), - isCompromised: vi.fn().mockReturnValue(false), - }), -})) - vi.mock("../utils/storage", () => ({ getStorageBasePath: vi.fn().mockResolvedValue("/tmp/test-storage"), })) diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 7b2dbd8a1f..d6c42e3280 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -7,14 +7,7 @@ import deepEqual from "fast-deep-equal" import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { - acquireFileLock, - assertLockUsable, - DESTRUCTIVE_LOCK_RETRIES, - LOCK_STALE_MS, - type AcquireFileLockOptions, -} from "../../utils/fileLock" -import { safeWriteJson } from "../../utils/safeWriteJson" +import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" @@ -273,10 +266,17 @@ export class TaskHistoryStore { */ async delete(taskId: string): Promise { return this.withLock(async () => { - await this.deleteTaskFile(taskId) this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) + // Remove per-task file (best-effort) + try { + const filePath = await this.getTaskFilePath(taskId) + await fs.unlink(filePath) + } catch { + // File may already be deleted + } + // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { await this.onWrite(this.getAll()) @@ -290,9 +290,15 @@ export class TaskHistoryStore { async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { for (const taskId of taskIds) { - await this.deleteTaskFile(taskId) this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) + + try { + const filePath = await this.getTaskFilePath(taskId) + await fs.unlink(filePath) + } catch { + // File may already be deleted + } } // Call onWrite callback inside the lock for serialized write-through @@ -810,21 +816,9 @@ export class TaskHistoryStore { await fs.access(filePath) // File already exists, skip (don't overwrite existing per-task files) } catch { - // File doesn't exist, write it under the same task guard as - // every other history write so a concurrent deletion cannot - // interleave with the removal of the task directory. - const guard = await this.acquireTaskIoGuard(item.id) - try { - assertLockUsable(guard, filePath, "write") - await safeWriteJson(filePath, item) - this.cache.set(item.id, item) - } finally { - try { - await guard.release() - } catch (error) { - console.error(`Failed to release task guard for ${item.id}:`, error) - } - } + // File doesn't exist, write it + await safeWriteJson(filePath, item) + this.cache.set(item.id, item) } } @@ -847,22 +841,6 @@ export class TaskHistoryStore { return { id, ...this.computeDelta(cached, incoming) } } - /** - * Cross-process guard for one task. It lives outside the removable task - * directory, so deletion can hold it across the history-file unlink and - * the recursive directory removal while history writers hold it around - * their writes. Neither side can then interleave with the other. - */ - private async acquireTaskIoGuard( - taskId: string, - options?: AcquireFileLockOptions, - ): Promise> { - const tasksDir = await this.getTasksDir() - const guardDir = path.join(tasksDir, ".guards") - await fs.mkdir(guardDir, { recursive: true }) - return acquireFileLock(path.join(guardDir, `${taskId}.guard`), options) - } - /** * Write a HistoryItem to its per-task `history_item.json` file. * @@ -873,30 +851,20 @@ export class TaskHistoryStore { */ private async writeTaskFile(item: HistoryItem, delta?: Partial): Promise { const filePath = await this.getTaskFilePath(item.id) - const guard = await this.acquireTaskIoGuard(item.id) - try { - assertLockUsable(guard, filePath, "write") - if (delta) { - let written: HistoryItem = item - const mergeFn = mergeWithDisk(delta) - await safeWriteJson(filePath, item, { - merge: (existing, incoming) => { - const result = mergeFn(existing, incoming) - written = result as HistoryItem - return result - }, - }) - return written - } else { - await safeWriteJson(filePath, item) - return item - } - } finally { - try { - await guard.release() - } catch (error) { - console.error(`Failed to release task guard for ${item.id}:`, error) - } + if (delta) { + let written: HistoryItem = item + const mergeFn = mergeWithDisk(delta) + await safeWriteJson(filePath, item, { + merge: (existing, incoming) => { + const result = mergeFn(existing, incoming) + written = result as HistoryItem + return result + }, + }) + return written + } else { + await safeWriteJson(filePath, item) + return item } } @@ -915,77 +883,6 @@ export class TaskHistoryStore { } } - /** - * Delete the history file and remove the task directory. - * - * The task guard lives outside the removable directory and spans the - * unlink and the recursive removal, so a concurrent history writer - * cannot interleave with either step. Both acquisitions use a retry - * budget that outlasts LOCK_STALE_MS, so a lock left by a crashed - * process is broken within this call instead of failing the deletion. - */ - private async deleteTaskFile(taskId: string): Promise { - const filePath = await this.getTaskFilePath(taskId) - const taskDir = path.dirname(filePath) - try { - await fs.access(taskDir) - } catch (error: unknown) { - const code = - error && typeof error === "object" && "code" in error ? (error as { code?: string }).code : undefined - if (code !== "ENOENT") { - throw error - } - // safeWriteJson creates the directory before taking this file lock. - // If it is absent, no writer has reached the shared protocol yet and - // deletion can linearize here without creating an empty task directory. - return - } - - const guard = await this.acquireTaskIoGuard(taskId, { retries: DESTRUCTIVE_LOCK_RETRIES }) - try { - let fileLock: Awaited> | undefined - try { - fileLock = await acquireFileLock(filePath, { retries: DESTRUCTIVE_LOCK_RETRIES }) - assertLockUsable(fileLock, filePath, "deletion") - // The guard can be lost while waiting for the file lock above. - assertLockUsable(guard, filePath, "deletion") - await fs.unlink(filePath) - } catch (error: unknown) { - const code = - error && typeof error === "object" && "code" in error - ? (error as { code?: string }).code - : undefined - if (code !== "ENOENT") { - throw error - } - } finally { - if (fileLock) { - try { - await fileLock.release() - } catch (error) { - console.error(`Failed to release lock for ${filePath}:`, error) - } - } - } - - // The guard must still hold exclusion before this second mutation. - assertLockUsable(guard, taskDir, "task directory removal") - try { - await fs.rm(taskDir, { recursive: true, force: true }) - } catch (error) { - // The record is already unlinked. Mirror the provider's historical - // tolerance for a directory that cannot be removed right now. - console.error(`[TaskHistoryStore] Failed to remove task directory ${taskDir}:`, error) - } - } finally { - try { - await guard.release() - } catch (error) { - console.error(`Failed to release task guard for ${taskId}:`, error) - } - } - } - // ────────────────────────────── Private: fs.watch ────────────────────────────── /** @@ -1188,32 +1085,22 @@ export class TaskHistoryStore { } const filePath = await this.getTaskFilePath(taskId) let authoritative: HistoryItem = cached - const guard = await this.acquireTaskIoGuard(taskId) - try { - assertLockUsable(guard, filePath, "write") - await safeWriteJson(filePath, cached, { - merge: (existing) => { - if (!existing || typeof existing !== "object" || !("id" in existing)) { - // Writing the cached record back would recreate a task - // another host deleted, so drop the stale entry first. - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - throw new Error( - `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, - ) - } - const disk = existing as HistoryItem - authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) - return authoritative - }, - }) - } finally { - try { - await guard.release() - } catch (error) { - console.error(`Failed to release task guard for ${taskId}:`, error) - } - } + await safeWriteJson(filePath, cached, { + merge: (existing) => { + if (!existing || typeof existing !== "object" || !("id" in existing)) { + // Writing the cached record back would recreate a task + // another host deleted, so drop the stale entry first. + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, + ) + } + const disk = existing as HistoryItem + authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) + return authoritative + }, + }) this.cache.set(taskId, authoritative) if (this.onWrite) { await this.onWrite(this.getAll()) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts index 40c555dbaa..cef9874e5f 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts @@ -142,25 +142,25 @@ describe("TaskHistoryStore cross-instance safety", () => { expect(storeB.get("shared-task")).toBeUndefined() }) - it("delete by instance A removes the history file, the task directory, and is detected by instance B", async () => { + it("delete by instance A is detected even when the task directory remains", async () => { await storeA.initialize() await storeB.initialize() - const item = makeHistoryItem({ id: "full-delete" }) + const item = makeHistoryItem({ id: "file-only-delete" }) await storeA.upsert(item) await storeB.reconcile() - expect(storeB.get("full-delete")).toBeDefined() + expect(storeB.get("file-only-delete")).toBeDefined() - // delete() unlinks history_item.json and removes the task directory - // under the task guard shared with history writes. - await storeA.delete("full-delete") + // delete() unlinks history_item.json but leaves the task directory. + await storeA.delete("file-only-delete") - const taskDir = path.join(tmpDir, "tasks", "full-delete") - await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) + // Directory still exists (other files like ui_messages.json may remain). + const taskDir = path.join(tmpDir, "tasks", "file-only-delete") + await expect(fs.access(taskDir)).resolves.toBeUndefined() await storeB.reconcile() - expect(storeB.get("full-delete")).toBeUndefined() + expect(storeB.get("file-only-delete")).toBeUndefined() }) it("per-task file updates by one instance are visible to another after invalidation", async () => { diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts deleted file mode 100644 index 3dfeec39cf..0000000000 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts +++ /dev/null @@ -1,149 +0,0 @@ -// pnpm --filter roo-cline test core/task-persistence/__tests__/TaskHistoryStore.guardCompromise.spec.ts - -import * as fs from "fs/promises" -import * as path from "path" -import * as os from "os" - -import type { HistoryItem } from "@roo-code/types" - -import { GlobalFileNames } from "../../../shared/globalFileNames" - -vi.mock("../../../utils/storage", () => ({ - getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => { - return defaultPath - }), -})) - -// Mock safeWriteJson with plain fs writes so only the outer task guard -// behavior is under test. -vi.mock("../../../utils/safeWriteJson", () => ({ - safeWriteJson: vi.fn().mockImplementation(async (filePath: string, data: unknown) => { - await fs.mkdir(path.dirname(filePath), { recursive: true }) - await fs.writeFile(filePath, JSON.stringify(data, null, "\t"), "utf8") - }), -})) - -function makeHistoryItem(overrides: Partial = {}): HistoryItem { - return { - id: `task-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, - number: 1, - ts: Date.now(), - task: "Test task", - tokensIn: 100, - tokensOut: 50, - totalCost: 0.01, - workspace: "/test/workspace", - ...overrides, - } -} - -async function seedTaskOnDisk(tmpDir: string, taskId: string): Promise<{ taskDir: string; historyFile: string }> { - const taskDir = path.join(tmpDir, "tasks", taskId) - const historyFile = path.join(taskDir, GlobalFileNames.historyItem) - await fs.mkdir(taskDir, { recursive: true }) - await fs.writeFile(historyFile, JSON.stringify(makeHistoryItem({ id: taskId }))) - return { taskDir, historyFile } -} - -describe("TaskHistoryStore task guard compromise", () => { - let tmpDir: string - - beforeEach(async () => { - tmpDir = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-guard-")) - }) - - afterEach(async () => { - vi.doUnmock("proper-lockfile") - vi.resetModules() - await fs.rm(tmpDir, { recursive: true, force: true }).catch(() => {}) - }) - - /** - * Simulate proper-lockfile reporting the guard lock (never the history - * file lock) as compromised the moment it is acquired. Guards live at - * `tasks/.guards/.guard`. - */ - function compromiseGuardsOnAcquire() { - vi.doMock("proper-lockfile", () => ({ - lock: (file: string, options: { onCompromised: (error: Error) => void }) => { - if (file.endsWith(".guard")) { - options.onCompromised(new Error("Guard no longer available")) - } - return Promise.resolve(async () => {}) - }, - })) - } - - async function importStoreClass(): Promise { - const module = await import("../TaskHistoryStore") - return module.TaskHistoryStore - } - - it("aborts a history write when the task guard is compromised", async () => { - compromiseGuardsOnAcquire() - const TaskHistoryStore = await importStoreClass() - const store = new TaskHistoryStore(tmpDir) - await store.initialize() - - await expect(store.upsert(makeHistoryItem({ id: "guard-write" }))).rejects.toThrow("was compromised") - - // The write aborted before reaching the history file. - const historyFile = path.join(tmpDir, "tasks", "guard-write", GlobalFileNames.historyItem) - await expect(fs.access(historyFile)).rejects.toThrow() - expect(store.get("guard-write")).toBeUndefined() - - store.dispose() - }) - - it("aborts a task deletion when the task guard is compromised before the unlink", async () => { - const { taskDir, historyFile } = await seedTaskOnDisk(tmpDir, "guard-delete") - - compromiseGuardsOnAcquire() - const TaskHistoryStore = await importStoreClass() - const store = new TaskHistoryStore(tmpDir) - await store.initialize() - expect(store.get("guard-delete")).toBeDefined() - - await expect(store.delete("guard-delete")).rejects.toThrow("was compromised") - - // Neither the history file nor the task directory was removed. - await expect(fs.access(historyFile)).resolves.toBeUndefined() - await expect(fs.access(taskDir)).resolves.toBeUndefined() - expect(store.get("guard-delete")).toBeDefined() - - store.dispose() - }) - - it("stops before removing the task directory when the guard is compromised mid-deletion", async () => { - const { taskDir, historyFile } = await seedTaskOnDisk(tmpDir, "guard-mid-delete") - - // The history-file lock's release runs after the unlink and before - // the directory removal, so it is the deterministic hook for losing - // the guard between the two mutations. In-flight operations cannot - // be cancelled; the store must re-check before each mutation. - let compromiseGuard: (() => void) | undefined - vi.doMock("proper-lockfile", () => ({ - lock: (file: string, options: { onCompromised: (error: Error) => void }) => { - if (file.endsWith(".guard")) { - compromiseGuard = () => options.onCompromised(new Error("Guard no longer available")) - return Promise.resolve(async () => {}) - } - return Promise.resolve(async () => { - compromiseGuard?.() - }) - }, - })) - - const TaskHistoryStore = await importStoreClass() - const store = new TaskHistoryStore(tmpDir) - await store.initialize() - - await expect(store.delete("guard-mid-delete")).rejects.toThrow("was compromised") - - // The unlink already landed, but the guarded directory removal did not. - await expect(fs.access(historyFile)).rejects.toThrow() - await expect(fs.access(taskDir)).resolves.toBeUndefined() - - store.dispose() - }) -}) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index b489f571e8..5040558fed 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -4,61 +4,8 @@ import * as path from "path" import type { HistoryItem } from "@roo-code/types" -import { acquireFileLock } from "../../../utils/fileLock" import { TaskHistoryStore } from "../TaskHistoryStore" -type WriteTaskFile = (item: HistoryItem, delta?: Partial) => Promise - -interface WriteBarrier { - arrivals(): number - dispose(): void -} - -function synchronizeNextWrites(stores: TaskHistoryStore[], timeoutMs = 2_000): WriteBarrier { - let arrivals = 0 - let release!: () => void - let rejectBarrier!: (error: Error) => void - let settled = false - let timer: ReturnType | undefined - const barrier = new Promise((resolve, reject) => { - rejectBarrier = reject - release = () => { - if (settled) return - settled = true - if (timer) clearTimeout(timer) - resolve() - } - timer = setTimeout(() => { - if (settled) return - settled = true - reject(new Error(`Only ${arrivals}/${stores.length} stores reached writeTaskFile within ${timeoutMs}ms`)) - }, timeoutMs) - }) - void barrier.catch(() => {}) - - for (const store of stores) { - const value: unknown = Reflect.get(store, "writeTaskFile") - if (typeof value !== "function") throw new Error("TaskHistoryStore.writeTaskFile is unavailable") - const original = value.bind(store) as WriteTaskFile - Reflect.set(store, "writeTaskFile", async (historyItem: HistoryItem, delta?: Partial) => { - arrivals++ - if (arrivals === stores.length) release() - await barrier - return original(historyItem, delta) - }) - } - - return { - arrivals: () => arrivals, - dispose: () => { - if (settled) return - settled = true - if (timer) clearTimeout(timer) - rejectBarrier(new Error("Write barrier disposed before all stores arrived")) - }, - } -} - function item(id: string): HistoryItem { return { id, @@ -286,187 +233,4 @@ describe("TaskHistoryStore real cross-host locking", () => { await fs.rm(storagePath, { recursive: true, force: true }) } }) - - it("serializes deletion with an in-flight settlement so the task stays deleted", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-during-settlement-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - const actionA = createAction("action-a", "action A") - const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") - const lockPath = `${filePath}.lock` - const largeTask = "x".repeat(16 * 1024 * 1024) - - try { - await storeA.initialize() - await storeA.upsert({ ...item("shared-task"), task: largeTask, pendingAction: actionA }) - await storeB.initialize() - - const settlement = storeA.clearPendingActionIfMatching("shared-task", actionA.actionId) - await vi.waitFor(() => expect(fs.stat(lockPath)).resolves.toBeDefined(), { interval: 1, timeout: 2_000 }) - const deletion = storeB.delete("shared-task") - - await expect(settlement).resolves.toMatchObject({ id: "shared-task", pendingAction: undefined }) - await expect(deletion).resolves.toBeUndefined() - await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) - expect(storeB.get("shared-task")).toBeUndefined() - await storeA.invalidate("shared-task") - expect(storeA.get("shared-task")).toBeUndefined() - } finally { - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("preserves independent stale-cache deltas through the real per-file lock", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-real-lock-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - let writeBarrier: WriteBarrier | undefined - - try { - await storeA.initialize() - await storeA.upsert(item("shared-task")) - await storeB.initialize() - writeBarrier = synchronizeNextWrites([storeA, storeB]) - - await Promise.all([ - storeA.atomicReadAndUpdate("shared-task", (current) => ({ ...current, mode: "architect" })), - storeB.atomicReadAndUpdate("shared-task", (current) => ({ ...current, totalCost: 42 })), - ]) - - expect(writeBarrier.arrivals()).toBe(2) - await storeA.invalidate("shared-task") - expect(storeA.get("shared-task")).toMatchObject({ mode: "architect", totalCost: 42 }) - } finally { - writeBarrier?.dispose() - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("reports a bounded error when one store never reaches the write barrier", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-missed-barrier-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - let writeBarrier: WriteBarrier | undefined - - try { - await storeA.initialize() - await storeA.upsert(item("shared-task")) - await storeB.initialize() - writeBarrier = synchronizeNextWrites([storeA, storeB], 50) - - await expect( - storeA.atomicReadAndUpdate("shared-task", (current) => ({ ...current, mode: "architect" })), - ).rejects.toThrow("Only 1/2 stores reached writeTaskFile within 50ms") - expect(writeBarrier.arrivals()).toBe(1) - } finally { - writeBarrier?.dispose() - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("serializes deletion with an external history-write guard holder and removes the task directory", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-guard-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - const tasksDir = path.join(storagePath, "tasks") - const guardPath = path.join(tasksDir, ".guards", "shared-task.guard") - const filePath = path.join(tasksDir, "shared-task", "history_item.json") - const taskDir = path.join(tasksDir, "shared-task") - - try { - await storeA.initialize() - await storeA.upsert(item("shared-task")) - await storeB.initialize() - - // Hold the same guard that history writes hold, so the deletion - // must wait instead of unlinking underneath a writer. - const guard = await acquireFileLock(guardPath) - const deletion = storeB.delete("shared-task") - await new Promise((resolve) => setTimeout(resolve, 50)) - await expect(fs.access(filePath)).resolves.toBeUndefined() - - await guard.release() - await expect(deletion).resolves.toBeUndefined() - - // The guard spans the history-file unlink and the recursive - // directory removal. - await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) - await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) - expect(storeB.get("shared-task")).toBeUndefined() - } finally { - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("lets a write that starts during deletion run only after the removal window", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-write-during-delete-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") - - try { - await storeA.initialize() - await storeA.upsert(item("shared-task")) - await storeB.initialize() - - const deletion = storeB.delete("shared-task") - await vi.waitFor(() => expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }), { - interval: 1, - timeout: 2_000, - }) - - // The write starts while the deletion still owns the task guard, - // so it cannot interleave with the removal of the task directory. - const write = storeA.upsert({ ...item("shared-task"), ts: 2000, task: "rewritten after deletion" }) - - await expect(deletion).resolves.toBeUndefined() - await expect(write).resolves.toBeDefined() - - const persisted = JSON.parse(await fs.readFile(filePath, "utf8")) as HistoryItem - expect(persisted).toMatchObject({ id: "shared-task", task: "rewritten after deletion" }) - } finally { - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("deletes a task whose file lock is held live beyond the legacy retry window", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-contention-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - const tasksDir = path.join(storagePath, "tasks") - const filePath = path.join(tasksDir, "shared-task", "history_item.json") - const taskDir = path.join(tasksDir, "shared-task") - - try { - await storeA.initialize() - await storeA.upsert(item("shared-task")) - await storeB.initialize() - - // A live holder keeps the lock mtime fresh, so staleness never - // breaks it. The deletion must wait out the hold instead of failing - // once the legacy retry budget of roughly 2.5 seconds is spent. - const holder = await acquireFileLock(filePath) - const deletion = storeB.delete("shared-task") - await new Promise((resolve) => setTimeout(resolve, 3_000)) - await holder.release() - - await expect(deletion).resolves.toBeUndefined() - await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) - await expect(fs.access(taskDir)).rejects.toMatchObject({ code: "ENOENT" }) - } finally { - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) }) diff --git a/src/core/task/__tests__/Task.persistence.spec.ts b/src/core/task/__tests__/Task.persistence.spec.ts index 03aab21ec5..bc6e9c3960 100644 --- a/src/core/task/__tests__/Task.persistence.spec.ts +++ b/src/core/task/__tests__/Task.persistence.spec.ts @@ -1459,6 +1459,48 @@ describe("Task persistence", () => { expect(ask).not.toHaveBeenCalled() }) + it("stops replay when settlement preserves a replacement create-subtask action", async () => { + const replacementAction = { + ...createSubtaskAction, + actionId: "create-action-b", + message: "Replacement child", + } + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + mockProvider.taskHistoryStore.clearPendingActionIfMatching = vi.fn().mockResolvedValue({ + id: "parent-1", + status: "interrupted", + pendingAction: replacementAction, + }) + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const ask = vi.spyOn(task, "ask") + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + + await expect(getTaskPersistenceAccess(task).resumeTaskFromHistory()).rejects.toThrow( + "[Task#settleInterruptedCreateSubtaskBeforeReplay] Task parent-1 still has a rejected create-subtask action", + ) + + expect(replay).not.toHaveBeenCalled() + expect(ask).not.toHaveBeenCalled() + }) + it("reconciles an already-persisted tool result before generic resume", async () => { mockReadTaskMessages.mockResolvedValue([{ ts: 1, type: "say", say: "text", text: "Child" }]) mockReadApiMessages.mockResolvedValue([ diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 0a45443ea3..86121bf3a4 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -2375,15 +2375,15 @@ export class ClineProvider } } - // Delete all tasks from state in one batch. The store removes each - // history file and its task directory under the shared task guard, - // so a concurrent history write cannot interleave with the removal. + // Delete all tasks from state in one batch await this.taskHistoryStore.deleteMany(allIdsToDelete) this.recentTasksCache = undefined - // Delete associated shadow repositories or branches + // Delete associated shadow repositories or branches and task directories const globalStorageDir = this.contextProxy.globalStorageUri.fsPath const workspaceDir = this.cwd + const { getTaskDirectoryPath } = await import("../../utils/storage") + const globalStoragePath = this.contextProxy.globalStorageUri.fsPath for (const taskId of allIdsToDelete) { try { @@ -2393,6 +2393,17 @@ export class ClineProvider `[deleteTaskWithId${taskId}] failed to delete associated shadow repository or branch: ${error instanceof Error ? error.message : String(error)}`, ) } + + // Delete the task directory + try { + const dirPath = await getTaskDirectoryPath(globalStoragePath, taskId) + await fs.rm(dirPath, { recursive: true, force: true }) + console.log(`[deleteTaskWithId${taskId}] removed task directory`) + } catch (error) { + console.error( + `[deleteTaskWithId${taskId}] failed to remove task directory: ${error instanceof Error ? error.message : String(error)}`, + ) + } } await this.postStateToWebview() diff --git a/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts b/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts index 8f04d36ae6..7d8493fba3 100644 --- a/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.sticky-profile.spec.ts @@ -88,17 +88,6 @@ vi.mock("../../prompts/sections/custom-instructions") vi.mock("../../../utils/safeWriteJson") -// The store's cross-process task guard and advisory file lock need a real -// filesystem, which this spec stubs out. The store behavior under test does -// not depend on them. -vi.mock("../../../utils/fileLock", async () => ({ - ...(await vi.importActual("../../../utils/fileLock")), - acquireFileLock: vi.fn().mockResolvedValue({ - release: vi.fn().mockResolvedValue(undefined), - isCompromised: vi.fn().mockReturnValue(false), - }), -})) - vi.mock("../../../api", () => ({ buildApiHandler: vi.fn().mockReturnValue({ getModel: vi.fn().mockReturnValue({ diff --git a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts index f8e1a06d77..2bbf0736c6 100644 --- a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts @@ -59,17 +59,6 @@ vi.mock("../../../utils/safeWriteJson", () => ({ safeWriteJson: vi.fn().mockResolvedValue(undefined), })) -// The store's cross-process task guard and advisory file lock need a real -// filesystem, which this spec stubs out. The store behavior under test does -// not depend on them. -vi.mock("../../../utils/fileLock", async () => ({ - ...(await vi.importActual("../../../utils/fileLock")), - acquireFileLock: vi.fn().mockResolvedValue({ - release: vi.fn().mockResolvedValue(undefined), - isCompromised: vi.fn().mockReturnValue(false), - }), -})) - vi.mock("@modelcontextprotocol/sdk/types.js", () => ({ CallToolResultSchema: {}, ListResourcesResultSchema: {}, diff --git a/src/eslint-suppressions.json b/src/eslint-suppressions.json index 5d8951d312..3f71416a12 100644 --- a/src/eslint-suppressions.json +++ b/src/eslint-suppressions.json @@ -1686,7 +1686,7 @@ }, "utils/__tests__/safeWriteJson.test.ts": { "@typescript-eslint/no-explicit-any": { - "count": 26 + "count": 27 } }, "utils/__tests__/shell.spec.ts": { @@ -1716,7 +1716,7 @@ }, "utils/safeWriteJson.ts": { "@typescript-eslint/no-explicit-any": { - "count": 2 + "count": 4 } }, "utils/tts.ts": { diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index b3f34022b9..79d08678a0 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -158,44 +158,57 @@ describe("safeWriteJson", () => { expect(content).toEqual({ initial: "content" }) }) - test("should leave the original file in place when the commit rename fails", async () => { + test("should handle failure when renaming filePath to tempBackupFilePath (filePath exists)", async () => { const initialData = { message: "Initial content, should remain" } const newData = { message: "New content, should not be written" } // Overwrite the pre-created file with specific initial data await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - // The replacement is a single rename, so its failure leaves the - // original file untouched — the target is never missing. + // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn vi.mocked(fs.rename).mockImplementationOnce(async () => { - throw new Error("Rename from temp to final failed") + throw new Error("Rename to backup failed") }) - await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename from temp to final failed") + await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename to backup failed") + // Verify the original file still exists with initial content const content = await readFileContent(currentTestFilePath) expect(content).toEqual(initialData) }) - test("should remove leftover temp and backup orphans for the target before writing", async () => { - const orphanNew = path.join(tempDir, ".test-file.json.new_123_abc.tmp") - const orphanBackup = path.join(tempDir, ".test-file.json.bak_456_def.tmp") - const orphanOtherTarget = path.join(tempDir, ".other-file.json.new_789_ghi.tmp") - await fs.writeFile(orphanNew, "{}") - await fs.writeFile(orphanBackup, "{}") - await fs.writeFile(orphanOtherTarget, "{}") + test("should handle failure when renaming tempNewFilePath to filePath (filePath exists, backup succeeded)", async () => { + const initialData = { message: "Initial content, should be restored" } + const newData = { message: "New content" } - const data = { message: "written after orphan recovery" } - await safeWriteJson(currentTestFilePath, data) + // Overwrite the pre-created file with specific initial data + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - // Orphans of this target are removed while holding the lock. - await expect(fileExists(orphanNew)).resolves.toBe(false) - await expect(fileExists(orphanBackup)).resolves.toBe(false) - // Temp files of other targets are left alone. - await expect(fileExists(orphanOtherTarget)).resolves.toBe(true) + // Track rename calls + let renameCallCount = 0 + + // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { + renameCallCount++ + if (renameCallCount === 1) { + // First call: filePath -> tempBackupFilePath (should succeed) + return fsPromisesActuals.rename!(oldPath, newPath) + } else if (renameCallCount === 2) { + // Second call: tempNewFilePath -> filePath (should fail) + throw new Error("Rename from temp to final failed") + } else if (renameCallCount === 3) { + // Third call: tempBackupFilePath -> filePath (rollback, should succeed) + return fsPromisesActuals.rename!(oldPath, newPath) + } + // Default: use original implementation + return fsPromisesActuals.rename!(oldPath, newPath) + }) + await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename from temp to final failed") + + // Verify the file was restored to initial content const content = await readFileContent(currentTestFilePath) - expect(content).toEqual(data) + expect(content).toEqual(initialData) }) // Tests for directory creation functionality @@ -279,6 +292,78 @@ describe("safeWriteJson", () => { expect(content).toEqual(data) }) + test("should handle failure when deleting tempBackupFilePath (filePath exists, all renames succeed)", async () => { + const initialData = { message: "Initial content" } + const newData = { message: "Successfully written new content" } + + // Overwrite the pre-created file with specific initial data + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + + // fs.unlink is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + vi.mocked(fs.unlink).mockImplementationOnce(async () => { + throw new Error("Failed to delete backup file") + }) + + // The write should succeed even if backup deletion fails + await safeWriteJson(currentTestFilePath, newData) + + // Verify the new content was written successfully + const content = await readFileContent(currentTestFilePath) + expect(content).toEqual(newData) + }) + + // Test for console error suppression during backup deletion + test("should suppress console.error when backup deletion fails", async () => { + const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) // Suppress console.error + const initialData = { message: "Initial" } + const newData = { message: "New" } + + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + + // fs.unlink is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + vi.mocked(fs.unlink).mockImplementation(async (filePath: any) => { + if (filePath.toString().includes(".bak_")) { + throw new Error("Backup deletion failed") + } + return fsPromisesActuals.unlink!(filePath) + }) + + await safeWriteJson(currentTestFilePath, newData) + + // Verify console.error was called with the expected message + expect(consoleErrorSpy).toHaveBeenCalledWith(expect.stringContaining("Successfully wrote"), expect.any(Error)) + + consoleErrorSpy.mockRestore() + vi.mocked(fs.unlink).mockRestore() + }) + + // The expected error message might need to change if the mock behaves differently. + test("should handle failure when renaming tempNewFilePath to filePath (filePath initially exists)", async () => { + // currentTestFilePath exists due to beforeEach. + const initialData = { message: "Initial content" } + const newData = { message: "New content" } + + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + + // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + let renameCallCount = 0 + vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { + renameCallCount++ + if (renameCallCount === 2) { + // Second call: tempNewFilePath -> filePath (should fail) + throw new Error("Rename failed") + } + // For all other calls, use the original implementation + return fsPromisesActuals.rename!(oldPath, newPath) + }) + + await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Rename failed") + + // The file should be restored to its initial content + const content = await readFileContent(currentTestFilePath) + expect(content).toEqual(initialData) + }) + test("should throw an error if an inter-process lock is already held for the filePath", async () => { vi.resetModules() // Clear module cache to ensure fresh imports for this test @@ -349,6 +434,41 @@ describe("safeWriteJson", () => { expect(vi.mocked(fs.access)).toHaveBeenCalled() }) + // Test for rollback failure scenario + test("should log error and re-throw original if rollback fails", async () => { + const initialData = { message: "Initial, should be lost if rollback fails" } + const newData = { message: "New content" } + + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + + const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) // Suppress console.error + + // fs.rename is already vi.fn() — use vi.mocked to avoid double-wrapping via vi.spyOn + let renameCallCount = 0 + vi.mocked(fs.rename).mockImplementation(async (oldPath, newPath) => { + renameCallCount++ + if (renameCallCount === 2) { + // Second call: tempNewFilePath -> filePath (fail) + throw new Error("Primary rename failed") + } else if (renameCallCount === 3) { + // Third call: tempBackupFilePath -> filePath (rollback, also fail) + throw new Error("Rollback rename failed") + } + return fsPromisesActuals.rename!(oldPath, newPath) + }) + + // Should throw the original error, not the rollback error + await expect(safeWriteJson(currentTestFilePath, newData)).rejects.toThrow("Primary rename failed") + + // Verify console.error was called for the rollback failure + expect(consoleErrorSpy).toHaveBeenCalledWith( + expect.stringContaining("Failed to restore backup"), + expect.objectContaining({ message: "Rollback rename failed" }), + ) + + consoleErrorSpy.mockRestore() + }) + // Merge option tests test("should merge incoming data with existing file content when merge callback is provided", async () => { const initial = { a: 1, b: 2 } @@ -422,68 +542,4 @@ describe("safeWriteJson", () => { const content = await readFileContent(currentTestFilePath) expect(content).toEqual({ c: 3 }) }) - - test("should remove no orphan temp files when the lock is compromised before orphan cleanup", async () => { - vi.resetModules() - - const compromisedPath = path.join(tempDir, "compromised-cleanup-file.json") - await fsPromisesActuals.writeFile!(compromisedPath, JSON.stringify({ initial: "content" })) - - const orphanNew = path.join(tempDir, ".compromised-cleanup-file.json.new_123_abc.tmp") - const orphanBackup = path.join(tempDir, ".compromised-cleanup-file.json.bak_456_def.tmp") - await fs.writeFile(orphanNew, "{}") - await fs.writeFile(orphanBackup, "{}") - - // Simulate proper-lockfile reporting the lost lock right after - // acquisition, before any file operation runs. - vi.doMock("proper-lockfile", () => ({ - lock: (_file: string, options: { onCompromised: (error: Error) => void }) => { - options.onCompromised(new Error("Lock no longer available")) - return Promise.resolve(async () => {}) - }, - })) - - const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") - - await expect(mockedSafeWriteJson(compromisedPath, { replaced: true })).rejects.toThrow("was compromised") - - // The compromised lock must not classify any temp file as an orphan: - // without exclusion it could belong to a live peer writer. - await expect(fileExists(orphanNew)).resolves.toBe(true) - await expect(fileExists(orphanBackup)).resolves.toBe(true) - - // The write aborted, so the target kept its original content. - const content = await readFileContent(compromisedPath) - expect(content).toEqual({ initial: "content" }) - - vi.doUnmock("proper-lockfile") - }) - - test("should abort before replacing the target when the lock is compromised", async () => { - vi.resetModules() - - const compromisedPath = path.join(tempDir, "compromised-lock-file.json") - await fsPromisesActuals.writeFile!(compromisedPath, JSON.stringify({ initial: "content" })) - - // Simulate proper-lockfile detecting the lost lock right after - // acquisition: it reports compromise from its update timer instead of - // rejecting the lock() promise. - vi.doMock("proper-lockfile", () => ({ - lock: (_file: string, options: { onCompromised: (error: Error) => void }) => { - options.onCompromised(new Error("Lock no longer available")) - return Promise.resolve(async () => {}) - }, - })) - - const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") - - await expect(mockedSafeWriteJson(compromisedPath, { replaced: true })).rejects.toThrow("was compromised") - - // The write aborted before the commit rename, so the target kept its - // original content instead of being replaced without exclusion. - const content = await readFileContent(compromisedPath) - expect(content).toEqual({ initial: "content" }) - - vi.unmock("proper-lockfile") - }) }) diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts deleted file mode 100644 index ded0e701e9..0000000000 --- a/src/utils/fileLock.ts +++ /dev/null @@ -1,73 +0,0 @@ -import * as path from "path" -import * as lockfile from "proper-lockfile" - -export const LOCK_STALE_MS = 31_000 - -/** - * Retry budget for destructive mutations. The backoff outlasts - * LOCK_STALE_MS, so a lock left behind by a crashed process is broken - * within the same acquisition instead of failing the mutation. - */ -export const DESTRUCTIVE_LOCK_RETRIES = { - retries: 36, - factor: 2, - minTimeout: 100, - maxTimeout: 1_000, -} - -export interface AcquiredFileLock { - release: () => Promise - /** - * True once proper-lockfile reported the held lock compromised, for - * example when another host's stale-lock recovery removed the lock - * directory. Callers must check this before mutating the protected - * target and fail the operation instead of writing without exclusion. - */ - isCompromised: () => boolean -} - -/** - * Fail with one deterministic error when the held lock was reported - * compromised, so the caller aborts before the next mutation of the - * protected target instead of writing or deleting without exclusion. - */ -export function assertLockUsable(lock: AcquiredFileLock, targetPath: string, action: string): void { - if (lock.isCompromised()) { - throw new Error(`Lock for ${targetPath} was compromised before ${action}`) - } -} - -export interface AcquireFileLockOptions { - retries?: lockfile.LockOptions["retries"] -} - -/** - * Acquire the advisory lock shared by JSON writes and destructive mutations. - * The target may not exist yet, so callers can use the same protocol for - * creation, replacement, and deletion. - */ -export function acquireFileLock(filePath: string, options?: AcquireFileLockOptions): Promise { - const absoluteFilePath = path.resolve(filePath) - let compromised = false - return lockfile - .lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10_000, - realpath: false, - retries: options?.retries ?? { - retries: 5, - factor: 2, - minTimeout: 100, - maxTimeout: 1_000, - }, - // proper-lockfile invokes this from a timer callback after - // acquisition, so a throw here never reaches the awaited operation. - // Record the state instead; callers check isCompromised before - // mutating and fail the operation themselves. - onCompromised: (error) => { - compromised = true - console.error(`Lock at ${absoluteFilePath} was compromised:`, error) - }, - }) - .then((release) => ({ release, isCompromised: () => compromised })) -} diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index c5be4940a4..957a0bb20f 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -1,10 +1,9 @@ import * as fs from "fs/promises" import * as fsSync from "fs" import * as path from "path" +import * as lockfile from "proper-lockfile" import { JsonStreamStringify } from "json-stream-stringify" -import { acquireFileLock, assertLockUsable, LOCK_STALE_MS } from "./fileLock" - /** * Options for safeWriteJson function */ @@ -32,13 +31,9 @@ export interface SafeWriteJsonOptions { * Safely writes JSON data to a file. * - Creates parent directories if they don't exist * - Uses 'proper-lockfile' for inter-process advisory locking to prevent concurrent writes to the same path. - * - Removes leftover temp files for this target left by a crashed writer. - * - Writes to a temporary file first, then replaces the target with one - * atomic rename, so the target is never missing between the two states. - * - Cleans up the temporary file in case of errors. - * - Aborts before orphan cleanup and before replacing the target when the - * advisory lock was compromised, so it never writes or deletes without - * exclusion. + * - Writes to a temporary file first. + * - If the target file exists, it's backed up before being replaced. + * - Attempts to roll back and clean up in case of errors. * - Supports pretty-printing with indentation while maintaining streaming efficiency. * * @param {string} filePath - The absolute path to the target file. @@ -49,35 +44,58 @@ export interface SafeWriteJsonOptions { async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJsonOptions): Promise { const absoluteFilePath = path.resolve(filePath) + let releaseLock = async () => {} // Initialized to a no-op // For directory creation const dirPath = path.dirname(absoluteFilePath) - // Ensure directory structure exists + // Ensure directory structure exists with improved reliability try { + // Create directory with recursive option await fs.mkdir(dirPath, { recursive: true }) + + // Verify directory exists after creation attempt await fs.access(dirPath) - } catch (dirError: unknown) { + } catch (dirError: any) { console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) throw dirError } - // Acquire the lock before any file operations. On failure the release - // helper stays unused and the error propagates. - let lock: Awaited> + // Acquire the lock before any file operations try { - lock = await acquireFileLock(absoluteFilePath) + releaseLock = await lockfile.lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long + realpath: false, // the file may not exist yet, which is acceptable + retries: { + // Configuration for retrying lock acquisition + retries: 5, // Number of retries after the initial attempt + factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) + minTimeout: 100, // Minimum time to wait before the first retry (in ms) + maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) + }, + onCompromised: (err) => { + console.error(`Lock at ${absoluteFilePath} was compromised:`, err) + throw err + }, + }) } catch (lockError) { + // If lock acquisition fails, we throw immediately. + // The releaseLock remains a no-op, so the finally block in the main file operations + // try-catch-finally won't try to release an unacquired lock if this path is taken. console.error(`Failed to acquire lock for ${absoluteFilePath}:`, lockError) + // Propagate the lock acquisition error throw lockError } - // Path of the temporary file while it exists, so the error path can clean it up. + // Variables to hold the actual paths of temp files if they are created. let actualTempNewFilePath: string | null = null + let actualTempBackupFilePath: string | null = null try { // If a merge callback was provided, read the current file under the lock - // and let the caller merge before we write. + // and let the caller merge before we write. Must be inside try/finally + // so a throwing merge still releases the lock. if (options?.merge) { let existing: unknown = null try { @@ -92,47 +110,108 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso data = options.merge(existing, data) } - // A compromised lock no longer excludes a peer writer, so its temp - // files are not orphans. Abort before discovery and again before - // each removal instead of deleting a live writer's temp files. - assertLockUsable(lock, absoluteFilePath, "orphan cleanup") - await removeLeftoverTempFiles(dirPath, path.basename(absoluteFilePath), () => - assertLockUsable(lock, absoluteFilePath, "orphan cleanup"), - ) - // Step 1: Write data to a new temporary file. actualTempNewFilePath = path.join( - dirPath, + path.dirname(absoluteFilePath), `.${path.basename(absoluteFilePath)}.new_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, ) await _streamDataToFile(actualTempNewFilePath, data, options?.prettyPrint) - // A compromised lock means another host may be mutating the target. - // Fail here instead of replacing the target without exclusion. - assertLockUsable(lock, absoluteFilePath, "commit") + // Step 2: Check if the target file exists. If so, rename it to a backup path. + try { + // Check for target file existence + await fs.access(absoluteFilePath) + // Target exists, create a backup path and rename. + actualTempBackupFilePath = path.join( + path.dirname(absoluteFilePath), + `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, + ) + await fs.rename(absoluteFilePath, actualTempBackupFilePath) + } catch (accessError: any) { + // Explicitly type accessError + if (accessError.code !== "ENOENT") { + // An error other than "file not found" occurred during access check. + throw accessError + } + // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. + } - // Step 2: Replace the target with one atomic rename. The target holds - // either the old content or the new content at every instant, so a - // crash cannot leave it missing. + // Step 3: Rename the new temporary file to the target file path. + // This is the main "commit" step. await fs.rename(actualTempNewFilePath, absoluteFilePath) + + // If we reach here, the new file is successfully in place. + // The original actualTempNewFilePath is now the main file, so we shouldn't try to clean it up as "temp". + // Mark as "used" or "committed" actualTempNewFilePath = null + + // Step 4: If a backup was created, attempt to delete it. + if (actualTempBackupFilePath) { + try { + await fs.unlink(actualTempBackupFilePath) + // Mark backup as handled + actualTempBackupFilePath = null + } catch (unlinkBackupError) { + // Log this error, but do not re-throw. The main operation was successful. + // actualTempBackupFilePath remains set, indicating an orphaned backup. + console.error( + `Successfully wrote ${absoluteFilePath}, but failed to clean up backup ${actualTempBackupFilePath}:`, + unlinkBackupError, + ) + } + } } catch (originalError) { - console.error(`Operation failed for ${absoluteFilePath}:`, originalError) + console.error(`Operation failed for ${absoluteFilePath}: [Original Error Caught]`, originalError) + + const newFileToCleanupWithinCatch = actualTempNewFilePath + const backupFileToRollbackOrCleanupWithinCatch = actualTempBackupFilePath + + // Attempt rollback if a backup was made + if (backupFileToRollbackOrCleanupWithinCatch) { + try { + await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) + // Mark as handled, prevent later unlink of this path + actualTempBackupFilePath = null + } catch (rollbackError) { + // actualTempBackupFilePath (outer scope) remains pointing to backupFileToRollbackOrCleanupWithinCatch + console.error( + `[Catch] Failed to restore backup ${backupFileToRollbackOrCleanupWithinCatch} to ${absoluteFilePath}:`, + rollbackError, + ) + } + } - if (actualTempNewFilePath) { + // Cleanup the .new file if it exists + if (newFileToCleanupWithinCatch) { try { - await fs.unlink(actualTempNewFilePath) + await fs.unlink(newFileToCleanupWithinCatch) } catch (cleanupError) { - console.error(`[Catch] Failed to clean up temporary new file ${actualTempNewFilePath}:`, cleanupError) + console.error( + `[Catch] Failed to clean up temporary new file ${newFileToCleanupWithinCatch}:`, + cleanupError, + ) } } - throw originalError + // Cleanup the .bak file if it still needs to be (i.e., wasn't successfully restored) + if (actualTempBackupFilePath) { + try { + await fs.unlink(actualTempBackupFilePath) + } catch (cleanupError) { + console.error( + `[Catch] Failed to clean up temporary backup file ${actualTempBackupFilePath}:`, + cleanupError, + ) + } + } + throw originalError // This MUST be the error that rejects the promise. } finally { // Release the lock in the main finally block. try { - await lock.release() + // releaseLock will be the actual unlock function if lock was acquired, + // or the initial no-op if acquisition failed. + await releaseLock() } catch (unlockError) { // Do not re-throw here, as the originalError from the try/catch (if any) is more important. console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) @@ -140,43 +219,6 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } -/** - * Remove leftover `..new_*.tmp` and `..bak_*.tmp` files. - * Safe while the caller holds the advisory lock for `targetBasename`. - * @param dirPath The directory holding the target file. - * @param targetBasename The target file's base name. - * @param assertLockUsable Called before each removal so the caller can - * abort while the lock is compromised. - */ -async function removeLeftoverTempFiles( - dirPath: string, - targetBasename: string, - assertLockUsable: () => void, -): Promise { - const newPrefix = `.${targetBasename}.new_` - const backupPrefix = `.${targetBasename}.bak_` - let entries: string[] - try { - entries = await fs.readdir(dirPath) - } catch { - return - } - for (const entry of entries) { - if (!entry.endsWith(".tmp")) { - continue - } - if (!entry.startsWith(newPrefix) && !entry.startsWith(backupPrefix)) { - continue - } - assertLockUsable() - try { - await fs.unlink(path.join(dirPath, entry)) - } catch (error) { - console.error(`Failed to clean up leftover temp file ${entry} for ${targetBasename}:`, error) - } - } -} - /** * Helper function to stream JSON data to a file. * @param targetPath The path to write the stream to. @@ -205,4 +247,6 @@ async function _streamDataToFile(targetPath: string, data: any, prettyPrint = fa }) } -export { LOCK_STALE_MS, safeWriteJson } +export const LOCK_STALE_MS = 31_000 + +export { safeWriteJson } From bbde189fab9e24d9924ee728e7a41caa13b1d909 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Thu, 24 Sep 2026 02:12:40 +0000 Subject: [PATCH 09/23] fix(lifecycle): handle authoritative settlement results --- .../ClineProvider.delegation.spec.ts | 66 +++++++++++++++++++ src/core/task-persistence/TaskHistoryStore.ts | 53 ++++++++++----- .../TaskHistoryStore.realConcurrency.spec.ts | 23 +++++++ .../task/__tests__/Task.persistence.spec.ts | 33 ++++++++++ src/core/webview/ClineProvider.ts | 8 ++- src/utils/__tests__/safeWriteJson.test.ts | 15 ++++- src/utils/safeWriteJson.ts | 28 +++++--- 7 files changed, 197 insertions(+), 29 deletions(-) diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index a741f1815d..cf3e9dfdc3 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -1016,6 +1016,72 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { expect(createTaskWithHistoryItem).not.toHaveBeenCalled() }) + it("does not restore a completed parent when settlement preserves the rejected action", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + const interruptedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const completedParent: HistoryItem = { + ...parentHistoryItem, + status: "completed", + pendingAction, + } + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const clearPendingActionIfMatching = vi.fn().mockResolvedValue(completedParent) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const provider = { + taskScheduler: new TaskScheduler(), + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId: vi.fn().mockResolvedValue({ historyItem: completedParent }), + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + deleteTaskWithId: vi.fn().mockResolvedValue(undefined), + createTaskWithHistoryItem, + log: vi.fn(), + isViewLaunched: false, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => interruptedParent), + atomicReadAndUpdate: vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + updater(interruptedParent) + return [] + }), + clearPendingActionIfMatching, + }, + } as unknown as ClineProvider + + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow("Invalid task status transition: interrupted → delegated") + + expect(clearPendingActionIfMatching).toHaveBeenCalledWith("parent-1", pendingAction.actionId) + expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) + expect(createTaskWithHistoryItem).not.toHaveBeenCalled() + }) + it("restores the authoritative parent record when settlement preserves a replacement action", async () => { const pendingAction = { kind: "create_subtask" as const, diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index d6c42e3280..5f70f47c11 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -1085,22 +1085,43 @@ export class TaskHistoryStore { } const filePath = await this.getTaskFilePath(taskId) let authoritative: HistoryItem = cached - await safeWriteJson(filePath, cached, { - merge: (existing) => { - if (!existing || typeof existing !== "object" || !("id" in existing)) { - // Writing the cached record back would recreate a task - // another host deleted, so drop the stale entry first. - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - throw new Error( - `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, - ) - } - const disk = existing as HistoryItem - authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) - return authoritative - }, - }) + let missingDiskRecord = false + try { + await safeWriteJson(filePath, cached, { + createParentDirectory: false, + merge: (existing) => { + if (!existing || typeof existing !== "object" || !("id" in existing)) { + // Writing the cached record back would recreate a task + // another host deleted, so drop the stale entry first. + missingDiskRecord = true + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, + ) + } + const disk = existing as HistoryItem + authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) + return authoritative + }, + }) + } catch (error) { + const missingLockPath = + error && + typeof error === "object" && + "code" in error && + error.code === "ENOENT" && + "path" in error && + error.path === `${filePath}.lock` + if (missingDiskRecord || missingLockPath) { + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, + ) + } + throw error + } this.cache.set(taskId, authoritative) if (this.onWrite) { await this.onWrite(this.getAll()) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index 5040558fed..fea7a1d805 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -215,6 +215,29 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("rejects and evicts stale cache without recreating artifacts after full directory deletion", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-deleted-directory-settlement-")) + const store = new TaskHistoryStore(storagePath) + const action = createAction("action-a", "action A") + const taskDirectory = path.join(storagePath, "tasks", "shared-task") + + try { + await store.initialize() + await store.upsert({ ...item("shared-task"), pendingAction: action }) + await fs.rm(taskDirectory, { recursive: true }) + + await expect(store.clearPendingActionIfMatching("shared-task", action.actionId)).rejects.toThrow( + "task shared-task not found", + ) + expect(store.get("shared-task")).toBeUndefined() + await expect(fs.access(taskDirectory)).rejects.toMatchObject({ code: "ENOENT" }) + expect(await fs.readdir(path.join(storagePath, "tasks"))).not.toContain("shared-task") + } finally { + store.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("rejects settlement for a task absent from the cache without creating it", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-cache-miss-settlement-")) const store = new TaskHistoryStore(storagePath) diff --git a/src/core/task/__tests__/Task.persistence.spec.ts b/src/core/task/__tests__/Task.persistence.spec.ts index bc6e9c3960..3dba6a42c9 100644 --- a/src/core/task/__tests__/Task.persistence.spec.ts +++ b/src/core/task/__tests__/Task.persistence.spec.ts @@ -1391,6 +1391,39 @@ describe("Task persistence", () => { expect(mockSaveTaskMessages).not.toHaveBeenCalled() }) + it("replays a create-subtask action for an active historical task without restart settlement", async () => { + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + const clearRejectedAction = vi.fn() + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "active", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const replay = vi + .spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + .mockResolvedValue(undefined) + + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + expect(clearRejectedAction).not.toHaveBeenCalled() + expect(replay).toHaveBeenCalledWith(createSubtaskAction) + }) + it("settles an interrupted create-subtask action before restart replay", async () => { mockReadTaskMessages.mockResolvedValue([ { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 86121bf3a4..0bf4848ecf 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -4067,7 +4067,13 @@ export class ClineProvider let settlementFailed = false if (pendingActionId && err instanceof LifecycleTransitionError) { try { - await this.taskHistoryStore.clearPendingActionIfMatching(parentTaskId, pendingActionId) + const authoritative = await this.taskHistoryStore.clearPendingActionIfMatching( + parentTaskId, + pendingActionId, + ) + settlementFailed = + authoritative.pendingAction?.kind === "create_subtask" && + authoritative.pendingAction.actionId === pendingActionId this.recentTasksCache = undefined } catch (settlementError) { settlementFailed = true diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index 79d08678a0..400ab97453 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -232,6 +232,17 @@ describe("safeWriteJson", () => { expect(content).toEqual(data) }) + test("should reject without creating the parent directory when parent creation is disabled", async () => { + const subDir = path.join(tempDir, "missing-parent") + const filePath = path.join(subDir, "file.json") + + await expect(safeWriteJson(filePath, { value: 1 }, { createParentDirectory: false })).rejects.toMatchObject({ + code: "ENOENT", + path: `${filePath}.lock`, + }) + await expect(fs.access(subDir)).rejects.toMatchObject({ code: "ENOENT" }) + }) + test("should handle multi-level directory creation", async () => { // Create a new non-existent subdirectory path with multiple levels const deepDir = path.join(tempDir, "level1", "level2", "level3") @@ -486,11 +497,11 @@ describe("safeWriteJson", () => { expect(content).toEqual({ a: 1, b: 3, c: 4 }) }) - test("should pass null to merge callback when file does not exist", async () => { + test("should pass null to merge callback under the lock when the parent exists but the file does not", async () => { const newFilePath = path.join(tempDir, "nonexistent.json") const mergeFn = vi.fn((existing, incoming) => incoming) - await safeWriteJson(newFilePath, { value: 42 }, { merge: mergeFn }) + await safeWriteJson(newFilePath, { value: 42 }, { createParentDirectory: false, merge: mergeFn }) expect(mergeFn).toHaveBeenCalledWith(null, { value: 42 }) const content = await readFileContent(newFilePath) diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 957a0bb20f..14dc92834f 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -8,6 +8,12 @@ import { JsonStreamStringify } from "json-stream-stringify" * Options for safeWriteJson function */ export interface SafeWriteJsonOptions { + /** + * Whether to create and verify the target file's parent directory. + * @default true + */ + createParentDirectory?: boolean + /** * Whether to pretty-print the JSON output with indentation. * When true, uses tab characters for indentation. @@ -49,16 +55,18 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // For directory creation const dirPath = path.dirname(absoluteFilePath) - // Ensure directory structure exists with improved reliability - try { - // Create directory with recursive option - await fs.mkdir(dirPath, { recursive: true }) - - // Verify directory exists after creation attempt - await fs.access(dirPath) - } catch (dirError: any) { - console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) - throw dirError + if (options?.createParentDirectory !== false) { + // Ensure directory structure exists with improved reliability + try { + // Create directory with recursive option + await fs.mkdir(dirPath, { recursive: true }) + + // Verify directory exists after creation attempt + await fs.access(dirPath) + } catch (dirError: any) { + console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) + throw dirError + } } // Acquire the lock before any file operations From 7018228027816d8797a987bc868706445282751a Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Sun, 27 Sep 2026 23:31:31 +0000 Subject: [PATCH 10/23] test(lifecycle): cover settlement lock errors --- .../TaskHistoryStore.crossInstance.spec.ts | 21 +++++++++++++++++++ 1 file changed, 21 insertions(+) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts index cef9874e5f..0de14134fc 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts @@ -204,6 +204,27 @@ describe("TaskHistoryStore cross-instance safety", () => { expect(storeB.getAll().length).toBe(10) }) + it("preserves the cached record when rejected-action settlement fails with an unrelated lock error", async () => { + await storeA.initialize() + const pendingAction = { + kind: "create_subtask" as const, + actionId: "action-a", + approvalText: "{}", + mode: "code", + message: "action A", + todos: [], + } + const cached = makeHistoryItem({ id: "settlement-error-task", pendingAction }) + await storeA.upsert(cached) + + const { safeWriteJson } = await import("../../../utils/safeWriteJson") + const lockError = Object.assign(new Error("lock acquisition failed"), { code: "ELOCKED" }) + vi.mocked(safeWriteJson).mockRejectedValueOnce(lockError) + + await expect(storeA.clearPendingActionIfMatching(cached.id, pendingAction.actionId)).rejects.toBe(lockError) + expect(storeA.get(cached.id)).toEqual(cached) + }) + /** * Host B completes a task on disk while host A's cache still has it * active. Host A's next save updates only totalCost (a full-object From 3ff37679bbbc1e86b4857007d1a7dc3e6945550b Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 00:08:40 +0000 Subject: [PATCH 11/23] fix(lifecycle): harden rejected action settlement --- .../ClineProvider.delegation.spec.ts | 2 + .../removeClineFromStack-delegation.spec.ts | 29 +++++++++++++ src/core/task-persistence/TaskHistoryStore.ts | 1 + .../TaskHistoryStore.crossInstance.spec.ts | 28 ++++++++++++ src/core/webview/ClineProvider.ts | 43 ++++++++++++++++--- src/utils/__tests__/safeWriteJson.test.ts | 12 ++++++ src/utils/safeWriteJson.ts | 42 +++++++++++------- 7 files changed, 136 insertions(+), 21 deletions(-) diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index cf3e9dfdc3..e337fcc2e3 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -792,6 +792,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { } const provider = { taskScheduler: new TaskScheduler(), + recentTasksCache: [parentHistoryItem], emit: vi.fn(), getCurrentTask, removeClineFromStack: vi.fn().mockResolvedValue(undefined), @@ -817,6 +818,7 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { expect(current.pendingAction).toBeUndefined() expect(current.status).toBe("interrupted") + expect((provider as unknown as { recentTasksCache?: HistoryItem[] }).recentTasksCache).toBeUndefined() expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) expect(provider.createTaskWithHistoryItem).toHaveBeenCalledWith( expect.objectContaining({ diff --git a/src/__tests__/removeClineFromStack-delegation.spec.ts b/src/__tests__/removeClineFromStack-delegation.spec.ts index e211a8f4ab..7ac6359df5 100644 --- a/src/__tests__/removeClineFromStack-delegation.spec.ts +++ b/src/__tests__/removeClineFromStack-delegation.spec.ts @@ -14,6 +14,7 @@ type MockTask = Pick & type PrivateClineProviderMethods = { removeClineFromStack: (this: ClineProvider) => ReturnType + cleanupFailedHistoryTask: (this: ClineProvider, task: Task) => Promise markDelegatedChildInterrupted: ( this: ClineProvider, ...args: Parameters @@ -176,6 +177,34 @@ describe("ClineProvider.removeClineFromStack() — pure lifecycle, no delegation }) }) +describe("ClineProvider failed history restoration cleanup", () => { + it("removes the failed task, its listeners, and its resources without saving stale history", async () => { + const cleanupListener = vi.fn() + const task = { + taskId: "failed-history-task", + instanceId: "inst-1", + emit: vi.fn(), + dispose: vi.fn().mockResolvedValue(undefined), + } as unknown as Task + const taskRegistry = new TaskRegistry() + taskRegistry.push(task) + const taskEventListeners = new Map([[task, [cleanupListener]]]) + const provider = { + taskRegistry, + taskEventListeners, + log: vi.fn(), + } as unknown as ClineProvider + + await privateClineProvider.cleanupFailedHistoryTask.call(provider, task) + + expect(taskRegistry.getById(task.taskId)).toBeUndefined() + expect(taskRegistry.current).toBeUndefined() + expect(cleanupListener).toHaveBeenCalledOnce() + expect(taskEventListeners.has(task)).toBe(false) + expect(task.dispose).toHaveBeenCalledOnce() + }) +}) + describe("ClineProvider.markDelegatedChildInterrupted() — live eviction path", () => { it("marks an active delegated child interrupted and leaves parent delegated", async () => { const childTaskId = "child-1" diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 5f70f47c11..3cfd0350a1 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -1089,6 +1089,7 @@ export class TaskHistoryStore { try { await safeWriteJson(filePath, cached, { createParentDirectory: false, + atomicReplace: true, merge: (existing) => { if (!existing || typeof existing !== "object" || !("id" in existing)) { // Writing the cached record back would recreate a task diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts index 0de14134fc..fe7846edd3 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts @@ -225,6 +225,34 @@ describe("TaskHistoryStore cross-instance safety", () => { expect(storeA.get(cached.id)).toEqual(cached) }) + it("evicts stale cache state when the disk record is invalid without recreating it", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "action-a", + approvalText: "{}", + mode: "code", + message: "action A", + todos: [], + } + const cached = makeHistoryItem({ id: "invalid-settlement-task", pendingAction }) + const filePath = path.join(tmpDir, "tasks", cached.id, GlobalFileNames.historyItem) + await fs.mkdir(path.dirname(filePath), { recursive: true }) + await fs.writeFile(filePath, "{invalid", "utf8") + const storeState = storeA as unknown as { + cache: Map + taskFileMtimes: Map + } + storeState.cache.set(cached.id, cached) + storeState.taskFileMtimes.set(cached.id, Date.now()) + + await expect(storeA.clearPendingActionIfMatching(cached.id, pendingAction.actionId)).rejects.toThrow( + `task ${cached.id} not found in cache`, + ) + expect(storeA.get(cached.id)).toBeUndefined() + expect(storeState.taskFileMtimes.has(cached.id)).toBe(false) + expect(await fs.readFile(filePath, "utf8")).toBe("{invalid") + }) + /** * Host B completes a task on disk while host A's cache still has it * active. Host A's next save updates only totalCost (a full-object diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 0bf4848ecf..90dd957854 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -177,10 +177,16 @@ function scheduleTask( task: Task, source: string, run: () => Promise = () => task.run(), + onError?: (error: unknown) => void | Promise, ): void { - void scheduler - .schedule(task, run) - .catch((error) => console.error(`[${source}] taskScheduler.schedule failed:`, error)) + void scheduler.schedule(task, run).catch(async (error) => { + console.error(`[${source}] taskScheduler.schedule failed:`, error) + try { + await onError?.(error) + } catch (cleanupError) { + console.error(`[${source}] task failure cleanup failed:`, cleanupError) + } + }) } type GetStateOptions = { @@ -650,6 +656,29 @@ export class ClineProvider } } + private async cleanupFailedHistoryTask(task: Task): Promise { + if (this.taskRegistry.getById(task.taskId) !== task) { + return + } + + this.taskRegistry.remove(task.taskId) + task.emit(RooCodeEventName.TaskUnfocused) + + const cleanupFunctions = this.taskEventListeners.get(task) + if (cleanupFunctions) { + cleanupFunctions.forEach((cleanup) => cleanup()) + this.taskEventListeners.delete(task) + } + + try { + await task.dispose() + } catch (error) { + this.log( + `[cleanupFailedHistoryTask] dispose() failed for ${task.taskId}.${task.instanceId}: ${error instanceof Error ? error.message : String(error)}`, + ) + } + } + /** * Evicts the current task from the stack and, if it was an active delegated child, * marks it interrupted so the parent stays delegated (rather than silently losing the link). @@ -1415,7 +1444,9 @@ export class ClineProvider ) if (options?.startTask !== false) { - scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem") + scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, () => + this.cleanupFailedHistoryTask(task), + ) } } else { await this.addClineToStack(task) @@ -1425,7 +1456,9 @@ export class ClineProvider ) if (options?.startTask !== false) { - scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem") + scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, () => + this.cleanupFailedHistoryTask(task), + ) } } diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index 400ab97453..bf244f4eb3 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -211,6 +211,18 @@ describe("safeWriteJson", () => { expect(content).toEqual(initialData) }) + test("should replace an existing file with one rename when atomicReplace is enabled", async () => { + const initialData = { message: "Initial content" } + const newData = { message: "New content" } + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + vi.mocked(fs.rename).mockClear() + + await safeWriteJson(currentTestFilePath, newData, { atomicReplace: true }) + + expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(1) + expect(await readFileContent(currentTestFilePath)).toEqual(newData) + }) + // Tests for directory creation functionality test("should create parent directory if it doesn't exist", async () => { // Create a path in a non-existent subdirectory of the temp dir diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 14dc92834f..b5ea09ea78 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -31,6 +31,14 @@ export interface SafeWriteJsonOptions { * cannot be parsed. */ merge?: (existing: unknown, incoming: unknown) => unknown + + /** + * Replace an existing target with one rename from the completed temporary file. + * Use this when an absent target during a process stop is less safe than losing + * the backup-based rollback that the default write path provides. + * @default false + */ + atomicReplace?: boolean } /** @@ -126,23 +134,25 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso await _streamDataToFile(actualTempNewFilePath, data, options?.prettyPrint) - // Step 2: Check if the target file exists. If so, rename it to a backup path. - try { - // Check for target file existence - await fs.access(absoluteFilePath) - // Target exists, create a backup path and rename. - actualTempBackupFilePath = path.join( - path.dirname(absoluteFilePath), - `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, - ) - await fs.rename(absoluteFilePath, actualTempBackupFilePath) - } catch (accessError: any) { - // Explicitly type accessError - if (accessError.code !== "ENOENT") { - // An error other than "file not found" occurred during access check. - throw accessError + if (!options?.atomicReplace) { + // Step 2: Check if the target file exists. If so, rename it to a backup path. + try { + // Check for target file existence + await fs.access(absoluteFilePath) + // Target exists, create a backup path and rename. + actualTempBackupFilePath = path.join( + path.dirname(absoluteFilePath), + `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, + ) + await fs.rename(absoluteFilePath, actualTempBackupFilePath) + } catch (accessError: any) { + // Explicitly type accessError + if (accessError.code !== "ENOENT") { + // An error other than "file not found" occurred during access check. + throw accessError + } + // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. } - // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. } // Step 3: Rename the new temporary file to the target file path. From 422166bb5abe9ead8d198e9cbb54127c25698aff Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 13:37:39 +0000 Subject: [PATCH 12/23] fix: narrow failed history cleanup --- .../removeClineFromStack-delegation.spec.ts | 36 ++++++++++++++-- src/core/task/Task.ts | 21 +++++++-- .../task/__tests__/Task.persistence.spec.ts | 12 +++++- src/core/webview/ClineProvider.ts | 16 ++++--- src/utils/__tests__/safeWriteJson.test.ts | 21 +++++++++ src/utils/safeWriteJson.ts | 43 ++++++++++--------- 6 files changed, 115 insertions(+), 34 deletions(-) diff --git a/src/__tests__/removeClineFromStack-delegation.spec.ts b/src/__tests__/removeClineFromStack-delegation.spec.ts index 7ac6359df5..bc5b8a4426 100644 --- a/src/__tests__/removeClineFromStack-delegation.spec.ts +++ b/src/__tests__/removeClineFromStack-delegation.spec.ts @@ -3,7 +3,7 @@ import { describe, it, expect, vi, type MockedFunction } from "vitest" import { ClineProvider } from "../core/webview/ClineProvider" import { TaskRegistry } from "../core/task/TaskRegistry" -import { type Task } from "../core/task/Task" +import { PendingActionSettlementError, type Task } from "../core/task/Task" import { makeProviderStub } from "./helpers/provider-stub" type MockTask = Pick & @@ -14,7 +14,7 @@ type MockTask = Pick & type PrivateClineProviderMethods = { removeClineFromStack: (this: ClineProvider) => ReturnType - cleanupFailedHistoryTask: (this: ClineProvider, task: Task) => Promise + cleanupFailedHistoryTask: (this: ClineProvider, task: Task, error: unknown) => Promise markDelegatedChildInterrupted: ( this: ClineProvider, ...args: Parameters @@ -195,7 +195,11 @@ describe("ClineProvider failed history restoration cleanup", () => { log: vi.fn(), } as unknown as ClineProvider - await privateClineProvider.cleanupFailedHistoryTask.call(provider, task) + await privateClineProvider.cleanupFailedHistoryTask.call( + provider, + task, + new PendingActionSettlementError("settlement failed"), + ) expect(taskRegistry.getById(task.taskId)).toBeUndefined() expect(taskRegistry.current).toBeUndefined() @@ -203,6 +207,32 @@ describe("ClineProvider failed history restoration cleanup", () => { expect(taskEventListeners.has(task)).toBe(false) expect(task.dispose).toHaveBeenCalledOnce() }) + + it("keeps the task active after an unrelated history resume failure", async () => { + const cleanupListener = vi.fn() + const task = { + taskId: "failed-history-task", + instanceId: "inst-1", + emit: vi.fn(), + dispose: vi.fn().mockResolvedValue(undefined), + } as unknown as Task + const taskRegistry = new TaskRegistry() + taskRegistry.push(task) + const taskEventListeners = new Map([[task, [cleanupListener]]]) + const provider = { + taskRegistry, + taskEventListeners, + log: vi.fn(), + } as unknown as ClineProvider + + await privateClineProvider.cleanupFailedHistoryTask.call(provider, task, new Error("history read failed")) + + expect(taskRegistry.getById(task.taskId)).toBe(task) + expect(taskRegistry.current).toBe(task) + expect(cleanupListener).not.toHaveBeenCalled() + expect(taskEventListeners.has(task)).toBe(true) + expect(task.dispose).not.toHaveBeenCalled() + }) }) describe("ClineProvider.markDelegatedChildInterrupted() — live eviction path", () => { diff --git a/src/core/task/Task.ts b/src/core/task/Task.ts index 7731944c08..e9ee655e22 100644 --- a/src/core/task/Task.ts +++ b/src/core/task/Task.ts @@ -216,6 +216,13 @@ type AssistantMessagePersistenceCancellation = { resolve: () => void } +export class PendingActionSettlementError extends Error { + constructor(message: string, options?: ErrorOptions) { + super(message, options) + this.name = "PendingActionSettlementError" + } +} + export class Task extends EventEmitter implements TaskLike { readonly taskId: string readonly rootTaskId?: string @@ -1008,18 +1015,26 @@ export class Task extends EventEmitter implements TaskLike { const provider = this.providerRef.deref() if (!provider) { - throw new Error( + throw new PendingActionSettlementError( `[Task#settleInterruptedCreateSubtaskBeforeReplay] Provider unavailable for task ${this.taskId}`, ) } - const authoritative = await provider.taskHistoryStore.clearPendingActionIfMatching(this.taskId, action.actionId) + let authoritative: HistoryItem + try { + authoritative = await provider.taskHistoryStore.clearPendingActionIfMatching(this.taskId, action.actionId) + } catch (error) { + throw new PendingActionSettlementError( + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Failed to settle rejected action for task ${this.taskId}`, + { cause: error }, + ) + } if (this.pendingAction?.actionId === action.actionId) { this.pendingAction = authoritative.pendingAction } if (this.pendingAction?.kind === "create_subtask") { - throw new Error( + throw new PendingActionSettlementError( `[Task#settleInterruptedCreateSubtaskBeforeReplay] Task ${this.taskId} still has a rejected create-subtask action`, ) } diff --git a/src/core/task/__tests__/Task.persistence.spec.ts b/src/core/task/__tests__/Task.persistence.spec.ts index 3dba6a42c9..1b25e814fa 100644 --- a/src/core/task/__tests__/Task.persistence.spec.ts +++ b/src/core/task/__tests__/Task.persistence.spec.ts @@ -14,7 +14,7 @@ import { import { TelemetryService } from "@roo-code/telemetry" import type { Anthropic } from "@anthropic-ai/sdk" -import { Task } from "../Task" +import { PendingActionSettlementError, Task } from "../Task" import { ClineProvider } from "../../webview/ClineProvider" import { ContextProxy } from "../../config/ContextProxy" import { providerIdentifiers } from "@roo-code/types/provider-identifiers" @@ -1486,7 +1486,15 @@ describe("Task persistence", () => { const ask = vi.spyOn(task, "ask") const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") - await expect(getTaskPersistenceAccess(task).resumeTaskFromHistory()).rejects.toThrow(settlementError) + const resumeError = await getTaskPersistenceAccess(task) + .resumeTaskFromHistory() + .catch((error: unknown) => error) + + expect(resumeError).toMatchObject({ + name: "PendingActionSettlementError", + cause: settlementError, + }) + expect(resumeError).toBeInstanceOf(PendingActionSettlementError) expect(replay).not.toHaveBeenCalled() expect(ask).not.toHaveBeenCalled() diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 90dd957854..6325435d54 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -112,7 +112,7 @@ import { forceFullModelDetailsLoad, hasLoadedFullDetails } from "../../api/provi import { ContextProxy } from "../config/ContextProxy" import { ProviderSettingsManager } from "../config/ProviderSettingsManager" import { CustomModesManager } from "../config/CustomModesManager" -import { Task } from "../task/Task" +import { PendingActionSettlementError, Task } from "../task/Task" import { webviewMessageHandler } from "./webviewMessageHandler" import type { ClineMessage, TodoItem } from "@roo-code/types" @@ -656,7 +656,11 @@ export class ClineProvider } } - private async cleanupFailedHistoryTask(task: Task): Promise { + private async cleanupFailedHistoryTask(task: Task, error: unknown): Promise { + if (!(error instanceof PendingActionSettlementError)) { + return + } + if (this.taskRegistry.getById(task.taskId) !== task) { return } @@ -1444,8 +1448,8 @@ export class ClineProvider ) if (options?.startTask !== false) { - scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, () => - this.cleanupFailedHistoryTask(task), + scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, (error) => + this.cleanupFailedHistoryTask(task, error), ) } } else { @@ -1456,8 +1460,8 @@ export class ClineProvider ) if (options?.startTask !== false) { - scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, () => - this.cleanupFailedHistoryTask(task), + scheduleTask(this.taskScheduler, task, "createTaskWithHistoryItem", undefined, (error) => + this.cleanupFailedHistoryTask(task, error), ) } } diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index bf244f4eb3..b91e52a4d1 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -10,6 +10,7 @@ import { safeWriteJson } from "../safeWriteJson" // test mockImplementation callbacks delegate to the real implementation. const fsPromisesActuals = vi.hoisted(() => ({ rename: undefined as (typeof import("fs/promises"))["rename"] | undefined, + copyFile: undefined as (typeof import("fs/promises"))["copyFile"] | undefined, unlink: undefined as (typeof import("fs/promises"))["unlink"] | undefined, writeFile: undefined as (typeof import("fs/promises"))["writeFile"] | undefined, })) @@ -17,6 +18,7 @@ const fsPromisesActuals = vi.hoisted(() => ({ vi.mock("fs/promises", async () => { const actual = await vi.importActual("fs/promises") fsPromisesActuals.rename = actual.rename + fsPromisesActuals.copyFile = actual.copyFile fsPromisesActuals.unlink = actual.unlink fsPromisesActuals.writeFile = actual.writeFile // Start with all actual implementations. @@ -28,6 +30,7 @@ vi.mock("fs/promises", async () => { mockedFs.writeFile = vi.fn(actual.writeFile) as any mockedFs.readFile = vi.fn(actual.readFile) as any mockedFs.rename = vi.fn(actual.rename) as any + mockedFs.copyFile = vi.fn(actual.copyFile) as typeof actual.copyFile mockedFs.unlink = vi.fn(actual.unlink) as any mockedFs.access = vi.fn(actual.access) as any mockedFs.mkdtemp = vi.fn(actual.mkdtemp) as any @@ -220,9 +223,27 @@ describe("safeWriteJson", () => { await safeWriteJson(currentTestFilePath, newData, { atomicReplace: true }) expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(1) + expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(1) expect(await readFileContent(currentTestFilePath)).toEqual(newData) }) + test("should restore the copied backup when atomic replacement fails", async () => { + const initialData = { message: "Initial content, should be restored" } + const newData = { message: "New content" } + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + vi.mocked(fs.copyFile).mockClear() + vi.mocked(fs.rename).mockClear() + vi.mocked(fs.rename).mockRejectedValueOnce(new Error("Atomic replacement failed")) + + await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( + "Atomic replacement failed", + ) + + expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(2) + expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(1) + expect(await readFileContent(currentTestFilePath)).toEqual(initialData) + }) + // Tests for directory creation functionality test("should create parent directory if it doesn't exist", async () => { // Create a path in a non-existent subdirectory of the temp dir diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index b5ea09ea78..d8b0b1963d 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -34,8 +34,8 @@ export interface SafeWriteJsonOptions { /** * Replace an existing target with one rename from the completed temporary file. - * Use this when an absent target during a process stop is less safe than losing - * the backup-based rollback that the default write path provides. + * A copied backup retains rollback support without removing the target before + * the replacement rename. * @default false */ atomicReplace?: boolean @@ -134,25 +134,23 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso await _streamDataToFile(actualTempNewFilePath, data, options?.prettyPrint) - if (!options?.atomicReplace) { - // Step 2: Check if the target file exists. If so, rename it to a backup path. - try { - // Check for target file existence - await fs.access(absoluteFilePath) - // Target exists, create a backup path and rename. - actualTempBackupFilePath = path.join( - path.dirname(absoluteFilePath), - `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, - ) + // Step 2: Check if the target file exists. If so, retain a rollback backup. + try { + await fs.access(absoluteFilePath) + actualTempBackupFilePath = path.join( + path.dirname(absoluteFilePath), + `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, + ) + if (options?.atomicReplace) { + await fs.copyFile(absoluteFilePath, actualTempBackupFilePath) + } else { await fs.rename(absoluteFilePath, actualTempBackupFilePath) - } catch (accessError: any) { - // Explicitly type accessError - if (accessError.code !== "ENOENT") { - // An error other than "file not found" occurred during access check. - throw accessError - } - // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. } + } catch (accessError: any) { + if (accessError.code !== "ENOENT") { + throw accessError + } + actualTempBackupFilePath = null } // Step 3: Rename the new temporary file to the target file path. @@ -188,7 +186,12 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Attempt rollback if a backup was made if (backupFileToRollbackOrCleanupWithinCatch) { try { - await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) + if (options?.atomicReplace) { + await fs.copyFile(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) + await fs.unlink(backupFileToRollbackOrCleanupWithinCatch) + } else { + await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) + } // Mark as handled, prevent later unlink of this path actualTempBackupFilePath = null } catch (rollbackError) { From ac0e9751caf81b0b13b6bfb1484526cb9ca3bf49 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 14:02:41 +0000 Subject: [PATCH 13/23] fix: preserve atomic rollback backup --- src/utils/__tests__/safeWriteJson.test.ts | 48 ++++++++++++++++++++++- src/utils/safeWriteJson.ts | 33 +++++++++++++--- 2 files changed, 74 insertions(+), 7 deletions(-) diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index b91e52a4d1..e5dcb49917 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -240,10 +240,56 @@ describe("safeWriteJson", () => { ) expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(2) - expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(1) + expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(2) expect(await readFileContent(currentTestFilePath)).toEqual(initialData) }) + test("should not roll back from an incomplete backup copy", async () => { + const initialData = { message: "Initial content" } + const newData = { message: "New content" } + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + vi.mocked(fs.copyFile).mockClear() + vi.mocked(fs.rename).mockClear() + vi.mocked(fs.copyFile).mockRejectedValueOnce(new Error("Backup copy failed")) + + await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( + "Backup copy failed", + ) + + expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(1) + expect(vi.mocked(fs.rename)).not.toHaveBeenCalled() + expect(await readFileContent(currentTestFilePath)).toEqual(initialData) + }) + + test("should preserve the completed backup and remove an incomplete rollback copy when atomic rollback fails", async () => { + const initialData = { message: "Initial content" } + const newData = { message: "New content" } + await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) + const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) + vi.mocked(fs.rename).mockRejectedValueOnce(new Error("Atomic replacement failed")) + vi.mocked(fs.copyFile) + .mockImplementationOnce(fsPromisesActuals.copyFile!) + .mockImplementationOnce(async (_source, target) => { + await fsPromisesActuals.writeFile!(target, "incomplete rollback") + throw new Error("Rollback copy failed") + }) + + await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( + "Atomic replacement failed", + ) + + const remainingFiles = await fs.readdir(tempDir) + const backupFiles = remainingFiles.filter((file) => file.includes(".bak_")) + expect(backupFiles).toHaveLength(1) + expect(await readFileContent(path.join(tempDir, backupFiles[0]))).toEqual(initialData) + expect(remainingFiles.some((file) => file.includes(".rollback_"))).toBe(false) + expect(await readFileContent(currentTestFilePath)).toEqual(initialData) + expect(consoleErrorSpy).toHaveBeenCalledWith( + expect.stringContaining("Failed to restore backup"), + expect.objectContaining({ message: "Rollback copy failed" }), + ) + }) + // Tests for directory creation functionality test("should create parent directory if it doesn't exist", async () => { // Create a path in a non-existent subdirectory of the temp dir diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index d8b0b1963d..daaca981c4 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -107,6 +107,7 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Variables to hold the actual paths of temp files if they are created. let actualTempNewFilePath: string | null = null let actualTempBackupFilePath: string | null = null + let actualTempRollbackFilePath: string | null = null try { // If a merge callback was provided, read the current file under the lock @@ -137,15 +138,16 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Step 2: Check if the target file exists. If so, retain a rollback backup. try { await fs.access(absoluteFilePath) - actualTempBackupFilePath = path.join( + const candidateBackupFilePath = path.join( path.dirname(absoluteFilePath), `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, ) if (options?.atomicReplace) { - await fs.copyFile(absoluteFilePath, actualTempBackupFilePath) + await fs.copyFile(absoluteFilePath, candidateBackupFilePath) } else { - await fs.rename(absoluteFilePath, actualTempBackupFilePath) + await fs.rename(absoluteFilePath, candidateBackupFilePath) } + actualTempBackupFilePath = candidateBackupFilePath } catch (accessError: any) { if (accessError.code !== "ENOENT") { throw accessError @@ -187,7 +189,13 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso if (backupFileToRollbackOrCleanupWithinCatch) { try { if (options?.atomicReplace) { - await fs.copyFile(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) + actualTempRollbackFilePath = path.join( + path.dirname(absoluteFilePath), + `.${path.basename(absoluteFilePath)}.rollback_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, + ) + await fs.copyFile(backupFileToRollbackOrCleanupWithinCatch, actualTempRollbackFilePath) + await fs.rename(actualTempRollbackFilePath, absoluteFilePath) + actualTempRollbackFilePath = null await fs.unlink(backupFileToRollbackOrCleanupWithinCatch) } else { await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) @@ -203,6 +211,19 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } + // A failed rollback can leave an incomplete rollback copy. The completed backup remains available for recovery. + if (actualTempRollbackFilePath) { + try { + await fs.unlink(actualTempRollbackFilePath) + actualTempRollbackFilePath = null + } catch (cleanupError) { + console.error( + `[Catch] Failed to clean up temporary rollback file ${actualTempRollbackFilePath}:`, + cleanupError, + ) + } + } + // Cleanup the .new file if it exists if (newFileToCleanupWithinCatch) { try { @@ -215,8 +236,8 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } - // Cleanup the .bak file if it still needs to be (i.e., wasn't successfully restored) - if (actualTempBackupFilePath) { + // A copied backup remains available for recovery when atomic rollback fails. + if (actualTempBackupFilePath && !options?.atomicReplace) { try { await fs.unlink(actualTempBackupFilePath) } catch (cleanupError) { From c842cd263f6a4f172f72734a72717eac661bc680 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 15:12:00 +0000 Subject: [PATCH 14/23] fix: hold advisory lock during task history deletion TaskHistoryStore.delete and deleteMany unlinked history_item.json outside the per-file advisory lock. A delete that landed between the settlement read and commit let clearPendingActionIfMatching recreate the deleted record. Extract acquireFileLock and withFileLock into src/utils/fileLock.ts and route safeWriteJson through the same helper, so writes and deletes serialize on one lock implementation with no nested acquisition. Unlink now runs under withFileLock. Cache and mtime entries change only after the file is absent or deletion succeeds, so a failed delete keeps the cached record and skips the write-through callback. Add a real-filesystem test that pauses settlement after its disk read, starts a deletion from another store, and verifies serialization, no recreation of history_item.json, and fail-closed settlement. Add unit tests for unlink and lock failure cache semantics. --- src/core/task-persistence/TaskHistoryStore.ts | 52 +++++++++------ .../TaskHistoryStore.realConcurrency.spec.ts | 62 ++++++++++++++++++ .../__tests__/TaskHistoryStore.spec.ts | 64 +++++++++++++++++++ src/utils/fileLock.ts | 50 +++++++++++++++ src/utils/safeWriteJson.ts | 24 ++----- 5 files changed, 212 insertions(+), 40 deletions(-) create mode 100644 src/utils/fileLock.ts diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 3cfd0350a1..2fbade81bb 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -7,7 +7,8 @@ import deepEqual from "fast-deep-equal" import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" +import { LOCK_STALE_MS, withFileLock } from "../../utils/fileLock" +import { safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" @@ -266,16 +267,7 @@ export class TaskHistoryStore { */ async delete(taskId: string): Promise { return this.withLock(async () => { - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - - // Remove per-task file (best-effort) - try { - const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) - } catch { - // File may already be deleted - } + await this.deleteTaskFile(taskId) // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { @@ -290,15 +282,7 @@ export class TaskHistoryStore { async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { for (const taskId of taskIds) { - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - - try { - const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) - } catch { - // File may already be deleted - } + await this.deleteTaskFile(taskId) } // Call onWrite callback inside the lock for serialized write-through @@ -883,6 +867,34 @@ export class TaskHistoryStore { } } + /** + * Delete one task file under the same advisory lock used by `safeWriteJson`. + * Cache state changes only after the file is absent or unlink succeeds. + */ + private async deleteTaskFile(taskId: string): Promise { + const filePath = await this.getTaskFilePath(taskId) + try { + await withFileLock(filePath, async (absoluteFilePath) => { + try { + await fs.unlink(absoluteFilePath) + } catch (error) { + if (!this.isFileNotFoundError(error)) { + throw error + } + } + }) + } catch (error) { + // A missing parent directory prevents lock creation and also proves + // that the history file is absent. + if (!this.isFileNotFoundError(error)) { + throw error + } + } + + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + } + // ────────────────────────────── Private: fs.watch ────────────────────────────── /** diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index fea7a1d805..39bfde667c 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -6,6 +6,11 @@ import type { HistoryItem } from "@roo-code/types" import { TaskHistoryStore } from "../TaskHistoryStore" +vi.mock("fs/promises", async () => { + const actual = await vi.importActual("fs/promises") + return { ...actual, readFile: vi.fn(actual.readFile) } +}) + function item(id: string): HistoryItem { return { id, @@ -215,6 +220,63 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("serializes deletion after settlement reads disk without recreating the record", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-during-settlement-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const action = createAction("action-a", "action A") + const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") + let signalReadComplete!: () => void + const readComplete = new Promise((resolve) => { + signalReadComplete = resolve + }) + let releaseSettlement!: () => void + const settlementCanContinue = new Promise((resolve) => { + releaseSettlement = resolve + }) + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: action }) + await storeB.initialize() + + const actualFs = await vi.importActual("fs/promises") + vi.mocked(fs.readFile).mockImplementation(async (...args: Parameters) => { + const result = await actualFs.readFile(...args) + if (args[0] === filePath) { + signalReadComplete() + await settlementCanContinue + } + return result + }) + + const settlement = storeA.clearPendingActionIfMatching("shared-task", action.actionId) + await readComplete + const deletion = storeB.delete("shared-task") + let deletionSettled = false + void deletion.finally(() => { + deletionSettled = true + }) + await new Promise((resolve) => setTimeout(resolve, 25)) + expect(deletionSettled).toBe(false) + + releaseSettlement() + await expect(settlement).resolves.toMatchObject({ id: "shared-task", pendingAction: undefined }) + await deletion + + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + await expect(storeA.clearPendingActionIfMatching("shared-task", action.actionId)).rejects.toThrow( + "task shared-task not found", + ) + expect(storeA.get("shared-task")).toBeUndefined() + expect(storeB.get("shared-task")).toBeUndefined() + } finally { + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("rejects and evicts stale cache without recreating artifacts after full directory deletion", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-deleted-directory-settlement-")) const store = new TaskHistoryStore(storagePath) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts index 3e277ac867..d6f8c58b7d 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts @@ -10,6 +10,11 @@ import { TaskHistoryStore, assertValidTransition } from "../TaskHistoryStore" import { GlobalFileNames } from "../../../shared/globalFileNames" import { ClineProvider } from "../../webview/ClineProvider" +vi.mock("fs/promises", async () => { + const actual = await vi.importActual("fs/promises") + return { ...actual, unlink: vi.fn(actual.unlink) } +}) + vi.mock("../../../utils/storage", () => ({ getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => { return defaultPath @@ -24,6 +29,15 @@ vi.mock("../../../utils/safeWriteJson", () => ({ }), })) +vi.mock("../../../utils/fileLock", () => ({ + LOCK_STALE_MS: 31_000, + withFileLock: vi + .fn() + .mockImplementation(async (filePath: string, operation: (filePath: string) => Promise) => + operation(filePath), + ), +})) + function makeHistoryItem(overrides: Partial = {}): HistoryItem { return { id: `task-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, @@ -244,6 +258,33 @@ describe("TaskHistoryStore", () => { await store.initialize() await expect(store.delete("non-existent")).resolves.not.toThrow() }) + + it("retains cache state when unlink fails", async () => { + await store.initialize() + const item = makeHistoryItem({ id: "unlink-failure" }) + await store.upsert(item) + const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) + vi.mocked(fs.unlink).mockRejectedValueOnce(unlinkError) + + await expect(store.delete(item.id)).rejects.toBe(unlinkError) + expect(store.get(item.id)).toEqual(item) + }) + + it("retains cache state when lock acquisition fails", async () => { + const onWrite = vi.fn().mockResolvedValue(undefined) + store = new TaskHistoryStore(tmpDir, { onWrite }) + await store.initialize() + const item = makeHistoryItem({ id: "lock-failure" }) + await store.upsert(item) + onWrite.mockClear() + const { withFileLock } = await import("../../../utils/fileLock") + const lockError = new Error("lock failed") + vi.mocked(withFileLock).mockRejectedValueOnce(lockError) + + await expect(store.delete(item.id)).rejects.toBe(lockError) + expect(store.get(item.id)).toEqual(item) + expect(onWrite).not.toHaveBeenCalled() + }) }) describe("deleteMany()", () => { @@ -259,6 +300,29 @@ describe("TaskHistoryStore", () => { expect(store.getAll()).toHaveLength(1) expect(store.get("batch-2")).toBeDefined() }) + + it("retains the failed item and later items when unlink fails", async () => { + await store.initialize() + const first = makeHistoryItem({ id: "batch-success" }) + const failed = makeHistoryItem({ id: "batch-failure" }) + const later = makeHistoryItem({ id: "batch-later" }) + await store.upsert(first) + await store.upsert(failed) + await store.upsert(later) + const actualFs = await vi.importActual("fs/promises") + const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) + vi.mocked(fs.unlink).mockImplementation(async (filePath) => { + if (filePath.toString().includes(failed.id)) { + throw unlinkError + } + return actualFs.unlink(filePath) + }) + + await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) + expect(store.get(first.id)).toBeUndefined() + expect(store.get(failed.id)).toEqual(failed) + expect(store.get(later.id)).toEqual(later) + }) }) describe("reconcile()", () => { diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts new file mode 100644 index 0000000000..2358845a94 --- /dev/null +++ b/src/utils/fileLock.ts @@ -0,0 +1,50 @@ +import * as path from "path" +import * as lockfile from "proper-lockfile" + +export const LOCK_STALE_MS = 31_000 + +export async function acquireFileLock(filePath: string): Promise<() => Promise> { + const absoluteFilePath = path.resolve(filePath) + try { + return await lockfile.lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10_000, + realpath: false, + retries: { + retries: 5, + factor: 2, + minTimeout: 100, + maxTimeout: 1_000, + }, + onCompromised: (error) => { + console.error(`Lock at ${absoluteFilePath} was compromised:`, error) + throw error + }, + }) + } catch (error) { + console.error(`Failed to acquire lock for ${absoluteFilePath}:`, error) + throw error + } +} + +/** + * Run one operation while holding the advisory lock for a file path. + * Callers must not acquire this lock again from inside `operation`. + */ +export async function withFileLock( + filePath: string, + operation: (absoluteFilePath: string) => Promise, +): Promise { + const absoluteFilePath = path.resolve(filePath) + const releaseLock = await acquireFileLock(absoluteFilePath) + + try { + return await operation(absoluteFilePath) + } finally { + try { + await releaseLock() + } catch (error) { + console.error(`Failed to release lock for ${absoluteFilePath}:`, error) + } + } +} diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index daaca981c4..56c05f84d5 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -1,9 +1,10 @@ import * as fs from "fs/promises" import * as fsSync from "fs" import * as path from "path" -import * as lockfile from "proper-lockfile" import { JsonStreamStringify } from "json-stream-stringify" +import { acquireFileLock, LOCK_STALE_MS } from "./fileLock" + /** * Options for safeWriteJson function */ @@ -79,22 +80,7 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Acquire the lock before any file operations try { - releaseLock = await lockfile.lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long - realpath: false, // the file may not exist yet, which is acceptable - retries: { - // Configuration for retrying lock acquisition - retries: 5, // Number of retries after the initial attempt - factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) - minTimeout: 100, // Minimum time to wait before the first retry (in ms) - maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) - }, - onCompromised: (err) => { - console.error(`Lock at ${absoluteFilePath} was compromised:`, err) - throw err - }, - }) + releaseLock = await acquireFileLock(absoluteFilePath) } catch (lockError) { // If lock acquisition fails, we throw immediately. // The releaseLock remains a no-op, so the finally block in the main file operations @@ -289,6 +275,4 @@ async function _streamDataToFile(targetPath: string, data: any, prettyPrint = fa }) } -export const LOCK_STALE_MS = 31_000 - -export { safeWriteJson } +export { LOCK_STALE_MS, safeWriteJson } From 8797436cab3538df7e650d25bf02822bda4e5d19 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 15:21:09 +0000 Subject: [PATCH 15/23] test: provide getStorageBasePath in provider storage mock TaskHistoryStore.deleteTaskFile resolves the task file path before the best-effort unlink block, so a storage mock without getStorageBasePath now fails the delete call instead of being swallowed. Add the export to the ClineProvider storage mock so the main-branch regression test runs the real delete semantics. --- src/core/webview/__tests__/ClineProvider.spec.ts | 1 + 1 file changed, 1 insertion(+) diff --git a/src/core/webview/__tests__/ClineProvider.spec.ts b/src/core/webview/__tests__/ClineProvider.spec.ts index 97c4dd877e..75885185bd 100644 --- a/src/core/webview/__tests__/ClineProvider.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.spec.ts @@ -87,6 +87,7 @@ vi.mock("../../../utils/storage", () => ({ getSettingsDirectoryPath: vi.fn().mockResolvedValue("/test/settings/path"), getTaskDirectoryPath: vi.fn().mockResolvedValue("/test/task/path"), getGlobalStoragePath: vi.fn().mockResolvedValue("/test/storage/path"), + getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => defaultPath), })) vi.mock("@modelcontextprotocol/sdk/types.js", () => ({ From 981b6688f4b7831d6327fd16129ae707e4bddae6 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 20:24:37 +0000 Subject: [PATCH 16/23] fix: flush write-through on partial deleteMany and reject on lock compromise - deleteMany now awaits the write-through callback with the current cache when a later deletion fails after an earlier success, so persisted globalState.taskHistory does not keep already-deleted tasks. The original deletion error stays the rejection, and a failed flush is logged without masking it. - acquireFileLock records a lock compromise instead of throwing from the proper-lockfile renewal timer. The release rejects with the recorded compromise error, and withFileLock and safeWriteJson reject a successful operation whose lock was compromised. Operation errors keep precedence. - Add regression tests for both behaviors. --- src/core/task-persistence/TaskHistoryStore.ts | 26 ++++- .../__tests__/TaskHistoryStore.spec.ts | 74 ++++++++++++ src/utils/__tests__/fileLock.spec.ts | 108 ++++++++++++++++++ src/utils/__tests__/safeWriteJson.test.ts | 61 ++++++++++ src/utils/fileLock.ts | 47 +++++++- src/utils/safeWriteJson.ts | 25 +++- 6 files changed, 327 insertions(+), 14 deletions(-) create mode 100644 src/utils/__tests__/fileLock.spec.ts diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 2fbade81bb..419ae44bcb 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -281,8 +281,30 @@ export class TaskHistoryStore { */ async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { - for (const taskId of taskIds) { - await this.deleteTaskFile(taskId) + let deletedCount = 0 + try { + for (const taskId of taskIds) { + await this.deleteTaskFile(taskId) + deletedCount++ + } + } catch (error) { + // Earlier deletions already removed their files and cache entries. + // Await the write-through with the current cache before the + // original deletion error rejects the call, so persisted + // globalState does not keep already-deleted tasks. + if (deletedCount > 0 && this.onWrite) { + try { + await this.onWrite(this.getAll()) + } catch (writeError) { + // The deletion failure is the primary error. Report the + // write-through failure without masking it. + console.error( + "[TaskHistoryStore] deleteMany write-through after partial deletion failed:", + writeError, + ) + } + } + throw error } // Call onWrite callback inside the lock for serialized write-through diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts index d6f8c58b7d..4ac087c56f 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts @@ -323,6 +323,80 @@ describe("TaskHistoryStore", () => { expect(store.get(failed.id)).toEqual(failed) expect(store.get(later.id)).toEqual(later) }) + + it("flushes the write-through with the current cache when a later deletion fails", async () => { + store.dispose() + const globalState: HistoryItem[] = [] + const onWrite = vi.fn(async (items: HistoryItem[]) => { + globalState.splice(0, globalState.length, ...items) + }) + store = new TaskHistoryStore(tmpDir, { onWrite }) + await store.initialize() + + const first = makeHistoryItem({ id: "batch-wt-success" }) + const failed = makeHistoryItem({ id: "batch-wt-failure" }) + const later = makeHistoryItem({ id: "batch-wt-later" }) + await store.upsert(first) + await store.upsert(failed) + await store.upsert(later) + + const actualFs = await vi.importActual("fs/promises") + const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) + vi.mocked(fs.unlink).mockImplementation(async (filePath) => { + if (filePath.toString().includes(failed.id)) { + throw unlinkError + } + return actualFs.unlink(filePath) + }) + + await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) + + // The earlier deleted task is evicted from the store cache. + expect(store.get(first.id)).toBeUndefined() + // The awaited write-through persisted that state before the original + // deletion error rejected the call, so stale globalState cannot + // return the deleted task. + expect(globalState.find((item) => item.id === first.id)).toBeUndefined() + expect(globalState.find((item) => item.id === failed.id)).toEqual(failed) + expect(globalState.find((item) => item.id === later.id)).toEqual(later) + + vi.mocked(fs.unlink).mockImplementation(async (filePath) => actualFs.unlink(filePath)) + }) + + it("keeps the original deletion error when the partial write-through also fails", async () => { + store.dispose() + const onWrite = vi.fn() + // Three upserts succeed, then the partial-failure flush rejects. + onWrite + .mockResolvedValueOnce(undefined) + .mockResolvedValueOnce(undefined) + .mockResolvedValueOnce(undefined) + .mockRejectedValueOnce(new Error("write-through failed")) + store = new TaskHistoryStore(tmpDir, { onWrite }) + await store.initialize() + + const first = makeHistoryItem({ id: "batch-wt2-success" }) + const failed = makeHistoryItem({ id: "batch-wt2-failure" }) + const later = makeHistoryItem({ id: "batch-wt2-later" }) + await store.upsert(first) + await store.upsert(failed) + await store.upsert(later) + + const actualFs = await vi.importActual("fs/promises") + const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) + vi.mocked(fs.unlink).mockImplementation(async (filePath) => { + if (filePath.toString().includes(failed.id)) { + throw unlinkError + } + return actualFs.unlink(filePath) + }) + + // The deletion failure is the primary error and must not be masked + // by the failed write-through. + await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) + + vi.mocked(fs.unlink).mockImplementation(async (filePath) => actualFs.unlink(filePath)) + }) }) describe("reconcile()", () => { diff --git a/src/utils/__tests__/fileLock.spec.ts b/src/utils/__tests__/fileLock.spec.ts new file mode 100644 index 0000000000..93d12bb2d4 --- /dev/null +++ b/src/utils/__tests__/fileLock.spec.ts @@ -0,0 +1,108 @@ +// pnpm --filter roo-cline test utils/__tests__/fileLock.spec.ts + +import * as path from "path" + +import { acquireFileLock, withFileLock } from "../fileLock" + +const { lockMock } = vi.hoisted(() => ({ lockMock: vi.fn() })) + +vi.mock("proper-lockfile", () => ({ lock: lockMock, default: { lock: lockMock } })) + +interface CapturedLockOptions { + onCompromised: (error: Error) => void +} + +function requireCapturedLockOptions(): CapturedLockOptions { + if (!capturedOptions) { + throw new Error("proper-lockfile options were not captured") + } + return capturedOptions +} + +let capturedOptions: CapturedLockOptions | undefined + +describe("fileLock", () => { + const absoluteFilePath = path.resolve("/virtual/dir/target.json") + let underlyingRelease: ReturnType + let consoleError: ReturnType + + beforeEach(() => { + consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) + underlyingRelease = vi.fn().mockResolvedValue(undefined) + capturedOptions = undefined + lockMock.mockImplementation(async (_filePath: string, options: CapturedLockOptions) => { + capturedOptions = options + return underlyingRelease + }) + }) + + afterEach(() => { + consoleError.mockRestore() + vi.clearAllMocks() + }) + + it("records the compromise without throwing from the renewal callback", () => { + const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) + + return acquireFileLock(absoluteFilePath).then((release) => { + expect(() => requireCapturedLockOptions().onCompromised(compromiseError)).not.toThrow() + return release().catch(() => {}) + }) + }) + + it("rejects the owning operation with the recorded compromise error", async () => { + const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) + const operation = vi.fn(async () => { + // Simulate the proper-lockfile renewal timer reporting a + // compromise while the operation still holds the lock. + requireCapturedLockOptions().onCompromised(compromiseError) + return "done" + }) + + await expect(withFileLock(absoluteFilePath, operation)).rejects.toBe(compromiseError) + expect(operation).toHaveBeenCalledTimes(1) + // proper-lockfile already marked the lock released, so the wrapper + // must not call the underlying release after a compromise. + expect(underlyingRelease).not.toHaveBeenCalled() + }) + + it("keeps the operation error when the operation fails after a compromise", async () => { + const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) + const operationError = new Error("operation failed") + const operation = vi.fn(async () => { + requireCapturedLockOptions().onCompromised(compromiseError) + throw operationError + }) + + await expect(withFileLock(absoluteFilePath, operation)).rejects.toBe(operationError) + }) + + it("rejects a direct release with the recorded compromise error", async () => { + const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) + const release = await acquireFileLock(absoluteFilePath) + + requireCapturedLockOptions().onCompromised(compromiseError) + + await expect(release()).rejects.toBe(compromiseError) + expect(underlyingRelease).not.toHaveBeenCalled() + }) + + it("releases normally without a compromise", async () => { + const operation = vi.fn(async () => "done") + + await expect(withFileLock(absoluteFilePath, operation)).resolves.toBe("done") + expect(underlyingRelease).toHaveBeenCalledTimes(1) + }) + + it("keeps reporting success when an unrelated release error occurs", async () => { + underlyingRelease.mockRejectedValue(new Error("unlock failed")) + + await expect( + withFileLock( + absoluteFilePath, + vi.fn(async () => "done"), + ), + ).resolves.toBe("done") + expect(consoleError).toHaveBeenCalled() + }) +}) diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index e5dcb49917..c58174a06c 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -477,6 +477,67 @@ describe("safeWriteJson", () => { await fs.unlink(lockTestFilePath).catch(() => {}) // Ignore errors if file doesn't exist vi.unmock("proper-lockfile") // Ensure the mock is removed after this test }) + + test("rejects with a lock compromise error after a successful write", async () => { + vi.resetModules() + + const data = { message: "compromise after success" } + const compromiseTestFilePath = path.join(tempDir, "compromise-test-file.json") + await fs.writeFile(compromiseTestFilePath, JSON.stringify({ initial: "content" })) + + const compromiseError = Object.assign(new Error("lock was compromised"), { code: "ECOMPROMISED" }) + vi.doMock("proper-lockfile", () => ({ + ...vi.importActual("proper-lockfile"), + lock: vi.fn().mockResolvedValue(vi.fn().mockRejectedValue(compromiseError)), + })) + + const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") + + // The write itself succeeds, but the release reports a compromised + // lock, so the operation must reject instead of reporting success. + await expect(mockedSafeWriteJson(compromiseTestFilePath, data)).rejects.toBe(compromiseError) + + await fs.unlink(compromiseTestFilePath).catch(() => {}) + vi.doUnmock("proper-lockfile") + }) + + test("keeps the original write error when the write fails and the lock was compromised", async () => { + vi.resetModules() + + const data = { message: "compromise after failure" } + const compromiseTestFilePath = path.join(tempDir, "compromise-failure-test-file.json") + await fs.writeFile(compromiseTestFilePath, JSON.stringify({ initial: "content" })) + + const compromiseError = Object.assign(new Error("lock was compromised"), { code: "ECOMPROMISED" }) + vi.doMock("proper-lockfile", () => ({ + ...vi.importActual("proper-lockfile"), + lock: vi.fn().mockResolvedValue(vi.fn().mockRejectedValue(compromiseError)), + })) + + const createWriteStreamSpy = vi.spyOn(fsSyncActual, "createWriteStream") + createWriteStreamSpy.mockImplementationOnce(() => { + // A plain Writable provides the pipe and error surface that + // `_streamDataToFile` uses, but not the full WriteStream + // interface, so this cast is deliberate. + const errorStream = new Writable({ + write: (_chunk, _encoding, callback) => { + callback(new Error("Stream write error")) + }, + }) as unknown as fsSyncActual.WriteStream + errorStream.close = vi.fn() + return errorStream + }) + + const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") + + // The write failure is the primary error and must not be masked by + // the compromise reported at release time. + await expect(mockedSafeWriteJson(compromiseTestFilePath, data)).rejects.toThrow("Stream write error") + + createWriteStreamSpy.mockRestore() + await fs.unlink(compromiseTestFilePath).catch(() => {}) + vi.doUnmock("proper-lockfile") + }) test("should release lock even if an error occurs mid-operation", async () => { const data = { message: "test lock release on error" } diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts index 2358845a94..54f9be732c 100644 --- a/src/utils/fileLock.ts +++ b/src/utils/fileLock.ts @@ -5,8 +5,13 @@ export const LOCK_STALE_MS = 31_000 export async function acquireFileLock(filePath: string): Promise<() => Promise> { const absoluteFilePath = path.resolve(filePath) + // proper-lockfile calls `onCompromised` from its renewal timer after the + // lock promise already resolved. A throw here becomes an uncaught + // exception in the host instead of a rejection, so record the error and + // fail the owning operation through its release call instead. + let compromisedError: Error | null = null try { - return await lockfile.lock(absoluteFilePath, { + const release = await lockfile.lock(absoluteFilePath, { stale: LOCK_STALE_MS, update: 10_000, realpath: false, @@ -18,9 +23,18 @@ export async function acquireFileLock(filePath: string): Promise<() => Promise { console.error(`Lock at ${absoluteFilePath} was compromised:`, error) - throw error + compromisedError = error }, }) + return async () => { + if (compromisedError) { + // proper-lockfile marks the lock released when it reports a + // compromise and its own release resolves silently, so reject + // with the recorded compromise error here. + throw compromisedError + } + await release() + } } catch (error) { console.error(`Failed to acquire lock for ${absoluteFilePath}:`, error) throw error @@ -38,13 +52,34 @@ export async function withFileLock( const absoluteFilePath = path.resolve(filePath) const releaseLock = await acquireFileLock(absoluteFilePath) + let result: T try { - return await operation(absoluteFilePath) - } finally { + result = await operation(absoluteFilePath) + } catch (operationError) { + // The operation error is the primary failure. Release without + // reporting a secondary release error over it. try { await releaseLock() - } catch (error) { - console.error(`Failed to release lock for ${absoluteFilePath}:`, error) + } catch (releaseError) { + console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) + } + throw operationError + } + + try { + await releaseLock() + } catch (releaseError) { + const code = + releaseError && typeof releaseError === "object" && "code" in releaseError + ? (releaseError as { code: unknown }).code + : undefined + if (code === "ECOMPROMISED") { + // The lock was compromised while this operation held it, so the + // operation ran without mutual exclusion. Reject the operation + // that owns the lock instead of reporting success. + throw releaseError } + console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) } + return result } diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 56c05f84d5..49ae87a06f 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -233,17 +233,30 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso ) } } - throw originalError // This MUST be the error that rejects the promise. - } finally { - // Release the lock in the main finally block. + // Release the lock before rejecting. The original write failure is the + // rejection and a release failure cannot mask it. try { - // releaseLock will be the actual unlock function if lock was acquired, - // or the initial no-op if acquisition failed. await releaseLock() } catch (unlockError) { - // Do not re-throw here, as the originalError from the try/catch (if any) is more important. console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) } + throw originalError // This MUST be the error that rejects the promise. + } + + // Release the lock on the success path. A compromised lock means this + // write ran without mutual exclusion, so reject this operation instead + // of reporting success. + try { + await releaseLock() + } catch (unlockError) { + const code = + unlockError && typeof unlockError === "object" && "code" in unlockError + ? (unlockError as { code: unknown }).code + : undefined + if (code === "ECOMPROMISED") { + throw unlockError + } + console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) } } From 748f8e346034e00a037349425f14c0eb2988362c Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 20:54:39 +0000 Subject: [PATCH 17/23] test: assert partial-failure write-through call and release error details - The partial write-through failure test now asserts the flush call, its item payload, and the logged write-through error next to the original unlink rejection. - The unrelated release error test asserts the exact log arguments and error object. --- .../task-persistence/__tests__/TaskHistoryStore.spec.ts | 8 ++++++++ src/utils/__tests__/fileLock.spec.ts | 5 +++-- 2 files changed, 11 insertions(+), 2 deletions(-) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts index 4ac087c56f..671b8168aa 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts @@ -393,7 +393,15 @@ describe("TaskHistoryStore", () => { // The deletion failure is the primary error and must not be masked // by the failed write-through. + const consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) + expect(onWrite).toHaveBeenCalledTimes(4) + expect(onWrite.mock.calls[3][0].map((item: HistoryItem) => item.id)).not.toContain(first.id) + expect(consoleError).toHaveBeenCalledWith( + "[TaskHistoryStore] deleteMany write-through after partial deletion failed:", + expect.objectContaining({ message: "write-through failed" }), + ) + consoleError.mockRestore() vi.mocked(fs.unlink).mockImplementation(async (filePath) => actualFs.unlink(filePath)) }) diff --git a/src/utils/__tests__/fileLock.spec.ts b/src/utils/__tests__/fileLock.spec.ts index 93d12bb2d4..66ef34a46b 100644 --- a/src/utils/__tests__/fileLock.spec.ts +++ b/src/utils/__tests__/fileLock.spec.ts @@ -95,7 +95,8 @@ describe("fileLock", () => { }) it("keeps reporting success when an unrelated release error occurs", async () => { - underlyingRelease.mockRejectedValue(new Error("unlock failed")) + const unlockError = new Error("unlock failed") + underlyingRelease.mockRejectedValue(unlockError) await expect( withFileLock( @@ -103,6 +104,6 @@ describe("fileLock", () => { vi.fn(async () => "done"), ), ).resolves.toBe("done") - expect(consoleError).toHaveBeenCalled() + expect(consoleError).toHaveBeenCalledWith(`Failed to release lock for ${absoluteFilePath}:`, unlockError) }) }) From eb7e9700ef82f2689fc5abd461c2c3c5d37359ac Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 22:16:47 +0000 Subject: [PATCH 18/23] fix(lifecycle): narrow PR to rejected-action settlement and dedupe Gemini call IDs Remove deletion coordination, the generic file-lock wrapper, atomicReplace backup machinery, lock-compromise propagation, and the partial deleteMany write-through. Issue #1714 requires none of them. Restore base safeWriteJson behavior and base store deletion semantics, and restore the pre-existing real-lock write-barrier tests. Reimplement clearPendingActionIfMatching on the base safeWriteJson merge callback. The merge callback reads the persisted record under the per-file lock, a throwing merge writes nothing, so a record deleted by another host still fails settlement closed without recreation. Add a request-unique component to synthesized Gemini tool-call IDs. The per-request counter restarts at zero, so the first new_task call of every request previously persisted the same pending-action ID. Exact-ID settlement matching is unchanged. Drop the tests that only covered removed behavior and add a focused test for request-unique Gemini call IDs. --- src/api/providers/__tests__/gemini.spec.ts | 37 ++++ src/api/providers/gemini.ts | 6 +- src/core/task-persistence/TaskHistoryStore.ts | 89 ++------ .../TaskHistoryStore.realConcurrency.spec.ts | 196 ++++++++++-------- .../__tests__/TaskHistoryStore.spec.ts | 146 ------------- .../webview/__tests__/ClineProvider.spec.ts | 1 - src/utils/__tests__/fileLock.spec.ts | 109 ---------- src/utils/__tests__/safeWriteJson.test.ts | 155 +------------- src/utils/fileLock.ts | 85 -------- src/utils/safeWriteJson.ts | 133 +++++------- 10 files changed, 218 insertions(+), 739 deletions(-) delete mode 100644 src/utils/__tests__/fileLock.spec.ts delete mode 100644 src/utils/fileLock.ts diff --git a/src/api/providers/__tests__/gemini.spec.ts b/src/api/providers/__tests__/gemini.spec.ts index 2f19028eb7..c7f1711ae7 100644 --- a/src/api/providers/__tests__/gemini.spec.ts +++ b/src/api/providers/__tests__/gemini.spec.ts @@ -332,6 +332,43 @@ describe("GeminiHandler", () => { }) }) + it("generates request-unique tool call IDs across requests (#1714)", async () => { + const metadata = { + taskId: "test-task", + tools: [{ type: "function", function: { name: "new_task", description: "", parameters: {} } }], + } satisfies ApiHandlerCreateMessageMetadata + const messages: Anthropic.Messages.MessageParam[] = [{ role: "user", content: "Delegate" }] + + // Each request restarts the tool-call counter at zero, so the + // synthesized ID must carry a request-unique component to keep + // persisted pending-action IDs distinct across requests. + const firstCallIds: string[] = [] + for (let request = 0; request < 2; request++) { + mockGenerateContentStream.mockResolvedValueOnce( + asyncStreamFrom([ + { + candidates: [{ content: { parts: [{ functionCall: { name: "new_task", args: {} } }] } }], + }, + ]), + ) + const chunks = await collectStream(handler.createMessage(systemPrompt, messages, metadata)) + const partialIds = chunks + .filter((chunk) => chunk.type === "tool_call_partial") + .map((chunk) => (chunk as { id: string }).id) + expect(partialIds.length).toBeGreaterThan(0) + // Both partial chunks of one synthesized call share one ID. + expect(new Set(partialIds).size).toBe(1) + firstCallIds.push(partialIds[0]) + } + + // The first call of each request must not collide. + expect(firstCallIds[0]).not.toBe(firstCallIds[1]) + for (const id of firstCallIds) { + expect(id).toMatch(/^new_task-.+-0$/) + expect(id).not.toBe("new_task-0") + } + }) + it("should handle API errors", async () => { const mockError = new Error("Gemini API error") ;(handler["client"].models.generateContentStream as any).mockRejectedValue(mockError) diff --git a/src/api/providers/gemini.ts b/src/api/providers/gemini.ts index ec0d14e4c9..7674521d13 100644 --- a/src/api/providers/gemini.ts +++ b/src/api/providers/gemini.ts @@ -354,6 +354,10 @@ export class GeminiHandler extends BaseProvider implements SingleCompletionHandl let finalResponse: { responseId?: string } | undefined let finishReason: string | undefined + // Gemini provides no call ID, so one is synthesized here. The + // request-unique component keeps persisted pending-action IDs (#1714) + // distinct across requests even though the counter restarts at zero. + const toolCallRequestId = crypto.randomUUID() let toolCallCounter = 0 let hasContent = false let hasReasoning = false @@ -398,7 +402,7 @@ export class GeminiHandler extends BaseProvider implements SingleCompletionHandl hasContent = true // Gemini sends complete function calls in a single chunk // Emit as partial chunks for consistent handling with NativeToolCallParser - const callId = `${part.functionCall.name}-${toolCallCounter}` + const callId = `${part.functionCall.name}-${toolCallRequestId}-${toolCallCounter}` const args = JSON.stringify(part.functionCall.args) // Emit name first diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 419ae44bcb..bb0012981a 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -7,8 +7,7 @@ import deepEqual from "fast-deep-equal" import type { HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { LOCK_STALE_MS, withFileLock } from "../../utils/fileLock" -import { safeWriteJson } from "../../utils/safeWriteJson" +import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" @@ -267,7 +266,16 @@ export class TaskHistoryStore { */ async delete(taskId: string): Promise { return this.withLock(async () => { - await this.deleteTaskFile(taskId) + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + + // Remove per-task file (best-effort) + try { + const filePath = await this.getTaskFilePath(taskId) + await fs.unlink(filePath) + } catch { + // File may already be deleted + } // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { @@ -281,30 +289,16 @@ export class TaskHistoryStore { */ async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { - let deletedCount = 0 - try { - for (const taskId of taskIds) { - await this.deleteTaskFile(taskId) - deletedCount++ - } - } catch (error) { - // Earlier deletions already removed their files and cache entries. - // Await the write-through with the current cache before the - // original deletion error rejects the call, so persisted - // globalState does not keep already-deleted tasks. - if (deletedCount > 0 && this.onWrite) { - try { - await this.onWrite(this.getAll()) - } catch (writeError) { - // The deletion failure is the primary error. Report the - // write-through failure without masking it. - console.error( - "[TaskHistoryStore] deleteMany write-through after partial deletion failed:", - writeError, - ) - } + for (const taskId of taskIds) { + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + + try { + const filePath = await this.getTaskFilePath(taskId) + await fs.unlink(filePath) + } catch { + // File may already be deleted } - throw error } // Call onWrite callback inside the lock for serialized write-through @@ -889,34 +883,6 @@ export class TaskHistoryStore { } } - /** - * Delete one task file under the same advisory lock used by `safeWriteJson`. - * Cache state changes only after the file is absent or unlink succeeds. - */ - private async deleteTaskFile(taskId: string): Promise { - const filePath = await this.getTaskFilePath(taskId) - try { - await withFileLock(filePath, async (absoluteFilePath) => { - try { - await fs.unlink(absoluteFilePath) - } catch (error) { - if (!this.isFileNotFoundError(error)) { - throw error - } - } - }) - } catch (error) { - // A missing parent directory prevents lock creation and also proves - // that the history file is absent. - if (!this.isFileNotFoundError(error)) { - throw error - } - } - - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) - } - // ────────────────────────────── Private: fs.watch ────────────────────────────── /** @@ -1122,12 +1088,12 @@ export class TaskHistoryStore { let missingDiskRecord = false try { await safeWriteJson(filePath, cached, { - createParentDirectory: false, - atomicReplace: true, merge: (existing) => { if (!existing || typeof existing !== "object" || !("id" in existing)) { // Writing the cached record back would recreate a task // another host deleted, so drop the stale entry first. + // A throwing merge writes nothing, so the deleted + // record stays deleted. missingDiskRecord = true this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) @@ -1141,16 +1107,7 @@ export class TaskHistoryStore { }, }) } catch (error) { - const missingLockPath = - error && - typeof error === "object" && - "code" in error && - error.code === "ENOENT" && - "path" in error && - error.path === `${filePath}.lock` - if (missingDiskRecord || missingLockPath) { - this.cache.delete(taskId) - this.taskFileMtimes.delete(taskId) + if (missingDiskRecord) { throw new Error( `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, ) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index 39bfde667c..7290b08757 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -6,10 +6,57 @@ import type { HistoryItem } from "@roo-code/types" import { TaskHistoryStore } from "../TaskHistoryStore" -vi.mock("fs/promises", async () => { - const actual = await vi.importActual("fs/promises") - return { ...actual, readFile: vi.fn(actual.readFile) } -}) +type WriteTaskFile = (item: HistoryItem, delta?: Partial) => Promise + +interface WriteBarrier { + arrivals(): number + dispose(): void +} + +function synchronizeNextWrites(stores: TaskHistoryStore[], timeoutMs = 2_000): WriteBarrier { + let arrivals = 0 + let release!: () => void + let rejectBarrier!: (error: Error) => void + let settled = false + let timer: ReturnType | undefined + const barrier = new Promise((resolve, reject) => { + rejectBarrier = reject + release = () => { + if (settled) return + settled = true + if (timer) clearTimeout(timer) + resolve() + } + timer = setTimeout(() => { + if (settled) return + settled = true + reject(new Error(`Only ${arrivals}/${stores.length} stores reached writeTaskFile within ${timeoutMs}ms`)) + }, timeoutMs) + }) + void barrier.catch(() => {}) + + for (const store of stores) { + const value: unknown = Reflect.get(store, "writeTaskFile") + if (typeof value !== "function") throw new Error("TaskHistoryStore.writeTaskFile is unavailable") + const original = value.bind(store) as WriteTaskFile + Reflect.set(store, "writeTaskFile", async (historyItem: HistoryItem, delta?: Partial) => { + arrivals++ + if (arrivals === stores.length) release() + await barrier + return original(historyItem, delta) + }) + } + + return { + arrivals: () => arrivals, + dispose: () => { + if (settled) return + settled = true + if (timer) clearTimeout(timer) + rejectBarrier(new Error("Write barrier disposed before all stores arrived")) + }, + } +} function item(id: string): HistoryItem { return { @@ -37,18 +84,63 @@ function createAction(actionId: string, message: string) { } describe("TaskHistoryStore real cross-host locking", () => { + it("preserves independent stale-cache deltas through the real per-file lock", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-real-lock-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + let writeBarrier: WriteBarrier | undefined + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + writeBarrier = synchronizeNextWrites([storeA, storeB]) + + await Promise.all([ + storeA.atomicReadAndUpdate("shared-task", (current) => ({ ...current, mode: "architect" })), + storeB.atomicReadAndUpdate("shared-task", (current) => ({ ...current, totalCost: 42 })), + ]) + + expect(writeBarrier.arrivals()).toBe(2) + await storeA.invalidate("shared-task") + expect(storeA.get("shared-task")).toMatchObject({ mode: "architect", totalCost: 42 }) + } finally { + writeBarrier?.dispose() + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + + it("reports a bounded error when one store never reaches the write barrier", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-missed-barrier-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + let writeBarrier: WriteBarrier | undefined + + try { + await storeA.initialize() + await storeA.upsert(item("shared-task")) + await storeB.initialize() + writeBarrier = synchronizeNextWrites([storeA, storeB], 50) + + await expect( + storeA.atomicReadAndUpdate("shared-task", (current) => ({ ...current, mode: "architect" })), + ).rejects.toThrow("Only 1/2 stores reached writeTaskFile within 50ms") + expect(writeBarrier.arrivals()).toBe(1) + } finally { + writeBarrier?.dispose() + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("preserves a replacement pending action when a stale store settles the prior action", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-stale-settlement-")) const storeA = new TaskHistoryStore(storagePath) const storeB = new TaskHistoryStore(storagePath) - const actionA = { - kind: "create_subtask" as const, - actionId: "action-a", - approvalText: "{}", - mode: "code", - message: "action A", - todos: [], - } + const actionA = createAction("action-a", "action A") const actionB = { ...actionA, actionId: "action-b", message: "action B" } try { @@ -220,86 +312,6 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) - it("serializes deletion after settlement reads disk without recreating the record", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-during-settlement-")) - const storeA = new TaskHistoryStore(storagePath) - const storeB = new TaskHistoryStore(storagePath) - const action = createAction("action-a", "action A") - const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") - let signalReadComplete!: () => void - const readComplete = new Promise((resolve) => { - signalReadComplete = resolve - }) - let releaseSettlement!: () => void - const settlementCanContinue = new Promise((resolve) => { - releaseSettlement = resolve - }) - - try { - await storeA.initialize() - await storeA.upsert({ ...item("shared-task"), pendingAction: action }) - await storeB.initialize() - - const actualFs = await vi.importActual("fs/promises") - vi.mocked(fs.readFile).mockImplementation(async (...args: Parameters) => { - const result = await actualFs.readFile(...args) - if (args[0] === filePath) { - signalReadComplete() - await settlementCanContinue - } - return result - }) - - const settlement = storeA.clearPendingActionIfMatching("shared-task", action.actionId) - await readComplete - const deletion = storeB.delete("shared-task") - let deletionSettled = false - void deletion.finally(() => { - deletionSettled = true - }) - await new Promise((resolve) => setTimeout(resolve, 25)) - expect(deletionSettled).toBe(false) - - releaseSettlement() - await expect(settlement).resolves.toMatchObject({ id: "shared-task", pendingAction: undefined }) - await deletion - - await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) - await expect(storeA.clearPendingActionIfMatching("shared-task", action.actionId)).rejects.toThrow( - "task shared-task not found", - ) - expect(storeA.get("shared-task")).toBeUndefined() - expect(storeB.get("shared-task")).toBeUndefined() - } finally { - storeA.dispose() - storeB.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - - it("rejects and evicts stale cache without recreating artifacts after full directory deletion", async () => { - const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-deleted-directory-settlement-")) - const store = new TaskHistoryStore(storagePath) - const action = createAction("action-a", "action A") - const taskDirectory = path.join(storagePath, "tasks", "shared-task") - - try { - await store.initialize() - await store.upsert({ ...item("shared-task"), pendingAction: action }) - await fs.rm(taskDirectory, { recursive: true }) - - await expect(store.clearPendingActionIfMatching("shared-task", action.actionId)).rejects.toThrow( - "task shared-task not found", - ) - expect(store.get("shared-task")).toBeUndefined() - await expect(fs.access(taskDirectory)).rejects.toMatchObject({ code: "ENOENT" }) - expect(await fs.readdir(path.join(storagePath, "tasks"))).not.toContain("shared-task") - } finally { - store.dispose() - await fs.rm(storagePath, { recursive: true, force: true }) - } - }) - it("rejects settlement for a task absent from the cache without creating it", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-cache-miss-settlement-")) const store = new TaskHistoryStore(storagePath) diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts index 671b8168aa..3e277ac867 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.spec.ts @@ -10,11 +10,6 @@ import { TaskHistoryStore, assertValidTransition } from "../TaskHistoryStore" import { GlobalFileNames } from "../../../shared/globalFileNames" import { ClineProvider } from "../../webview/ClineProvider" -vi.mock("fs/promises", async () => { - const actual = await vi.importActual("fs/promises") - return { ...actual, unlink: vi.fn(actual.unlink) } -}) - vi.mock("../../../utils/storage", () => ({ getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => { return defaultPath @@ -29,15 +24,6 @@ vi.mock("../../../utils/safeWriteJson", () => ({ }), })) -vi.mock("../../../utils/fileLock", () => ({ - LOCK_STALE_MS: 31_000, - withFileLock: vi - .fn() - .mockImplementation(async (filePath: string, operation: (filePath: string) => Promise) => - operation(filePath), - ), -})) - function makeHistoryItem(overrides: Partial = {}): HistoryItem { return { id: `task-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, @@ -258,33 +244,6 @@ describe("TaskHistoryStore", () => { await store.initialize() await expect(store.delete("non-existent")).resolves.not.toThrow() }) - - it("retains cache state when unlink fails", async () => { - await store.initialize() - const item = makeHistoryItem({ id: "unlink-failure" }) - await store.upsert(item) - const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) - vi.mocked(fs.unlink).mockRejectedValueOnce(unlinkError) - - await expect(store.delete(item.id)).rejects.toBe(unlinkError) - expect(store.get(item.id)).toEqual(item) - }) - - it("retains cache state when lock acquisition fails", async () => { - const onWrite = vi.fn().mockResolvedValue(undefined) - store = new TaskHistoryStore(tmpDir, { onWrite }) - await store.initialize() - const item = makeHistoryItem({ id: "lock-failure" }) - await store.upsert(item) - onWrite.mockClear() - const { withFileLock } = await import("../../../utils/fileLock") - const lockError = new Error("lock failed") - vi.mocked(withFileLock).mockRejectedValueOnce(lockError) - - await expect(store.delete(item.id)).rejects.toBe(lockError) - expect(store.get(item.id)).toEqual(item) - expect(onWrite).not.toHaveBeenCalled() - }) }) describe("deleteMany()", () => { @@ -300,111 +259,6 @@ describe("TaskHistoryStore", () => { expect(store.getAll()).toHaveLength(1) expect(store.get("batch-2")).toBeDefined() }) - - it("retains the failed item and later items when unlink fails", async () => { - await store.initialize() - const first = makeHistoryItem({ id: "batch-success" }) - const failed = makeHistoryItem({ id: "batch-failure" }) - const later = makeHistoryItem({ id: "batch-later" }) - await store.upsert(first) - await store.upsert(failed) - await store.upsert(later) - const actualFs = await vi.importActual("fs/promises") - const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) - vi.mocked(fs.unlink).mockImplementation(async (filePath) => { - if (filePath.toString().includes(failed.id)) { - throw unlinkError - } - return actualFs.unlink(filePath) - }) - - await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) - expect(store.get(first.id)).toBeUndefined() - expect(store.get(failed.id)).toEqual(failed) - expect(store.get(later.id)).toEqual(later) - }) - - it("flushes the write-through with the current cache when a later deletion fails", async () => { - store.dispose() - const globalState: HistoryItem[] = [] - const onWrite = vi.fn(async (items: HistoryItem[]) => { - globalState.splice(0, globalState.length, ...items) - }) - store = new TaskHistoryStore(tmpDir, { onWrite }) - await store.initialize() - - const first = makeHistoryItem({ id: "batch-wt-success" }) - const failed = makeHistoryItem({ id: "batch-wt-failure" }) - const later = makeHistoryItem({ id: "batch-wt-later" }) - await store.upsert(first) - await store.upsert(failed) - await store.upsert(later) - - const actualFs = await vi.importActual("fs/promises") - const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) - vi.mocked(fs.unlink).mockImplementation(async (filePath) => { - if (filePath.toString().includes(failed.id)) { - throw unlinkError - } - return actualFs.unlink(filePath) - }) - - await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) - - // The earlier deleted task is evicted from the store cache. - expect(store.get(first.id)).toBeUndefined() - // The awaited write-through persisted that state before the original - // deletion error rejected the call, so stale globalState cannot - // return the deleted task. - expect(globalState.find((item) => item.id === first.id)).toBeUndefined() - expect(globalState.find((item) => item.id === failed.id)).toEqual(failed) - expect(globalState.find((item) => item.id === later.id)).toEqual(later) - - vi.mocked(fs.unlink).mockImplementation(async (filePath) => actualFs.unlink(filePath)) - }) - - it("keeps the original deletion error when the partial write-through also fails", async () => { - store.dispose() - const onWrite = vi.fn() - // Three upserts succeed, then the partial-failure flush rejects. - onWrite - .mockResolvedValueOnce(undefined) - .mockResolvedValueOnce(undefined) - .mockResolvedValueOnce(undefined) - .mockRejectedValueOnce(new Error("write-through failed")) - store = new TaskHistoryStore(tmpDir, { onWrite }) - await store.initialize() - - const first = makeHistoryItem({ id: "batch-wt2-success" }) - const failed = makeHistoryItem({ id: "batch-wt2-failure" }) - const later = makeHistoryItem({ id: "batch-wt2-later" }) - await store.upsert(first) - await store.upsert(failed) - await store.upsert(later) - - const actualFs = await vi.importActual("fs/promises") - const unlinkError = Object.assign(new Error("unlink failed"), { code: "EACCES" }) - vi.mocked(fs.unlink).mockImplementation(async (filePath) => { - if (filePath.toString().includes(failed.id)) { - throw unlinkError - } - return actualFs.unlink(filePath) - }) - - // The deletion failure is the primary error and must not be masked - // by the failed write-through. - const consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) - await expect(store.deleteMany([first.id, failed.id, later.id])).rejects.toBe(unlinkError) - expect(onWrite).toHaveBeenCalledTimes(4) - expect(onWrite.mock.calls[3][0].map((item: HistoryItem) => item.id)).not.toContain(first.id) - expect(consoleError).toHaveBeenCalledWith( - "[TaskHistoryStore] deleteMany write-through after partial deletion failed:", - expect.objectContaining({ message: "write-through failed" }), - ) - consoleError.mockRestore() - - vi.mocked(fs.unlink).mockImplementation(async (filePath) => actualFs.unlink(filePath)) - }) }) describe("reconcile()", () => { diff --git a/src/core/webview/__tests__/ClineProvider.spec.ts b/src/core/webview/__tests__/ClineProvider.spec.ts index 75885185bd..97c4dd877e 100644 --- a/src/core/webview/__tests__/ClineProvider.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.spec.ts @@ -87,7 +87,6 @@ vi.mock("../../../utils/storage", () => ({ getSettingsDirectoryPath: vi.fn().mockResolvedValue("/test/settings/path"), getTaskDirectoryPath: vi.fn().mockResolvedValue("/test/task/path"), getGlobalStoragePath: vi.fn().mockResolvedValue("/test/storage/path"), - getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => defaultPath), })) vi.mock("@modelcontextprotocol/sdk/types.js", () => ({ diff --git a/src/utils/__tests__/fileLock.spec.ts b/src/utils/__tests__/fileLock.spec.ts deleted file mode 100644 index 66ef34a46b..0000000000 --- a/src/utils/__tests__/fileLock.spec.ts +++ /dev/null @@ -1,109 +0,0 @@ -// pnpm --filter roo-cline test utils/__tests__/fileLock.spec.ts - -import * as path from "path" - -import { acquireFileLock, withFileLock } from "../fileLock" - -const { lockMock } = vi.hoisted(() => ({ lockMock: vi.fn() })) - -vi.mock("proper-lockfile", () => ({ lock: lockMock, default: { lock: lockMock } })) - -interface CapturedLockOptions { - onCompromised: (error: Error) => void -} - -function requireCapturedLockOptions(): CapturedLockOptions { - if (!capturedOptions) { - throw new Error("proper-lockfile options were not captured") - } - return capturedOptions -} - -let capturedOptions: CapturedLockOptions | undefined - -describe("fileLock", () => { - const absoluteFilePath = path.resolve("/virtual/dir/target.json") - let underlyingRelease: ReturnType - let consoleError: ReturnType - - beforeEach(() => { - consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) - underlyingRelease = vi.fn().mockResolvedValue(undefined) - capturedOptions = undefined - lockMock.mockImplementation(async (_filePath: string, options: CapturedLockOptions) => { - capturedOptions = options - return underlyingRelease - }) - }) - - afterEach(() => { - consoleError.mockRestore() - vi.clearAllMocks() - }) - - it("records the compromise without throwing from the renewal callback", () => { - const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) - - return acquireFileLock(absoluteFilePath).then((release) => { - expect(() => requireCapturedLockOptions().onCompromised(compromiseError)).not.toThrow() - return release().catch(() => {}) - }) - }) - - it("rejects the owning operation with the recorded compromise error", async () => { - const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) - const operation = vi.fn(async () => { - // Simulate the proper-lockfile renewal timer reporting a - // compromise while the operation still holds the lock. - requireCapturedLockOptions().onCompromised(compromiseError) - return "done" - }) - - await expect(withFileLock(absoluteFilePath, operation)).rejects.toBe(compromiseError) - expect(operation).toHaveBeenCalledTimes(1) - // proper-lockfile already marked the lock released, so the wrapper - // must not call the underlying release after a compromise. - expect(underlyingRelease).not.toHaveBeenCalled() - }) - - it("keeps the operation error when the operation fails after a compromise", async () => { - const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) - const operationError = new Error("operation failed") - const operation = vi.fn(async () => { - requireCapturedLockOptions().onCompromised(compromiseError) - throw operationError - }) - - await expect(withFileLock(absoluteFilePath, operation)).rejects.toBe(operationError) - }) - - it("rejects a direct release with the recorded compromise error", async () => { - const compromiseError = Object.assign(new Error("lock renewal failed"), { code: "ECOMPROMISED" }) - const release = await acquireFileLock(absoluteFilePath) - - requireCapturedLockOptions().onCompromised(compromiseError) - - await expect(release()).rejects.toBe(compromiseError) - expect(underlyingRelease).not.toHaveBeenCalled() - }) - - it("releases normally without a compromise", async () => { - const operation = vi.fn(async () => "done") - - await expect(withFileLock(absoluteFilePath, operation)).resolves.toBe("done") - expect(underlyingRelease).toHaveBeenCalledTimes(1) - }) - - it("keeps reporting success when an unrelated release error occurs", async () => { - const unlockError = new Error("unlock failed") - underlyingRelease.mockRejectedValue(unlockError) - - await expect( - withFileLock( - absoluteFilePath, - vi.fn(async () => "done"), - ), - ).resolves.toBe("done") - expect(consoleError).toHaveBeenCalledWith(`Failed to release lock for ${absoluteFilePath}:`, unlockError) - }) -}) diff --git a/src/utils/__tests__/safeWriteJson.test.ts b/src/utils/__tests__/safeWriteJson.test.ts index c58174a06c..79d08678a0 100644 --- a/src/utils/__tests__/safeWriteJson.test.ts +++ b/src/utils/__tests__/safeWriteJson.test.ts @@ -10,7 +10,6 @@ import { safeWriteJson } from "../safeWriteJson" // test mockImplementation callbacks delegate to the real implementation. const fsPromisesActuals = vi.hoisted(() => ({ rename: undefined as (typeof import("fs/promises"))["rename"] | undefined, - copyFile: undefined as (typeof import("fs/promises"))["copyFile"] | undefined, unlink: undefined as (typeof import("fs/promises"))["unlink"] | undefined, writeFile: undefined as (typeof import("fs/promises"))["writeFile"] | undefined, })) @@ -18,7 +17,6 @@ const fsPromisesActuals = vi.hoisted(() => ({ vi.mock("fs/promises", async () => { const actual = await vi.importActual("fs/promises") fsPromisesActuals.rename = actual.rename - fsPromisesActuals.copyFile = actual.copyFile fsPromisesActuals.unlink = actual.unlink fsPromisesActuals.writeFile = actual.writeFile // Start with all actual implementations. @@ -30,7 +28,6 @@ vi.mock("fs/promises", async () => { mockedFs.writeFile = vi.fn(actual.writeFile) as any mockedFs.readFile = vi.fn(actual.readFile) as any mockedFs.rename = vi.fn(actual.rename) as any - mockedFs.copyFile = vi.fn(actual.copyFile) as typeof actual.copyFile mockedFs.unlink = vi.fn(actual.unlink) as any mockedFs.access = vi.fn(actual.access) as any mockedFs.mkdtemp = vi.fn(actual.mkdtemp) as any @@ -214,82 +211,6 @@ describe("safeWriteJson", () => { expect(content).toEqual(initialData) }) - test("should replace an existing file with one rename when atomicReplace is enabled", async () => { - const initialData = { message: "Initial content" } - const newData = { message: "New content" } - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - vi.mocked(fs.rename).mockClear() - - await safeWriteJson(currentTestFilePath, newData, { atomicReplace: true }) - - expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(1) - expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(1) - expect(await readFileContent(currentTestFilePath)).toEqual(newData) - }) - - test("should restore the copied backup when atomic replacement fails", async () => { - const initialData = { message: "Initial content, should be restored" } - const newData = { message: "New content" } - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - vi.mocked(fs.copyFile).mockClear() - vi.mocked(fs.rename).mockClear() - vi.mocked(fs.rename).mockRejectedValueOnce(new Error("Atomic replacement failed")) - - await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( - "Atomic replacement failed", - ) - - expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(2) - expect(vi.mocked(fs.rename)).toHaveBeenCalledTimes(2) - expect(await readFileContent(currentTestFilePath)).toEqual(initialData) - }) - - test("should not roll back from an incomplete backup copy", async () => { - const initialData = { message: "Initial content" } - const newData = { message: "New content" } - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - vi.mocked(fs.copyFile).mockClear() - vi.mocked(fs.rename).mockClear() - vi.mocked(fs.copyFile).mockRejectedValueOnce(new Error("Backup copy failed")) - - await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( - "Backup copy failed", - ) - - expect(vi.mocked(fs.copyFile)).toHaveBeenCalledTimes(1) - expect(vi.mocked(fs.rename)).not.toHaveBeenCalled() - expect(await readFileContent(currentTestFilePath)).toEqual(initialData) - }) - - test("should preserve the completed backup and remove an incomplete rollback copy when atomic rollback fails", async () => { - const initialData = { message: "Initial content" } - const newData = { message: "New content" } - await fsPromisesActuals.writeFile!(currentTestFilePath, JSON.stringify(initialData)) - const consoleErrorSpy = vi.spyOn(console, "error").mockImplementation(() => {}) - vi.mocked(fs.rename).mockRejectedValueOnce(new Error("Atomic replacement failed")) - vi.mocked(fs.copyFile) - .mockImplementationOnce(fsPromisesActuals.copyFile!) - .mockImplementationOnce(async (_source, target) => { - await fsPromisesActuals.writeFile!(target, "incomplete rollback") - throw new Error("Rollback copy failed") - }) - - await expect(safeWriteJson(currentTestFilePath, newData, { atomicReplace: true })).rejects.toThrow( - "Atomic replacement failed", - ) - - const remainingFiles = await fs.readdir(tempDir) - const backupFiles = remainingFiles.filter((file) => file.includes(".bak_")) - expect(backupFiles).toHaveLength(1) - expect(await readFileContent(path.join(tempDir, backupFiles[0]))).toEqual(initialData) - expect(remainingFiles.some((file) => file.includes(".rollback_"))).toBe(false) - expect(await readFileContent(currentTestFilePath)).toEqual(initialData) - expect(consoleErrorSpy).toHaveBeenCalledWith( - expect.stringContaining("Failed to restore backup"), - expect.objectContaining({ message: "Rollback copy failed" }), - ) - }) - // Tests for directory creation functionality test("should create parent directory if it doesn't exist", async () => { // Create a path in a non-existent subdirectory of the temp dir @@ -311,17 +232,6 @@ describe("safeWriteJson", () => { expect(content).toEqual(data) }) - test("should reject without creating the parent directory when parent creation is disabled", async () => { - const subDir = path.join(tempDir, "missing-parent") - const filePath = path.join(subDir, "file.json") - - await expect(safeWriteJson(filePath, { value: 1 }, { createParentDirectory: false })).rejects.toMatchObject({ - code: "ENOENT", - path: `${filePath}.lock`, - }) - await expect(fs.access(subDir)).rejects.toMatchObject({ code: "ENOENT" }) - }) - test("should handle multi-level directory creation", async () => { // Create a new non-existent subdirectory path with multiple levels const deepDir = path.join(tempDir, "level1", "level2", "level3") @@ -477,67 +387,6 @@ describe("safeWriteJson", () => { await fs.unlink(lockTestFilePath).catch(() => {}) // Ignore errors if file doesn't exist vi.unmock("proper-lockfile") // Ensure the mock is removed after this test }) - - test("rejects with a lock compromise error after a successful write", async () => { - vi.resetModules() - - const data = { message: "compromise after success" } - const compromiseTestFilePath = path.join(tempDir, "compromise-test-file.json") - await fs.writeFile(compromiseTestFilePath, JSON.stringify({ initial: "content" })) - - const compromiseError = Object.assign(new Error("lock was compromised"), { code: "ECOMPROMISED" }) - vi.doMock("proper-lockfile", () => ({ - ...vi.importActual("proper-lockfile"), - lock: vi.fn().mockResolvedValue(vi.fn().mockRejectedValue(compromiseError)), - })) - - const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") - - // The write itself succeeds, but the release reports a compromised - // lock, so the operation must reject instead of reporting success. - await expect(mockedSafeWriteJson(compromiseTestFilePath, data)).rejects.toBe(compromiseError) - - await fs.unlink(compromiseTestFilePath).catch(() => {}) - vi.doUnmock("proper-lockfile") - }) - - test("keeps the original write error when the write fails and the lock was compromised", async () => { - vi.resetModules() - - const data = { message: "compromise after failure" } - const compromiseTestFilePath = path.join(tempDir, "compromise-failure-test-file.json") - await fs.writeFile(compromiseTestFilePath, JSON.stringify({ initial: "content" })) - - const compromiseError = Object.assign(new Error("lock was compromised"), { code: "ECOMPROMISED" }) - vi.doMock("proper-lockfile", () => ({ - ...vi.importActual("proper-lockfile"), - lock: vi.fn().mockResolvedValue(vi.fn().mockRejectedValue(compromiseError)), - })) - - const createWriteStreamSpy = vi.spyOn(fsSyncActual, "createWriteStream") - createWriteStreamSpy.mockImplementationOnce(() => { - // A plain Writable provides the pipe and error surface that - // `_streamDataToFile` uses, but not the full WriteStream - // interface, so this cast is deliberate. - const errorStream = new Writable({ - write: (_chunk, _encoding, callback) => { - callback(new Error("Stream write error")) - }, - }) as unknown as fsSyncActual.WriteStream - errorStream.close = vi.fn() - return errorStream - }) - - const { safeWriteJson: mockedSafeWriteJson } = await import("../safeWriteJson") - - // The write failure is the primary error and must not be masked by - // the compromise reported at release time. - await expect(mockedSafeWriteJson(compromiseTestFilePath, data)).rejects.toThrow("Stream write error") - - createWriteStreamSpy.mockRestore() - await fs.unlink(compromiseTestFilePath).catch(() => {}) - vi.doUnmock("proper-lockfile") - }) test("should release lock even if an error occurs mid-operation", async () => { const data = { message: "test lock release on error" } @@ -637,11 +486,11 @@ describe("safeWriteJson", () => { expect(content).toEqual({ a: 1, b: 3, c: 4 }) }) - test("should pass null to merge callback under the lock when the parent exists but the file does not", async () => { + test("should pass null to merge callback when file does not exist", async () => { const newFilePath = path.join(tempDir, "nonexistent.json") const mergeFn = vi.fn((existing, incoming) => incoming) - await safeWriteJson(newFilePath, { value: 42 }, { createParentDirectory: false, merge: mergeFn }) + await safeWriteJson(newFilePath, { value: 42 }, { merge: mergeFn }) expect(mergeFn).toHaveBeenCalledWith(null, { value: 42 }) const content = await readFileContent(newFilePath) diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts deleted file mode 100644 index 54f9be732c..0000000000 --- a/src/utils/fileLock.ts +++ /dev/null @@ -1,85 +0,0 @@ -import * as path from "path" -import * as lockfile from "proper-lockfile" - -export const LOCK_STALE_MS = 31_000 - -export async function acquireFileLock(filePath: string): Promise<() => Promise> { - const absoluteFilePath = path.resolve(filePath) - // proper-lockfile calls `onCompromised` from its renewal timer after the - // lock promise already resolved. A throw here becomes an uncaught - // exception in the host instead of a rejection, so record the error and - // fail the owning operation through its release call instead. - let compromisedError: Error | null = null - try { - const release = await lockfile.lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10_000, - realpath: false, - retries: { - retries: 5, - factor: 2, - minTimeout: 100, - maxTimeout: 1_000, - }, - onCompromised: (error) => { - console.error(`Lock at ${absoluteFilePath} was compromised:`, error) - compromisedError = error - }, - }) - return async () => { - if (compromisedError) { - // proper-lockfile marks the lock released when it reports a - // compromise and its own release resolves silently, so reject - // with the recorded compromise error here. - throw compromisedError - } - await release() - } - } catch (error) { - console.error(`Failed to acquire lock for ${absoluteFilePath}:`, error) - throw error - } -} - -/** - * Run one operation while holding the advisory lock for a file path. - * Callers must not acquire this lock again from inside `operation`. - */ -export async function withFileLock( - filePath: string, - operation: (absoluteFilePath: string) => Promise, -): Promise { - const absoluteFilePath = path.resolve(filePath) - const releaseLock = await acquireFileLock(absoluteFilePath) - - let result: T - try { - result = await operation(absoluteFilePath) - } catch (operationError) { - // The operation error is the primary failure. Release without - // reporting a secondary release error over it. - try { - await releaseLock() - } catch (releaseError) { - console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) - } - throw operationError - } - - try { - await releaseLock() - } catch (releaseError) { - const code = - releaseError && typeof releaseError === "object" && "code" in releaseError - ? (releaseError as { code: unknown }).code - : undefined - if (code === "ECOMPROMISED") { - // The lock was compromised while this operation held it, so the - // operation ran without mutual exclusion. Reject the operation - // that owns the lock instead of reporting success. - throw releaseError - } - console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) - } - return result -} diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 49ae87a06f..957a0bb20f 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -1,20 +1,13 @@ import * as fs from "fs/promises" import * as fsSync from "fs" import * as path from "path" +import * as lockfile from "proper-lockfile" import { JsonStreamStringify } from "json-stream-stringify" -import { acquireFileLock, LOCK_STALE_MS } from "./fileLock" - /** * Options for safeWriteJson function */ export interface SafeWriteJsonOptions { - /** - * Whether to create and verify the target file's parent directory. - * @default true - */ - createParentDirectory?: boolean - /** * Whether to pretty-print the JSON output with indentation. * When true, uses tab characters for indentation. @@ -32,14 +25,6 @@ export interface SafeWriteJsonOptions { * cannot be parsed. */ merge?: (existing: unknown, incoming: unknown) => unknown - - /** - * Replace an existing target with one rename from the completed temporary file. - * A copied backup retains rollback support without removing the target before - * the replacement rename. - * @default false - */ - atomicReplace?: boolean } /** @@ -64,23 +49,36 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // For directory creation const dirPath = path.dirname(absoluteFilePath) - if (options?.createParentDirectory !== false) { - // Ensure directory structure exists with improved reliability - try { - // Create directory with recursive option - await fs.mkdir(dirPath, { recursive: true }) - - // Verify directory exists after creation attempt - await fs.access(dirPath) - } catch (dirError: any) { - console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) - throw dirError - } + // Ensure directory structure exists with improved reliability + try { + // Create directory with recursive option + await fs.mkdir(dirPath, { recursive: true }) + + // Verify directory exists after creation attempt + await fs.access(dirPath) + } catch (dirError: any) { + console.error(`Failed to create or access directory for ${absoluteFilePath}:`, dirError) + throw dirError } // Acquire the lock before any file operations try { - releaseLock = await acquireFileLock(absoluteFilePath) + releaseLock = await lockfile.lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long + realpath: false, // the file may not exist yet, which is acceptable + retries: { + // Configuration for retrying lock acquisition + retries: 5, // Number of retries after the initial attempt + factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) + minTimeout: 100, // Minimum time to wait before the first retry (in ms) + maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) + }, + onCompromised: (err) => { + console.error(`Lock at ${absoluteFilePath} was compromised:`, err) + throw err + }, + }) } catch (lockError) { // If lock acquisition fails, we throw immediately. // The releaseLock remains a no-op, so the finally block in the main file operations @@ -93,7 +91,6 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Variables to hold the actual paths of temp files if they are created. let actualTempNewFilePath: string | null = null let actualTempBackupFilePath: string | null = null - let actualTempRollbackFilePath: string | null = null try { // If a merge callback was provided, read the current file under the lock @@ -121,24 +118,23 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso await _streamDataToFile(actualTempNewFilePath, data, options?.prettyPrint) - // Step 2: Check if the target file exists. If so, retain a rollback backup. + // Step 2: Check if the target file exists. If so, rename it to a backup path. try { + // Check for target file existence await fs.access(absoluteFilePath) - const candidateBackupFilePath = path.join( + // Target exists, create a backup path and rename. + actualTempBackupFilePath = path.join( path.dirname(absoluteFilePath), `.${path.basename(absoluteFilePath)}.bak_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, ) - if (options?.atomicReplace) { - await fs.copyFile(absoluteFilePath, candidateBackupFilePath) - } else { - await fs.rename(absoluteFilePath, candidateBackupFilePath) - } - actualTempBackupFilePath = candidateBackupFilePath + await fs.rename(absoluteFilePath, actualTempBackupFilePath) } catch (accessError: any) { + // Explicitly type accessError if (accessError.code !== "ENOENT") { + // An error other than "file not found" occurred during access check. throw accessError } - actualTempBackupFilePath = null + // Target file does not exist, so no backup is made. actualTempBackupFilePath remains null. } // Step 3: Rename the new temporary file to the target file path. @@ -174,18 +170,7 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso // Attempt rollback if a backup was made if (backupFileToRollbackOrCleanupWithinCatch) { try { - if (options?.atomicReplace) { - actualTempRollbackFilePath = path.join( - path.dirname(absoluteFilePath), - `.${path.basename(absoluteFilePath)}.rollback_${Date.now()}_${Math.random().toString(36).substring(2)}.tmp`, - ) - await fs.copyFile(backupFileToRollbackOrCleanupWithinCatch, actualTempRollbackFilePath) - await fs.rename(actualTempRollbackFilePath, absoluteFilePath) - actualTempRollbackFilePath = null - await fs.unlink(backupFileToRollbackOrCleanupWithinCatch) - } else { - await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) - } + await fs.rename(backupFileToRollbackOrCleanupWithinCatch, absoluteFilePath) // Mark as handled, prevent later unlink of this path actualTempBackupFilePath = null } catch (rollbackError) { @@ -197,19 +182,6 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } - // A failed rollback can leave an incomplete rollback copy. The completed backup remains available for recovery. - if (actualTempRollbackFilePath) { - try { - await fs.unlink(actualTempRollbackFilePath) - actualTempRollbackFilePath = null - } catch (cleanupError) { - console.error( - `[Catch] Failed to clean up temporary rollback file ${actualTempRollbackFilePath}:`, - cleanupError, - ) - } - } - // Cleanup the .new file if it exists if (newFileToCleanupWithinCatch) { try { @@ -222,8 +194,8 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso } } - // A copied backup remains available for recovery when atomic rollback fails. - if (actualTempBackupFilePath && !options?.atomicReplace) { + // Cleanup the .bak file if it still needs to be (i.e., wasn't successfully restored) + if (actualTempBackupFilePath) { try { await fs.unlink(actualTempBackupFilePath) } catch (cleanupError) { @@ -233,30 +205,17 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso ) } } - // Release the lock before rejecting. The original write failure is the - // rejection and a release failure cannot mask it. + throw originalError // This MUST be the error that rejects the promise. + } finally { + // Release the lock in the main finally block. try { + // releaseLock will be the actual unlock function if lock was acquired, + // or the initial no-op if acquisition failed. await releaseLock() } catch (unlockError) { + // Do not re-throw here, as the originalError from the try/catch (if any) is more important. console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) } - throw originalError // This MUST be the error that rejects the promise. - } - - // Release the lock on the success path. A compromised lock means this - // write ran without mutual exclusion, so reject this operation instead - // of reporting success. - try { - await releaseLock() - } catch (unlockError) { - const code = - unlockError && typeof unlockError === "object" && "code" in unlockError - ? (unlockError as { code: unknown }).code - : undefined - if (code === "ECOMPROMISED") { - throw unlockError - } - console.error(`Failed to release lock for ${absoluteFilePath}:`, unlockError) } } @@ -288,4 +247,6 @@ async function _streamDataToFile(targetPath: string, data: any, prettyPrint = fa }) } -export { LOCK_STALE_MS, safeWriteJson } +export const LOCK_STALE_MS = 31_000 + +export { safeWriteJson } From 69f761b1657c25bd096cfab886da372cd8c02332 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 22:29:41 +0000 Subject: [PATCH 19/23] test: require both partial chunks in the Gemini call-ID test The handler emits exactly one name partial and one arguments partial per synthesized call. Require both so the test fails when either chunk is missing. --- src/api/providers/__tests__/gemini.spec.ts | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/src/api/providers/__tests__/gemini.spec.ts b/src/api/providers/__tests__/gemini.spec.ts index c7f1711ae7..8658e835bd 100644 --- a/src/api/providers/__tests__/gemini.spec.ts +++ b/src/api/providers/__tests__/gemini.spec.ts @@ -355,8 +355,9 @@ describe("GeminiHandler", () => { const partialIds = chunks .filter((chunk) => chunk.type === "tool_call_partial") .map((chunk) => (chunk as { id: string }).id) - expect(partialIds.length).toBeGreaterThan(0) - // Both partial chunks of one synthesized call share one ID. + // The handler emits one name partial and one arguments partial + // for the synthesized call. + expect(partialIds).toHaveLength(2) expect(new Set(partialIds).size).toBe(1) firstCallIds.push(partialIds[0]) } From 29765dcf48f45a4dd789a6f027876f4ec22cd7b3 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Mon, 28 Sep 2026 23:10:13 +0000 Subject: [PATCH 20/23] fix: serialize history deletion with settlement and validate the locked record Address PR #1726 review findings with minimal changes scoped to the new settlement writer. Deletion versus settlement (#1726 finding 1) - src/utils/fileLock.ts exports acquireFileLock and withFileLock. They use the exact advisory lock protocol that safeWriteJson uses, so one implementation owns the per-file lock. safeWriteJson acquires through acquireFileLock and keeps its write, backup, and release behavior. - TaskHistoryStore.delete and deleteMany unlink each history_item.json inside withFileLock. Cache eviction and write-through semantics stay identical to base. Lock ordering stays store lock first, then one per-file lock, matching the write path, so no path nests the file lock. - The real-filesystem race test pauses settlement inside its locked disk read, starts deletion from a second store, verifies the deletion stays blocked, releases settlement, and verifies the file stays deleted without resurrection. Locked-record validation (#1726 finding 3) - The settlement merge callback now validates the disk record with the canonical historyItemSchema and requires parsed.data.id === taskId. - Malformed or mismatched records evict the stale cache entry and fail settlement closed without rewriting the record. Missing-record eviction and error behavior stay unchanged. Gemini partial assertions (#1726 finding 2) - The request-unique tool-call ID test now asserts exactly one partial with name new_task and no arguments and exactly one partial with serialized empty arguments and no name, then keeps the shared-ID and cross-request uniqueness assertions. The pre-commit hook was skipped because the full-workspace eslint in this environment was killed for memory (exit 137). Targeted eslint --prune-suppressions --max-warnings=0 passed on every edited file with no suppression count increase, and CI runs the full lint. Validation: focused Vitest suites pass, tsc --noEmit passes, all seven bounded lifecycle model checkers pass, and full pnpm test passes (13 tasks, 488 files, 9021 tests). --- docs/architecture/task-lifecycle-model.md | 2 +- src/api/providers/__tests__/gemini.spec.ts | 17 +++- src/core/task-persistence/TaskHistoryStore.ts | 69 +++++++++++++--- .../TaskHistoryStore.crossInstance.spec.ts | 63 +++++++++++++++ .../TaskHistoryStore.realConcurrency.spec.ts | 73 +++++++++++++++++ src/utils/fileLock.ts | 78 +++++++++++++++++++ src/utils/safeWriteJson.ts | 39 +++------- 7 files changed, 297 insertions(+), 44 deletions(-) create mode 100644 src/utils/fileLock.ts diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index 38324af9ba..ed186961f7 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -83,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear. Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear. History-file deletion serializes on the same per-file advisory lock as `safeWriteJson`, and the real-filesystem suite covers a deletion that targets the settlement read-to-commit window plus schema and task-ID validation of the locked disk record before settlement (#1726). Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. ## Task cleanup protocol model diff --git a/src/api/providers/__tests__/gemini.spec.ts b/src/api/providers/__tests__/gemini.spec.ts index 8658e835bd..0920042586 100644 --- a/src/api/providers/__tests__/gemini.spec.ts +++ b/src/api/providers/__tests__/gemini.spec.ts @@ -352,12 +352,21 @@ describe("GeminiHandler", () => { ]), ) const chunks = await collectStream(handler.createMessage(systemPrompt, messages, metadata)) - const partialIds = chunks + const partials = chunks .filter((chunk) => chunk.type === "tool_call_partial") - .map((chunk) => (chunk as { id: string }).id) + .map((chunk) => chunk as { id: string; name?: string; arguments?: string }) // The handler emits one name partial and one arguments partial - // for the synthesized call. - expect(partialIds).toHaveLength(2) + // for the synthesized call. Assert both semantic halves so a + // duplicate name partial cannot satisfy the ID comparison. + expect(partials).toHaveLength(2) + expect( + partials.filter(({ name, arguments: args }) => name === "new_task" && args === undefined), + ).toHaveLength(1) + expect( + partials.filter(({ name, arguments: args }) => name === undefined && args === "{}"), + ).toHaveLength(1) + + const partialIds = partials.map(({ id }) => id) expect(new Set(partialIds).size).toBe(1) firstCallIds.push(partialIds[0]) } diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index bb0012981a..9c70d01a0d 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -4,10 +4,11 @@ import * as path from "path" import crypto from "crypto" import deepEqual from "fast-deep-equal" -import type { HistoryItem } from "@roo-code/types" +import { historyItemSchema, type HistoryItem } from "@roo-code/types" import { GlobalFileNames } from "../../shared/globalFileNames" -import { LOCK_STALE_MS, safeWriteJson } from "../../utils/safeWriteJson" +import { LOCK_STALE_MS, withFileLock } from "../../utils/fileLock" +import { safeWriteJson } from "../../utils/safeWriteJson" import { getStorageBasePath } from "../../utils/storage" import { assertValidTransition, settleRejectedCreateSubtaskAction, type HistoryItemStatus } from "./taskLifecycle" import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./taskStoreConcurrency" @@ -269,12 +270,25 @@ export class TaskHistoryStore { this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) - // Remove per-task file (best-effort) + // Remove per-task file (best-effort). The unlink runs under the + // same per-file advisory lock that `safeWriteJson` holds, so a + // deletion cannot interleave with another host's locked + // read-modify-write (for example the settlement in + // `clearPendingActionIfMatching`). Lock ordering stays store + // lock first, then one per-file lock, matching the write path. try { const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) + await withFileLock(filePath, async (absoluteFilePath) => { + try { + await fs.unlink(absoluteFilePath) + } catch { + // File may already be deleted + } + }) } catch { - // File may already be deleted + // A missing task directory also proves the history file is + // absent, so remaining failures keep the base best-effort + // deletion semantics. } // Call onWrite callback inside the lock for serialized write-through @@ -293,11 +307,21 @@ export class TaskHistoryStore { this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) + // Serialize each unlink with `safeWriteJson` under the same + // per-file advisory lock. See `delete` for the lock order. try { const filePath = await this.getTaskFilePath(taskId) - await fs.unlink(filePath) + await withFileLock(filePath, async (absoluteFilePath) => { + try { + await fs.unlink(absoluteFilePath) + } catch { + // File may already be deleted + } + }) } catch { - // File may already be deleted + // A missing task directory also proves the history file + // is absent, so remaining failures keep the base + // best-effort deletion semantics. } } @@ -1073,9 +1097,13 @@ export class TaskHistoryStore { * * Deletion by another host is authoritative (#1726): when no persisted * record exists, the merge callback removes the stale cache entry and - * throws instead of writing the cached record back to disk. + * throws instead of writing the cached record back to disk. A persisted + * record must also match the canonical task-history schema and carry the + * requested task ID; malformed or mismatched records drop the stale + * cache entry and fail settlement closed without rewriting the record. * - * @throws If the task ID is not present in the cache or no persisted record remains on disk. + * @throws If the task ID is not present in the cache, no persisted record remains on disk, + * or the persisted record is invalid or belongs to a different task ID. */ public async clearPendingActionIfMatching(taskId: string, expectedActionId: string): Promise { return this.withLock(async () => { @@ -1089,7 +1117,7 @@ export class TaskHistoryStore { try { await safeWriteJson(filePath, cached, { merge: (existing) => { - if (!existing || typeof existing !== "object" || !("id" in existing)) { + if (existing === null || existing === undefined) { // Writing the cached record back would recreate a task // another host deleted, so drop the stale entry first. // A throwing merge writes nothing, so the deleted @@ -1101,6 +1129,27 @@ export class TaskHistoryStore { `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} not found in cache`, ) } + // Validate the locked disk record with the canonical + // task-history schema (#1726). Settlement must not + // rewrite malformed data and must not clear the action + // on a record persisted under a different task ID, so + // both cases drop the stale cache entry and fail + // closed without touching the disk record. + const parsed = historyItemSchema.safeParse(existing) + if (!parsed.success) { + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} has an invalid disk record`, + ) + } + if (parsed.data.id !== taskId) { + this.cache.delete(taskId) + this.taskFileMtimes.delete(taskId) + throw new Error( + `[TaskHistoryStore] clearPendingActionIfMatching: task ${taskId} has a disk record with mismatched id ${parsed.data.id}`, + ) + } const disk = existing as HistoryItem authoritative = settleRejectedCreateSubtaskAction(disk, expectedActionId) return authoritative diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts index fe7846edd3..5b85c9b4e2 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.crossInstance.spec.ts @@ -253,6 +253,69 @@ describe("TaskHistoryStore cross-instance safety", () => { expect(await fs.readFile(filePath, "utf8")).toBe("{invalid") }) + it("fails settlement closed for a malformed disk record without rewriting it", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "action-a", + approvalText: "{}", + mode: "code", + message: "action A", + todos: [], + } + const cached = makeHistoryItem({ id: "malformed-settlement-task", pendingAction }) + const filePath = path.join(tmpDir, "tasks", cached.id, GlobalFileNames.historyItem) + await fs.mkdir(path.dirname(filePath), { recursive: true }) + // Valid JSON that fails the canonical task-history schema: id must + // be a string and the required history fields are absent. + const malformed = JSON.stringify({ id: 123, pendingAction }) + await fs.writeFile(filePath, malformed, "utf8") + const storeState = storeA as unknown as { + cache: Map + taskFileMtimes: Map + } + storeState.cache.set(cached.id, cached) + storeState.taskFileMtimes.set(cached.id, Date.now()) + + await expect(storeA.clearPendingActionIfMatching(cached.id, pendingAction.actionId)).rejects.toThrow( + `task ${cached.id} has an invalid disk record`, + ) + expect(storeA.get(cached.id)).toBeUndefined() + expect(storeState.taskFileMtimes.has(cached.id)).toBe(false) + // The malformed record stays on disk untouched. + expect(await fs.readFile(filePath, "utf8")).toBe(malformed) + }) + + it("fails settlement closed when the disk record carries a different task id", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "action-a", + approvalText: "{}", + mode: "code", + message: "action A", + todos: [], + } + const cached = makeHistoryItem({ id: "mismatch-settlement-task", pendingAction }) + const other = makeHistoryItem({ id: "other-task" }) + const filePath = path.join(tmpDir, "tasks", cached.id, GlobalFileNames.historyItem) + await fs.mkdir(path.dirname(filePath), { recursive: true }) + const mismatched = JSON.stringify(other, null, "\t") + await fs.writeFile(filePath, mismatched, "utf8") + const storeState = storeA as unknown as { + cache: Map + taskFileMtimes: Map + } + storeState.cache.set(cached.id, cached) + storeState.taskFileMtimes.set(cached.id, Date.now()) + + await expect(storeA.clearPendingActionIfMatching(cached.id, pendingAction.actionId)).rejects.toThrow( + `task ${cached.id} has a disk record with mismatched id other-task`, + ) + expect(storeA.get(cached.id)).toBeUndefined() + expect(storeState.taskFileMtimes.has(cached.id)).toBe(false) + // The other task's record stays on disk untouched. + expect(await fs.readFile(filePath, "utf8")).toBe(mismatched) + }) + /** * Host B completes a task on disk while host A's cache still has it * active. Host A's next save updates only totalCost (a full-object diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts index 7290b08757..d35dfc49cf 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.realConcurrency.spec.ts @@ -6,6 +6,13 @@ import type { HistoryItem } from "@roo-code/types" import { TaskHistoryStore } from "../TaskHistoryStore" +// Wrap the real fs/promises so a test can pause inside the per-file lock +// while every call still runs against the real filesystem. +vi.mock("fs/promises", async () => { + const actual = await vi.importActual("fs/promises") + return { ...actual, readFile: vi.fn(actual.readFile) } +}) + type WriteTaskFile = (item: HistoryItem, delta?: Partial) => Promise interface WriteBarrier { @@ -312,6 +319,72 @@ describe("TaskHistoryStore real cross-host locking", () => { } }) + it("serializes deletion after settlement reads disk without recreating the record", async () => { + const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-during-settlement-")) + const storeA = new TaskHistoryStore(storagePath) + const storeB = new TaskHistoryStore(storagePath) + const action = createAction("action-a", "action A") + const filePath = path.join(storagePath, "tasks", "shared-task", "history_item.json") + const actualFs = await vi.importActual("fs/promises") + let signalReadComplete!: () => void + const readComplete = new Promise((resolve) => { + signalReadComplete = resolve + }) + let releaseSettlement!: () => void + const settlementCanContinue = new Promise((resolve) => { + releaseSettlement = resolve + }) + + try { + await storeA.initialize() + await storeA.upsert({ ...item("shared-task"), pendingAction: action }) + await storeB.initialize() + + // Pause settlement inside its locked disk read, then start a + // deletion from another store so it targets the settlement's + // read-to-commit window. + vi.mocked(fs.readFile).mockImplementation(async (...args: Parameters) => { + const result = await actualFs.readFile(...args) + if (args[0] === filePath) { + signalReadComplete() + await settlementCanContinue + } + return result + }) + + const settlement = storeA.clearPendingActionIfMatching("shared-task", action.actionId) + await readComplete + const deletion = storeB.delete("shared-task") + let deletionSettled = false + void deletion.finally(() => { + deletionSettled = true + }) + await new Promise((resolve) => setTimeout(resolve, 25)) + // The deletion must stay blocked while settlement holds the + // per-file lock across its read-to-commit window. + expect(deletionSettled).toBe(false) + + releaseSettlement() + await expect(settlement).resolves.toMatchObject({ id: "shared-task", pendingAction: undefined }) + await deletion + + // The settlement committed first and the locked deletion + // removed the file afterwards, so the record stays deleted + // instead of being resurrected. + await expect(fs.access(filePath)).rejects.toMatchObject({ code: "ENOENT" }) + await expect(storeA.clearPendingActionIfMatching("shared-task", action.actionId)).rejects.toThrow( + "task shared-task not found", + ) + expect(storeA.get("shared-task")).toBeUndefined() + expect(storeB.get("shared-task")).toBeUndefined() + } finally { + vi.mocked(fs.readFile).mockImplementation(actualFs.readFile) + storeA.dispose() + storeB.dispose() + await fs.rm(storagePath, { recursive: true, force: true }) + } + }) + it("rejects settlement for a task absent from the cache without creating it", async () => { const storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-cache-miss-settlement-")) const store = new TaskHistoryStore(storagePath) diff --git a/src/utils/fileLock.ts b/src/utils/fileLock.ts new file mode 100644 index 0000000000..9f7cad7653 --- /dev/null +++ b/src/utils/fileLock.ts @@ -0,0 +1,78 @@ +import * as path from "path" +import * as lockfile from "proper-lockfile" + +/** + * Shared staleness window for per-file advisory locks. This module owns the + * single advisory lock protocol used by `safeWriteJson` and by callers that + * must serialize with it, such as task-history deletion. + */ +export const LOCK_STALE_MS = 31_000 + +/** + * Acquire the advisory lock for one file path using the exact protocol + * `safeWriteJson` uses, so operations that hold this lock serialize with + * every `safeWriteJson` write to the same path. Callers must release the + * returned function exactly once and must not acquire the same lock again + * while holding it. + */ +export async function acquireFileLock(filePath: string): Promise<() => Promise> { + const absoluteFilePath = path.resolve(filePath) + try { + return await lockfile.lock(absoluteFilePath, { + stale: LOCK_STALE_MS, + update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long + realpath: false, // the file may not exist yet, which is acceptable + retries: { + // Configuration for retrying lock acquisition + retries: 5, // Number of retries after the initial attempt + factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) + minTimeout: 100, // Minimum time to wait before the first retry (in ms) + maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) + }, + onCompromised: (err) => { + console.error(`Lock at ${absoluteFilePath} was compromised:`, err) + throw err + }, + }) + } catch (lockError) { + console.error(`Failed to acquire lock for ${absoluteFilePath}:`, lockError) + throw lockError + } +} + +/** + * Run one operation while holding the advisory lock for a file path. + * Callers must not acquire this lock again from inside `operation`, and + * they must keep the documented lock order when combining this helper + * with other locks to prevent deadlock. + */ +export async function withFileLock( + filePath: string, + operation: (absoluteFilePath: string) => Promise, +): Promise { + const absoluteFilePath = path.resolve(filePath) + const releaseLock = await acquireFileLock(absoluteFilePath) + + let result: T + try { + result = await operation(absoluteFilePath) + } catch (operationError) { + // The operation error is the primary failure. Release without + // reporting a secondary release error over it. + try { + await releaseLock() + } catch (releaseError) { + console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) + } + throw operationError + } + + try { + await releaseLock() + } catch (releaseError) { + // The operation already succeeded, so a release failure is only + // logged, matching how `safeWriteJson` handles release failures. + console.error(`Failed to release lock for ${absoluteFilePath}:`, releaseError) + } + return result +} diff --git a/src/utils/safeWriteJson.ts b/src/utils/safeWriteJson.ts index 957a0bb20f..7da68b2a7a 100644 --- a/src/utils/safeWriteJson.ts +++ b/src/utils/safeWriteJson.ts @@ -1,9 +1,10 @@ import * as fs from "fs/promises" import * as fsSync from "fs" import * as path from "path" -import * as lockfile from "proper-lockfile" import { JsonStreamStringify } from "json-stream-stringify" +import { acquireFileLock } from "./fileLock" + /** * Options for safeWriteJson function */ @@ -61,32 +62,14 @@ async function safeWriteJson(filePath: string, data: any, options?: SafeWriteJso throw dirError } - // Acquire the lock before any file operations - try { - releaseLock = await lockfile.lock(absoluteFilePath, { - stale: LOCK_STALE_MS, - update: 10000, // Update mtime every 10 seconds to prevent staleness if operation is long - realpath: false, // the file may not exist yet, which is acceptable - retries: { - // Configuration for retrying lock acquisition - retries: 5, // Number of retries after the initial attempt - factor: 2, // Exponential backoff factor (e.g., 100ms, 200ms, 400ms, ...) - minTimeout: 100, // Minimum time to wait before the first retry (in ms) - maxTimeout: 1000, // Maximum time to wait for any single retry (in ms) - }, - onCompromised: (err) => { - console.error(`Lock at ${absoluteFilePath} was compromised:`, err) - throw err - }, - }) - } catch (lockError) { - // If lock acquisition fails, we throw immediately. - // The releaseLock remains a no-op, so the finally block in the main file operations - // try-catch-finally won't try to release an unacquired lock if this path is taken. - console.error(`Failed to acquire lock for ${absoluteFilePath}:`, lockError) - // Propagate the lock acquisition error - throw lockError - } + // Acquire the lock before any file operations. `acquireFileLock` owns the + // shared advisory lock protocol, so callers that lock the same path with + // it (for example task-history deletion) serialize with this write. + // If lock acquisition fails, it throws immediately. The releaseLock + // remains a no-op, so the finally block in the main file operations + // try-catch-finally won't try to release an unacquired lock if this + // path is taken. + releaseLock = await acquireFileLock(absoluteFilePath) // Variables to hold the actual paths of temp files if they are created. let actualTempNewFilePath: string | null = null @@ -247,6 +230,4 @@ async function _streamDataToFile(targetPath: string, data: any, prettyPrint = fa }) } -export const LOCK_STALE_MS = 31_000 - export { safeWriteJson } From edcf937e2b55f64e2fe0fa956bcf0b10cd34cf1b Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Tue, 29 Sep 2026 01:18:27 +0000 Subject: [PATCH 21/23] fix: fail task-history deletion closed and verify the disk outcome delete and deleteMany evicted the cache before the locked unlink and swallowed all lock and unlink errors. onWrite could persist a removal while the history file stayed on disk. Callers then removed the task directory. - delete removes the file under the per-file lock first. Only a locked unlink success or a confirmed ENOENT evicts the cache and runs the write-through. Lock and unlink errors propagate. - deleteMany stops at the first failed item. Failed and unattempted items keep their cache entries. Completed deletions get a write-through before the rejection. The caller receives TaskHistoryPartialDeleteError with the original deletion error as cause. A failure-path write-through error is logged only. - deleteTaskWithId aborts before checkpoint and task directory removal when history deletion fails. Tests cover lock acquisition failure, non-ENOENT unlink failure, confirmed ENOENT, partial deleteMany cache and write-through state, and caller fail-closed directory handling. --- src/core/task-persistence/TaskHistoryStore.ts | 148 ++++++++--- .../TaskHistoryStore.deleteSemantics.spec.ts | 245 ++++++++++++++++++ src/core/webview/ClineProvider.ts | 5 +- .../ClineProvider.taskHistory.spec.ts | 24 ++ 4 files changed, 383 insertions(+), 39 deletions(-) create mode 100644 src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index 9c70d01a0d..a96ecb409a 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -16,6 +16,33 @@ import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./ta export { assertValidTransition, type HistoryItemStatus } from "./taskLifecycle" export { DeltaRejectedError } from "./taskStoreConcurrency" +/** True when an fs error reports that the target path does not exist. */ +function isEnoentError(error: unknown): boolean { + return typeof error === "object" && error !== null && (error as NodeJS.ErrnoException).code === "ENOENT" +} + +/** + * Reported by `deleteMany` when one or more history files could not be + * deleted. Items deleted before the failure stay deleted; the failed item + * and every unattempted item keep their cache entries. `cause` holds the + * first original deletion error. + */ +export class TaskHistoryPartialDeleteError extends Error { + public readonly failures: ReadonlyArray<{ readonly taskId: string; readonly reason: unknown }> + + constructor(failures: ReadonlyArray<{ taskId: string; reason: unknown }>) { + const firstReason = failures[0]?.reason + const firstMessage = + firstReason instanceof Error ? firstReason.message : String(firstReason ?? "unknown deletion error") + super( + `Failed to delete ${failures.length} task history item${failures.length === 1 ? "" : "s"}: ${firstMessage}`, + { cause: firstReason }, + ) + this.name = "TaskHistoryPartialDeleteError" + this.failures = failures + } +} + /** * Build a `safeWriteJson` merge callback that applies only `delta` to the * current disk state, preserving fields written by another process. @@ -264,33 +291,20 @@ export class TaskHistoryStore { /** * Delete a single task's history item. + * + * The deletion fails closed: a lock acquisition failure or a non-ENOENT + * unlink failure propagates to the caller and keeps the cache entry, so + * the history item stays readable. Only a verified deletion (a locked + * unlink success or a confirmed ENOENT) evicts in-memory state and runs + * the `onWrite` write-through. */ async delete(taskId: string): Promise { return this.withLock(async () => { + await this.removeHistoryFile(taskId) + this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) - // Remove per-task file (best-effort). The unlink runs under the - // same per-file advisory lock that `safeWriteJson` holds, so a - // deletion cannot interleave with another host's locked - // read-modify-write (for example the settlement in - // `clearPendingActionIfMatching`). Lock ordering stays store - // lock first, then one per-file lock, matching the write path. - try { - const filePath = await this.getTaskFilePath(taskId) - await withFileLock(filePath, async (absoluteFilePath) => { - try { - await fs.unlink(absoluteFilePath) - } catch { - // File may already be deleted - } - }) - } catch { - // A missing task directory also proves the history file is - // absent, so remaining failures keep the base best-effort - // deletion semantics. - } - // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { await this.onWrite(this.getAll()) @@ -300,38 +314,96 @@ export class TaskHistoryStore { /** * Delete multiple tasks' history items in a batch. + * + * The batch stops at the first item whose deletion fails. Items deleted + * before the failure stay deleted and are written through; the failed + * item and every unattempted item keep their cache entries. The original + * deletion error is reported through `TaskHistoryPartialDeleteError`. A + * write-through failure on the partial-failure path is logged and never + * replaces the original deletion error; a write-through failure on a + * fully successful batch propagates. */ async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { + const failures: Array<{ taskId: string; reason: unknown }> = [] + let deletedCount = 0 + for (const taskId of taskIds) { + try { + await this.removeHistoryFile(taskId) + } catch (reason) { + failures.push({ taskId, reason }) + break + } + this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) + deletedCount++ + } - // Serialize each unlink with `safeWriteJson` under the same - // per-file advisory lock. See `delete` for the lock order. + // Write through the deletions that completed, so persisted state + // matches the cache even when the batch fails partway. + if (deletedCount > 0 && this.onWrite) { try { - const filePath = await this.getTaskFilePath(taskId) - await withFileLock(filePath, async (absoluteFilePath) => { - try { - await fs.unlink(absoluteFilePath) - } catch { - // File may already be deleted - } - }) - } catch { - // A missing task directory also proves the history file - // is absent, so remaining failures keep the base - // best-effort deletion semantics. + await this.onWrite(this.getAll()) + } catch (writeError) { + if (failures.length === 0) { + throw writeError + } + // The original deletion failure stays the reported error. + console.error("[TaskHistoryStore] Write-through after partial deleteMany failed:", writeError) } } - // Call onWrite callback inside the lock for serialized write-through - if (this.onWrite) { - await this.onWrite(this.getAll()) + if (failures.length > 0) { + throw new TaskHistoryPartialDeleteError(failures) } }) } + /** + * Remove one history file under the same per-file advisory lock that + * `safeWriteJson` holds, so a deletion cannot interleave with another + * host's locked read-modify-write (for example the settlement in + * `clearPendingActionIfMatching`). Lock ordering stays store lock first, + * then one per-file lock, matching the write path. + * + * A locked unlink success or a confirmed ENOENT completes the deletion. + * Any other failure leaves the disk outcome unknown, so the file itself + * decides: an absent file is a completed deletion, a present file fails + * the deletion and its error propagates. + */ + private async removeHistoryFile(taskId: string): Promise { + const filePath = await this.getTaskFilePath(taskId) + try { + await withFileLock(filePath, async (absoluteFilePath) => { + try { + await fs.unlink(absoluteFilePath) + } catch (error) { + if (!isEnoentError(error)) { + throw error + } + } + }) + } catch (error) { + // Verify the disk state before deciding the semantics. A missing + // task directory also proves the history file is absent. + if (!(await this.isFileAbsent(filePath))) { + throw error + } + } + } + + /** True when the path does not exist. Other stat errors stay errors. */ + private async isFileAbsent(filePath: string): Promise { + try { + await fs.stat(filePath) + return false + } catch (error) { + return isEnoentError(error) + } + } + // ────────────────────────────── Reconciliation ────────────────────────────── /** diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts new file mode 100644 index 0000000000..9be7625c42 --- /dev/null +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts @@ -0,0 +1,245 @@ +// pnpm --filter roo-cline test core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts +// +// Focused fail-closed deletion semantics for `delete()` and `deleteMany()`. +// Lock, unlink, and fs behavior stay real by default; individual tests force +// one lock or unlink failure through the wrappers below. + +import * as fs from "fs/promises" +import * as os from "os" +import * as path from "path" + +import type { HistoryItem } from "@roo-code/types" + +import { TaskHistoryStore, TaskHistoryPartialDeleteError } from "../TaskHistoryStore" +import { withFileLock } from "../../../utils/fileLock" +import { GlobalFileNames } from "../../../shared/globalFileNames" + +vi.mock("../../../utils/storage", () => ({ + getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => defaultPath), +})) + +// The default implementation stays the real one, so `safeWriteJson` writes +// and per-file locking behave exactly as in production unless a test forces +// a failure. +vi.mock("../../../utils/fileLock", async () => { + const actual = await vi.importActual("../../../utils/fileLock") + return { ...actual, withFileLock: vi.fn(actual.withFileLock) } +}) + +vi.mock("fs/promises", async () => { + const actual = await vi.importActual("fs/promises") + return { ...actual, unlink: vi.fn(actual.unlink) } +}) + +const actualFs = await vi.importActual("fs/promises") +const actualFileLock = await vi.importActual("../../../utils/fileLock") + +function makeHistoryItem(overrides: Partial = {}): HistoryItem { + return { + id: `task-${Date.now()}-${Math.random().toString(36).substring(2, 8)}`, + number: 1, + ts: Date.now(), + task: "Test task", + tokensIn: 100, + tokensOut: 50, + totalCost: 0.01, + workspace: "/test/workspace", + ...overrides, + } +} + +function historyFilePath(storagePath: string, taskId: string): string { + return path.join(storagePath, "tasks", taskId, GlobalFileNames.historyItem) +} + +function epermLike(message: string): NodeJS.ErrnoException { + return Object.assign(new Error(message), { code: "EACCES" }) +} + +describe("TaskHistoryStore fail-closed deletion semantics", () => { + let storagePath: string + let stores: TaskHistoryStore[] + let onWrite: ReturnType + + beforeEach(async () => { + storagePath = await fs.mkdtemp(path.join(os.tmpdir(), "task-history-delete-semantics-")) + stores = [] + onWrite = vi.fn().mockResolvedValue(undefined) + vi.mocked(withFileLock).mockImplementation(actualFileLock.withFileLock) + vi.mocked(fs.unlink).mockImplementation(actualFs.unlink) + }) + + afterEach(async () => { + for (const store of stores) { + store.dispose() + } + await fs.rm(storagePath, { recursive: true, force: true }).catch(() => {}) + }) + + function createStore(): TaskHistoryStore { + const store = new TaskHistoryStore(storagePath, { + onWrite: onWrite as (items: HistoryItem[]) => Promise, + }) + stores.push(store) + return store + } + + describe("delete()", () => { + it("propagates a lock acquisition failure, keeps the cache entry and file, and skips write-through", async () => { + const store = createStore() + await store.initialize() + await store.upsert(makeHistoryItem({ id: "lock-fail" })) + onWrite.mockClear() + + const lockError = new Error("lock acquisition timed out") + vi.mocked(withFileLock).mockRejectedValueOnce(lockError) + + await expect(store.delete("lock-fail")).rejects.toBe(lockError) + + // The file stayed on disk, the cache entry survived, and no + // removal was persisted through onWrite. + await expect(fs.access(historyFilePath(storagePath, "lock-fail"))).resolves.toBeUndefined() + expect(store.get("lock-fail")).toBeDefined() + expect(onWrite).not.toHaveBeenCalled() + }) + + it("propagates a non-ENOENT unlink failure, keeps the cache entry and file, and stays deletable afterwards", async () => { + const store = createStore() + await store.initialize() + await store.upsert(makeHistoryItem({ id: "perm-fail" })) + onWrite.mockClear() + + const unlinkError = epermLike("EACCES: permission denied, unlink") + vi.mocked(fs.unlink).mockRejectedValueOnce(unlinkError) + + await expect(store.delete("perm-fail")).rejects.toBe(unlinkError) + + await expect(fs.access(historyFilePath(storagePath, "perm-fail"))).resolves.toBeUndefined() + expect(store.get("perm-fail")).toBeDefined() + expect(onWrite).not.toHaveBeenCalled() + + // The per-file lock was released: a retry without the injected + // failure deletes the file and writes through. + await expect(store.delete("perm-fail")).resolves.toBeUndefined() + await expect(fs.access(historyFilePath(storagePath, "perm-fail"))).rejects.toMatchObject({ + code: "ENOENT", + }) + expect(store.get("perm-fail")).toBeUndefined() + expect(onWrite).toHaveBeenCalledTimes(1) + }) + + it("treats a confirmed ENOENT as a successful deletion", async () => { + const store = createStore() + await store.initialize() + + // A task that never existed resolves without throwing. + await expect(store.delete("never-existed")).resolves.toBeUndefined() + + // A task whose file a peer already removed also deletes cleanly. + await store.upsert(makeHistoryItem({ id: "peer-removed" })) + await fs.unlink(historyFilePath(storagePath, "peer-removed")) + await expect(store.delete("peer-removed")).resolves.toBeUndefined() + expect(store.get("peer-removed")).toBeUndefined() + expect(onWrite).toHaveBeenCalled() + }) + }) + + describe("deleteMany()", () => { + it("stops at the first failure, preserves failed and unattempted cache entries, and writes through completed deletions before rejecting", async () => { + const store = createStore() + await store.initialize() + await store.upsert(makeHistoryItem({ id: "batch-a", ts: 1000 })) + await store.upsert(makeHistoryItem({ id: "batch-b", ts: 2000 })) + await store.upsert(makeHistoryItem({ id: "batch-c", ts: 3000 })) + onWrite.mockClear() + + const unlinkError = epermLike("EACCES: permission denied, unlink") + const events: string[] = [] + onWrite.mockImplementation(async () => { + events.push("write-through") + }) + + vi.mocked(fs.unlink).mockImplementation(async (p) => { + if (p === historyFilePath(storagePath, "batch-b")) { + throw unlinkError + } + return actualFs.unlink(p) + }) + + const pending = store.deleteMany(["batch-a", "batch-b", "batch-c"]).catch((error) => { + events.push("rejected") + return error + }) + const partialError = (await pending) as TaskHistoryPartialDeleteError + + // The write-through of completed deletions is awaited before the + // rejection reaches the caller. + expect(events).toEqual(["write-through", "rejected"]) + + expect(partialError).toBeInstanceOf(TaskHistoryPartialDeleteError) + expect(partialError.failures).toEqual([{ taskId: "batch-b", reason: unlinkError }]) + expect(partialError.cause).toBe(unlinkError) + + // batch-a completed: file gone, cache evicted. + await expect(fs.access(historyFilePath(storagePath, "batch-a"))).rejects.toMatchObject({ code: "ENOENT" }) + expect(store.get("batch-a")).toBeUndefined() + + // batch-b failed: file present, cache entry preserved. + await expect(fs.access(historyFilePath(storagePath, "batch-b"))).resolves.toBeUndefined() + expect(store.get("batch-b")).toBeDefined() + + // batch-c was never attempted: file present, cache entry preserved. + await expect(fs.access(historyFilePath(storagePath, "batch-c"))).resolves.toBeUndefined() + expect(store.get("batch-c")).toBeDefined() + + // The write-through saw the cache after completed deletions only. + const writtenIds = (onWrite.mock.calls[0][0] as HistoryItem[]).map((item) => item.id).sort() + expect(writtenIds).toEqual(["batch-b", "batch-c"]) + }) + + it("preserves the original deletion error when the failure-path write-through also fails", async () => { + const store = createStore() + await store.initialize() + await store.upsert(makeHistoryItem({ id: "batch-a", ts: 1000 })) + await store.upsert(makeHistoryItem({ id: "batch-b", ts: 2000 })) + onWrite.mockClear() + + const unlinkError = epermLike("EACCES: permission denied, unlink") + const consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) + + vi.mocked(fs.unlink).mockImplementation(async (p) => { + if (p === historyFilePath(storagePath, "batch-b")) { + throw unlinkError + } + return actualFs.unlink(p) + }) + onWrite.mockRejectedValue(new Error("write-through boom")) + + try { + const partialError = (await store + .deleteMany(["batch-a", "batch-b"]) + .catch((error) => error)) as TaskHistoryPartialDeleteError + + // The original deletion failure stays the reported error. + expect(partialError).toBeInstanceOf(TaskHistoryPartialDeleteError) + expect(partialError.failures).toEqual([{ taskId: "batch-b", reason: unlinkError }]) + expect(partialError.cause).toBe(unlinkError) + + // The secondary write-through failure is logged, not thrown. + expect(consoleError).toHaveBeenCalledWith( + "[TaskHistoryStore] Write-through after partial deleteMany failed:", + expect.any(Error), + ) + + // batch-a stays deleted and batch-b stays intact. + await expect(fs.access(historyFilePath(storagePath, "batch-a"))).rejects.toMatchObject({ + code: "ENOENT", + }) + await expect(fs.access(historyFilePath(storagePath, "batch-b"))).resolves.toBeUndefined() + expect(store.get("batch-b")).toBeDefined() + } finally { + consoleError.mockRestore() + } + }) + }) +}) diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index 6325435d54..037e6f56ae 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -2412,7 +2412,10 @@ export class ClineProvider } } - // Delete all tasks from state in one batch + // Delete all tasks from state in one batch. A failed history + // deletion fails closed: the error leaves this method before any + // checkpoint or task directory removal below, so a directory is + // never removed while its history file could still exist. await this.taskHistoryStore.deleteMany(allIdsToDelete) this.recentTasksCache = undefined diff --git a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts index 2bbf0736c6..797bdf948d 100644 --- a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts @@ -1,5 +1,6 @@ // pnpm --filter roo-cline test core/webview/__tests__/ClineProvider.taskHistory.spec.ts +import * as fs from "fs/promises" import * as vscode from "vscode" import type { HistoryItem, ExtensionMessage } from "@roo-code/types" import { providerIdentifiers, RooCodeEventName } from "@roo-code/types" @@ -25,6 +26,9 @@ vi.mock("fs/promises", () => ({ rmdir: vi.fn().mockResolvedValue(undefined), access: vi.fn().mockResolvedValue(undefined), rm: vi.fn().mockResolvedValue(undefined), + // Deletion verifies disk state after a failed lock or unlink, so the + // harness reports every probed file as absent. + stat: vi.fn().mockRejectedValue(Object.assign(new Error("no such file or directory"), { code: "ENOENT" })), })) vi.mock("axios", () => ({ @@ -816,6 +820,26 @@ describe("ClineProvider Task History Synchronization", () => { expect(ids).not.toContain("remove-me") }) + it("fails closed: does not remove task directories when history deletion fails", async () => { + await provider.resolveWebviewView(mockWebviewView) + + const item = createHistoryItem({ id: "dir-keep", task: "Keep directory" }) + await provider.updateTaskHistory(item, { broadcast: false }) + + const failure = new Error("lock acquisition failed") + const deleteManySpy = vi.spyOn(provider.taskHistoryStore, "deleteMany").mockRejectedValue(failure) + + try { + // The failed history deletion aborts deleteTaskWithId before + // any checkpoint or task directory removal. + await expect(provider.deleteTaskWithId("dir-keep")).rejects.toBe(failure) + } finally { + deleteManySpy.mockRestore() + } + + expect(fs.rm).not.toHaveBeenCalled() + }) + it("does not block subsequent writes when a previous store write errors", async () => { await provider.resolveWebviewView(mockWebviewView) From c326dad1cadb99bb1610fb6ba84c574b09426f36 Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Tue, 29 Sep 2026 02:02:16 +0000 Subject: [PATCH 22/23] test: provide getStorageBasePath in the ClineProvider.spec storage mock TaskHistoryStore deletion now resolves the tasks directory before it removes a history file, so the storage mock must supply the getStorageBasePath passthrough. Without it, a deletion inside the file-backed history tests throws from the mock instead of exercising the real path. --- src/core/webview/__tests__/ClineProvider.spec.ts | 3 +++ 1 file changed, 3 insertions(+) diff --git a/src/core/webview/__tests__/ClineProvider.spec.ts b/src/core/webview/__tests__/ClineProvider.spec.ts index feef95d870..8ba9144f1c 100644 --- a/src/core/webview/__tests__/ClineProvider.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.spec.ts @@ -87,6 +87,9 @@ vi.mock("../../../utils/storage", () => ({ getSettingsDirectoryPath: vi.fn().mockResolvedValue("/test/settings/path"), getTaskDirectoryPath: vi.fn().mockResolvedValue("/test/task/path"), getGlobalStoragePath: vi.fn().mockResolvedValue("/test/storage/path"), + // Deletion resolves the tasks directory before it removes a history + // file, so the harness must provide the passthrough base path. + getStorageBasePath: vi.fn().mockImplementation((defaultPath: string) => defaultPath), })) vi.mock("@modelcontextprotocol/sdk/types.js", () => ({ From 7c2528c51bccf9b008ee210e07e7cd24e697a03d Mon Sep 17 00:00:00 2001 From: Elliott de Launay Date: Wed, 30 Sep 2026 01:04:03 +0000 Subject: [PATCH 23/23] fix: refresh task history before restart settlement and restore best-effort deletion settleInterruptedCreateSubtaskBeforeReplay forces a TaskHistoryStore reconcile, settles the refreshed create_subtask action through the compare-and-clear, and adopts the authoritative pending action. Refresh, lookup, and settlement failures convert to PendingActionSettlementError and stop replay. TaskHistoryStore delete and deleteMany restore base ordering and semantics: evict cache and mtime first, best-effort locked unlink, swallow lock and unlink failures, continue every batch item, and write through once after processing. The per-file lock stays, so deletion still serializes with settlement. Remove TaskHistoryPartialDeleteError, outcome verification helpers, first-failure batch behavior, partial write-through, and the fail-closed deleteTaskWithId comment. --- docs/architecture/task-lifecycle-model.md | 2 +- .../ClineProvider.delegation.spec.ts | 96 +++++++ src/core/task-persistence/TaskHistoryStore.ts | 133 ++------- .../TaskHistoryStore.deleteSemantics.spec.ts | 204 +++++++------ src/core/task/Task.ts | 41 ++- .../task/__tests__/Task.persistence.spec.ts | 271 +++++++++++++++++- src/core/webview/ClineProvider.ts | 5 +- .../ClineProvider.taskHistory.spec.ts | 24 -- 8 files changed, 518 insertions(+), 258 deletions(-) diff --git a/docs/architecture/task-lifecycle-model.md b/docs/architecture/task-lifecycle-model.md index ed186961f7..744ac8eb81 100644 --- a/docs/architecture/task-lifecycle-model.md +++ b/docs/architecture/task-lifecycle-model.md @@ -83,7 +83,7 @@ CI fails if either exact causal witness or violation class changes, a witness di The known-unsafe witnesses currently compare exact shortest action sequences. This is intentionally simple and reviewable, but brittle to harmless action renames or serialization refactors. A causal partial-order comparator would reduce that brittleness but would add a second trace-equivalence protocol to maintain. Until that complexity is justified, update an exact witness only after confirming the terminal violation class and required causal ordering are unchanged. -`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear. History-file deletion serializes on the same per-file advisory lock as `safeWriteJson`, and the real-filesystem suite covers a deletion that targets the settlement read-to-commit window plus schema and task-ID validation of the locked disk record before settlement (#1726). Restart recovery is covered by `Task.persistence.spec.ts`: it treats an interrupted task's pending `create_subtask` action as rejected, settles that exact action before replay, and a failed recovery write stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. +`TaskHistoryStore.realConcurrency.spec.ts` complements the abstract interleavings with real-filesystem checks through the real `proper-lockfile` and filesystem rename path, including stale-settlement compare-and-clear. History-file deletion holds the same per-file advisory lock as `safeWriteJson` around each unlink, and the real-filesystem suite covers a deletion that targets the settlement read-to-commit window plus schema and task-ID validation of the locked disk record before settlement (#1726). That lock serialization is best-effort: a lock acquisition or unlink failure is swallowed, so a deletion can proceed without the lock and the two sides serialize only when both acquire it. Restart recovery is covered by `Task.persistence.spec.ts`: it refreshes the persisted record first, treats a refreshed interrupted `create_subtask` action as rejected, settles that exact action before replay, adopts the authoritative pending action, and a failed refresh, lookup, or settlement stops replay rather than creating another child. Broader VS Code E2E remains reserved for other restart and extension-host behavior. ## Task cleanup protocol model diff --git a/src/__tests__/ClineProvider.delegation.spec.ts b/src/__tests__/ClineProvider.delegation.spec.ts index e337fcc2e3..68f39c6517 100644 --- a/src/__tests__/ClineProvider.delegation.spec.ts +++ b/src/__tests__/ClineProvider.delegation.spec.ts @@ -1,5 +1,9 @@ // npx vitest run __tests__/provider-delegation.spec.ts +import * as fs from "fs/promises" +import * as os from "os" +import * as path from "path" + import { describe, it, expect, vi } from "vitest" import type { HistoryItem } from "@roo-code/types" import { providerIdentifiers, RooCodeEventName } from "@roo-code/types" @@ -1151,4 +1155,96 @@ describe("ClineProvider.delegateParentAndOpenChild()", () => { expect(provider.deleteTaskWithId).toHaveBeenCalledWith("child-1", false) expect(createTaskWithHistoryItem).toHaveBeenCalledWith(replacedParent) }) + + it("keeps directory cleanup and parent restoration when the child history lock failure is swallowed", async () => { + const pendingAction = { + kind: "create_subtask" as const, + actionId: "create-action", + approvalText: "{}", + mode: "code", + message: "Do something", + todos: [], + } + const interruptedParent: HistoryItem = { + ...parentHistoryItem, + status: "interrupted", + pendingAction, + } + const settledParent: HistoryItem = { ...interruptedParent, pendingAction: undefined } + const childItem: HistoryItem = { ...parentHistoryItem, id: "child-1", task: "Child" } + + const globalStorageDir = await fs.mkdtemp(path.join(os.tmpdir(), "delegation-rollback-cleanup-")) + const childDir = path.join(globalStorageDir, "tasks", "child-1") + await fs.mkdir(childDir, { recursive: true }) + await fs.writeFile(path.join(childDir, "ui_messages.json"), "[]") + + const parentTask = makeParentTask() + const child = { taskId: "child-1", run: vi.fn().mockResolvedValue(undefined) } + const getCurrentTask = vi.fn().mockReturnValue(parentTask) + const createTask = vi.fn(async () => { + getCurrentTask.mockReturnValue(child) + return child + }) + const createTaskWithHistoryItem = vi.fn().mockResolvedValue(undefined) + const getTaskWithId = vi.fn(async (taskId: string) => { + if (taskId === "parent-1") { + return { historyItem: settledParent } + } + return { taskDirPath: childDir, historyItem: childItem } + }) + const clearPendingActionIfMatching = vi.fn(async () => settledParent) + + const provider = { + taskScheduler: new TaskScheduler(), + recentTasksCache: [parentHistoryItem], + emit: vi.fn(), + getCurrentTask, + removeClineFromStack: vi.fn().mockResolvedValue(undefined), + createTask, + getTaskWithId, + handleModeSwitch: vi.fn().mockResolvedValue(undefined), + // The real deletion path: the store swallows a per-file history + // lock failure, so deleteMany resolves and cleanup continues. + deleteTaskWithId: ClineProvider.prototype.deleteTaskWithId, + createTaskWithHistoryItem, + log: vi.fn(), + postStateToWebview: vi.fn().mockResolvedValue(undefined), + isViewLaunched: false, + contextProxy: { globalStorageUri: { fsPath: globalStorageDir } }, + cwd: globalStorageDir, + taskHistoryStore: { + invalidate: vi.fn().mockResolvedValue(undefined), + get: vi.fn(() => interruptedParent), + atomicReadAndUpdate: vi.fn(async (_taskId: string, updater: (item: HistoryItem) => HistoryItem) => { + updater(interruptedParent) + return [] + }), + clearPendingActionIfMatching, + deleteMany: vi.fn().mockResolvedValue(undefined), + }, + } as unknown as ClineProvider + + try { + await expect( + ClineProvider.prototype.delegateParentAndOpenChild.call(provider, { + parentTaskId: "parent-1", + message: pendingAction.message, + initialTodos: pendingAction.todos, + mode: pendingAction.mode, + pendingActionId: pendingAction.actionId, + }), + ).rejects.toThrow("Invalid task status transition: interrupted → delegated") + + // The child task directory was removed even though the child's + // history lock failed and was swallowed, and the settled parent + // was restored before the original rejection surfaced. + await expect(fs.access(childDir)).rejects.toMatchObject({ code: "ENOENT" }) + expect(clearPendingActionIfMatching).toHaveBeenCalledWith("parent-1", "create-action") + expect(createTaskWithHistoryItem).toHaveBeenCalledWith( + expect.objectContaining({ status: "interrupted", pendingAction: undefined }), + ) + } finally { + await fs.rm(globalStorageDir, { recursive: true, force: true }).catch(() => {}) + } + }) }) diff --git a/src/core/task-persistence/TaskHistoryStore.ts b/src/core/task-persistence/TaskHistoryStore.ts index a96ecb409a..346d697d9a 100644 --- a/src/core/task-persistence/TaskHistoryStore.ts +++ b/src/core/task-persistence/TaskHistoryStore.ts @@ -16,33 +16,6 @@ import { computeHistoryDelta, DeltaRejectedError, mergeHistoryDelta } from "./ta export { assertValidTransition, type HistoryItemStatus } from "./taskLifecycle" export { DeltaRejectedError } from "./taskStoreConcurrency" -/** True when an fs error reports that the target path does not exist. */ -function isEnoentError(error: unknown): boolean { - return typeof error === "object" && error !== null && (error as NodeJS.ErrnoException).code === "ENOENT" -} - -/** - * Reported by `deleteMany` when one or more history files could not be - * deleted. Items deleted before the failure stay deleted; the failed item - * and every unattempted item keep their cache entries. `cause` holds the - * first original deletion error. - */ -export class TaskHistoryPartialDeleteError extends Error { - public readonly failures: ReadonlyArray<{ readonly taskId: string; readonly reason: unknown }> - - constructor(failures: ReadonlyArray<{ taskId: string; reason: unknown }>) { - const firstReason = failures[0]?.reason - const firstMessage = - firstReason instanceof Error ? firstReason.message : String(firstReason ?? "unknown deletion error") - super( - `Failed to delete ${failures.length} task history item${failures.length === 1 ? "" : "s"}: ${firstMessage}`, - { cause: firstReason }, - ) - this.name = "TaskHistoryPartialDeleteError" - this.failures = failures - } -} - /** * Build a `safeWriteJson` merge callback that applies only `delta` to the * current disk state, preserving fields written by another process. @@ -292,19 +265,26 @@ export class TaskHistoryStore { /** * Delete a single task's history item. * - * The deletion fails closed: a lock acquisition failure or a non-ENOENT - * unlink failure propagates to the caller and keeps the cache entry, so - * the history item stays readable. Only a verified deletion (a locked - * unlink success or a confirmed ENOENT) evicts in-memory state and runs - * the `onWrite` write-through. + * Deletion is best-effort: the unlink runs under the same per-file + * advisory lock as `safeWriteJson`, so a locked read-modify-write (for + * example the settlement in `clearPendingActionIfMatching`) cannot + * interleave with it. A lock or unlink failure is swallowed because the + * file may already be deleted; the in-memory eviction and the write + * through still complete. */ async delete(taskId: string): Promise { return this.withLock(async () => { - await this.removeHistoryFile(taskId) - this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) + // Remove per-task file (best-effort) + try { + const filePath = await this.getTaskFilePath(taskId) + await withFileLock(filePath, (absoluteFilePath) => fs.unlink(absoluteFilePath)) + } catch { + // File may already be deleted + } + // Call onWrite callback inside the lock for serialized write-through if (this.onWrite) { await this.onWrite(this.getAll()) @@ -315,95 +295,32 @@ export class TaskHistoryStore { /** * Delete multiple tasks' history items in a batch. * - * The batch stops at the first item whose deletion fails. Items deleted - * before the failure stay deleted and are written through; the failed - * item and every unattempted item keep their cache entries. The original - * deletion error is reported through `TaskHistoryPartialDeleteError`. A - * write-through failure on the partial-failure path is logged and never - * replaces the original deletion error; a write-through failure on a - * fully successful batch propagates. + * Every item follows the `delete` semantics and is attempted even when an + * earlier unlink fails. The single write-through runs once after the + * whole batch. */ async deleteMany(taskIds: string[]): Promise { return this.withLock(async () => { - const failures: Array<{ taskId: string; reason: unknown }> = [] - let deletedCount = 0 - for (const taskId of taskIds) { - try { - await this.removeHistoryFile(taskId) - } catch (reason) { - failures.push({ taskId, reason }) - break - } - this.cache.delete(taskId) this.taskFileMtimes.delete(taskId) - deletedCount++ - } - // Write through the deletions that completed, so persisted state - // matches the cache even when the batch fails partway. - if (deletedCount > 0 && this.onWrite) { + // Remove per-task file (best-effort) try { - await this.onWrite(this.getAll()) - } catch (writeError) { - if (failures.length === 0) { - throw writeError - } - // The original deletion failure stays the reported error. - console.error("[TaskHistoryStore] Write-through after partial deleteMany failed:", writeError) + const filePath = await this.getTaskFilePath(taskId) + await withFileLock(filePath, (absoluteFilePath) => fs.unlink(absoluteFilePath)) + } catch { + // File may already be deleted } } - if (failures.length > 0) { - throw new TaskHistoryPartialDeleteError(failures) + // Call onWrite callback inside the lock for serialized write-through + if (this.onWrite) { + await this.onWrite(this.getAll()) } }) } - /** - * Remove one history file under the same per-file advisory lock that - * `safeWriteJson` holds, so a deletion cannot interleave with another - * host's locked read-modify-write (for example the settlement in - * `clearPendingActionIfMatching`). Lock ordering stays store lock first, - * then one per-file lock, matching the write path. - * - * A locked unlink success or a confirmed ENOENT completes the deletion. - * Any other failure leaves the disk outcome unknown, so the file itself - * decides: an absent file is a completed deletion, a present file fails - * the deletion and its error propagates. - */ - private async removeHistoryFile(taskId: string): Promise { - const filePath = await this.getTaskFilePath(taskId) - try { - await withFileLock(filePath, async (absoluteFilePath) => { - try { - await fs.unlink(absoluteFilePath) - } catch (error) { - if (!isEnoentError(error)) { - throw error - } - } - }) - } catch (error) { - // Verify the disk state before deciding the semantics. A missing - // task directory also proves the history file is absent. - if (!(await this.isFileAbsent(filePath))) { - throw error - } - } - } - - /** True when the path does not exist. Other stat errors stay errors. */ - private async isFileAbsent(filePath: string): Promise { - try { - await fs.stat(filePath) - return false - } catch (error) { - return isEnoentError(error) - } - } - // ────────────────────────────── Reconciliation ────────────────────────────── /** diff --git a/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts b/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts index 9be7625c42..5071c9de73 100644 --- a/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts +++ b/src/core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts @@ -1,6 +1,6 @@ // pnpm --filter roo-cline test core/task-persistence/__tests__/TaskHistoryStore.deleteSemantics.spec.ts // -// Focused fail-closed deletion semantics for `delete()` and `deleteMany()`. +// Best-effort deletion semantics for `delete()` and `deleteMany()`. // Lock, unlink, and fs behavior stay real by default; individual tests force // one lock or unlink failure through the wrappers below. @@ -10,7 +10,7 @@ import * as path from "path" import type { HistoryItem } from "@roo-code/types" -import { TaskHistoryStore, TaskHistoryPartialDeleteError } from "../TaskHistoryStore" +import { TaskHistoryStore } from "../TaskHistoryStore" import { withFileLock } from "../../../utils/fileLock" import { GlobalFileNames } from "../../../shared/globalFileNames" @@ -52,11 +52,17 @@ function historyFilePath(storagePath: string, taskId: string): string { return path.join(storagePath, "tasks", taskId, GlobalFileNames.historyItem) } -function epermLike(message: string): NodeJS.ErrnoException { - return Object.assign(new Error(message), { code: "EACCES" }) +function storeInternals(store: TaskHistoryStore): { + cache: Map + taskFileMtimes: Map +} { + return { + cache: store["cache"], + taskFileMtimes: store["taskFileMtimes"], + } } -describe("TaskHistoryStore fail-closed deletion semantics", () => { +describe("TaskHistoryStore best-effort deletion semantics", () => { let storagePath: string let stores: TaskHistoryStore[] let onWrite: ReturnType @@ -85,67 +91,92 @@ describe("TaskHistoryStore fail-closed deletion semantics", () => { } describe("delete()", () => { - it("propagates a lock acquisition failure, keeps the cache entry and file, and skips write-through", async () => { + it("unlinks under the shared per-file lock, evicts cache and mtime, and writes through once", async () => { + const store = createStore() + await store.initialize() + await store.upsert(makeHistoryItem({ id: "locked-delete" })) + onWrite.mockClear() + + await expect(store.delete("locked-delete")).resolves.toBeUndefined() + + expect(vi.mocked(withFileLock)).toHaveBeenCalledWith( + historyFilePath(storagePath, "locked-delete"), + expect.any(Function), + ) + await expect(fs.access(historyFilePath(storagePath, "locked-delete"))).rejects.toMatchObject({ + code: "ENOENT", + }) + + const { cache, taskFileMtimes } = storeInternals(store) + expect(cache.has("locked-delete")).toBe(false) + expect(taskFileMtimes.has("locked-delete")).toBe(false) + expect(store.get("locked-delete")).toBeUndefined() + + expect(onWrite).toHaveBeenCalledTimes(1) + const writtenIds = (onWrite.mock.calls[0][0] as HistoryItem[]).map((item) => item.id) + expect(writtenIds).not.toContain("locked-delete") + }) + + it("swallows a lock acquisition failure, evicts cache and mtime, and still writes through once", async () => { const store = createStore() await store.initialize() await store.upsert(makeHistoryItem({ id: "lock-fail" })) onWrite.mockClear() - const lockError = new Error("lock acquisition timed out") - vi.mocked(withFileLock).mockRejectedValueOnce(lockError) + vi.mocked(withFileLock).mockRejectedValueOnce(new Error("lock acquisition timed out")) - await expect(store.delete("lock-fail")).rejects.toBe(lockError) + await expect(store.delete("lock-fail")).resolves.toBeUndefined() - // The file stayed on disk, the cache entry survived, and no - // removal was persisted through onWrite. + // The unlink never ran, but the in-memory eviction and the single + // write-through still completed. await expect(fs.access(historyFilePath(storagePath, "lock-fail"))).resolves.toBeUndefined() - expect(store.get("lock-fail")).toBeDefined() - expect(onWrite).not.toHaveBeenCalled() + const { cache, taskFileMtimes } = storeInternals(store) + expect(cache.has("lock-fail")).toBe(false) + expect(taskFileMtimes.has("lock-fail")).toBe(false) + expect(store.get("lock-fail")).toBeUndefined() + expect(onWrite).toHaveBeenCalledTimes(1) }) - it("propagates a non-ENOENT unlink failure, keeps the cache entry and file, and stays deletable afterwards", async () => { + it("swallows a non-ENOENT unlink failure, evicts cache and mtime, and stays deletable afterwards", async () => { const store = createStore() await store.initialize() await store.upsert(makeHistoryItem({ id: "perm-fail" })) onWrite.mockClear() - const unlinkError = epermLike("EACCES: permission denied, unlink") - vi.mocked(fs.unlink).mockRejectedValueOnce(unlinkError) + vi.mocked(fs.unlink).mockRejectedValueOnce( + Object.assign(new Error("EACCES: permission denied, unlink"), { code: "EACCES" }), + ) - await expect(store.delete("perm-fail")).rejects.toBe(unlinkError) + await expect(store.delete("perm-fail")).resolves.toBeUndefined() await expect(fs.access(historyFilePath(storagePath, "perm-fail"))).resolves.toBeUndefined() - expect(store.get("perm-fail")).toBeDefined() - expect(onWrite).not.toHaveBeenCalled() + const { cache, taskFileMtimes } = storeInternals(store) + expect(cache.has("perm-fail")).toBe(false) + expect(taskFileMtimes.has("perm-fail")).toBe(false) + expect(onWrite).toHaveBeenCalledTimes(1) // The per-file lock was released: a retry without the injected - // failure deletes the file and writes through. + // failure deletes the file and writes through again. await expect(store.delete("perm-fail")).resolves.toBeUndefined() await expect(fs.access(historyFilePath(storagePath, "perm-fail"))).rejects.toMatchObject({ code: "ENOENT", }) - expect(store.get("perm-fail")).toBeUndefined() - expect(onWrite).toHaveBeenCalledTimes(1) + expect(onWrite).toHaveBeenCalledTimes(2) }) - it("treats a confirmed ENOENT as a successful deletion", async () => { + it("treats a missing file as a completed deletion", async () => { const store = createStore() await store.initialize() + onWrite.mockClear() - // A task that never existed resolves without throwing. await expect(store.delete("never-existed")).resolves.toBeUndefined() - - // A task whose file a peer already removed also deletes cleanly. - await store.upsert(makeHistoryItem({ id: "peer-removed" })) - await fs.unlink(historyFilePath(storagePath, "peer-removed")) - await expect(store.delete("peer-removed")).resolves.toBeUndefined() - expect(store.get("peer-removed")).toBeUndefined() - expect(onWrite).toHaveBeenCalled() + expect(store.get("never-existed")).toBeUndefined() + expect(onWrite).toHaveBeenCalledTimes(1) }) }) describe("deleteMany()", () => { - it("stops at the first failure, preserves failed and unattempted cache entries, and writes through completed deletions before rejecting", async () => { + it("continues the batch after a failed unlink, evicts every entry, and writes through exactly once", async () => { const store = createStore() await store.initialize() await store.upsert(makeHistoryItem({ id: "batch-a", ts: 1000 })) @@ -153,93 +184,60 @@ describe("TaskHistoryStore fail-closed deletion semantics", () => { await store.upsert(makeHistoryItem({ id: "batch-c", ts: 3000 })) onWrite.mockClear() - const unlinkError = epermLike("EACCES: permission denied, unlink") - const events: string[] = [] - onWrite.mockImplementation(async () => { - events.push("write-through") - }) - vi.mocked(fs.unlink).mockImplementation(async (p) => { if (p === historyFilePath(storagePath, "batch-b")) { - throw unlinkError + throw Object.assign(new Error("EACCES: permission denied, unlink"), { code: "EACCES" }) } return actualFs.unlink(p) }) - const pending = store.deleteMany(["batch-a", "batch-b", "batch-c"]).catch((error) => { - events.push("rejected") - return error - }) - const partialError = (await pending) as TaskHistoryPartialDeleteError - - // The write-through of completed deletions is awaited before the - // rejection reaches the caller. - expect(events).toEqual(["write-through", "rejected"]) - - expect(partialError).toBeInstanceOf(TaskHistoryPartialDeleteError) - expect(partialError.failures).toEqual([{ taskId: "batch-b", reason: unlinkError }]) - expect(partialError.cause).toBe(unlinkError) + await expect(store.deleteMany(["batch-a", "batch-b", "batch-c"])).resolves.toBeUndefined() - // batch-a completed: file gone, cache evicted. + // batch-b failed but the batch continued around it. await expect(fs.access(historyFilePath(storagePath, "batch-a"))).rejects.toMatchObject({ code: "ENOENT" }) - expect(store.get("batch-a")).toBeUndefined() - - // batch-b failed: file present, cache entry preserved. await expect(fs.access(historyFilePath(storagePath, "batch-b"))).resolves.toBeUndefined() - expect(store.get("batch-b")).toBeDefined() - - // batch-c was never attempted: file present, cache entry preserved. - await expect(fs.access(historyFilePath(storagePath, "batch-c"))).resolves.toBeUndefined() - expect(store.get("batch-c")).toBeDefined() - - // The write-through saw the cache after completed deletions only. - const writtenIds = (onWrite.mock.calls[0][0] as HistoryItem[]).map((item) => item.id).sort() - expect(writtenIds).toEqual(["batch-b", "batch-c"]) + await expect(fs.access(historyFilePath(storagePath, "batch-c"))).rejects.toMatchObject({ code: "ENOENT" }) + + const { cache, taskFileMtimes } = storeInternals(store) + expect(cache.has("batch-a")).toBe(false) + expect(cache.has("batch-b")).toBe(false) + expect(cache.has("batch-c")).toBe(false) + expect(taskFileMtimes.has("batch-a")).toBe(false) + expect(taskFileMtimes.has("batch-b")).toBe(false) + expect(taskFileMtimes.has("batch-c")).toBe(false) + expect(store.get("batch-b")).toBeUndefined() + + // One write-through, after the whole batch, seeing the final cache. + expect(onWrite).toHaveBeenCalledTimes(1) + const writtenIds = (onWrite.mock.calls[0][0] as HistoryItem[]).map((item) => item.id) + expect(writtenIds).toEqual([]) }) - it("preserves the original deletion error when the failure-path write-through also fails", async () => { + it("swallows a lock failure for one item and continues the batch with one write-through", async () => { const store = createStore() await store.initialize() - await store.upsert(makeHistoryItem({ id: "batch-a", ts: 1000 })) - await store.upsert(makeHistoryItem({ id: "batch-b", ts: 2000 })) + await store.upsert(makeHistoryItem({ id: "lock-a", ts: 1000 })) + await store.upsert(makeHistoryItem({ id: "lock-b", ts: 2000 })) onWrite.mockClear() - const unlinkError = epermLike("EACCES: permission denied, unlink") - const consoleError = vi.spyOn(console, "error").mockImplementation(() => {}) + vi.mocked(withFileLock).mockRejectedValueOnce(new Error("lock acquisition timed out")) - vi.mocked(fs.unlink).mockImplementation(async (p) => { - if (p === historyFilePath(storagePath, "batch-b")) { - throw unlinkError - } - return actualFs.unlink(p) - }) - onWrite.mockRejectedValue(new Error("write-through boom")) - - try { - const partialError = (await store - .deleteMany(["batch-a", "batch-b"]) - .catch((error) => error)) as TaskHistoryPartialDeleteError - - // The original deletion failure stays the reported error. - expect(partialError).toBeInstanceOf(TaskHistoryPartialDeleteError) - expect(partialError.failures).toEqual([{ taskId: "batch-b", reason: unlinkError }]) - expect(partialError.cause).toBe(unlinkError) - - // The secondary write-through failure is logged, not thrown. - expect(consoleError).toHaveBeenCalledWith( - "[TaskHistoryStore] Write-through after partial deleteMany failed:", - expect.any(Error), - ) - - // batch-a stays deleted and batch-b stays intact. - await expect(fs.access(historyFilePath(storagePath, "batch-a"))).rejects.toMatchObject({ - code: "ENOENT", - }) - await expect(fs.access(historyFilePath(storagePath, "batch-b"))).resolves.toBeUndefined() - expect(store.get("batch-b")).toBeDefined() - } finally { - consoleError.mockRestore() - } + await expect(store.deleteMany(["lock-a", "lock-b"])).resolves.toBeUndefined() + + // The first item's lock failed before its unlink; the second item + // still completed. + await expect(fs.access(historyFilePath(storagePath, "lock-a"))).resolves.toBeUndefined() + await expect(fs.access(historyFilePath(storagePath, "lock-b"))).rejects.toMatchObject({ code: "ENOENT" }) + + const { cache, taskFileMtimes } = storeInternals(store) + expect(cache.has("lock-a")).toBe(false) + expect(cache.has("lock-b")).toBe(false) + expect(taskFileMtimes.has("lock-a")).toBe(false) + expect(taskFileMtimes.has("lock-b")).toBe(false) + + expect(onWrite).toHaveBeenCalledTimes(1) + const writtenIds = (onWrite.mock.calls[0][0] as HistoryItem[]).map((item) => item.id) + expect(writtenIds).toEqual([]) }) }) }) diff --git a/src/core/task/Task.ts b/src/core/task/Task.ts index 4633a54098..143c6b886c 100644 --- a/src/core/task/Task.ts +++ b/src/core/task/Task.ts @@ -1003,14 +1003,15 @@ export class Task extends EventEmitter implements TaskLike { } /** - * An interrupted task cannot legally delegate, so its staged create-subtask - * action is a durable rejection marker rather than replayable work. Reconcile - * it before restart replay; if persistence is still unavailable, propagate the - * error and leave the task stopped instead of creating another doomed child. + * An interrupted task cannot legally delegate, so a staged create-subtask + * action is a durable rejection marker rather than replayable work. The + * constructor-injected history item can be stale, so the persisted record + * is refreshed first and the refreshed action is the one settled. A failed + * refresh, lookup, or settlement stops replay instead of risking another + * doomed child. */ private async settleInterruptedCreateSubtaskBeforeReplay(): Promise { - const action = this.pendingAction - if (this.initialStatus !== "interrupted" || action?.kind !== "create_subtask") { + if (this.initialStatus !== "interrupted") { return } @@ -1021,24 +1022,38 @@ export class Task extends EventEmitter implements TaskLike { ) } - let authoritative: HistoryItem try { - authoritative = await provider.taskHistoryStore.clearPendingActionIfMatching(this.taskId, action.actionId) + await provider.taskHistoryStore.reconcile({ forceRefresh: true }) } catch (error) { throw new PendingActionSettlementError( - `[Task#settleInterruptedCreateSubtaskBeforeReplay] Failed to settle rejected action for task ${this.taskId}`, + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Failed to refresh task history for task ${this.taskId}`, { cause: error }, ) } - if (this.pendingAction?.actionId === action.actionId) { - this.pendingAction = authoritative.pendingAction + + const refreshedItem = provider.taskHistoryStore.get(this.taskId) + if (!refreshedItem) { + throw new PendingActionSettlementError( + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Task ${this.taskId} not found in refreshed task history`, + ) } + this.pendingAction = refreshedItem.pendingAction - if (this.pendingAction?.kind === "create_subtask") { + const action = refreshedItem.pendingAction + if (action?.kind !== "create_subtask") { + return + } + + let authoritative: HistoryItem + try { + authoritative = await provider.taskHistoryStore.clearPendingActionIfMatching(this.taskId, action.actionId) + } catch (error) { throw new PendingActionSettlementError( - `[Task#settleInterruptedCreateSubtaskBeforeReplay] Task ${this.taskId} still has a rejected create-subtask action`, + `[Task#settleInterruptedCreateSubtaskBeforeReplay] Failed to settle rejected action for task ${this.taskId}`, + { cause: error }, ) } + this.pendingAction = authoritative.pendingAction } private handleQueuedAskResponse(message: QueuedMessage, resolution: QueuedAskResolution): string | undefined { diff --git a/src/core/task/__tests__/Task.persistence.spec.ts b/src/core/task/__tests__/Task.persistence.spec.ts index 1b25e814fa..77cb813e51 100644 --- a/src/core/task/__tests__/Task.persistence.spec.ts +++ b/src/core/task/__tests__/Task.persistence.spec.ts @@ -1244,6 +1244,13 @@ describe("Task persistence", () => { }, ]) + // Restart settlement needs the refreshed record to exist, and a + // record without a pending action keeps the replay path unchanged. + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "interrupted-subtask", + status: "interrupted", + }) + await getTaskPersistenceAccess(task).resumeTaskFromHistory() expect(initiateTaskLoopSpy).toHaveBeenCalledTimes(1) @@ -1309,6 +1316,13 @@ describe("Task persistence", () => { }, ]) + // Restart settlement needs the refreshed record to exist, and a + // record without a pending action keeps the replay path unchanged. + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "interrupted-subtask-2", + status: "interrupted", + }) + await getTaskPersistenceAccess(task).resumeTaskFromHistory() expect(initiateTaskLoopSpy).toHaveBeenCalledTimes(1) @@ -1421,6 +1435,7 @@ describe("Task persistence", () => { await getTaskPersistenceAccess(task).resumeTaskFromHistory() expect(clearRejectedAction).not.toHaveBeenCalled() + expect(mockProvider.taskHistoryStore.reconcile).not.toHaveBeenCalled() expect(replay).toHaveBeenCalledWith(createSubtaskAction) }) @@ -1429,6 +1444,12 @@ describe("Task persistence", () => { { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, ]) mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: createSubtaskAction, + }) const clearRejectedAction = vi.fn().mockResolvedValue({ id: "parent-1", status: "interrupted", @@ -1457,6 +1478,8 @@ describe("Task persistence", () => { await getTaskPersistenceAccess(task).resumeTaskFromHistory() + expect(mockProvider.taskHistoryStore.reconcile).toHaveBeenCalledWith({ forceRefresh: true }) + expect(mockProvider.taskHistoryStore.get).toHaveBeenCalledWith("parent-1") expect(clearRejectedAction).toHaveBeenCalledWith("parent-1", "create-action") expect(replay).not.toHaveBeenCalled() expect(task.ask).toHaveBeenCalledWith("resume_task") @@ -1466,6 +1489,12 @@ describe("Task persistence", () => { const settlementError = new Error("settlement unavailable") mockReadTaskMessages.mockResolvedValue([]) mockReadApiMessages.mockResolvedValue([]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: createSubtaskAction, + }) mockProvider.taskHistoryStore.clearPendingActionIfMatching = vi.fn().mockRejectedValue(settlementError) const task = new Task({ provider: mockProvider, @@ -1500,7 +1529,7 @@ describe("Task persistence", () => { expect(ask).not.toHaveBeenCalled() }) - it("stops replay when settlement preserves a replacement create-subtask action", async () => { + it("adopts the authoritative replacement pending action returned by settlement", async () => { const replacementAction = { ...createSubtaskAction, actionId: "create-action-b", @@ -1510,11 +1539,70 @@ describe("Task persistence", () => { { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, ]) mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) - mockProvider.taskHistoryStore.clearPendingActionIfMatching = vi.fn().mockResolvedValue({ + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: createSubtaskAction, + }) + const clearRejectedAction = vi.fn().mockResolvedValue({ id: "parent-1", status: "interrupted", pendingAction: replacementAction, }) + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const taskState = task as unknown as { pendingAction?: PendingTaskAction } + const ask = vi.spyOn(task, "ask") + const replay = vi + .spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + .mockResolvedValue(undefined) + + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + expect(clearRejectedAction).toHaveBeenCalledWith("parent-1", "create-action") + expect(taskState.pendingAction).toEqual(replacementAction) + expect(replay).toHaveBeenCalledWith(replacementAction) + expect(ask).not.toHaveBeenCalled() + }) + + it("settles the refreshed action when the staged action was replaced before restart", async () => { + const refreshedAction = { + ...createSubtaskAction, + actionId: "create-action-b", + message: "Refreshed child", + } + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: refreshedAction, + }) + const clearRejectedAction = vi.fn().mockResolvedValue({ + id: "parent-1", + status: "interrupted", + pendingAction: undefined, + }) + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction const task = new Task({ provider: mockProvider, apiConfiguration: mockApiConfig, @@ -1531,13 +1619,186 @@ describe("Task persistence", () => { }, startTask: false, }) + const taskState = task as unknown as { pendingAction?: PendingTaskAction } const ask = vi.spyOn(task, "ask") const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") - await expect(getTaskPersistenceAccess(task).resumeTaskFromHistory()).rejects.toThrow( - "[Task#settleInterruptedCreateSubtaskBeforeReplay] Task parent-1 still has a rejected create-subtask action", - ) + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + // The stale staged action never reaches settlement; the refreshed + // one does, and the cleared authoritative result ends the replay. + expect(clearRejectedAction).toHaveBeenCalledTimes(1) + expect(clearRejectedAction).toHaveBeenCalledWith("parent-1", "create-action-b") + expect(taskState.pendingAction).toBeUndefined() + expect(replay).not.toHaveBeenCalled() + expect(ask).toHaveBeenCalledWith("resume_task") + }) + + it("replays without settlement when the refreshed record has no pending action", async () => { + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: undefined, + }) + const clearRejectedAction = vi.fn() + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const taskState = task as unknown as { pendingAction?: PendingTaskAction } + const ask = vi.spyOn(task, "ask") + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + expect(mockProvider.taskHistoryStore.reconcile).toHaveBeenCalledWith({ forceRefresh: true }) + expect(clearRejectedAction).not.toHaveBeenCalled() + expect(taskState.pendingAction).toBeUndefined() + expect(replay).not.toHaveBeenCalled() + expect(ask).toHaveBeenCalledWith("resume_task") + }) + + it("replays a refreshed non-create-subtask action without settlement", async () => { + const refreshedFinishAction: PendingTaskAction = { + ...pendingAction, + actionId: "finish-action-b", + } + mockReadTaskMessages.mockResolvedValue([ + { ts: 1, type: "ask", ask: "tool", text: createSubtaskAction.approvalText }, + ]) + mockReadApiMessages.mockResolvedValue([{ role: "assistant", content: "Previous response" }]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue({ + id: "parent-1", + status: "interrupted", + pendingAction: refreshedFinishAction, + }) + const clearRejectedAction = vi.fn() + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const taskState = task as unknown as { pendingAction?: PendingTaskAction } + const ask = vi.spyOn(task, "ask") + const replay = vi + .spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + .mockResolvedValue(undefined) + + await getTaskPersistenceAccess(task).resumeTaskFromHistory() + + expect(clearRejectedAction).not.toHaveBeenCalled() + expect(taskState.pendingAction).toEqual(refreshedFinishAction) + expect(replay).toHaveBeenCalledWith(refreshedFinishAction) + expect(ask).not.toHaveBeenCalled() + }) + + it("stops replay when the refreshed task history cannot be read", async () => { + const refreshError = new Error("reconcile unavailable") + mockReadTaskMessages.mockResolvedValue([]) + mockReadApiMessages.mockResolvedValue([]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockRejectedValue(refreshError) + const clearRejectedAction = vi.fn() + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const ask = vi.spyOn(task, "ask") + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + const resumeError = await getTaskPersistenceAccess(task) + .resumeTaskFromHistory() + .catch((error: unknown) => error) + + expect(resumeError).toBeInstanceOf(PendingActionSettlementError) + expect(resumeError).toMatchObject({ + name: "PendingActionSettlementError", + cause: refreshError, + }) + expect(clearRejectedAction).not.toHaveBeenCalled() + expect(replay).not.toHaveBeenCalled() + expect(ask).not.toHaveBeenCalled() + }) + + it("stops replay when the task is missing from the refreshed task history", async () => { + mockReadTaskMessages.mockResolvedValue([]) + mockReadApiMessages.mockResolvedValue([]) + mockProvider.taskHistoryStore.reconcile = vi.fn().mockResolvedValue(undefined) + mockProvider.taskHistoryStore.get = vi.fn().mockReturnValue(undefined) + const clearRejectedAction = vi.fn() + mockProvider.taskHistoryStore.clearPendingActionIfMatching = clearRejectedAction + const task = new Task({ + provider: mockProvider, + apiConfiguration: mockApiConfig, + historyItem: { + id: "parent-1", + number: 1, + ts: 1, + task: "Parent", + tokensIn: 0, + tokensOut: 0, + totalCost: 0, + status: "interrupted", + pendingAction: createSubtaskAction, + }, + startTask: false, + }) + const ask = vi.spyOn(task, "ask") + const replay = vi.spyOn(getTaskPersistenceAccess(task), "resumePendingTaskAction") + + const resumeError = await getTaskPersistenceAccess(task) + .resumeTaskFromHistory() + .catch((error: unknown) => error) + + expect(resumeError).toBeInstanceOf(PendingActionSettlementError) + expect(resumeError).toMatchObject({ + name: "PendingActionSettlementError", + message: expect.stringContaining("not found in refreshed task history"), + }) + expect(clearRejectedAction).not.toHaveBeenCalled() expect(replay).not.toHaveBeenCalled() expect(ask).not.toHaveBeenCalled() }) diff --git a/src/core/webview/ClineProvider.ts b/src/core/webview/ClineProvider.ts index b7f41c1711..39f19130d7 100644 --- a/src/core/webview/ClineProvider.ts +++ b/src/core/webview/ClineProvider.ts @@ -2366,10 +2366,7 @@ export class ClineProvider } } - // Delete all tasks from state in one batch. A failed history - // deletion fails closed: the error leaves this method before any - // checkpoint or task directory removal below, so a directory is - // never removed while its history file could still exist. + // Delete all tasks from state in one batch await this.taskHistoryStore.deleteMany(allIdsToDelete) this.recentTasksCache = undefined diff --git a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts index bd71fadee6..753332dd5b 100644 --- a/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts +++ b/src/core/webview/__tests__/ClineProvider.taskHistory.spec.ts @@ -1,6 +1,5 @@ // pnpm --filter roo-cline test core/webview/__tests__/ClineProvider.taskHistory.spec.ts -import * as fs from "fs/promises" import * as vscode from "vscode" import type { HistoryItem, ExtensionMessage } from "@roo-code/types" import { providerIdentifiers, RooCodeEventName } from "@roo-code/types" @@ -26,9 +25,6 @@ vi.mock("fs/promises", () => ({ rmdir: vi.fn().mockResolvedValue(undefined), access: vi.fn().mockResolvedValue(undefined), rm: vi.fn().mockResolvedValue(undefined), - // Deletion verifies disk state after a failed lock or unlink, so the - // harness reports every probed file as absent. - stat: vi.fn().mockRejectedValue(Object.assign(new Error("no such file or directory"), { code: "ENOENT" })), })) vi.mock("axios", () => ({ @@ -847,26 +843,6 @@ describe("ClineProvider Task History Synchronization", () => { expect(ids).not.toContain("remove-me") }) - it("fails closed: does not remove task directories when history deletion fails", async () => { - await provider.resolveWebviewView(mockWebviewView) - - const item = createHistoryItem({ id: "dir-keep", task: "Keep directory" }) - await provider.updateTaskHistory(item, { broadcast: false }) - - const failure = new Error("lock acquisition failed") - const deleteManySpy = vi.spyOn(provider.taskHistoryStore, "deleteMany").mockRejectedValue(failure) - - try { - // The failed history deletion aborts deleteTaskWithId before - // any checkpoint or task directory removal. - await expect(provider.deleteTaskWithId("dir-keep")).rejects.toBe(failure) - } finally { - deleteManySpy.mockRestore() - } - - expect(fs.rm).not.toHaveBeenCalled() - }) - it("does not block subsequent writes when a previous store write errors", async () => { await provider.resolveWebviewView(mockWebviewView)