Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
@@ -1,3 +1,4 @@
node_modules/
*.log
.DS_Store
*.swp
71 changes: 66 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,7 @@ defaults:
"options": {
"maxTurns": 20,
"stallLimit": 2,
"pollLimit": 3,
"judgeModel": { "providerID": "openrouter", "id": "google/gemini-3-flash-preview" }
}
}
Expand All @@ -106,6 +107,7 @@ defaults:
| --- | --- | --- |
| `maxTurns` | `20` | Automatic continuation turns before the loop auto-pauses. A whole number, or `"unlimited"` / `null` for no ceiling. The initial `/goal` turn is not counted, so `20` allows 21 agent executions in total. See below. |
| `stallLimit` | `2` | Consecutive turns that ran **no tools** before the loop is declared stalled. |
| `pollLimit` | `3` | Consecutive turns that re-read an unchanged result before the loop is called polling. Distinct from `stallLimit`: a polling turn *does* use tools, it just learns nothing. |
| `judgeModel` | session model | Model used for the `done` / `continue` / `blocked` verdict. |
| `quiet` | `false` | Stop posting the loop's turn banner and completion notices into the transcript while the panel is open. |

Expand Down Expand Up @@ -143,6 +145,7 @@ until the session ends — it never rewrites your config.
| --- | --- | --- |
| `/goal budget <n\|inf\|default>` | `maxTurns` | See [Removing the turn limit](#removing-the-turn-limit) |
| `/goal stall <n\|default>` | `stallLimit` | Turns with **no tool calls** before giving up |
| `/goal poll <n\|default>` | `pollLimit` | Turns reading the **same unchanged result** before calling it polling |
| `/goal quiet <on\|off\|default>` | `quiet` | When on, the panel replaces the loop's transcript notices |
| `/goal judge <provider/model[#variant]\|default>` | `judgeModel` | Cheaper and sharper models judge better and cost less |
| `/goal settings` | — | Shows all four, and whether each is a session override or the config default |
Expand Down Expand Up @@ -206,9 +209,11 @@ It applies to the running goal immediately and to every goal set afterwards in t
| Still stops it | |
| --- | --- |
| The judge says `done` | the goal is met, with evidence |
| The judge says `blocked` | the goal is unreachable, or the agent is going in circles |
| The judge says `blocked` | the goal is unreachable, on goal-level evidence |
| Stall guard | `stallLimit` consecutive turns with no tool calls |
| Polling guard | `pollLimit` consecutive turns with an unchanged result |
| Repetition guard | the same reply twice running |
| Polling guard | 3 turns that re-read an unchanged result |
| You | `/goal pause`, `/goal clear`, or <kbd>esc</kbd> |

So an agent that keeps making small, genuine-looking progress can now run indefinitely. Nothing
Expand Down Expand Up @@ -299,6 +304,23 @@ non-trivial. Comma or space separated, with an optional trailing `on`/`off`:
/goal display footer,sidebar off turn both off, leave the rest alone
```

A change appears immediately, without reopening anything.

That took a fix worth recording, because the cause was not obvious from the symptom.
`context.storage.store` is durable but **not reactive**, and a slot's `render` runs once, so
`<Show when={display.panel}>` read a snapshot of the placements and never re-read it. The
visible effect was that a placement change only appeared after something forced the slot to
re-render — in practice, toggling the panel with `/goal panel`, which is exactly the manual
workaround that should not be needed.

So the stored value is mirrored into a signal, which is the mechanism the goal state already
used and which demonstrably re-renders in this host. The store stays the single durable
record; the signal is only what the UI reads, and it leads the write so a change shows on
the frame it happens rather than after the storage round trip. Because the UI shows four
placements but only ever changes the named one, `applyMutation` returns the **whole** value
rather than the keys that changed — otherwise an untouched placement would silently revert on
the next write. Both properties are pinned in `test/display.test.ts`.

With an explicit `on`/`off` the named placements are set and the rest untouched; without one,
each named placement toggles.

Expand Down Expand Up @@ -366,26 +388,65 @@ Only these exact prefixes are recognised, so an ordinary goal containing a colon

## How the loop stops

The point of this plugin is that it terminates. Five conditions, checked in order:
The point of this plugin is that it terminates. Six conditions, checked in order:

| # | Condition | Kind |
| --- | --- | --- |
| 1 | Judge returns `done` — the reply carries concrete evidence, such as a passing command and its output | model |
| 2 | Judge returns `blocked` — impossible, out of scope, needs credentials or hardware you do not have, or the agent is going in circles | model |
| 2 | Judge returns `blocked` — impossible, out of scope, needs credentials or hardware you do not have | model |
| 3 | **Stall** — `stallLimit` consecutive turns ran no tools at all, so nothing changed however confident the prose | deterministic |
| 4 | **Repetition** — the agent produced the same reply twice running | deterministic |
| 5 | **Budget** — `maxTurns` continuation turns spent (21 executions by default) | deterministic |
| 5 | **Polling** — `pollLimit` consecutive turns ran tools and read back the same unchanged result | deterministic |
| 6 | **Budget** — `maxTurns` continuation turns spent (21 executions by default) | deterministic |

Conditions 3–5 do not consult the model. This is deliberate: in testing, a weak judge model
Conditions 3–6 do not consult the model. This is deliberate: in testing, a weak judge model
answered `continue` to twenty byte-identical replies and happily spent the entire budget. The
deterministic guards are what actually stopped it, in two turns and about $0.004 instead of
twenty turns and roughly $0.02. Treat the judge's `blocked` verdict as a useful fourth opinion,
not as the safety net.

### `blocked` means the goal, not the turn

The one rule worth stating on its own, because getting it wrong stops work that would have
finished. An earlier version told the judge to answer `blocked` when "the loop is going in
circles". That mapped a **turn-level** observation onto a **goal-level** verdict, and the
failure mode was concrete: an agent waiting on a five-minute build spends each turn reading
an unchanged log, and the judge — seeing a turn that changed nothing — declared the whole
goal unachievable and paused. The work was not unachievable; the agent was waiting.

So repetition is now explicitly **not** a reason to answer `blocked`, however many turns it
has happened, because repetition is recoverable: the next turn can do something different. A
turn that checks on work already in flight is **waiting**, not circling, and the judge is told
so and told not to assume work did *not* happen off-screen.

### The polling guard

The stall and repetition guards miss a specific shape. An agent waiting on a long job says
something different every turn (so the reply digest moves) and calls a tool every turn (so
the tool count is non-zero), while learning nothing. Only the model noticed, and it reached
for the terminal verdict.

So the loop also digests **what the tool calls reported**, ignoring the commands that
produced them, and pauses after `pollLimit` turns whose observation is unchanged. Its message
says the work is not moving, says to wait for a running command rather than re-read it, and
says plainly that this is not a statement about whether the goal is reachable.

`pollLimit` is configurable exactly like `stallLimit` — `pollLimit` in `opencode.json`, or
`/goal poll <n|default>` for the session — because the right threshold is a property of the
work: a build that takes five minutes wants a higher limit than a test that takes five
seconds, and a hardcoded 3 was the one number here a user could not argue with.

It fails open: an unrecognised tool-state shape yields an empty digest and the guard stays
quiet, because a heuristic that pauses a loop on a guess is worse than one that misses.

## Caveats

- **The judge is only as good as its model.** A weak judge is permissive. Set `judgeModel` to
something small and sharp, and expect to use `/goal status` to sanity-check its verdicts.
- **A `blocked` verdict is worth reading twice.** It means the goal looks unreachable, which
is a strong claim resting on one model call. If the agent was mid-way through a long
command, a large batch, or a build, the more likely story is that it was waiting and the
judge read a quiet turn as a dead end. `/goal resume` costs nothing but the turn.
- **Your agent model must actually use tools.** If the session model narrates intentions
without calling tools, the stall guard fires after two turns. That is the guard working, but
it means the goal will not get done.
Expand Down
14 changes: 14 additions & 0 deletions src/display.ts
Original file line number Diff line number Diff line change
Expand Up @@ -99,3 +99,17 @@ export function applyAll(enabled: boolean): Display {
return { panel: enabled, footer: enabled, composer: enabled, sidebar: enabled }
}

/**
* Apply a mutation to a display and return the complete next value.
*
* The UI shows all four placements but only ever changes the named ones, so persisting
* has to write the *whole* value rather than the keys the caller happened to touch.
* Returning a fresh object rather than editing in place also gives the UI something new
* to compare, which is what makes a change visible on the frame it happens.
*/
export function applyMutation(current: Display, mutate: (draft: Display) => void): Display {
const next: Display = { ...current }
mutate(next)
return { panel: next.panel, footer: next.footer, composer: next.composer, sidebar: next.sidebar }
}

116 changes: 116 additions & 0 deletions src/listen.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,116 @@
/**
* A subscription that survives.
*
* The client's own `rpc.events.on` is one `for await` loop with a catch at the
* end, and nothing that restarts it. Measured against the real library:
*
* - a handler that throws once ends the loop, so every later event is dropped
* and the error goes to `console.error`;
* - a stream that simply ends ends the loop with no error at all, and nothing
* opens a second connection.
*
* Either way the listener stays registered and looks alive while being
* permanently deaf, and the UI it feeds stops updating until something else
* refreshes it - here, reopening the panel, which reads state directly.
*
* `listen` fixes both. A handler error is reported and the next event is still
* delivered; when the stream ends or fails it is reopened after a growing delay,
* and `onResume` runs first so the caller can re-read whatever it missed while
* nothing was listening.
*
* No dependency on the client, Solid or JSX, so it can be tested with plain
* async iterables.
*/

export type Stop = () => void

/**
* Where a failure came from. A dropped stream is routine and heals itself, while
* a throwing handler is a bug, so callers usually want to treat them differently.
*/
export type Phase = "handler" | "stream" | "resume"

export interface ListenOptions {
/**
* Runs after a gap in delivery, before the stream is reopened. Events sent
* during the gap are gone, so this is where the caller re-reads state. It is
* not called before the first connection.
*/
onResume?: () => void | Promise<void>
/** Told about a failure and where it came from. May throw; that is contained. */
onError?: (error: unknown, phase: Phase) => void
/** Injectable so tests need not wait. */
sleep?: (ms: number, signal: AbortSignal) => Promise<void>
/** First reconnect delay. Default 500ms. */
minDelay?: number
/** Ceiling for the doubling delay. Default 10s. */
maxDelay?: number
}

const sleepUnlessAborted = (ms: number, signal: AbortSignal) =>
new Promise<void>((resolve) => {
if (signal.aborted) return resolve()
const finish = () => {
clearTimeout(timer)
signal.removeEventListener("abort", finish)
resolve()
}
const timer = setTimeout(finish, ms)
signal.addEventListener("abort", finish, { once: true })
})

export function listen<E>(
subscribe: (signal: AbortSignal) => AsyncIterable<E>,
handler: (event: E) => void | Promise<void>,
options: ListenOptions = {},
): Stop {
const controller = new AbortController()
const { signal } = controller
const sleep = options.sleep ?? sleepUnlessAborted
const minDelay = options.minDelay ?? 500
const maxDelay = options.maxDelay ?? 10_000
const report = (error: unknown, phase: Phase) => {
try {
options.onError?.(error, phase)
} catch {
// A broken error sink must not take the listener down with it.
}
}

void (async () => {
let delay = minDelay
let reconnecting = false

while (!signal.aborted) {
if (reconnecting) {
await sleep(delay, signal)
if (signal.aborted) return
delay = Math.min(delay * 2, maxDelay)
try {
await options.onResume?.()
} catch (error) {
report(error, "resume")
}
if (signal.aborted) return
}
reconnecting = true

try {
for await (const event of subscribe(signal)) {
if (signal.aborted) return
// Something arrived, so the connection is healthy: forget the backoff.
delay = minDelay
try {
await handler(event)
} catch (error) {
report(error, "handler")
}
}
} catch (error) {
if (!signal.aborted) report(error, "stream")
}
}
})()

return () => controller.abort()
}
55 changes: 55 additions & 0 deletions src/observation.ts
Original file line number Diff line number Diff line change
@@ -0,0 +1,55 @@
/**
* Turning a turn's tool calls into one stable observation string.
*
* The point of this is narrow and specific: an agent waiting on a long build produces a
* turn that calls tools, says something new, and learns nothing. That defeats both of the
* loop's existing no-progress checks — the reply digest moves, so the reply-repetition
* guard never fires, and the tool count is non-zero, so the stall guard never fires. Only
* the judge noticed, and its answer was to declare the *goal* unachievable from
* turn-level evidence.
*
* So this digests what the tool calls *reported*, not what the agent said or asked for.
* Two deliberate choices:
*
* - Tool inputs are skipped. An agent that reworded the same command each turn was still
* reading the same unchanged result, and that is precisely the case worth catching.
* - It fails open. An unrecognised state shape yields an empty digest, and the guard
* stays quiet, because a heuristic that pauses a loop on a guess is worse than one
* that misses.
*/

/** Keys that describe *how* the tool was called rather than what it found. */
const CALL_SHAPE = new Set(["input", "args", "description", "title", "callID", "id"])

/** Everything a turn's tool calls reported, as one stable string. */
export function observationOf(messages: readonly unknown[]): string {
const parts: string[] = []
const walk = (value: unknown, depth = 0): void => {
if (depth > 6 || value === null || value === undefined) return
if (typeof value === "string") {
parts.push(value)
return
}
if (typeof value === "number" || typeof value === "boolean") {
parts.push(String(value))
return
}
if (Array.isArray(value)) {
for (const item of value) walk(item, depth + 1)
return
}
if (typeof value === "object")
for (const [key, nested] of Object.entries(value as Record<string, unknown>))
if (!CALL_SHAPE.has(key)) walk(nested, depth + 1)
}

for (const message of messages) {
const record = message as { type?: unknown; content?: unknown } | null
if (record?.type !== "assistant") continue
for (const part of (record.content ?? []) as readonly unknown[]) {
const entry = part as { type?: unknown; state?: unknown } | null
if (entry?.type === "tool") walk(entry.state)
}
}
return parts.join("\n")
}
5 changes: 5 additions & 0 deletions src/rpc.ts
Original file line number Diff line number Diff line change
Expand Up @@ -35,6 +35,7 @@ export const HELP_TEXT = `/goal <text> set the goal and start working
/goal budget inf no turn limit (also: unlim, unlimited, none, ∞)
/goal budget default back to the configured maxTurns
/goal stall <n> turns with no tools before the loop gives up
/goal poll <n> turns re-reading an unchanged result before it counts as polling
/goal quiet <on|off> whether the panel replaces the transcript notices
/goal judge <model> provider/model[#variant] used to judge each turn
/goal settings show every setting, and where it came from
Expand All @@ -57,6 +58,9 @@ export type GoalView = {
maxTurns: number | null
stalled: number
repeats: number
/** Turns that re-read an unchanged result. Optional: added after v1 shipped, so it is
* deliberately not in `required` and older views simply do not carry it. */
observing?: number
/** The judge's most recent one-sentence reason, or "". */
reason: string
/** The goal's own proof condition, or "". */
Expand All @@ -74,6 +78,7 @@ const state = {
maxTurns: { anyOf: [{ type: "number" }, { type: "null" }] },
stalled: { type: "number" },
repeats: { type: "number" },
observing: { type: "number" },
reason: { type: "string" },
verification: { type: "string" },
updatedAt: { type: "number" },
Expand Down
Loading
Loading