fix: blocked must mean the goal, not that a turn was unproductive - #1
Merged
Merged
Conversation
The judge prompt told the model to answer "blocked" when the loop was going in circles. That mapped a turn-level observation onto a goal-level verdict, and the failure mode was concrete: an agent waiting on a five-minute build spends each turn reading an unchanged log, and the judge — seeing a turn that changed nothing — declared the goal unachievable and paused the loop. The work was not unachievable; the agent was waiting. Repetition is recoverable, because the next turn can do something different, so it can no longer produce "blocked" at any count. `blocked` now requires goal-level evidence, `continue` explicitly covers the unproductive turn, and a turn checking on work already in flight is defined as waiting rather than circling. The judge is also told not to assume work did *not* happen off-screen, which was the assumption that made a long command look like a dead end. The two deterministic guards could not catch that shape, which is why only the model saw it: the reply digest moves when an agent varies its prose, and the tool count is non-zero when it polls, so `repeats` and `stall` both stay at zero. A third guard now digests what the tool calls *reported*, ignoring the commands that produced them, and pauses after three turns whose observation is unchanged. It fails open on an unrecognised tool-state shape, because a heuristic that pauses a loop on a guess is worse than one that misses, and its message says the work is not moving, says to wait for a running command rather than re-read it, and says plainly that this is not a statement about whether the goal is reachable. The repetition guard's message also claimed "the goal is not reachable in this session", which was an inference stated as a fact. Softened. `observing` is added to the goal view as an optional field rather than a required one, so views written against the previous schema still validate.
`/goal display` set the placements correctly and the change simply did not
show. Toggling the panel with `/goal panel` made it appear, which is the
workaround this should never have needed.
The cause is not obvious from the symptom, so it is worth stating: the
placements were read straight off `context.storage.store`, which is durable
but not reactive, and a slot's `render` runs once. So `<Show when=
{display.panel}>` read a snapshot of the placements and never re-read it. The
goal state right beside it reads `states()`, a signal accessor, which is why
that updates live and these did not.
So the stored value is now mirrored into a signal — the same mechanism, in
the same file, that already re-renders here. The store stays the single
durable record; the signal is only what the UI reads, and it is updated
before the storage write so a change lands on the frame it happens rather
than after the round trip.
`applyMutation` is extracted into display.ts for that write path, which
keeps it testable in a module with no JSX. It returns the whole value, not
just the keys the caller changed: the UI shows four placements but only
changes the named one, so writing only the difference would silently revert
an untouched placement on the next save. Both that and the no-edit of the
input are pinned.
The polling guard shipped with its threshold hardcoded at 3 while `stallLimit` and `maxTurns` were both configurable, so it was the one number in this plugin a user could not argue with. The right threshold is a property of the work: a build that takes five minutes wants a higher limit than a test that takes five seconds, and the person paying for the tokens is the one who should set it. So `pollLimit` joins `stallLimit` everywhere `stallLimit` already was, rather than only in the place it is used: - the plugin option, defaulting to 3 and ignoring anything below 1 - a per-session override, read back with the rest of the stored settings - `/goal poll <n|default>`, reusing `parseCount` so "default", padding and nonsense behave exactly as they do for `/goal stall` - `/goal settings`, which reports the value and whether it came from this session or the config - the help text and the README, including the settings summary Four existing tests changed shape rather than behaviour, and are worth noting as such: the resolved object gained a `poll` and an `overridden.poll`, and one test counted the parenthesised rows in the settings summary, which is now five instead of four. The count is of row markers, not the words themselves, precisely so adding a row is a visible change rather than a silent one.
The panel, footer, composer and sidebar stopped updating on their own, and
reopening the panel brought them current again.
The client's rpc.events.on is a single for-await loop with a catch at the end
and nothing that restarts it. Measured against the real library:
- a handler that throws once ends the loop, so every later event is dropped
and the error goes to console.error;
- a stream that ends ends the loop with no error at all, and nothing opens
a second connection.
Either way the listener stays registered and looks alive while being deaf.
Reopening the panel appeared to fix it only because that path calls rpc.get
directly, bypassing events. The footer, composer and sidebar never fetched
anything themselves, so they had nothing to fall back on.
listen() replaces rpc.events.on. A handler error is reported and the next event
is still delivered. A stream that ends or fails is reopened after a doubling
delay, and onResume runs first so state missed during the gap is re-read.
Handler errors surface as a toast, rate limited; a dropped stream stays quiet
because it heals itself.
The TUI now also:
- fetches a session's state when a compact area first shows it, so a goal
already running when the session opens is visible;
- re-reads every known session after a gap, and announces a goal that ended
while nothing was listening;
- polls active goals every 5s, since a half-open stream reports no error to
reconnect on;
- guards the terminal toast, so a failing toast cannot end the listener.
listen() has 16 checks. Breaking it in scratch copies (letting handler errors
escape, or never reopening) fails 3 and 7 of them respectively, so they detect
the two failures they exist for. What is not verified is what triggers the
silent stop in a live TUI; only the library behaviour was reproduced.
Add goal_wait, which holds a turn open until a background command finishes (detected by its output file being closed, since the completion notice only arrives after the turn ends). Independently, the repeat, stall and polling guards stand down while a background command is pending, with a back-off and a 60 minute bound.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The judge prompt told the model to answer "blocked" when the loop was going in circles. That mapped a turn-level observation onto a goal-level verdict, and the failure mode was concrete: an agent waiting on a five-minute build spends each turn reading an unchanged log, and the judge — seeing a turn that changed nothing — declared the goal unachievable and paused the loop. The work was not unachievable; the agent was waiting.
Repetition is recoverable, because the next turn can do something different, so it can no longer produce "blocked" at any count.
blockednow requires goal-level evidence,continueexplicitly covers the unproductive turn, and a turn checking on work already in flight is defined as waiting rather than circling. The judge is also told not to assume work did not happen off-screen, which was the assumption that made a long command look like a dead end.The two deterministic guards could not catch that shape, which is why only the model saw it: the reply digest moves when an agent varies its prose, and the tool count is non-zero when it polls, so
repeatsandstallboth stay at zero. A third guard now digests what the tool calls reported, ignoring the commands that produced them, and pauses after three turns whose observation is unchanged. It fails open on an unrecognised tool-state shape, because a heuristic that pauses a loop on a guess is worse than one that misses, and its message says the work is not moving, says to wait for a running command rather than re-read it, and says plainly that this is not a statement about whether the goal is reachable.The repetition guard's message also claimed "the goal is not reachable in this session", which was an inference stated as a fact. Softened.
observingis added to the goal view as an optional field rather than a required one, so views written against the previous schema still validate.