Skip to content

fix: blocked must mean the goal, not that a turn was unproductive - #1

Merged
zignd merged 5 commits into
mainfrom
fix/blocked-means-goal-not-turn
Sep 29, 2026
Merged

zignd merged 5 commits into
mainfrom
fix/blocked-means-goal-not-turn

Conversation

@zignd

@zignd zignd commented Sep 29, 2026

Copy link
Copy Markdown
Owner

The judge prompt told the model to answer "blocked" when the loop was going in circles. That mapped a turn-level observation onto a goal-level verdict, and the failure mode was concrete: an agent waiting on a five-minute build spends each turn reading an unchanged log, and the judge — seeing a turn that changed nothing — declared the goal unachievable and paused the loop. The work was not unachievable; the agent was waiting.

Repetition is recoverable, because the next turn can do something different, so it can no longer produce "blocked" at any count. blocked now requires goal-level evidence, continue explicitly covers the unproductive turn, and a turn checking on work already in flight is defined as waiting rather than circling. The judge is also told not to assume work did not happen off-screen, which was the assumption that made a long command look like a dead end.

The two deterministic guards could not catch that shape, which is why only the model saw it: the reply digest moves when an agent varies its prose, and the tool count is non-zero when it polls, so repeats and stall both stay at zero. A third guard now digests what the tool calls reported, ignoring the commands that produced them, and pauses after three turns whose observation is unchanged. It fails open on an unrecognised tool-state shape, because a heuristic that pauses a loop on a guess is worse than one that misses, and its message says the work is not moving, says to wait for a running command rather than re-read it, and says plainly that this is not a statement about whether the goal is reachable.

The repetition guard's message also claimed "the goal is not reachable in this session", which was an inference stated as a fact. Softened.

observing is added to the goal view as an optional field rather than a required one, so views written against the previous schema still validate.

The judge prompt told the model to answer "blocked" when the loop was going
in circles. That mapped a turn-level observation onto a goal-level verdict,
and the failure mode was concrete: an agent waiting on a five-minute build
spends each turn reading an unchanged log, and the judge — seeing a turn that
changed nothing — declared the goal unachievable and paused the loop. The
work was not unachievable; the agent was waiting.

Repetition is recoverable, because the next turn can do something different,
so it can no longer produce "blocked" at any count. `blocked` now requires
goal-level evidence, `continue` explicitly covers the unproductive turn, and a
turn checking on work already in flight is defined as waiting rather than
circling. The judge is also told not to assume work did *not* happen
off-screen, which was the assumption that made a long command look like a
dead end.

The two deterministic guards could not catch that shape, which is why only
the model saw it: the reply digest moves when an agent varies its prose, and
the tool count is non-zero when it polls, so `repeats` and `stall` both stay
at zero. A third guard now digests what the tool calls *reported*, ignoring
the commands that produced them, and pauses after three turns whose
observation is unchanged. It fails open on an unrecognised tool-state shape,
because a heuristic that pauses a loop on a guess is worse than one that
misses, and its message says the work is not moving, says to wait for a
running command rather than re-read it, and says plainly that this is not a
statement about whether the goal is reachable.

The repetition guard's message also claimed "the goal is not reachable in
this session", which was an inference stated as a fact. Softened.

`observing` is added to the goal view as an optional field rather than a
required one, so views written against the previous schema still validate.
`/goal display` set the placements correctly and the change simply did not
show. Toggling the panel with `/goal panel` made it appear, which is the
workaround this should never have needed.

The cause is not obvious from the symptom, so it is worth stating: the
placements were read straight off `context.storage.store`, which is durable
but not reactive, and a slot's `render` runs once. So `<Show when=
{display.panel}>` read a snapshot of the placements and never re-read it. The
goal state right beside it reads `states()`, a signal accessor, which is why
that updates live and these did not.

So the stored value is now mirrored into a signal — the same mechanism, in
the same file, that already re-renders here. The store stays the single
durable record; the signal is only what the UI reads, and it is updated
before the storage write so a change lands on the frame it happens rather
than after the round trip.

`applyMutation` is extracted into display.ts for that write path, which
keeps it testable in a module with no JSX. It returns the whole value, not
just the keys the caller changed: the UI shows four placements but only
changes the named one, so writing only the difference would silently revert
an untouched placement on the next save. Both that and the no-edit of the
input are pinned.
The polling guard shipped with its threshold hardcoded at 3 while
`stallLimit` and `maxTurns` were both configurable, so it was the one number
in this plugin a user could not argue with. The right threshold is a property
of the work: a build that takes five minutes wants a higher limit than a test
that takes five seconds, and the person paying for the tokens is the one who
should set it.

So `pollLimit` joins `stallLimit` everywhere `stallLimit` already was, rather
than only in the place it is used:

- the plugin option, defaulting to 3 and ignoring anything below 1
- a per-session override, read back with the rest of the stored settings
- `/goal poll <n|default>`, reusing `parseCount` so "default", padding and
  nonsense behave exactly as they do for `/goal stall`
- `/goal settings`, which reports the value and whether it came from this
  session or the config
- the help text and the README, including the settings summary

Four existing tests changed shape rather than behaviour, and are worth noting
as such: the resolved object gained a `poll` and an `overridden.poll`, and one
test counted the parenthesised rows in the settings summary, which is now five
instead of four. The count is of row markers, not the words themselves,
precisely so adding a row is a visible change rather than a silent one.
The panel, footer, composer and sidebar stopped updating on their own, and
reopening the panel brought them current again.

The client's rpc.events.on is a single for-await loop with a catch at the end
and nothing that restarts it. Measured against the real library:

  - a handler that throws once ends the loop, so every later event is dropped
    and the error goes to console.error;
  - a stream that ends ends the loop with no error at all, and nothing opens
    a second connection.

Either way the listener stays registered and looks alive while being deaf.
Reopening the panel appeared to fix it only because that path calls rpc.get
directly, bypassing events. The footer, composer and sidebar never fetched
anything themselves, so they had nothing to fall back on.

listen() replaces rpc.events.on. A handler error is reported and the next event
is still delivered. A stream that ends or fails is reopened after a doubling
delay, and onResume runs first so state missed during the gap is re-read.
Handler errors surface as a toast, rate limited; a dropped stream stays quiet
because it heals itself.

The TUI now also:
  - fetches a session's state when a compact area first shows it, so a goal
    already running when the session opens is visible;
  - re-reads every known session after a gap, and announces a goal that ended
    while nothing was listening;
  - polls active goals every 5s, since a half-open stream reports no error to
    reconnect on;
  - guards the terminal toast, so a failing toast cannot end the listener.

listen() has 16 checks. Breaking it in scratch copies (letting handler errors
escape, or never reopening) fails 3 and 7 of them respectively, so they detect
the two failures they exist for. What is not verified is what triggers the
silent stop in a live TUI; only the library behaviour was reproduced.
Add goal_wait, which holds a turn open until a background command finishes
(detected by its output file being closed, since the completion notice only
arrives after the turn ends). Independently, the repeat, stall and polling
guards stand down while a background command is pending, with a back-off
and a 60 minute bound.
@zignd
zignd merged commit f29f9fe into main Sep 29, 2026
@zignd
zignd deleted the fix/blocked-means-goal-not-turn branch September 29, 2026 23:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant