Skip to content

fix(voice): anchor stopped_speaking_at on the provider's word timestamps - #7653

Closed
aayushbaluni wants to merge 3 commits into
livekit:mainfrom
aayushbaluni:fix/7651-stt-stopped-speaking-at
Closed

aayushbaluni wants to merge 3 commits into
livekit:mainfrom
aayushbaluni:fix/7651-stt-stopped-speaking-at

Conversation

@aayushbaluni

Copy link
Copy Markdown

Summary

Fixes two of the three causes in #7651. The report is unusually precise — root causes with line numbers and a repro that needs no API keys — so this PR takes the two that are self-contained logic bugs and leaves the third, which needs a design decision.

Both make stopped_speaking_at report when an event arrived instead of when the user stopped talking, and transcription_delay, end_of_turn_delay and the agent's e2e_latency are all derived from that anchor.

Cause 1 — END_OF_SPEECH discarded the word timestamp

The final transcript sets _last_speaking_time from the last word's end_time. END_OF_SPEECH then overwrote it: Soniox, Deepgram and AssemblyAI send that event with no alternatives and no speech_end_time, so stt_last_speaking_time collapses to time.time().

if ev.speech_end_time is not None:
    self._last_speaking_time = min(ev.speech_end_time, now)
+elif self._last_stt_word_end_time is not None:
+    self._last_speaking_time = self._last_stt_word_end_time
else:
    self._last_speaking_time = stt_last_speaking_time

_last_stt_word_end_time holds the newest word-derived end for the current turn. It is set wherever has_stt_end_time is already computed, and cleared per turn — in _clear_user_turn and on START_OF_SPEECH — so a turn whose transcript carries no word times cannot inherit the previous turn's word end and commit against an anchor that is already seconds in the past.

This is the case #4388 reported; that fix only helped when END_OF_SPEECH carries a timestamp, which these providers do not send. It is also why transcription_delay was always 0 in stt mode: the anchor and now were the same instant.

Cause 2 — a new STT stream used a different time base than the reader

The recognition loop resolves a word time as _stt_pipeline.input_started_at + end_time, and that anchor is stamped on the first frame to reach the pipeline. stt_node has to measure start_time_offset from the same anchor, and it did not:

_audio_input_started_at: float = (
    activity._audio_recognition._input_started_at
    if ... is not None
    else (recording_started_at or session._started_at or time.time())   # <- different anchor
)
stream.start_time_offset = time.time() - _audio_input_started_at

A pipeline with no anchor yet is about to receive the frame that sets it, so the offset is zero. Falling back to the session start counted the time since then twice — once in the offset, again when the new pipeline stamped its own anchor — putting every timestamp on that stream in the future, where it was clamped to now.

This is not a corner case: a new pipeline is created on every handoff to an agent that overrides stt_node (the pipeline is only reused for the default node), so from the second agent onwards every turn reported the arrival time even with a provider that does timestamp END_OF_SPEECH.

A reused pipeline still gets the gap between its audio start and now, which is the case the offset exists for.

One thing I could not establish

I could not find why the recording/session-start fallback was introduced — my clone is shallow and git log -S did not reach it. No test pins it (every test sets start_time_offset directly), and start_time_offset is only ever used by plugins to offset their own word times, so the recognition loop is its only consumer. But if that fallback was there for a consumer I have not found, please say so and I will rework it rather than argue the point.

Not in this PR

Cause 3, gaps in the mic audio. Provider time only advances while audio is pushed, so after a silent gap end_time + input_started_at is early by the length of the gap — sometimes before the turn's own started_speaking_at, which makes _compute_end_of_turn_metrics drop all four metrics. Fixing it means tracking how much audio the pipeline has actually received and anchoring on that instead of wall clock. That is a larger change with its own trade-offs, and it seemed better to keep it out of a PR that already changes two timing paths. Happy to take it as a follow-up.

Tests

tests/test_stt_turn_speaking_time.py (5) drives _process_stt_event directly, in the style of test_audio_recognition_turn_detection.py, and tests/test_stt_node_start_time_offset.py (2) drives Agent.default.stt_node with an STT fake that records the offset it is given.

Three of the seven fail without the change, which is how I checked they are measuring it:

Removing Result
the elif branch in cause 1 test_end_of_speech_keeps_the_word_timestamp and test_later_word_in_the_same_turn_moves_the_anchor_forward fail — anchor comes back as now instead of the word end
the offset change in cause 2 test_new_pipeline_offsets_from_its_own_first_frame fails with 30.000926971435547 == 0.0 ± 0.5 — exactly the session age, the double count

The other four pass in both directions on purpose — they pin behaviour that must not change: an explicit speech_end_time still wins, a turn with no word timestamps anywhere still falls back to now, a new turn does not inherit the previous turn's word end, and a reused pipeline still gets a non-zero offset.

Green: 238 passed across test_audio_recognition_*, test_end_of_turn_metrics, test_e2e_latency_handoff, test_agent_stt_node, test_eou_wait_span, test_audio_turn_detector_fallback, test_agent_session and the two new files. ruff check and ruff format --check clean on all four files. I did not run the plugin suites — they need plugin packages I do not have installed — but nothing here touches a plugin.

Closes #7651 partially (causes 1 and 2).

Two of the three causes in livekit#7651. Both are in the core turn handling, not in
a plugin, and both make `stopped_speaking_at` -- and the
`transcription_delay`, `end_of_turn_delay` and `e2e_latency` derived from it
-- report when an event arrived rather than when the user stopped talking.

END_OF_SPEECH no longer discards the word timestamp. Soniox, Deepgram and
AssemblyAI send it with no alternatives and no `speech_end_time`, so the
fallback collapsed to `time.time()`, overwriting the anchor the final
transcript had just set from the last word's `end_time`. That is late by the
provider's endpointing delay (~0.6s with Soniox) and by the whole pause when
the mic goes quiet right after the last word, and it made
`transcription_delay` always 0 in stt mode. The newest word end for the turn
is now kept and preferred over the arrival time; an explicit
`speech_end_time` still wins, a turn with no word timestamps anywhere still
falls back to `now`, and the anchor is cleared per turn so a turn without a
timestamped transcript cannot inherit the previous turn's word end.

`stt_node` now measures `start_time_offset` from the same anchor the
recognition loop adds back. The loop resolves a word time as
`_stt_pipeline.input_started_at + end_time`, and that anchor is stamped on
the first frame to reach the pipeline. When the pipeline had no anchor yet --
a new pipeline, which is created on every handoff to an agent that overrides
`stt_node` -- the node fell back to the recording or session start, so the
time since then was counted twice and every timestamp on that stream landed
in the future, where it was clamped to `now`. A pipeline with no anchor yet
is about to receive the frame that sets it, so the offset is zero; a reused
pipeline still gets the gap between its audio start and now.

The third cause in the report, gaps in the mic audio moving provider time
out of step with wall clock, is not addressed here. It needs the pipeline to
track how much audio it has actually received, which is a larger change than
these two.
@aayushbaluni
aayushbaluni requested a review from a team as a code owner October 7, 2026 09:32

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 2 potential issues.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

# and by the whole pause if the mic went quiet after the last word.
# The transcript for this turn already carried word timestamps,
# so use that rather than when the message happened to arrive.
self._last_speaking_time = self._last_stt_word_end_time

@devin-ai-integration devin-ai-integration Bot Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Short utterances lose their word-end anchor

When an untimestamped speech-start signal arrives after a short utterance's last word, _last_stt_word_end_time falls before _speech_start_time. The guard rejects that valid word time, so stopped_speaking_at and delay metrics use endpoint arrival instead.

Learn more

Speech-start events without an onset timestamp use their arrival time in _process_stt_event. Both Soniox and Deepgram emit these events without an onset timestamp. On a short utterance, the first signal can arrive after the last word has ended, even with a continuous audio stream and correct provider word timestamps. The comparison treats that normal recognition delay as a clock gap and discards the word-end anchor. The end event then falls back to its later arrival time.

Example: Audio starts at 100.0; a one-word utterance ends at 100.4. The provider's speech-start signal arrives at 100.6, its final word reports end_time=0.4, and END_OF_SPEECH arrives at 101.0. The guard reports 101.0 instead of 100.4.

Recommended fix: Track whether _speech_start_time came from a measured onset or event arrival. Do not use an arrival-based onset to reject a final word timestamp; handle true audio-stream clock gaps separately.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

Comment thread livekit-agents/livekit/agents/voice/audio_recognition.py Outdated
`test_transcription_delay_anchor.py` builds AudioRecognition with `__new__`
and sets the attributes the event handlers touch, so the new
`_last_stt_word_end_time` has to be listed there too -- 13 tests failed on
AttributeError without it.

No expectation changed. Those tests set `_last_speaking_time` directly and
never process a transcript carrying a word `end_time`, so the new branch
cannot fire in any of them and END_OF_SPEECH still moves the anchor to the
arrival time there. Added a note to
`test_stt_end_of_speech_without_timestamps_still_anchors_the_turn`, whose
rationale ("arrival time is the only estimate") is exactly what makes it
different from the case in livekit#7651, so the two comments do not read as
contradicting each other.
@aayushbaluni

Copy link
Copy Markdown
Author

unit-tests was red on the first push and that one was mine — pushed 34a24f8.

tests/test_transcription_delay_anchor.py builds AudioRecognition with __new__ and sets the attributes the event handlers touch, so the new _last_stt_word_end_time had to be listed there as well. 13 tests failed on AttributeError, not on an expectation.

No expectation changed. Every one of those tests sets _last_speaking_time directly and none processes a transcript carrying a word end_time, so the new branch cannot fire in any of them — END_OF_SPEECH still moves the anchor to the arrival time there.

That includes test_stt_end_of_speech_without_timestamps_still_anchors_the_turn, which is the test closest to this change and deliberately pins "an explicit endpointing signal moves the anchor". I want to be clear that I did not quietly weaken it: its premise is that arrival time is the only estimate available, which is true when no transcript in the turn was timestamped. The change only applies when one was. I added a note to its docstring pointing at the new file so the two rationales do not read as contradicting each other, and I would rather hear it now if you think the ordering should be different — my ordering is speech_end_time > word end from this turn > arrival time.

On how I missed it: I picked test files by name and ran those, which left out the one file whose name does not contain audio_recognition, stt_node or end_of_turn. Re-ran properly this time — every non-plugin test file that mentions any symbol this PR touches, 25 files, 392 passed. (6 failures in test_chat_item_proto_coverage.py are from a missing telemetry dependency in my environment: they fail identically on unmodified HEAD~1 and they are green in your CI.)

… finals

Two problems with the previous commit, both found in review.

The word anchor is now only preferred when it does not predate the turn's
own `_speech_start_time`. After a gap in the mic audio, provider time is
behind the wall clock by the length of the gap, so the word end resolves to
before the turn started -- the third cause in livekit#7651, which this PR does not
fix. Preferring that anchor made `_compute_end_of_turn_metrics` drop all
four metrics for predating the turn start, which is worse than the arrival
time it replaced. Falling back to `now` there keeps the previous behaviour
in a case this change was never meant to touch.

The anchor is also recorded from FINAL_TRANSCRIPT only, not from any event
carrying word times. A streaming provider revises its interim words, so an
interim that reached further than the final it was revised into would win
the max and report speech ending later than it did. Finals are additive
segments of one turn, so the max over them is also the newest; it stays to
tolerate out-of-order delivery.

The reset on START_OF_SPEECH is now unconditional -- that event carries no
alternatives, so the `has_stt_end_time` arm was dead.

Two tests added, and each of the three behaviours is caught by its own test:
removing the preference branch fails 3, dropping the onset bound fails
`test_word_end_predating_the_turn_falls_back_to_arrival`, and recording from
interims fails `test_retracted_interim_word_does_not_move_the_anchor`.

Timestamp assertions now pass an explicit `abs=` tolerance. `pytest.approx`
defaults to a relative 1e-6, which on a unix timestamp is about +/- 1790
seconds, so the equality checks in the first version barely constrained
anything.
@aayushbaluni

Copy link
Copy Markdown
Author

Both Devin flags were right and both are fixed in e212a34. The first one in particular was a regression I introduced and had not thought through, so thank you for it.

1. Audio gaps erasing the metrics — a real regression, now bounded

This is the third cause in the issue, which I had declared out of scope. What I missed is that my change made it worse rather than leaving it alone.

Before: after a mic gap the final transcript set a gap-skewed anchor, and END_OF_SPEECH overwrote it with now — late, but after started_speaking_at, so _compute_end_of_turn_metrics still produced numbers. After my first commit: the gap-skewed word end survived to the commit, predated the turn start, and all four metrics were dropped. Wrong numbers became no numbers, in exactly the scenario the reporter's input_gap repro covers.

The word end is now only preferred when it does not predate the turn's own onset:

elif self._last_stt_word_end_time is not None and (
    self._speech_start_time is None
    or self._last_stt_word_end_time >= self._speech_start_time
):

When provider time and wall clock have drifted apart, that is a signal the stream's clock is behind, and now is the better of two bad answers. Cause 3 is still unfixed, but it is no longer made worse here.

2. Retracted interim words — the bookkeeping was in the wrong place

Correct: I recorded the word end wherever has_stt_end_time held, which includes interim and preflight events, and max() then preserved a retracted later timestamp. Moved into the FINAL_TRANSCRIPT branch so only committed words count. Finals are additive segments of one turn, so the max over them is also the newest — it stays to tolerate out-of-order delivery, not to merge guesses.

Also dropped a dead arm: START_OF_SPEECH carries no alternatives, so the has_stt_end_time condition on that reset could never be true.

Checking the guards actually do something

Each of the three behaviours is caught by its own test — I removed them one at a time:

Removed Fails
the preference branch test_end_of_speech_keeps_the_word_timestamp, test_later_word_in_the_same_turn_moves_the_anchor_forward, test_retracted_interim_word_does_not_move_the_anchor
the _speech_start_time bound test_word_end_predating_the_turn_falls_back_to_arrival
finals-only recording test_retracted_interim_word_does_not_move_the_anchor

One thing I got wrong in the first version of the tests

My timestamp assertions used bare pytest.approx, whose default tolerance is a relative 1e-6 — on a unix timestamp that is about ±1790 seconds, so assert anchor == approx(expected) was passing on practically any value. The only assertion with teeth was an incidental < comparison. Every timestamp assertion now carries an explicit abs=0.05, and the fixture's clock is anchored to the real time.time() rather than a fake epoch, so the onset bound is exercised on a realistic timeline.

394 passed across the 25 non-plugin test files mentioning any symbol this PR touches; ruff check and ruff format --check clean.

@chenghao-mou chenghao-mou left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I am going to close this one for now (see my comment). You are welcome to open another one to address each plugin's reporting.

# newest word-timestamp-derived end of speech seen in the current turn.
# END_OF_SPEECH arrives with no alternatives, so this is the only way to
# keep the provider's own timing instead of the arrival time.
self._last_stt_word_end_time: float | None = None

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the PR, but I don't think we really need this at framework level. Each STT has the option to report speech_end_time on each SpeechEvent. This is what we do with the StreamAdapter.

The decision on how and when to report it is up to each plugin implementation, so the issue with Soniox, Deepgram, and AssemblyAI not reporting them should be solved in those plugins.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

STT turn detection: stopped_speaking_at ignores word timestamps, and is wrong after a handoff or a gap in mic audio

2 participants