Repository navigation
fix(voice): anchor stopped_speaking_at on the provider's word timestamps - #7653
aayushbaluni wants to merge 3 commits into
Conversation
Two of the three causes in livekit#7651. Both are in the core turn handling, not in a plugin, and both make `stopped_speaking_at` -- and the `transcription_delay`, `end_of_turn_delay` and `e2e_latency` derived from it -- report when an event arrived rather than when the user stopped talking. END_OF_SPEECH no longer discards the word timestamp. Soniox, Deepgram and AssemblyAI send it with no alternatives and no `speech_end_time`, so the fallback collapsed to `time.time()`, overwriting the anchor the final transcript had just set from the last word's `end_time`. That is late by the provider's endpointing delay (~0.6s with Soniox) and by the whole pause when the mic goes quiet right after the last word, and it made `transcription_delay` always 0 in stt mode. The newest word end for the turn is now kept and preferred over the arrival time; an explicit `speech_end_time` still wins, a turn with no word timestamps anywhere still falls back to `now`, and the anchor is cleared per turn so a turn without a timestamped transcript cannot inherit the previous turn's word end. `stt_node` now measures `start_time_offset` from the same anchor the recognition loop adds back. The loop resolves a word time as `_stt_pipeline.input_started_at + end_time`, and that anchor is stamped on the first frame to reach the pipeline. When the pipeline had no anchor yet -- a new pipeline, which is created on every handoff to an agent that overrides `stt_node` -- the node fell back to the recording or session start, so the time since then was counted twice and every timestamp on that stream landed in the future, where it was clamped to `now`. A pipeline with no anchor yet is about to receive the frame that sets it, so the offset is zero; a reused pipeline still gets the gap between its audio start and now. The third cause in the report, gaps in the mic audio moving provider time out of step with wall clock, is not addressed here. It needs the pipeline to track how much audio it has actually received, which is a larger change than these two.
There was a problem hiding this comment.
Devin Review found 2 potential issues.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
| # and by the whole pause if the mic went quiet after the last word. | ||
| # The transcript for this turn already carried word timestamps, | ||
| # so use that rather than when the message happened to arrive. | ||
| self._last_speaking_time = self._last_stt_word_end_time |
There was a problem hiding this comment.
🟡 Short utterances lose their word-end anchor
When an untimestamped speech-start signal arrives after a short utterance's last word, _last_stt_word_end_time falls before _speech_start_time. The guard rejects that valid word time, so stopped_speaking_at and delay metrics use endpoint arrival instead.
Learn more
Speech-start events without an onset timestamp use their arrival time in _process_stt_event. Both Soniox and Deepgram emit these events without an onset timestamp. On a short utterance, the first signal can arrive after the last word has ended, even with a continuous audio stream and correct provider word timestamps. The comparison treats that normal recognition delay as a clock gap and discards the word-end anchor. The end event then falls back to its later arrival time.
Example: Audio starts at 100.0; a one-word utterance ends at 100.4. The provider's speech-start signal arrives at 100.6, its final word reports end_time=0.4, and END_OF_SPEECH arrives at 101.0. The guard reports 101.0 instead of 100.4.
Recommended fix: Track whether _speech_start_time came from a measured onset or event arrival. Do not use an arrival-based onset to reject a final word timestamp; handle true audio-stream clock gaps separately.
Was this helpful? React with 👍 or 👎 to provide feedback.
`test_transcription_delay_anchor.py` builds AudioRecognition with `__new__`
and sets the attributes the event handlers touch, so the new
`_last_stt_word_end_time` has to be listed there too -- 13 tests failed on
AttributeError without it.
No expectation changed. Those tests set `_last_speaking_time` directly and
never process a transcript carrying a word `end_time`, so the new branch
cannot fire in any of them and END_OF_SPEECH still moves the anchor to the
arrival time there. Added a note to
`test_stt_end_of_speech_without_timestamps_still_anchors_the_turn`, whose
rationale ("arrival time is the only estimate") is exactly what makes it
different from the case in livekit#7651, so the two comments do not read as
contradicting each other.
|
No expectation changed. Every one of those tests sets That includes On how I missed it: I picked test files by name and ran those, which left out the one file whose name does not contain |
… finals Two problems with the previous commit, both found in review. The word anchor is now only preferred when it does not predate the turn's own `_speech_start_time`. After a gap in the mic audio, provider time is behind the wall clock by the length of the gap, so the word end resolves to before the turn started -- the third cause in livekit#7651, which this PR does not fix. Preferring that anchor made `_compute_end_of_turn_metrics` drop all four metrics for predating the turn start, which is worse than the arrival time it replaced. Falling back to `now` there keeps the previous behaviour in a case this change was never meant to touch. The anchor is also recorded from FINAL_TRANSCRIPT only, not from any event carrying word times. A streaming provider revises its interim words, so an interim that reached further than the final it was revised into would win the max and report speech ending later than it did. Finals are additive segments of one turn, so the max over them is also the newest; it stays to tolerate out-of-order delivery. The reset on START_OF_SPEECH is now unconditional -- that event carries no alternatives, so the `has_stt_end_time` arm was dead. Two tests added, and each of the three behaviours is caught by its own test: removing the preference branch fails 3, dropping the onset bound fails `test_word_end_predating_the_turn_falls_back_to_arrival`, and recording from interims fails `test_retracted_interim_word_does_not_move_the_anchor`. Timestamp assertions now pass an explicit `abs=` tolerance. `pytest.approx` defaults to a relative 1e-6, which on a unix timestamp is about +/- 1790 seconds, so the equality checks in the first version barely constrained anything.
|
Both Devin flags were right and both are fixed in e212a34. The first one in particular was a regression I introduced and had not thought through, so thank you for it. 1. Audio gaps erasing the metrics — a real regression, now boundedThis is the third cause in the issue, which I had declared out of scope. What I missed is that my change made it worse rather than leaving it alone. Before: after a mic gap the final transcript set a gap-skewed anchor, and The word end is now only preferred when it does not predate the turn's own onset: elif self._last_stt_word_end_time is not None and (
self._speech_start_time is None
or self._last_stt_word_end_time >= self._speech_start_time
):When provider time and wall clock have drifted apart, that is a signal the stream's clock is behind, and 2. Retracted interim words — the bookkeeping was in the wrong placeCorrect: I recorded the word end wherever Also dropped a dead arm: Checking the guards actually do somethingEach of the three behaviours is caught by its own test — I removed them one at a time:
One thing I got wrong in the first version of the testsMy timestamp assertions used bare 394 passed across the 25 non-plugin test files mentioning any symbol this PR touches; |
chenghao-mou
left a comment
There was a problem hiding this comment.
I am going to close this one for now (see my comment). You are welcome to open another one to address each plugin's reporting.
| # newest word-timestamp-derived end of speech seen in the current turn. | ||
| # END_OF_SPEECH arrives with no alternatives, so this is the only way to | ||
| # keep the provider's own timing instead of the arrival time. | ||
| self._last_stt_word_end_time: float | None = None |
There was a problem hiding this comment.
Thanks for the PR, but I don't think we really need this at framework level. Each STT has the option to report speech_end_time on each SpeechEvent. This is what we do with the StreamAdapter.
The decision on how and when to report it is up to each plugin implementation, so the issue with Soniox, Deepgram, and AssemblyAI not reporting them should be solved in those plugins.
Summary
Fixes two of the three causes in #7651. The report is unusually precise — root causes with line numbers and a repro that needs no API keys — so this PR takes the two that are self-contained logic bugs and leaves the third, which needs a design decision.
Both make
stopped_speaking_atreport when an event arrived instead of when the user stopped talking, andtranscription_delay,end_of_turn_delayand the agent'se2e_latencyare all derived from that anchor.Cause 1 —
END_OF_SPEECHdiscarded the word timestampThe final transcript sets
_last_speaking_timefrom the last word'send_time.END_OF_SPEECHthen overwrote it: Soniox, Deepgram and AssemblyAI send that event with no alternatives and nospeech_end_time, sostt_last_speaking_timecollapses totime.time()._last_stt_word_end_timeholds the newest word-derived end for the current turn. It is set whereverhas_stt_end_timeis already computed, and cleared per turn — in_clear_user_turnand onSTART_OF_SPEECH— so a turn whose transcript carries no word times cannot inherit the previous turn's word end and commit against an anchor that is already seconds in the past.This is the case #4388 reported; that fix only helped when
END_OF_SPEECHcarries a timestamp, which these providers do not send. It is also whytranscription_delaywas always 0 in stt mode: the anchor andnowwere the same instant.Cause 2 — a new STT stream used a different time base than the reader
The recognition loop resolves a word time as
_stt_pipeline.input_started_at + end_time, and that anchor is stamped on the first frame to reach the pipeline.stt_nodehas to measurestart_time_offsetfrom the same anchor, and it did not:A pipeline with no anchor yet is about to receive the frame that sets it, so the offset is zero. Falling back to the session start counted the time since then twice — once in the offset, again when the new pipeline stamped its own anchor — putting every timestamp on that stream in the future, where it was clamped to
now.This is not a corner case: a new pipeline is created on every handoff to an agent that overrides
stt_node(the pipeline is only reused for the default node), so from the second agent onwards every turn reported the arrival time even with a provider that does timestampEND_OF_SPEECH.A reused pipeline still gets the gap between its audio start and now, which is the case the offset exists for.
One thing I could not establish
I could not find why the recording/session-start fallback was introduced — my clone is shallow and
git log -Sdid not reach it. No test pins it (every test setsstart_time_offsetdirectly), andstart_time_offsetis only ever used by plugins to offset their own word times, so the recognition loop is its only consumer. But if that fallback was there for a consumer I have not found, please say so and I will rework it rather than argue the point.Not in this PR
Cause 3, gaps in the mic audio. Provider time only advances while audio is pushed, so after a silent gap
end_time + input_started_atis early by the length of the gap — sometimes before the turn's ownstarted_speaking_at, which makes_compute_end_of_turn_metricsdrop all four metrics. Fixing it means tracking how much audio the pipeline has actually received and anchoring on that instead of wall clock. That is a larger change with its own trade-offs, and it seemed better to keep it out of a PR that already changes two timing paths. Happy to take it as a follow-up.Tests
tests/test_stt_turn_speaking_time.py(5) drives_process_stt_eventdirectly, in the style oftest_audio_recognition_turn_detection.py, andtests/test_stt_node_start_time_offset.py(2) drivesAgent.default.stt_nodewith an STT fake that records the offset it is given.Three of the seven fail without the change, which is how I checked they are measuring it:
elifbranch in cause 1test_end_of_speech_keeps_the_word_timestampandtest_later_word_in_the_same_turn_moves_the_anchor_forwardfail — anchor comes back asnowinstead of the word endtest_new_pipeline_offsets_from_its_own_first_framefails with30.000926971435547 == 0.0 ± 0.5— exactly the session age, the double countThe other four pass in both directions on purpose — they pin behaviour that must not change: an explicit
speech_end_timestill wins, a turn with no word timestamps anywhere still falls back tonow, a new turn does not inherit the previous turn's word end, and a reused pipeline still gets a non-zero offset.Green: 238 passed across
test_audio_recognition_*,test_end_of_turn_metrics,test_e2e_latency_handoff,test_agent_stt_node,test_eou_wait_span,test_audio_turn_detector_fallback,test_agent_sessionand the two new files.ruff checkandruff format --checkclean on all four files. I did not run the plugin suites — they need plugin packages I do not have installed — but nothing here touches a plugin.Closes #7651 partially (causes 1 and 2).