Skip to content

feat: add multi-turn answering-machine detection - #6202

Open
chenghao-mou wants to merge 31 commits into
mainfrom
chenghao/feat/AGT-3038-call-screening-support-in-amd
Open

chenghao-mou wants to merge 31 commits into
mainfrom
chenghao/feat/AGT-3038-call-screening-support-in-amd

Conversation

@chenghao-mou

@chenghao-mou chenghao-mou commented Jun 24, 2026 •

Copy link
Copy Markdown
Member

AMD handles screening, voicemail, and IVR across turns using rolling history and successful DTMF calls. WAIT suppresses replies during announcements; stale predictions are discarded.

A category FSM recommends next stages; another category corrects the stage. AMD owns deadlines and reply authorization and supports realtime replies. The classifier prompt is tuned on scripted and real calls across several models, and the default classifier is now google/gemma-4-31b-it. See design and usage.

Addresses AGT-3038. Adopts @karan-dhir's screening proposal (source issue) and @eashwar-mp's rolling-history guidance.

@chenghao-mou
chenghao-mou force-pushed the chenghao/feat/AGT-3038-call-screening-support-in-amd branch from 3173918 to 14d36f4 Compare July 24, 2026 12:00
@chenghao-mou
chenghao-mou force-pushed the chenghao/feat/AGT-3038-call-screening-support-in-amd branch from 94f4e05 to 7dfc05f Compare August 8, 2026 21:51

@eashwar-mp eashwar-mp left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice work, this lands right where we need it. Sharing a few things from running outbound AMD (screeners + voicemail) in production that might save some pain, especially around trusting the LLM for the screening category. Happy to dig into any of these.

Input: "Please state your name and why you're calling, and I will check if the person is available"
Output: machine-ivr
Note: this should apply for any call screening prompts.
Output: machine-screening

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@chenghao-mou from running outbound screening in production, this category is genuinely non-deterministic on the LLM alone. We saw the same Google Call Screen greeting classified as machine-vm once and machine-ivr twice across 11 calls, and a single example won't cover the real spread. Two things that helped us: a lot more examples, and a deterministic transcript substring pre-gate before trusting the LLM (matched at the start of the greeting so STT truncation doesn't break it). The failure if it's wrong: a screener tagged machine-vm plays the voicemail message and ends, so you leave a voicemail on a call screen instead of reaching the human.

Screener markers we substring-match on:
call assist, by google, automatically screening, record your name, reason for calling, asking for more information, if this person is available, see if this person is available

And the phrases we use to detect a screener rolling over to message-taking:
leave a message, leave me a message, leave an additional message, leave your message, record your message, after the tone, after the beep, at the tone, at the beep, forwarded to voicemail, reached the voicemail (plus combos like take the call + add anything else). Happy to share the full lists as examples to fold into the prompt.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the feedback. Just to make sure I understand this scope when you say "the same Google Call Screen greeting classified as machine-vm once and machine-ivr twice across 11 calls, and a single example won't cover the real spread." do you mean running the existing no-screening version with the default AMD LLM?

@eashwar-mp eashwar-mp Aug 24, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, exactly, this was the pre-#6202 build with no machine-screening category, running the default auto-selected AMD LLM (gemini-3.1-flash-lite + ink-whisper, we pass no explicit llm/stt). So a screener could only surface as machine-ivr (which we then sub-classify screener-vs-keypad by substring) or misfire to machine-vm.

What we saw: across a small sample, the identical "Colossus by Google" greeting came back machine-ivr twice and machine-vm once, with reason=llm on the verdict, and roughly 1 in 11 screener calls misfired to a non-screener category. Confirmed against the CallState telemetry and the STT transcripts, same greeting text each time.

Fair caveat: this was without your new dedicated category and its prompt example, so that should help. My concern is just that the same greeting doesn't always get the same verdict, so one example won't fully fix it on its own.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for confirming. We now include Google call screening support in scope and will evaluate its performance before release.

inference gateway; otherwise it falls back to the session's own
LLM.
instance or an inference model string (e.g. ``"openai/gpt-4.1-mini"``).
When omitted, AMD reuses the session's own LLM (``session.llm``).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is dropping the auto-selected gemini-flash-lite / ink-whisper defaults intended as part of this PR? It's a behavior change for anyone relying on auto-select, and reusing the session's main LLM for classification can be slower and, in our experience, less reliable on the vm-vs-screening call specifically. Might be worth calling out in the description.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That was a detour during the initial design. We now have the following rank:

  • explicit values
  • omitted and credentials present: our chosen models
  • None or otherwise: fall back to session or agent settings

classifier.arm_no_speech_timer()
# remember the screening playback so a later human verdict reports it
self._last_playback = playback
continue

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a cap on screening turns here? Each machine-screening verdict replays screening_message and re-arms the detection timer, so a chatty screener or a flaky re-classification could loop the message with no upper bound. A max-iterations guard might be worth adding.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Current design: AMD has a hard timeout from the start of listening, defaulting to 120 seconds, which new turns cannot extend. There is no screening-specific turn cap, so extended exchanges remain possible within that time limit (there is also no fixed screening message now that Google call screening support is added)

self._arm_eot_timer()
# pick session stt when it arrives before AMD stt
if self._source == "amd_stt" and source == "stt" and not self._amd_stt_seen:
logger.warning("amd: session STT won the transcript race, using session STT")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: this fires on the normal dual-STT path whenever session STT wins the race, which is expected. Might read better as debug so it's not warning-level noise.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. That warning is gone. When session STT wins, AMD selects it and closes the dedicated stream without logging a warning. An actual AMD STT failure still logs a warning before falling back.

logger.info("playing screening message")
await classifier.reset()
classifier.start_listening()
self._session._on_aec_warmup_expired()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the classifier keeps listening while screening_message plays, is there a risk AMD transcribes the agent's own message as the callee's next greeting? Looks like _on_aec_warmup_expired is meant to cover that, just want to confirm the playout can't poison the next turn's classification.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD receives participant input audio, and its classifier history excludes the Agent’s replies. That prevents direct inclusion of our output, but it does not rule out echo returning through the participant’s audio. The same can be said for normal agent sessions as well. This also depends on device, environment, and many other factors where typically you already have echo cancellation embedded in the browser or microphones.

AMDCategory.HUMAN: agent_pb.AmdCategory.AMD_HUMAN,
AMDCategory.MACHINE_IVR: agent_pb.AmdCategory.AMD_MACHINE_IVR,
# TODO: @chenghao-mou The remote-session protocol does not yet represent screening.
AMDCategory.MACHINE_SCREENING: agent_pb.AmdCategory.AMD_UNKNOWN,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Worth flagging for remote-session users: screening collapses to AMD_UNKNOWN here, and screening_detected/message_playback aren't on the proto, so the whole signal is invisible over the wire. Fine as a follow-up given the TODO, just noting the gap.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the gap won't be closed until we ship core protobuf changes. It will be ready when releasing.

if verdict.is_machine and self._interrupt_on_machine:
await self._session.interrupt(force=True)

if verdict.category == AMDCategory.MACHINE_VM:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One case from our screener handling worth considering: Google Call Screen often rolls over to message-taking with the decline split across several short STT turns ("okay", "they can't take the call", "feel free to leave a message"). Since each turn is classified independently after reset(), a rollover spread across turns may never produce a single transcript the LLM tags machine-vm, so the rollover-to-voicemail gets missed and the call just times out. We catch it by matching a voicemail-takeover over a rolling multi-turn window.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new design does include recent messages so the classification has more context about the current stage.

if self._silence_timer is not None:
self._silence_timer.cancel()
self._silence_timer = asyncio.get_running_loop().call_later(
self._human_silence_threshold,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Heads-up from a prod bug (we ended up widening human_silence from 0.5 to 1.0): a voicemail greeting whose start VAD misses runs through this synthesized bracket on the human_silence window. With the shorter default there's a risk the silence timer fires before a slow classification lands. Probably fine since this waits for the verdict rather than pre-baking human, but flagging since the short window bit us.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The synthesized human-silence is dropped. AMD now classifies at committed end of turn and waits for the prediction or its inference deadline.

Classify at client-side EOT across screening, voicemail, and IVR turns.
Keep reply control in AgentSession and expose separate prediction and
completion events. Race optional STT for AMD without changing Agent text.

Addresses AGT-3038

Co-authored-by: karan-dhir <karan@woflow.com>
Co-authored-by: eashwar-mp <eashwar.perumal@salesai.com>
@chenghao-mou
chenghao-mou force-pushed the chenghao/feat/AGT-3038-call-screening-support-in-amd branch from 7dfc05f to 7d3110f Compare September 11, 2026 11:59
@chenghao-mou chenghao-mou added the review_effort:high Needs broad context and substantial validation label Sep 11, 2026
@chenghao-mou chenghao-mou changed the title feat(amd): support call screening feat: add multi-turn answering-machine detection Sep 11, 2026
Machine stages require 1.5 seconds of continuous participant silence before predictions or replies can be released. Human and initial uncertain results keep normal EOT timing. Apply the same rule to machine-stage fallbacks and cancel pending releases when speech resumes.

Addresses AGT-3038
Keep AMD reply gating intact after the turn-completion tracing refactor. Preserve AMD turn IDs and trace cleanup.

Addresses AGT-3038
Select the default AMD models when llm or stt is omitted and LiveKit Cloud credentials are available. Explicit None inherits the active Agent LLM or session transcript.

Addresses AGT-3038
The detector mixed stage policy, turn bookkeeping, and asyncio timers in
one class, so every rule needed an event loop to test. Move policy into
_AMDFSM with injected clocks and returned events; the detector keeps I/O.

Delete the unused legacy classifier, move classification to a tool call,
and give AgentActivity _is_busy and _cancel_pending_replies so AMD stops
reading its private fields.
Turn state was a set of nullable fields whose legal combinations were
implicit, so deadline handling matched floats and idle checks probed
three fields. Each turn now has one phase (idle, inferring, holding) and
one deadline, and the FSM applies due deadlines in order.

Replace the six-way category handler table with one stage transition,
return a ReplyDecision instead of overloading AMDCategory, type reasons
as AMDReason, and build the classifier request as a model.

Remove the unused should_wait and detection_delay fields.
The detector held its agent, model, completion future, and open transcript
as Optionals that were None only before __aenter__, so every use asserted.
Group them in a _Run created on enter, with one guard at the boundary.

Task and listener failures now complete the run as internal_error instead
of inference_error, so callers can tell a model failure from a bug.

Split transcript_ready so superseding older work and building the new
request are separate steps.
Reduce coupling between AMD policy and speech scheduling. Keep each turn's
reply guard valid after AMD detaches, including while a customer hook runs.

Addresses AGT-3038
Remove the public user_turn_committed event and session turn counter.
AMD returns a guard bound to each turn before the customer hook runs, so
pending replies keep their original decision after newer turns arrive.

Addresses AGT-3038
Route model and fallback predictions through the same release path.
Inline turn handling and document the FSM and prediction reasons.
Keep saved decisions intact across late results and silent supersession.

Addresses AGT-3038
Use TurnHooks for committed user input, reply authorization, and accepted agent output. Keep reply tool preparation separate so preemptive attempts do not commit an agent turn.

Return turn-bound hooks at commit so delayed customer callbacks retain their AMD decision.

Addresses AGT-3038
The FSM kept a phase per turn and patched predictions across turns to
wake waiters, so stale work lingered and empty turns needed special
cases. Keep one inference and one hold. A new transcript supersedes
both, an empty turn reuses them through a pointer, and a late model
result retargets a held REUSED prediction instead of being dropped.

Speech after the last commit freezes a hold until the next speech end
or commit. finish() no longer fabricates predictions; waiters exit on
completion. The detector arms one timer for the FSM's next deadline.
Menu extraction and voicemail tracking are FSM decisions.

UserStateChangedEvent gains speech_timestamp so created_at keeps the
event-creation convention. Rename AMD.notify_dtmf_sent to on_dtmf_event.

Addresses AGT-3038
Keep category transitions in the FSM and manage inference, deadlines,
and reply authorization in AMD. Add WAIT, bounded classifier history,
and successful DTMF context. Discard stale predictions and restore
interruption settings after detection.

Preserve speech timestamps, roll back partial setup, and cover pipeline
and realtime replies with regression tests.

Addresses AGT-3038
@chenghao-mou
chenghao-mou marked this pull request as ready for review September 18, 2026 10:15
@chenghao-mou
chenghao-mou requested a review from a team as a code owner September 18, 2026 10:15
devin-ai-integration[bot]

This comment was marked as resolved.

Create owned model clients after entry validation and close them when setup
fails. Keep provider and listener exception content out of AMD logs and
DTMF tool errors.

Addresses AGT-3038
Keep event-creation time in created_at. Direct speech-timing consumers to
speech_timestamp, with created_at as the fallback when the boundary is unknown.
This applies to sessions with or without AMD.

Addresses AGT-3038
Successful DTMF sends keep their result in history and wait for the next prompt. Use ToolResult from #7340 so the tool controls reply generation.
devin-ai-integration[bot]

This comment was marked as resolved.

Hide end_call while AMD is active. Wait 10 seconds after voicemail playback for a post-message menu, then end the call from the completed AMD result unless a human has answered.
devin-ai-integration[bot]

This comment was marked as resolved.

Let later transcripts correct machine classifications while preserving completed voicemail and DTMF actions.

Document the tradeoff between preserving user hooks and delaying realtime audio commits.
devin-ai-integration[bot]

This comment was marked as resolved.

The classifier often returned the right category with a malformed
correction flag or an empty quote, and AMD rejected the answer. In eval
runs, every rejected answer had the correct category.

The model now returns only a category. A category outside the stage's
recommended next categories corrects the stage, and
AMDPredictionEvent.corrects_stage is derived from the FSM.
correction_evidence is removed.
CLASSIFY_PROMPT was tuned with GEPA on scripted calls and real recorded
calls. It decides person or machine from signs of automation, treats
holds and notices as wait, treats a request for a name or reason as
screening, and keeps carrier "not available" openers from ending the
call.

google/gemma-4-31b-it was the most accurate and the fastest classifier
in these evals, so AMD now auto-selects it.
The Gemma-tuned prompt from the previous commit did not hold up on the
pooled test splits (107 scripts, paired). Restore the earlier tuned
prompt: Gemma is about as accurate (88% vs 86%), flash-lite gains 6.5
points (81% vs 74%), and the prompt is 260 words shorter.

@devin-ai-integration devin-ai-integration Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note

Newer findings are available below. Devin Review posted a newer report on this PR, in addition to the findings presented here.

Devin Review found 3 new potential issues.

2 flags not posted on this PR by your GitHub settings β€” view them in Devin Review. (Configure)

Devin Review

Comment on lines +58 to +60
When the latest words ask nothing, such as one moment, please hold, stay on the line,
let me check, a thank you or confirmation after a sent digit, a recording notice, or
routing, return wait inside a machine stage, not machine-screening or machine-ivr. Before any machine stage, use wait only for an unmistakable

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

πŸ”΄ Voicemail message lost at the tone

When a voicemail request continues with a separate tone announcement, wait suppresses the message if playback has not started. The agent never records its voicemail message.

Learn more

A voicemail request puts AMD in the voicemail stage, where a reply can deliver one recorded message. If the greeting is split into turns, a subsequent tone or recording instruction can arrive before the reply clears the silence threshold. The next turn cancels the pending reply in _on_user_turn_committed. Classifying that continuation as wait makes _authorize_reply refuse to speak, although no message has played.

Example: The recording says β€œPlease leave a message,” pauses, then says β€œSpeak after the tone.” If the second turn arrives before the agent starts speaking, AMD retains the voicemail stage but predicts wait, and sends no message.

Recommended fix: Distinguish voicemail continuation and tone instructions from holds when the voicemail message has not played. Preserve machine-vm for the remaining greeting unless a required menu choice or a real person takes over.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

Comment on lines +58 to +60
When the latest words ask nothing, such as one moment, please hold, stay on the line,
let me check, a thank you or confirmation after a sent digit, a recording notice, or
routing, return wait inside a machine stage, not machine-screening or machine-ivr. Before any machine stage, use wait only for an unmistakable

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

πŸ”΄ Human greeting silenced after machine handoff

After a screener or IVR transfer, a person's plain β€œHello” asks nothing, so wait can suppress the reply. The call remains in the machine stage instead of recognizing the new speaker.

Learn more

A wait prediction retains the current machine stage and prevents an agent reply through _authorize_reply. The instruction to return wait for the latest words asking nothing does not exempt a plain greeting from a new human. A transfer to a person often begins with just β€œHello,” so the classifier can keep waiting even though the caller can now converse.

Example: An IVR says β€œTransferring you now” and then an employee answers β€œHello.” The second turn asks nothing, and a wait prediction leaves the agent silent instead of responding to the employee.

Recommended fix: Explicitly exempt conversational greetings and human takeover from the machine-stage wait rule. For a plain hello with insufficient evidence, prefer uncertain or human, not wait.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

Comment on lines +53 to +55
Return machine-screening, machine-vm, or machine-ivr only when an automated system asks
the caller for something: a name or reason, a message, or a menu choice. Menu
instructions after voicemail can be machine-ivr. A turn that asks for a name or reason

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟑 Optional voicemail key hijacks message delivery

A voicemail greeting offering β€œpress 1 for more options” can become machine-ivr, although the key is optional. The agent receives menu instructions instead of delivering its voicemail message.

Learn more

The classifier can move directly from a voicemail stage to IVR when it sees a menu choice. An optional key advertised in a voicemail greeting is not a required step for leaving a message. Moving to IVR gives the reply _DEFAULT_IVR_INSTRUCTIONS and enables the DTMF tool in _on_reply_generation, rather than supplying voicemail instructions.

Example: A recording says β€œLeave a message after the tone; press 1 for more options.” The optional 1 can be interpreted as an IVR choice, and the agent navigates options rather than recording its message.

Recommended fix: Restrict voicemail-to-IVR transitions to choices required to proceed, such as re-recording or sending a message. Keep optional keys within the voicemail stage.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

AMD now returns the reply decision when the turn's prediction arrives.
The stage instructions go into the reply context at that time.
Reply generation starts during the machine silence hold.
Playback and tool execution still wait until the hold ends.
The AMD timer resumes playback authorization when the hold ends.
New participant speech extends the hold, as before.
A newer committed turn cancels a reply that still waits.

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 new potential issue.

2 flags not posted on this PR by your GitHub settings β€” view them in Devin Review. (Configure)

Devin Review

Comment on lines +850 to +853
if self._reply_held_at(time.monotonic()):
self._held_turn_id = turn_id
else:
activity._resume_authorization()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟑 Realtime replies miss early generation

For realtime agents, _should_reply keeps authorization closed during the machine silence hold. _realtime_reply_task waits on that gate before generating, so model latency starts after silence ends.

Learn more

The authorization event gates different operations in the two agent paths. The pipeline starts its LLM and TTS work before waiting for authorization in _pipeline_reply_task_impl. The realtime path instead waits on authorization before calling generate_reply in _realtime_reply_task. Keeping the event closed protects playback but also prevents realtime generation from overlapping the machine silence threshold.

Example: A screening prediction arrives 0.2 seconds after speech ends with 1.3 seconds of silence remaining. A pipeline reply starts its LLM immediately; a realtime reply waits 1.3 seconds before asking the provider to generate anything.

Recommended fix: Separate realtime generation permission from playback authorization, or arrange for realtime responses to be generated during the hold and buffered until release. Preserve the existing cancellation of superseded replies and the no-playback-before-silence constraint.

Devin Review


Was this helpful? React with πŸ‘ or πŸ‘Ž to provide feedback.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

review_effort:high Needs broad context and substantial validation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants