Hi from what you’ve shared, this looks much more like a long-session media/stream stability issue than an LLM latency problem. When calls run past 10–15 minutes, WebRTC or audio sockets can degrade, partially reconnect, or desync, which leads to delayed playback, mismatched TTS, and choppy recordings even though the model and TTS are responding on time. The fact that Cartesia and Deepgram are still firing normally suggests generation is working, but delivery to the call stream isn’t staying stable. I’d focus first on checking for silent reconnects, buffer overflows, keepalive timeouts, or stream resets in your logs, and also verify your max call duration, idle timeouts, and whether your ASR/TTS sessions are being recycled mid-call. Handoffs can help as a workaround, but stabilizing the underlying media session will give you much better reliability long-term. Could you share your current call timeout limits, TTS/ASR providers, and whether you see any reconnect or transport warnings around the 10–15 minute mark?
@jan