Hey folks - quick infra question, not a support issue.
We’re building a voice AI speech layer focused on high-quality output, low perceived latency, and stable cost behavior under concurrency for real-time agents.
Curious if teams here ever evaluate or swap voice / TTS layers underneath agent orchestration as usage scales, and whether short technical pilots make sense when voice quality or latency becomes a bottleneck.
Happy to clarify context if useful.