yuki
04/02/2026, 6:01 AMDo you have a few minutes to chat about how we might be able to help your business?
- Output (transcribed audio): Do you have a few minutes to *to* chat about how we might be able to help your business?
It literally repeats words mid-sentence to simulate natural human disfluency.
Is this implemented as:
1. A rule-based text transformation applied to the LLM output before TTS?
2. An additional LLM layer that rewrites the text to add disfluencies?
3. Something handled at the TTS level directly?