Yeah, I’ve run into this before. The delay is usually caused by the system waiting for speech-finalization or the tool pipeline before playing the filler response, even with that setting enabled.
In most cases, the better fix is triggering the filler message immediately when the intent/tool detection starts rather than waiting for the actual tool execution. I can help you narrow it down if you share the platform or voice stack you’re using.
@Isha Malik