If you set it to assistant-waits-for-user it should speak quickly after the user.
Might be slightly delayed on the first call as the first message waits to be generated.
All calls after that, it will use the saved first message output to minimize latency.
If you make 2 calls you may notice the first message sounds identical. It doesn't get regenerated