Hey! GPT-5.x is naturally slow for voice, so that ~2s first token is expected.
For real-time, switch to GPT-4o mini, Claude 3.5 Haiku, or Llama 3—they’re much faster and better suited for low-latency voice agents.
If you want, I can help you get sub-1s response time with the right stack
@Cheswick