I don't think vapi is streaming the audio from gpt-4o yet because if changing the Voice to open ai for TTS it adds a further 500ms of latency over what I previously had with payht vs if it were receiving an audio stream instead of a text stream and thus did not need to do TTS, then it would be a bunch less