any word if vapi can stream the audio output from ...
# general-english
s
any word if vapi can stream the audio output from gpt-4o instead of text that then uses a different model to create the audio? in my setup with deepgram, gpt-4o, and playht, if playht wasn't needed and the latency wasn't too different for receiving audio instead of text from gpt-4o, then that could shave 350ms off the latency and have total latency down to 700ms