Issue Type Assistant-Related Issue → Latency / Mod...
# office-hours
h
Issue Type Assistant-Related Issue → Latency / Model Behavior Observed Behavior Using GPT-5 mini on Azure OpenAI within VAPI leads to noticeably higher latency compared to previous models Responses are delayed, especially in real-time voice interactions No visible configuration option in VAPI to control or reduce reasoning effort (e.g., minimal/low reasoning) Latency varies across calls at different times of the day even with identical prompts and assistant configuration Expected Behavior The model should support a configurable low or minimal reasoning mode to reduce latency for real-time voice use cases Latency should be consistent and optimized for conversational workflows at all times of the day.