Issue Type
Assistant-Related Issue → Latency / Model Behavior
Observed Behavior
Using GPT-5 mini on Azure OpenAI within VAPI leads to noticeably higher latency compared to previous models
Responses are delayed, especially in real-time voice interactions
No visible configuration option in VAPI to control or reduce reasoning effort (e.g., minimal/low reasoning)
Latency varies across calls at different times of the day even with identical prompts and assistant configuration
Expected Behavior
The model should support a configurable low or minimal reasoning mode to reduce latency for real-time voice use cases
Latency should be consistent and optimized for conversational workflows at all times of the day.