vj
12/19/2025, 7:46 AMVapi Support Bot
12/19/2025, 7:47 AMstartSpeakingPlan as follows:
json
{
"startSpeakingPlan": {
"waitSeconds": 0.1,
"smartEndpointingPlan": {
"provider": "livekit",
"waitFunction": "200 + 1000 * x"
}
}
}
This should help your assistant respond faster to simple confirmations and yes/no answers. For more details, see the [voice pipeline configuration documentation](https://docs.vapi.ai/customization/voice-pipeline-configuration).
Source:
- [Voice pipeline configuration documentation](https://docs.vapi.ai/customization/voice-pipeline-configuration)vj
12/19/2025, 7:57 AMvj
12/23/2025, 7:42 AMvj
12/26/2025, 6:48 AMvj
01/05/2026, 7:46 AMvj
01/19/2026, 7:15 AMChiranjeet Mishra
01/19/2026, 1:35 PMvj
01/23/2026, 6:18 AMvj
01/23/2026, 6:19 AMvj
01/23/2026, 6:20 AMvj
01/23/2026, 6:29 AMvj
01/27/2026, 10:22 AMChiranjeet Mishra
01/27/2026, 10:41 AM019b2624-577c-7334-b05b-8b1cff191f1c
• 019b262b-532a-7778-bf69-aee5e90eabd5
• 019b262e-1501-788d-9cff-22cea616e251
Root Cause Analysis
Based on Vapi's latency architecture and similar documented incidents, the delays likely stem from:
1. Endpointing Latency Spikes
The default onNoPunctuationSeconds is 1500ms, which can cause perceivable delays while waiting to confirm the user finished speaking.
2. LLM Model Latency
Model latency (time to first token) can spike to 1200-1700ms during peak load or provider issues.
3. Knowledge Base Retrieval Timeouts
If KB was configured, the default 15-second timeout can cause silence-timeout triggers.
4. Compounding Effects
The inconsistency of delays (4s, 7s, 8s, 13s, 16s) suggests multiple factors compounding at different points in the pipeline.
Recommendations
1. **Optimize Endpointing**: Use aggressive wait function, reduce onNoPunctuationSeconds to 500-800ms
2. **Set KB Timeout**: Reduce from 15s default to < 5 seconds
3. **Monitor LLM Latency**: Set alerts for >1000ms spikes
4. **Configure Fallback Providers**: For both LLM and TTS
5. **Review Assistant Config**: Check endpointing settings, wait function, custom rules
Next Steps
To fully diagnose, we need to:
1. Review the assistant's configuration for endpointing settings
2. Check infrastructure metrics for Jan 27 around 1:22-1:33 PM
3. Verify Knowledge Base configuration and timeouts if applicable
Let me know if you'd like me to dig deeper into any specific area or if you can share the assistant configuration for these calls.vj
01/27/2026, 1:41 PM