To achieve the lowest latency for your Custom LLM completions endpoint, deploy it in a region geographically closest to your primary users or to the Vapi infrastructure.
For example, if most users are in the US East, choose a US East data center.
>
Tip: Regional deployment is key for reducing latency and ensuring fast response times ([DeepInfra documentation](
https://docs.vapi.ai/providers/model/deepinfra)).
Source:
- [DeepInfra documentation](
https://docs.vapi.ai/providers/model/deepinfra)