Hi,
In terms of RPM, a typical call generates around 3–5 requests per minute. If your calls involve frequent interruptions, this can go up to 8–10 RPM per call.
For TPM, short or simple calls generally consume around 5,000–10,000 tokens per minute, standard calls around 10,000–15,000, and longer or more complex calls can reach 15,000–25,000 TPM. If you're using structured outputs, expect to add roughly 50,000–70,000 TPM on top of those figures per call - it's a significant multiplier.
As a rule of thumb, multiply your per-call estimates by your expected concurrency and add a 1.5x safety buffer. For production campaigns we'd recommend at least OpenAI Tier 2–3, and Tier 5 if you're running 100+ concurrent calls.
These are ballpark estimates - actual usage will vary based on call length and complexity, so it's worth monitoring your OpenAI usage dashboard after go-live and adjusting limits from there.
Let us know if you have any other questions!
Best regards,
Vapi Support