We’ve noticed that with each message in the conversation, the number of tokens keeps increasing. For example, if the initial prompt is around 250 tokens, then each new message adds about 20 tokens for user input and 20 for the assistant’s response. So by the 5th or 6th turn, the total token count becomes quite large.
Is there any way in Vapi to avoid sending the full conversation history every time? Maybe some kind of prompt caching, trimming, or limiting?
We’re just looking for a way to manage the growing token size better and avoid hitting limits or unnecessary costs in longer conversations.