Hey there, could I please get some feedback from the Vapi team (or community) about an issue we’ve started encountering.
Please feel free to point me to another channel if appropriate.
Im not sure if its been a recent change, or if we’ve only just noticed it because we’re playing around with a non-english assistant, but it looks like the platform is using the STT tooling as the point of truth for a call’s active transcript.
This obviously makes sense for the user’s input, but it seems to be applied to the whole turn, and includes the assistant’s output as well.
Converting from the LLM output, to voice, and then back to text is introducing compounding errors.
This is creating a large number of headaches in our non English assistant, but I’ve also noticed it effects our English assistants as well.