Hi everyone,
I have a couple of questions:
Can VAPI detect emotions or perform tonal analysis of voice input? Or does it rely solely on converting speech to text and then using a language model for inference based on the text?
Are you using a native language model for multimodal speech-to-speech processing, or not?
Thanks in advance for your help!