But it can understand stuff like guitar notes. That wouldn't be possible if it wasn't taking in pure audio input and pure audio output
Unless their speech to text models just knows how to interpret those sounds and pass a certain type of syntax that the intelligence model & TTS model can understand but that seems harder to pull off then audio to audio. All eleven labs ssml does is affect things like speed and it doesn't even work that well at that