Henry Black
03/17/2025, 6:47 PMspeech-update tells us when the assistant starts/stops speaking and model-output gives the text before it is spoken. We could theoretically estimate which word the assistant was speaking but it likely wouldn't be that accurate. The transcript message is received after the sentence is spoken so that wouldn't work either.
I appreciate the detailed responses though