Also is there any plan for the ability to choose Gemini for generating structured outputs from the audio rather than transcript?
I've built a separate workflow for this, but would be great to have it built in so it can better extract things like tone, flow, sentiment etc compared to just a transcript.