This is actually a big problem for my use case. The main issue is that real-time transcription models are just not very performant.
I've been doing post-conversation data extraction with GPT 4.1 and it's been working like a charm. Unfortunately this means that I can't confirm information with the user on the call, but I just specify in my instructions for them to spell information out slowly.