That's right yeah. I would love to see a feature in the future that can "escalate" audio segments to more powerful transcription models and update the conversation context.
The way this would work here is you ask the user for their email, they respond with something pseudo-reasonable and you log it, then later on after the more powerful model has completed transcription and extraction, you would inject a tool message or something into the context and generate a response to confirm the value with the user.