I think structuring works, like telling it to confirm letters one by one since the llm will put commas/spaces between the letters and the voice STT can recognize that and pronounce letters differently but for actual word pronunciation I believe it doesn’t do anything