Hey Favour, good to know I'm not alone here.
To answer your question: no, i'm not storing raw audio or model parameters per call right now. I just store the final transcripts and outputs. Which is probably exactly the problem. 🙁
But, how are you handling this? Are you building something custom to capture that raw audio + context, or using an existing tool? And how much effort did it take to get to replayable test cases?