Yeah so we use voice activated detection (basicailly the equivalent of "hey siri"), that triggers the realtime api, once it's done with responding, we close the connection, then send the context of what just happened to our database. The next time someone uses VAD, we connect to realtime again and pull the context