Use gemini-2.0-flash for VAPI custom LLM
# support
b
💡 Question: I’m trying to integrate gemini-2.0-flash as a streaming model with Vapi, but I’m hitting a blocker. Vapi’s custom LLM example uses OpenAI’s ChatCompletionChunk for StreamingResponse, but Gemini streams using google.genai.types.GenerateContentResponse. To work around this, I hacked the OpenAI proto to repurpose delta.content with Gemini’s output, but that feels hacky and wrong. Does anyone have a clean example or guidance on how to properly stream gemini-2.0-flash as a StreamingResponse in a custom LLM using Vapi?
More details, I hack the following chunk piece: { "id": "chatcmpl-abc123", "object": "chat.completion.chunk", "created": 1712345678, "model": "gpt-4o-2024-05-13", "choices": [ { "index": 0, "delta": { "content": }, "finish_reason": None } ] }
c
Use Gemini’s native WebSocket (BidiGenerateContent) and wrap its responses into a custom StreamingResponse<GeminiChunk> for proper handling..
b
Thanks Kings. I will try this method at the end of the week and let you know the result. It looks like a reasonable solution. Many appreciate.
c
Hey birdfly7259, Thanks for reaching out! We're currently swamped with a bunch of tickets and it's taking us longer than usual to get back to everyone. I know waiting isn't fun, but we want to make sure we give your issue the attention it deserves rather than rushing through it. We'll circle back with you soon to get this sorted out.
b
Thanks SHubham and KIngs. I just check the internet and it looks like there is no GeminiChunk. Can you provide me some example of what the chunk looks like? Probably you can provide me one is the first chunk and another is the final chunk (mark the steaming ending).
c
Hey birdfly7259, I wanted to let you know that we're managing a high volume of support requests at the moment, so our response time might be a bit slower than usual. I truly appreciate your understanding and will get back to you as soon as possible!  Thanks again for your patience!
Hi, First off, we want to sincerely apologize for the delay in getting back to you. We understand how frustrating it is to wait - especially when you're counting on us - and we owe you a clear explanation of what’s been happening and how we’re addressing it. Over the past few weeks, we've seen a significant increase in support requests. While this reflects exciting growth, it has also stretched our small team and exposed some real challenges in scaling our support operations. To improve your experience, we’ve taken a step back to reassess our approach. Here’s what we’re implementing: - Smarter Support Through Automation: We’re investing in our AI support systems to help you resolve issues more efficiently. Soon, our support bot will offer expanded capabilities, making it easier to access accurate, instant help—particularly for common or repetitive queries. - Expanding the Support Team: To meet growing demand, we’re adding 2–3 new team members focused on managing support volume and improving response times. - Prioritized SLAs for High-Usage Accounts: We’re introducing service level improvements for users who are growing with us: - Accounts with usage over 1,000 minutes/month will receive prioritized support. - For all general inquiries, we’re establishing a standardized 48-hour response time. We’re confident these steps will lead to faster, more reliable support and help us better serve you as you grow with us. Also, in case you still need help with this ongoing ticket, do let us know, and we will help you get this resolved as soon as possible. Thank you for your continued patience and for being part of our journey. Warm regards, Vapi Team