Leo
04/09/2026, 6:49 PMLeo
04/09/2026, 7:01 PMAdam
04/10/2026, 12:41 AMLeo
04/10/2026, 12:40 PMAdam
04/10/2026, 12:42 PMLeo
04/10/2026, 12:48 PMLeo
04/10/2026, 12:51 PMLeo
04/10/2026, 12:52 PMLeo
04/10/2026, 3:30 PMLeo
04/10/2026, 3:37 PMLeo
04/10/2026, 3:38 PMLeo
04/10/2026, 4:24 PMLeo
04/10/2026, 4:25 PMLeo
04/10/2026, 4:30 PMLeo
04/10/2026, 4:30 PMAllen
04/11/2026, 5:46 PMAllen
04/11/2026, 5:54 PMAllen
04/11/2026, 6:23 PMcustomer-did-not-answer with 0 messages, no recording, no transcript, Start Time = N/A, end-of-call-report webhook arrives ~6 min later.
- The stream URL in the TwiML VAPI hands back to us points at phone-call-websocket.aws-us-west-2-backend-production-**weekly2**.vapi.ai/{callId}/transport. Can you check yours? I'm curious whether you're on weekly2 too, or on a different cluster instance. If you're on weekly2, that's a strong pointer at a specific bad cluster.
- We were moved from weekly1 → weekly2 on the morning of Apr 7, ~24–36h before our failures started. Your ramp (Apr 5: 18 → Apr 7: 179) would fit a similar migration happening ~1–2 days earlier for you. Worth asking VAPI support when your numbers were last moved.
- Last night (Fri ~midnight UTC-4) I tested on both Daily and weekly2 in rapid succession from two phones — everything worked on both. But I can't tell if weekly2 was actually fixed or if it's just low weekend load masking the bug. Very much hoping it holds for Monday.
- Your "concurrency limit reached" notifications showing up even on dropped calls is a big deal — it's consistent with what I've been seeing, where failures are much worse when calls are close together. Smells like a per-cluster resource pool that's either leaking or under-provisioned on weekly2.
Happy to DM and compare notes on setup — if we can rule in/out anything common between our configs (size of prompt, keyterm list length, transient vs imported assistants, etc.) that'd help VAPI narrow it down.Allen
04/11/2026, 6:25 PMcustomer-did-not-answer with cost: 0, 0 messages, no logs/transcript/recording, Start Time = N/A.
- **Cluster**: our failing stream URL host is phone-call-websocket.aws-us-west-2-backend-production-weekly2.vapi.ai. Leo is checking his.
- **Ramp**: Leo saw 18 → 81 → 179 → 138 failures/day Apr 5–8. Ours started Apr 8 ~16:15 UTC. We were moved weekly1 → weekly2 on Apr 7 morning, ~36h before failures began.
- **Concurrency signal**: Leo is seeing "concurrency limit reached" notifications on dropped calls. We see failures cluster tightly under concurrent load. Points at a per-cluster resource pool, not a per-customer quota.
- **Architecture**: both of us POST /call with transient config and return VAPI's TwiML to Twilio — the WS handshake is Twilio ↔ your edge, nothing customer-side is in the audio path.
- **Misclassification**: these calls are being bucketed as customer-did-not-answer when the actual failure is the WS upgrade never getting a response. That's probably why it hasn't tripped your internal alerts — they're sitting in a "customer's fault" bucket.
Failing call: 019d7875-daac-7000-a141-394278ca6bbd (2026-04-10 17:34:42 UTC). Composer confirmed cost: 0, no logs, no transcript, no recording — but checking whether Twilio's WS upgrade reached your edge needs SRE.
**What we need**: someone with infra-log access to check weekly2 handshake success rates Apr 5–11, and tell us whether it's (a) a known issue, (b) fixed overnight Apr 10→11 (which would explain weekend tests passing), or (c) expected to come back Monday. I have a customer demo Monday on this infra. Can share more call IDs if useful.Leo
04/13/2026, 1:08 PMLeo
04/13/2026, 7:00 PMkyle
04/14/2026, 3:50 AMLeo
04/14/2026, 2:00 PMLeo
04/14/2026, 2:02 PMLeo
04/14/2026, 6:07 PMLeo
04/16/2026, 3:13 PMLeo
04/27/2026, 7:11 PMAdam
04/28/2026, 12:49 AMLeo
04/30/2026, 2:28 PMChiranjeet Mishra
05/20/2026, 5:13 AMLeo
05/21/2026, 2:48 PM