Join Discord
Powered by
either use it or don't; it's not random spam. here...
# projects-showcase
p
Pavan Katta
06/04/2026, 10:26 PM
either use it or don't; it's not random spam. here's a community showcase version. hey everyone, noticed in most of my calls, gpt4.1 was eating 60%+ of the latency budget with it's 700ms ttft. wanted to try out some open source models and hosting the latest qwen on 8*B200 so you guys can try can it. Model: Qwen/Qwen3.6-35B-A3B Endpoint URL:
https://llm.monosemantic.com/v1
with this config, seeing
sub 100ms in live vapi calls
.
https://cdn.discordapp.com/attachments/1211483571454218261/1512221006716997663/Screenshot_2026-06-04_at_2.46.43_PM.png?ex=6a234d0f&is=6a21fb8f&hm=d2d561436cc621f2755408c2fa4d5c58707ebaece2665733b06d0a6b3a3c14c5&
Previous
Next