
Building with VLMs? Check out the Overshoot API
I’ve been working with real-time vision/VLM inference lately and I’m curious what stacks people here are using.
We’ve been building around the Overshoot API, especially for applications where latency matters ( < 200 ms). Would be interested to hear what others are using for hosted VLM inference vs self-hosting.