Does anyone actually know what your AI features cost per request?
Curious how teams are handling this. Between OpenAI/Anthropic API bills, GPU instances on RunPod or EC2, and vector DB costs, it seems like most places have one big number and no idea which feature or model is driving it.
- Do you know your cost per request, or per user, for anything AI-powered?
- Is anyone tracking token spend by feature, or is it all one line item?
- If you self-host, do you know your actual GPU utilisation, or is it "the box is up"?
- Has anyone gone through and actually cut this — what worked?