u/mosquito1459

▲ 9 r/ollama

Has anyone else noticed DeepSeek-V4-Flash:0731 using way more Ollama Cloud usage lately?

I’ve been using deepseek-v4-flash:0731 on Ollama Cloud, and recently it feels like my usage is getting consumed much faster than before.

My prompts and general workflow haven’t changed much, but the amount of usage being deducted seems noticeably higher. I’m not sure if something changed on the backend, the model’s token accounting changed, or if the model itself was updated to use more compute/tokens.

Has anyone else noticed the same thing recently?

Would be especially interested to hear from people who have been using this model regularly and can compare its usage before vs. now.

reddit.com
u/mosquito1459 — 1 day ago