▲ 110 r/DeepSeek

Flash is on fire today

Not sure what’s going on, and I really hate these low value posts, but I gotta say: Flash is insanely good today.

I’ve been working on some skills and observability stuff and I used to have to go with Pro to actually get the job done but today tried with Flash (via OpenCode Go sub) and it is a monster. If that’s the GA version, I’m gonna start using this for so much more.

The thing just done a full trace investigation with my OTel observability, from getting the error logs to investigating src code, infra, AKS, running probes, and pointing to the exact issue with a fix suggested in 2min, and a total cost of $0.034. Zero tool calls failed. ZERO.

I’d be ashamed to work for Scam Altman or Scamthropic today. Damn!

reddit.com
u/somerussianbear — 1 month ago
▲ 0 r/ollama

Ollama $20 plan, what’s the quant used?

Mostly interested in GLM, Kimi, and maybe Minimax models. What’s your experience? (Coming from OpenCode Go, where GLM most certainly run on low quants)

Thanks!

reddit.com
u/somerussianbear — 1 month ago
▲ 172 r/DeepSeek

The "OpenCode Go is cheaper than DeepSeek API because you pay $10 and get $60" falacy

As a software engineer, my day-to-day setup is basically Claude and OpenAI subscriptions, around $100 each.

Since DeepSeek V4 came out, I’ve also been using its API (on Pi) for smaller investigations, observability work, and anything where I want a quick answer without waiting a full minute for Claude to think before doing a few greps and file reads.

V4 Pro, and sometimes even Flash, are great for this. After a few back-and-forth turns, once I’m happy with the result, I ask it to write down some notes. I then feed those notes into Claude or GPT as the starting point for the actual implementation.

The API is great and ridiculously cheap. You can get a lot of work done this way for less than $10 a month.

Then I kept hearing about the OpenCode Go subscription. It looked interesting because it gives access to other models like GLM 5.2 and now Kimi K3. Since DeepSeek was also available there, I started using it through OpenCode Go. I was already paying for the subscription, so I thought I might as well save my official DeepSeek API credits.

I’ve lost count of how many times I’ve seen people say OpenCode Go is a much better deal because you pay $10 and get $60 in usage, supposedly a 5x benefit. So I decided to test it.

The test was simple:
- Same prompt in both sessions
- Same code investigation
- A codebase with more than 100 repositories
- The goal was to understand and gather knowledge about one specific area of the system
- Two terminals running at the same time, one using OpenCode Go and one using the official DeepSeek API

The result: the OpenCode Go session was around 4x more expensive than the official DeepSeek API session.

Considering that the main selling point is paying $10 for $60 of usage, that 5x benefit suddenly doesn’t look like much of a deal.

This wasn’t a one-off test either. I ran several different sessions with other prompts and follow ups, and the average for a short session was consistently around 4x. Some longer sessions went above 10x the API cost, while a few were closer to 2x, so it balanced out around that number.

The main difference seems to be cache hits. The official DeepSeek API appears to handle caching much better than OpenCode Go. The longer the session goes, the more cache misses OpenCode Go seems to accumulate, and the larger the cost multiplier becomes.

For short tasks, OpenCode Go may still be convenient. But for longer sessions, where the combined model time across turns goes beyond 5 minutes or so, you may end up paying significantly more through OpenCode Go than you would through the official API.

One final detail: with GLM, I was able to improve cache hits through model-specific settings, similar to configuring temperature or other request parameters. In particular, there are settings that prevent the model from rewriting previous messages or stripping reasoning from earlier turns.

That matters because changing anything in the previous conversation changes the prompt prefix and breaks the cache. The next request then becomes a much more expensive cache miss instead of a cache hit.

As far as I can tell, OpenCode does not expose equivalent model settings for DeepSeek. Has anyone found a way to configure OpenCode Go so DeepSeek preserves the previous conversation and reasoning exactly as-is, or otherwise improves cache-hit rates?

https://preview.redd.it/63wtbl73ardh1.png?width=1011&format=png&auto=webp&s=f1e251af5d811b218a04cd833e749c69806ffe41

reddit.com
u/somerussianbear — 1 month ago
▲ 74 r/ZaiGLM

Neuralwatt price "adjustment": PAYG doubles the price and the $20 sub gets 200% more expensive

So many things are wrong in this email, to start with the lenght of it and the fact that he tries to sugarcoat an incredible 100% price hike (2x) on the price per token and a 200 to 214% increase (3x) on the plans.

The audacity of this dude to hide the % increase in a long email like that. I mean, the target audience is SWE, do you think people won't do the math?

IMO it'd be insane to take this price. Their service is wobbly, sometimes you get predictable TTFT and 40-50 TPS on GLM 5.2, but some times of the day it just times out or replies at 1 TPS.

I'm using OpenCode Go too, but it doesn't handle my needs in capacity, I'd need some 3 or 4 of those plans, which would still be cheaper than Neuralwatt by a huge margin. What are you guys using? I only hear bad stuff about the official API, so not sure I'm going there.

Email in full.

>Hi {name}

>Thank you for being part of Neuralwatt Cloud. The response to energy-based pricing and flex inference has exceeded our expectations, and we're grateful for your trust and engagement as we've grown.

>We're writing to let you know about an upcoming pricing change and what it means for your account.

>What's changing on July 16, 2026:

>- Base energy rate: $5/kWh → $10/kWh
- PAYG packs: Starting July 16, these will be flexible — purchase any amount from $5 to $1,000 at the flat rate
- Subscription tiers: Monthly subscriptions will be adjusted for the new rate and rolled over automatically; no action needed on your part
- New pricing tiers coming soon: Flex (latency-tolerant, discounted), Standard (improved SLA), Express (priority lanes, tighter TTFT), and Enterprise (dedicated capacity)

>Updated subscription kWh allocations:
- Basic: 2 kWh included
- Standard: 5.2 kWh included, 5 concurrent requests
- Pro: 10.5 kWh included, 10 concurrent requests

>How this affects you:

>- Credits purchased before this announcement will continue to be deducted at the old $5/kWh rate until your balance is used up — you won't see a surprise jump on existing credits.
- New purchases after July 16 will be billed at the new $10/kWh rate.
- Auto top-ups will need to be re-enabled after July 16, as existing auto top-up configurations will not carry over to the new pricing structure.
- Subscriptions will automatically roll onto the updated rates — your plan stays active, with the included energy allocation adjusted for the new pricing.
- Annual subscriptions are temporarily unavailable until July 16 while we prepare the new pricing structure.

>Energy-based pricing will still be substantially cheaper than token-based pricing for most workloads — and the efficiency compounds further with the new Flex tier if your use case can tolerate some latency.

>We know pricing changes aren't fun. We wouldn't be making this adjustment if it weren't necessary to sustain the level of reliability, capacity, and model availability you rely on. We're committed to making this transition as smooth as possible. You can see the full pricing details at https://portal.neuralwatt.com/pricing.

>If you have any questions, reach out in our Discord (https://discord.gg/ZJEfU2BZw2) or email info@neuralwatt.com.

>Thanks!
Scott Chamberlin
Neuralwatt CTO

u/somerussianbear — 1 month ago

pi agent with DeepSeek v4 Pro is the beast: $0.45 for a heavy 90min coding session with hundreds of tool calls

The image is just to show the stats (green box), nothing important in the last turn there, it's more about what follows here.

Not sure if this post is more about the pi agent or DeepSeek really, but gotta tell you that them coupled do a hell of a job, and all that fast and cheap.

I have a bunch of microservices (25 to be precise) running in different tech stacks, with different logging frameworks/observability patterns set up and troubleshooting them is always a shitty job cause you gotta start kubectling their logs, figuring out their patterns, checking a baseline distribution for HTTP Status Codes, then grouping log messages, getting a few random log entries to then "feel" if the app is behaving correctly or if there's some little issue happening since last deployment. Sometimes these issues are quite small and hard to trace, takes days to realize something is wrong (or even weeks). Of course there's much more to this than just grepping logs, but I don't want to go deep into that here as it's not the point of the post.

The point is that I usually use Opus or GPT latest on my day to day activities and DeepSeek with OpenCode for some side quests, but it's been a while that I'm not happy with how things go there. OpenCode is strongly biased into spinning sub-agents for anything and in my experience sub-agents (with DS) just don't have enough context to do a good run on the task they're given. Most of the time what happens is that the main agent comes back with the answer of the sub-agents summarized and spits out some lie/incomplete picture of the thing, then I push back and the main agent goes back to reviewing what the sub-agents said and figures out mistakes and basically the entire thing falls apart with the main agent saying some variation of "these results can't be trusted, gotta do the job myself" and I just wasted time/tokens on that. I know we can disable sub-agents in OpenCode but I just gave it a shot with pi at the task described below and the speed of Pro (not Flash!) combined with how light pi is AND its cache-friendliness surprised me a hell lot. I also tried Claude Code with DeepSeek already, but to me it felt like DeepSeek trying to drive a car using the instructions of how to ride a horse. Totally uncomfortable, tool calls failing all the time, no "learning" during the session, no benefiting on the tools that were designed for Claude models.

Coming back to pi, this session lasted around 80-90 minutes, had a shit ton of turns and some 250-400 tool calls (at least!), zero compaction during the job and I was still at 32.8% context window. Very rarely tool calls failed, I could say something like less than 2% of the tool calls failed, and these tool calls were quite complex. kubectl logs, piping to rg, piping to awk, and lots more.

The task was basically something on the lines of "you've got these apps, access their logs, figure out patterns, recipes, how-tos, gotchas, and write an md documenting it all, one doc per app. all commands documented must be executed/tested/validated so to ensure quality of the documentation". Bit more complex than that, but you get the gist of it.

Some of the things that surprised me the most were:

  • huge amount of tool call output and very low context size increase: working with Codex/Claude Code all the time, my feeling is that the same work would have had some 3 or 4 compactions already with those tool calls, and here we ended the whole thing with 0 compactions and 2/3 of the context window still free.
  • cache hit: 99% across the whole session. You gotta think that DS is very (but VERY) cheap for cache hit (~3 million tokens = $0.01) so an agent/harness that can ensure append-only behavior will be extremely good at token economics.

Main stats you can see in the green area of the image but I'll translate here to get it clearer:

  • input tokens (cache miss): 223K
  • input tokens (cache hit): 62 million (that's a 99% cache hit across the entire session!)
  • output tokens: 142K
  • cost: $0.45

I know lots here is about "vibe" but honestly, I'm coding with these tools for almost two years now (Copilot/Cursor/Claude Code/OpenCode/Codex and now pi) - and 18 years without them before that - and these vibe checks are important to get confidence in the tool/model, so I hope this is useful to anybody thinking about using DeepSeek v4 Pro for something serious and wondering about capability, harness, pricing etc.

u/somerussianbear — 2 months ago
▲ 74 r/codex

End of 2x promo: On my 100$ plan, there's WAY less juice than just half of that orange

Pretty simple: today is the first day without the 2x consumption promo, and I’m already noticing a huge difference.

I’m using Codex normally. For you to have an idea, I kept it in /fast for the last 4-5 weeks, but today I had to stop using fast mode entirely because just a few prompts burned through around 10% of my 5-hour limit in a matter of minutes, which would have been unthinkable last week.

And even without fast mode, it still feels like the limit is draining way, way faster than 2x. At this rate, I’ll probably hit 100% of my 5-hour limit pretty soon, which means I’m back to Claude.

I really don’t think this feels like a 50% reduction. It feels closer to 70%, which would be the first very Claude-looking move from the Codex team.

Hopefully this is just a temporary issue, and Tibo comes out soon with a tweet and a reset to save the day.

reddit.com
u/somerussianbear — 3 months ago