Unlimited DeepSeek for $0.20/hr — with a guaranteed 97 tok/s lane. Would you use it?

We ran a beta of a new AI inference pricing model last week, and our post here kind of blew up:

https://www.reddit.com/r/opencodeCLI/s/UKc0ErRp0U

There was a lot of interest, but also a lot of questions, doubts, and confusion because we didn't explain it well. So this is the follow-up that clears it all up, and we're opening slots for the next beta.

The one-liner: for $0.20hr, you get a dedicated lane on a GPU running the full-weight DeepSeek V4 Flash 0731 — not a quant — for one hour. Your own guaranteed slice, no shared rate limits.

Before you start doing the math, let me lay some groundwork.

Right now you have two ways to run inference -

  1. Pay-per-token APIs

Fine until you're a heavy user — then it gets expensive fast, and DeepSeek's price hike made it worse. If you're spending $100+/mo on tokens, you're exactly who this is for.

  1. Host on your own GPU

What most big teams do — full privacy, zero data retention, and once your workload is big enough, the monthly GPU cost beats per-token pricing.

But for solo builders and small teams this is a dead end: GPUs start around $12–30/hr and rack up $7k+/mo, and you'll never keep one saturated. You're paying for a whole GPU to use a sliver of it.

So we're building the middle ground: Shared Reserved Inference

We host the model, 30–60 people split the GPU cost for an hour, and each person gets a dedicated lane on it.

You get self-hosted-style dedicated inference for a fraction of the price — without renting the whole box.

And like self-hosting: we log zero prompts and zero completions. Only aggregate metrics like latency, throughput, tokens, and cache-hit rate. Your code never leaves your session.

The numbers

Full breakdown: https://www.singularityapi.dev/benchmark

From our last live run, on a lane costing $0.20/user/hr (rough estimate — don't hold me to the exact figure):

97% cache-hit rate under real agentic coding load

Each lane pushed 14M input tokens, 97% cached, and hundreds of thousands of output tokens in the hour

That worked out to 1.7×–3.4× the token value you'd get spending the same on DeepSeek, from off-peak to peak pricing

And we only ran the node at 30% capacity — there was a lot of headroom left

Clearing up the confusion from last time

  1. On the tok/s numbers

The per-second figures we quote are floors — measured with everyone hammering the node at the exact same time.

Real agent sessions interleave: different prompts, different timing, tool calls, waiting, etc. So in practice your effective throughput runs 2–4× above the floor.

The floor is the worst case, not the normal case.

  1. It only works on fully reserved, saturated GPUs

That means you reserve your hour in advance. If there's no node slot available in your timezone, we simply can't offer the lane.

This isn't an always-on API.

  1. It's for focused coding, not agent swarms

You get 1–2 concurrent requests + a few in-flight — plenty for a normal coding session with a subagent or two.

If you're running 5+ subagents hammering the API at once, this is not for you.

  1. It's a fixed hourly reservation — for now

You book a lane for a full hour.

If your session runs 40 minutes, you still reserve and pay for the hour. If it runs 1h20, you book a second hour.

That's the tradeoff of a guaranteed reserved lane today.

As demand grows and our node occupancy fills out, we want to move toward pay-for-what-you-use — billed for the 20 or 40 minutes you're actually on the lane — but that's down the road, not now.

It's also why we're being picky about matching beta slots to when you'll actually use them.

We're opening the next beta — free

A free 1-hour run, 64 seats.

You get a key + base URL, point your tools — Claude Code, opencode, Cline, Cursor, or direct API — at it, and code on your real project.

Pick a slot that fits your timezone:

Landing page: https://www.singularityapi.dev/beta

Signup form (60 sec): https://tally.so/r/EkoJkN

Benchmark: https://www.singularityapi.dev/benchmark

Now hammer me with questions — ask away.

reddit.com

Unlimited DeepSeek for $0.20/hr — with a guaranteed 97 tok/s lane. Would you use it?

We ran a beta of a new AI inference pricing model last week, and our post here kind of blew up:

https://www.reddit.com/r/DeepSeek/s/eFUlYOMpqS

There was a lot of interest, but also a lot of questions, doubts, and confusion because we didn't explain it well. So this is the follow-up that clears it all up, and we're opening slots for the next beta.

The one-liner: for ~$0.20/hr, you get a dedicated lane on a GPU running the full-weight DeepSeek V4 Flash 0731 — not a quant — for one hour. Your own guaranteed slice, no shared rate limits.

Before you start doing the math, let me lay some groundwork.

Right now you have two ways to run inference -

  1. Pay-per-token APIs

Fine until you're a heavy user — then it gets expensive fast, and DeepSeek's price hike made it worse. If you're spending $100+/mo on tokens, you're exactly who this is for.

  1. Host on your own GPU

What most big teams do — full privacy, zero data retention, and once your workload is big enough, the monthly GPU cost beats per-token pricing.

But for solo builders and small teams this is a dead end: GPUs start around $12–30/hr and rack up $7k+/mo, and you'll never keep one saturated. You're paying for a whole GPU to use a sliver of it.

So we're building the middle ground: Shared Reserved Inference

We host the model, 30–60 people split the GPU cost for an hour, and each person gets a dedicated lane on it.

You get self-hosted-style dedicated inference for a fraction of the price — without renting the whole box.

And like self-hosting: we log zero prompts and zero completions. Only aggregate metrics like latency, throughput, tokens, and cache-hit rate. Your code never leaves your session.

The numbers

Full breakdown: https://www.singularityapi.dev/benchmark

From our last live run, on a lane costing $0.20/user/hr (rough estimate — don't hold me to the exact figure):

97% cache-hit rate under real agentic coding load

Each lane pushed 14M input tokens, 97% cached, and hundreds of thousands of output tokens in the hour

That worked out to 1.7×–3.4× the token value you'd get spending the same on DeepSeek, from off-peak to peak pricing

And we only ran the node at 30% capacity — there was a lot of headroom left

Clearing up the confusion from last time

  1. On the tok/s numbers

The per-second figures we quote are floors — measured with everyone hammering the node at the exact same time.

Real agent sessions interleave: different prompts, different timing, tool calls, waiting, etc. So in practice your effective throughput runs ~2–4× above the floor.

The floor is the worst case, not the normal case.

  1. It only works on fully reserved, saturated GPUs

That means you reserve your hour in advance. If there's no node slot available in your timezone, we simply can't offer the lane.

This isn't an always-on API.

  1. It's for focused coding, not agent swarms

You get 1–2 concurrent requests + a few in-flight — plenty for a normal coding session with a subagent or two.

If you're running 5+ subagents hammering the API at once, this is not for you.

  1. It's a fixed hourly reservation — for now

You book a lane for a full hour.

If your session runs 40 minutes, you still reserve and pay for the hour. If it runs 1h20, you book a second hour.

That's the tradeoff of a guaranteed reserved lane today.

As demand grows and our node occupancy fills out, we want to move toward pay-for-what-you-use — billed for the 20 or 40 minutes you're actually on the lane — but that's down the road, not now.

It's also why we're being picky about matching beta slots to when you'll actually use them.

We're opening the next beta — free

A free 1-hour run, 64 seats.

You get a key + base URL, point your tools — Claude Code, opencode, Cline, Cursor, or direct API — at it, and code on your real project.

Pick a slot that fits your timezone:

Landing page: https://www.singularityapi.dev/beta

Signup form (60 sec): https://tally.so/r/EkoJkN

Benchmark: https://www.singularityapi.dev/benchmark

Now hammer me with questions — ask away.

reddit.com

Unlimited DeepSeek for $0.49/hr — with a guaranteed 160 tok/s lane. Would you use it?

We’ve been experimenting with a different way to price hosted inference at Singularity API.

Instead of charging per token or locking people into a subscription, we’re testing reserved inference slots at $0.49 per slot-hour. One slot = one guaranteed concurrent lane. If you need parallel agents, you can reserve multiple slots and each gets its own lane.

We’re currently serving DeepSeek-V4-Flash-0731 at full weights, with the full 1M context window. The service is built around reserved capacity rather than a shared best-effort pool, so each booked slot has a defined throughput floor regardless of how busy the rest of the service is.

These are measurements from the live deployment:

- $0.49 per slot-hour

- 160 tok/s guaranteed generation floor

- Typically 200–340 tok/s when spare capacity is available

- ~205k+ output tokens per slot-hour

- 5M fresh input tokens per slot-hour

- Unlimited cached input

- 98.1% measured prefix-cache hit rate across load levels

- Full 1M context

- One slot = one guaranteed concurrent lane

The reason we started exploring this is that DeepSeek changed API pricing significantly on August 16, while API throughput is still best-effort and can slow down during busy periods.

For workloads like agentic coding, parallel agent swarms, RAG over stable corpora, or anything repeatedly sending large warm contexts, we think hourly reserved capacity may make more sense than constantly paying again for the same cached tokens.

The important limitation is that this isn’t really meant for light or occasional API usage. You’re reserving a slot for the hour, so if you only make a few requests, normal per-token APIs will probably make more sense.

We’re still small and this is an interest check, not a GA launch. If there’s enough interest, we’ll open a waitlist on singularityapi.dev and start letting people in gradually.

Would you actually pay $0.49/hour for a guaranteed DeepSeek lane instead of paying per token?

What generation-speed floor would matter to you: 100, 160, 200+ tok/s?

If you currently use DeepSeek directly or through OpenRouter, what would make you switch?

Edit: Quick clarification since this confused a few people — the token numbers in the post are minimum floor values, not maximum limits.

If the system has spare capacity, it automatically flows to whoever is generating, so in normal coding/agent usage you'll generally see 2–4x higher throughput than the floor.

reddit.com
u/Individual_Team_2344 — 3 days ago

Unlimited DeepSeek for $0.49/hr — with a guaranteed 160 tok/s lane. Would you use it?

We’ve been experimenting with a different way to price hosted inference at Singularity API.

Instead of charging per token or locking people into a subscription, we’re testing reserved inference slots at $0.49 per slot-hour. One slot = one guaranteed concurrent lane. If you need parallel agents, you can reserve multiple slots and each gets its own lane.

We’re currently serving DeepSeek-V4-Flash-0731 at full weights, with the full 1M context window. The service is built around reserved capacity rather than a shared best-effort pool, so each booked slot has a defined throughput floor regardless of how busy the rest of the service is.

These are measurements from the live deployment:

- $0.49 per slot-hour

- 160 tok/s guaranteed generation floor

- Typically 200–340 tok/s when spare capacity is available

- ~205k+ output tokens per slot-hour

- 5M fresh input tokens per slot-hour

- Unlimited cached input

- 98.1% measured prefix-cache hit rate across load levels

- Full 1M context

- One slot = one guaranteed concurrent lane

The reason we started exploring this is that DeepSeek changed API pricing significantly on August 16, while API throughput is still best-effort and can slow down during busy periods.

For workloads like agentic coding, parallel agent swarms, RAG over stable corpora, or anything repeatedly sending large warm contexts, we think hourly reserved capacity may make more sense than constantly paying again for the same cached tokens.

The important limitation is that this isn’t really meant for light or occasional API usage. You’re reserving a slot for the hour, so if you only make a few requests, normal per-token APIs will probably make more sense.

We’re still small and this is an interest check, not a GA launch. If there’s enough interest, we’ll open a waitlist on singularityapi.dev and start letting people in gradually.

Would you actually pay $0.49/hour for a guaranteed DeepSeek lane instead of paying per token?

What generation-speed floor would matter to you: 100, 160, 200+ tok/s?

If you currently use DeepSeek directly or through OpenRouter, what would make you switch?

Edit: Quick clarification since this confused a few people — the token numbers in the post are minimum floor values, not maximum limits.

If the system has spare capacity, it automatically flows to whoever is generating, so in normal coding/agent usage you'll generally see 2–4x higher throughput than the floor.

Edit2: If this interests you please fill out this form - https://tally.so/r/ob8bj1

u/Individual_Team_2344 — 3 days ago

Indian AI startups: DeepSeek V4 Flash is completely free for the next few days

Inference bills can start hurting long before an AI startup has found product-market fit.

We’re building SingularityAPI from India—one OpenAI-compatible API for accessing open-source and frontier models, with some of the lowest inference pricing in the market.

For the next few days, DeepSeek V4 Flash is completely free to use, with no credit card required.

We’re also giving free inference credits to Indian startups building AI products, including agents, developer tools, workflow automation, voice AI, SaaS products, and other model-heavy applications.

Apart from pricing, we’re working on two problems we repeatedly faced with inference providers:

Model Contracts: Your application uses a permanent model ID, allowing you to switch the underlying model without changing your code or redeploying your application.

Verified Routing: Every request generates an inference receipt showing the exact model used, token consumption, latency, and cost. This prevents models from being silently swapped for lower-quality or lower-quantized variants without visibility.

To apply for free credits:

  1. Create an account at https://www.singularityapi.dev

  2. Preferably register using your business email

  3. Join our Discord and open a Startup Ticket with a short description of what you’re building

Approved startups will receive credits within 24 hours.

Would love to hear what Indian founders here are building and which models are currently consuming most of their inference budget.

reddit.com
u/Individual_Team_2344 — 15 days ago

Free API credits for DeepSeek V4, Kimi K2, and other open-weight LLMs — unified API, closed beta, no card required

Hey folks 👋

I’ve been heads-down building SingularityAPI—a unified API for accessing some of the strongest open-weight models available right now, including:

- DeepSeek V3.2

- DeepSeek V4 Pro

- DeepSeek V4 Flash

- Kimi K2.6

- Kimi K2.7 Code

Everything works through a single API key and base URL.

The API is fully OpenAI-compatible, including "/v1/chat/completions" and "/v1/responses", so you can drop it into the OpenAI SDK or most OpenAI-compatible tools with little to no code changes. End-to-end streaming is supported as well.

I’m currently giving beta users free API credits while I stress-test the routing and inference infrastructure. No credit card is required, and there’s no catch.

To be completely transparent, I know that “one API for multiple models” is already a solved problem. That’s the foundation, not the main product.

The bigger problem I’m trying to solve is the lack of transparency across inference providers.

When you build on most providers, you often have no reliable way to know exactly what served your request. Models can be swapped, quantized, modified, or deprecated without clear notice. Your bill is whatever the dashboard says it is. And when output quality suddenly drops, you have very little evidence showing what actually changed.

You’re essentially building your product on a black box.

I’m building an inference layer where those changes cannot happen silently.

I’m not ready to publicly reveal all of those features yet, but they’ll begin shipping during the beta. Early users will get access first, use them for free, and help shape how they work.

What’s available today:

- Multiple leading open-weight models behind one endpoint

- Switch models by changing only the model name in your request

- OpenAI-compatible requests and responses

- Support for "/v1/chat/completions" and "/v1/responses"

- End-to-end streaming

- One API key and base URL for every supported model

I’m keeping the beta small for now so I can closely monitor usage, fix issues, and stress-test the routing layer before opening it more broadly.

Accounts are currently being provisioned manually, so I can only onboard a limited number of people at a time.

Interested in testing it?

Send me a DM with a little information about what you’re building—or simply say that you want to experiment with it—and I’ll set you up.

I’d genuinely appreciate feedback from people building real products: bugs, missing features, unexpected behaviour, rough edges, or anything else you notice.

reddit.com
u/Individual_Team_2344 — 24 days ago