▲ 7 r/StrixHalo+1 crossposts

Qwen3.8 Q4_k_m 1M context Strix Halo / 3080 -> 45 tps

Been playing qwen3.8 for a couple of days, on a AMD 128GB strix machine, with oculink eGPU (3080ti) 12GB.

Running Q6, 262K context

on just strix halo without MTP 10tps, with MTP with n=4 24 tps, with FastMTP offloaded to 3080 with n=4 about 28 tps

  1. Q4_k_m 262K context
    Splitting layers 5GB weights & FastMPT on 3080 and rest on iGPU~ 53 tps

  2. Q4_k_m, 1M context, k-q8, v-q4

Splitting layers 2GB weights & FastMPT on 3080 and rest on iGPU~ 45 tps

Still isn’t as good as a 5090 but AMD still show 60-70GB free so I can still load qwen3 embedding and a qwen3 Reranker to go with everything

Now need to see how good this model really is compared to qwen3.6-35b that I have running on 4x3090 and runs at 450 tps….

reddit.com
u/TrifleHopeful5418 — 3 days ago
▲ 0 r/claude

Claude Code, it was good before the enshitification

The problem isn't that my AI got worse. It's that I can't prove it — and neither can you.

I canceled two premium AI subscriptions this week because the product degraded under me. But the degradation isn't the interesting part. The interesting part is that I have almost no way to prove it happened, and that should worry every single person paying for a hosted model.

Here is the structural problem consumer-protection law was never built for.

Everything about the product is invisible except the price.

When a candy bar shrinks, you can see it. When a car gets slower, you can measure it. When a hosted model gets worse, you get vibes. The outputs are stochastic, so every individual failure is deniable — "that's just variance." You can't see the weights. You can't see the quantization. You can't see the routing, the reasoning-effort setting, or the system prompt sitting in front of the model. The vendor controls every lever that determines quality, and not one of them is visible to the person paying the bill. So the customer is left trying to prove a regression using the one instrument the vendor has made unreadable on purpose. The burden of proof is inverted, and it is unmeetable by design.

"We never intentionally degrade" is a non-answer.

It is the standard reassurance, and notice what it does: it moves the entire argument onto intent — the one thing a customer can never see inside. But intent is irrelevant to the harm. Whether the reasoning was deliberately cut to save compute, "optimized" for latency, or broken by a bug nobody caught, the outcome is the same: I paid for X, I received less than X, and nobody told me. The damage is identical in all three cases. Stop letting the fight be about whether they meant it. The consumer question is whether they told you — and they didn't.

You don't actually know what you bought.

The model has a name — "Opus 4.8" — but a name is not a specification. The thing behind it can be re-tuned, re-quantized, re-routed, and re-prompted at any time, under the same label, at the same price, with no version you can pin to and no changelog for the changes that affect quality. Ask an API customer: they can pin a dated snapshot. Ask a subscriber: you get whatever is live today. Status pages track uptime and errors, not capability — so a model can quietly get dumber while every dashboard stays green. You are buying a product whose definition floats, and you are not allowed to watch it move.

The cost of noticing is dumped on you.

Someone has to catch the regression. It is never the vendor who shipped it. It is the customer, who burns hours on debugging, guardrail-building, and forensic log analysis to discover a change the vendor already knew about. That is unpaid QA on a defect you didn't introduce. And "just cancel" is glib once you've built workflows, tooling, and real dependencies on the thing. The switching cost is the lock-in, and the lock-in is exactly what makes silent alteration profitable.

Strip all of it down and you land on the oldest principle in consumer protection: you cannot materially change a product someone is paying for, keep the price and the label identical, and shift the entire burden of noticing onto the buyer — least of all when you control every piece of information required to notice. (Whether that is legally actionable is a separate question, and I'm not a lawyer. The principle holds regardless.)

Here is what a fair version looks like, and none of it is exotic:

- Pinnable, immutable model snapshots for subscribers, not just for API users.
- A public changelog for quality-affecting serving changes — reasoning-effort defaults, quantization, routing, system-prompt edits — not just an uptime page.
- Advance notice for material changes to a paid tier, the same as you'd give for a price change or a change of terms.
- Independent, continuous, third-party capability monitoring, so detection doesn't rest on the one party motivated not to find problems, or on individual users who can't prove anything alone.
- Credits for confirmed degraded periods. If the product I rented wasn't the product for two weeks, that isn't on me.

A market cannot function on a good whose quality is unobservable and whose definition is mutable at the seller's sole discretion. That is not a premium AI subscription. That is a slot machine with a monthly fee.

I would love to be told I'm wrong about this. In public. With specifics.

#AI #ConsumerProtection #LLMs #AIagents #ProductManagement

u/TrifleHopeful5418 — 1 month ago
▲ 2 r/gpu

RTX 3090 EBay Pricing is Crazy!!

Couple of years ago, before Local LLMs were in vogue, I bought 8 RTX 3090 @ $700 each to build a AI rig, it been working great and I was looking to build another to increase my capacity but looking at EBay those are now selling for 1,300 -1,500 range!

That price seems totally crazy because on my main machine I have 3090 Ti that I bought new 5 years ago for about 1,400.

Needless to say, I was in shock and started looking for other GPUs. Then I went to Amazon and can buy a brand spanking new 3090 for 1,550!

Please tell me if you can buy a new GPU with great thermals why are people buying 5 years old used GPUs with degraded thermals for 1,400+ and keeping the EBay prices so high. What am I missing here?

reddit.com
u/TrifleHopeful5418 — 2 months ago
▲ 175 r/pcpartsales+2 crossposts

RTX 3090 EBay Pricing is Crazy!!

Couple of years ago, before Local LLMs were in vogue, I bought 8 RTX 3090 @ $700 each to build a AI rig, it been working great and I was looking to build another to increase my capacity but looking at EBay those are now selling for 1,300 -1,500 range!

That price seems totally crazy because on my main machine I have 3090 Ti that I bought new 5 years ago for about 1,400.

Needless to say, I was in shock and started looking for other GPUs. Then I went to Amazon and can buy a brand spanking new 3090 for 1,550!

Please tell me if you can buy a new GPU with great thermals why are people buying 5 years old used GPUs with degraded thermals for 1,400+ and keeping the EBay prices so high. What am I missing here?

reddit.com
u/TrifleHopeful5418 — 3 months ago