
CDW has bumped the MSRP of the RTX Pro 6000 from $16,000 to $19,999
Did they slip up and leak future pricing?
Live link: https://www.cdw.com/product/pny-nvidia-rtx-pro-6000-graphic-card-96-gb-gddr7/8326705

Did they slip up and leak future pricing?
Live link: https://www.cdw.com/product/pny-nvidia-rtx-pro-6000-graphic-card-96-gb-gddr7/8326705
I'm looking for the best open-weight LLM I can realistically run locally for OpenCode, with a target of at least 40+ output tokens/sec while using the model's full context window.
My workstation:
The main use case is agentic coding through OpenCode, so I'm prioritizing coding ability, tool use, long-context reliability, instruction following, and avoiding repetition/loops.
I'm fine with FP8, NVFP4/MXFP4, GGUF, etc. if the quality trade-off is reasonable. The model does not necessarily have to fit entirely in VRAM; CPU/RAM offloading is also an option, but I still want 40+ tok/s generation speed at full context.
I'm basically looking for the smartest model this machine can run at that speed, rather than the fastest small model.
What would you pick today?
I'd especially appreciate actual benchmarks from similar dual-96GB Blackwell setups rather than theoretical estimates.
I have 2 rtx pro 6000 max qs in a non threadripper computer with only 64gb of ram. This has been fine for a lot of what I do but it's not ideal for video generation.
I'm reluctant to buy more ram even though it's the clear weakness of my computer because I think prices will eventually go down a lot, though it might take 3 years. Gpus I don't mind buying now because prices aren't going down anytime soon or could even be worth more later on.
So what are you all doing? I could sell a 5090 and then buy a bunch of ram, but I would hate to do that.
I always backed up my computers with UPS before but now that my computer is creeping up in power I'm being forced into even more expensive UPS than I was expecting. The 2000w UPS is very expensive. Just curious what everyone else is doing. Are you plugging into the wall with surge protection or are you backing up your computers with UPS?
For context: All tests Unsloth Q8_K_XL, F16KV, on a single RTX Pro 6000 Blackwell, Max-Q (the other one was sticking its tongue out at me)
Oh boy... this entire Qwen madness has been driving me nuts. So i put together my own benchmark. Small, honest. testing all sorts of combinations, and because of ONE big glitch, went down an actual rabbit hole and proved that this thing can INDEED rip. batches and batches of 300, 500 even 800+ tokens per second. But funny enough, that's not the most important bit. The Quality of the results mattered more. R$ed bars are failed generations (JS errors, doesn't matter how small), Yellow are positioning hiccups but otherwise ok), Green are fully successful.
While i still do more testing - now moving from raw curl tests to an actual harness, to evaluate my findings on long context horizon, tool calling, etc, i'll leave you with these images.
A note on token counts: these aren’t short-completion benchmarks. Prompt 1 was deliberately open-ended and the model routinely generated ~48k–70k tokens while reasoning and building the artifact. Prompt 2 is different: it feeds a previously generated artifact back into context and asks Qwen to work from/reproduce that existing structure, so the actual generated completion is typically ~20k–23k tokens. That’s where ngram speculation gets enough reusable material to become completely ridiculous.
The throughput numbers in the chart are full-run generation averages reported by llama.cpp, not the 3-second peaks. During the fastest P2 runs, tg_3s repeatedly reached 300–700+ t/s; those bursts are interesting, but the headline result is still 234.41 t/s averaged across the entire ~20k-token generation. And Even P1 runs (the one marked p1, s666) had bursts of 300-812 (peak)
The last 5 runs are the important bits:
P1 / seed 666: 86.72 t/s
P2 baseline / 666: 116.15 t/s
P2 baseline / 667: 125.84 t/s
P2 ngram-mod / 666: 163.11 t/s
P2 ngram-mod / 667: 234.41 t/s
Because prompt 1 (create a simple dashboard, one file, self contained), at some point generated some ludicrous segments of 300-800 tokens per second, i went down investigating that. Tested cold vs hot runs (funny enough, that had nothing to do with it), to reproduce, and it did. Then, on a hunch, i changed the prompt that ngram had enough reproducible content to absolutely RIP through content. see the last 2 runs.
"qwen38-ngram-mod":
env:
- "CUDA_VISIBLE_DEVICES=0"
cmd: >
/app/llama-server
-m /models/qwen3.8-27b/Qwen3.8-27B-UD-Q8_K_XL.gguf
--mmproj /models/qwen-mmproj/mmproj-F16.gguf
--port ${PORT}
-ngl 99
--ctx-size 131072
--batch-size 2048
--ubatch-size 512
--cache-type-k f16
--cache-type-v f16
--spec-type ngram-mod
--spec-draft-n-max 8
--spec-draft-p-min 0.8
--flash-attn on
-np 1
--jinja
--chat-template-file /models/qwen-mmproj/chat_template.jinja
--reasoning-format deepseek
--reasoning-preserve
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0
Chat template from here: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
Github repo here: https://github.com/mihaiCroitoru/qwen38-bench
The repo includes the exact prompts/configs, raw benchmark harness, timing methodology, Chrome artifact validation, and full per-run reporting. I’ll push all preserved run artifacts/results once I finish the harness-level tests.
Feel free to run your own tests, tell me what i missed about mine!
I’ve been loving running deepseek v4 flash so much that I haven’t wanted to take the downtime to test the pp/tg on my rtx rig… anyone got numbers yet?
I will own one RTX PRO 6000 (Workstation Edition) by the end of the month and place it inside a North XL case (192mm GPU height clearance - since the GPU is ~138mm tall, I will have ~54mm of cable clearance). Also, temperature-wise (both for the GPU and myself, having the workstation right under my desk), I will power limit the GPU to 300W (50% of its TDP).
Given the above, should I be worried about the cable (and/or GPU connector) melting like with other 50xx GPUs? Are there instances of cable melting for RTX PRO GPUs (I could not find even one instance)?
The PSU will be the Corsair RM1000X (ATX 3.1, PCIe 5.1) with a 12V-2x6 cable included.
I already have 1 rtx pro 6000 and regret not getting the 2nd at the time.
Is it worth getting a 2nd 6000? I want to be able to run deep seek flash and see that I can run it on 1 card but I haven't seen much discussion on how performance is 1 card vs 2.
Hi,
I'm just curious if there is someone in this group who recently bought or knows someone who recently bought an RTX 5090 or 6000 PRO.
Why did you do it? Why were you willing to spend so much money on it, and what value will this provide for you/your company?
I just want to know why people are paying these prices.
I got some serious FOMO for the TP 4 and orderd one more 5090 for 4300€ - it was the Xtreme waterforce WB - but still very expensive.
Thank you.
Hey everyone,
After ~5 days of tuning and ~$1.2k in compute I released a specialized Neuron-pruned IQ1_S quant of Kimi K3.
Key details:
- Size: ~308GB (vs the common ~594GB baseline)
- Every one of the 82,432 routed experts is present — no experts dropped
- HumanEval: 94.5% (matches full model 1:1)
- AIME: 92.5% (full ~96.1%)
- GSM8K: 95%
- MMLU: 79.49% (full ~85%)
- Throughput: 12.5492 tok/s average on 3× DGX Sparks (SparkInfer TP3 + custom speculative decoding patches; from ~2 t/s baseline)
This is my work (self-promo disclosure).
Links:
- HF (gated, request access): https://huggingface.co/vcruz305/Kimi-K3-GGUF
- SparkInfer patches + TP3 recipe + one-command deployment + benchmark receipts: https://github.com/vcruz305/kimi-k3-neuron-tp3-dgxspark-recipe
- Full announcement thread with more details: https://x.com/ViC305/status/2087609292442751209
Hardware notes: The reported 12.5 t/s is on 3 DGX Sparks via SparkInfer + my speculative decoding patches. It also has a llama.cpp fallback path. Optimized for multi-GPU / Blackwell-class setups.
If you have 3 Sparks (or 4× Blackwell / equivalent), please try it and share logs/results. Happy to answer questions on the neuron pruning, calibration, or multi-node setup. Feedback and PRs welcome!
Did a search everything was for 2x 6k pros. I'm looking to try some different models than qwen3.6 and gemma.
I got this rtx pro 6000 Blackwell yesterday and the first thing I noticed is that under load the coil whine is extremely noticeable. I did some research and it seems like some people think it's normal and some people have quiet cards, so I'm wondering what you guys think.
I recorded the video in a quiet room, it's in the DEG2 dock, I included keyboard and mouse clicks as reference. Coil whine starts around 15 seconds into the video.
I have been eyeing the RTX PRO 6000s for a while, and with the recent price increase announcement I'm really considering buying one quickly (although it's still a LOT of money).
I'm digging around and noticed that the HP variant is about 1000 EUR cheaper than the NVIDIA or PNY variants.
What is the difference between those cards? Can I use the HP in a non-HP workstation? Or does it require special HP drivers?
Which brand would you recommend?
Edit: thank you guys for the replies! Finally, I have ordered the HP. I will post an update once I get it (probably 1-2 weeks)