Best local LLM for OpenCode at 40+ tok/s with 2× RTX PRO 6000 Blackwell?
I'm looking for the best open-weight LLM I can realistically run locally for OpenCode, with a target of at least 40+ output tokens/sec while using the model's full context window.
My workstation:
- AMD Threadripper PRO 9985WX, 64C/128T
- 512 GB DDR5-5600 ECC RDIMM, 8-channel
- 2× NVIDIA RTX PRO 6000 Blackwell 96 GB (192 GB total VRAM)
- 1× Workstation Edition
- 1× Workstation Max-Q
- ASUS Pro WS WRX90E-SAGE SE
- Linux
- Mainly using llama.cpp / LM Studio, but I'm also open to vLLM or SGLang if they make more sense
The main use case is agentic coding through OpenCode, so I'm prioritizing coding ability, tool use, long-context reliability, instruction following, and avoiding repetition/loops.
I'm fine with FP8, NVFP4/MXFP4, GGUF, etc. if the quality trade-off is reasonable. The model does not necessarily have to fit entirely in VRAM; CPU/RAM offloading is also an option, but I still want 40+ tok/s generation speed at full context.
I'm basically looking for the smartest model this machine can run at that speed, rather than the fastest small model.
What would you pick today?
I'd especially appreciate actual benchmarks from similar dual-96GB Blackwell setups rather than theoretical estimates.