Qwen 3.8 27b on blackwell anyone?
I’ve been loving running deepseek v4 flash so much that I haven’t wanted to take the downtime to test the pp/tg on my rtx rig… anyone got numbers yet?
u/ObviouzFigure — 4 days ago
I’ve been loving running deepseek v4 flash so much that I haven’t wanted to take the downtime to test the pp/tg on my rtx rig… anyone got numbers yet?
Heyyo everyone -- wondering who out there has a Dual RTX Pro 6000 rig and if anyone is running Ds4F 0731? -- I had success running it through LM Studio/Llama.cpp (Windows) and was getting ~40 t/s... after many hours and many anthropic credits I was able to get vllm serving ds4f and I'm getting 100+ t/s but my contact is limited to about 140k with one concurrent session. Planning on setting up Linux this weekend... Opus says that if I set it up on linux/bare metal I'll see over 200 t/ks. Anybody else having success?