Image 1 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
Image 2 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
Image 3 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
Image 4 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
Image 5 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
Image 6 — Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram
▲ 10 r/LocalAIServers+1 crossposts

Quadro RTX 5000 16 GB + ​Dual GeForce RTX 3060 12GB = 40GB of vram

This is my local AI Server. It runs the best local model around qwen3.8:27b via ollama. I was inspired by Digital Spaceport on YouTube to make a 8-bit style arcade suite in a html file so i can host it on my website. I used hermes for my agent and it worked great, after a few update prompts it was finished - PIXELARCADE.

Ollama question:
The system has 40GB of Vram. qwen3.8:27b uses 23gb of vram in my setup. When I run gemma4:12b while qwen3.8:27b is loaded, 9.7gb is used. but the CPU is being used with a 16%:CPU 84%:GPU split. Why dose this happen? how can i fix it? will llama.cpp solve my issues?
This server only supports 1-2 users and I would like to run qwen3.8:27b and one more smaller model.

https://sikiru-ekunsumi.xyz/Projects.html

https://digitalspaceport.com/qwen-3-8-27b-review-prompts-and-vllm-settings/

u/Sik-Server — 5 hours ago

Got an OnexFly F1 from FB Marketplace for $390 (plus a fun repair story)

​

AMD Ryzen 7 7840U

AMD Radeon 780M

32GB LPDDR5X

2TB M.2 PCIe 4.0 NVMe SSD

Got this OnexFly F1 on Facebook marketplace for $390. (talked them down another $10 for showing up late).

The right trigger wasn't working at first. I contacted tech support, and they actually helped me even though I didn't buy from them! I got instructions to Open up the device and find a misplaced magnet. I found it, replaced it in the trigger, tested, saw that it was upside down, then flipped it, tested, and boom it worked.

I feel like I got a great deal, this handheld is more powerful than my laptop that I got for $500 6 years ago.

The cyberpunk benchmark is from the OnexFly F1.

I use sunshine/moonlight to stream from my main RTX 5070 PC.

u/Sik-Server — 1 month ago