u/Zestyclose-Bison-340

▲ 25 r/OdysseusAI+1 crossposts

8GB VRAM + 32GB RAM - how far can I realistically go with local LLMs?

Hi, I'm new to running LLMs locally and I'm trying to figure out what my laptop can actually handle.

Specs:

  • RTX 5060 Laptop GPU — 8GB VRAM
  • 32GB RAM (31.2GB usable)
  • Ryzen 7 260 — 8 cores / 16 threads
  • Radeon 780M iGPU

I'm a bit confused about how far I can go with bigger models.

I know something like a 30B model won't fit in 8GB VRAM, but from what I understand I can use system RAM as well and offload part of the model to the CPU/RAM.

So what would be the realistic limit with this setup? Could I use something around 20B–30B at Q4, or does it become painfully slow once too much of it is running from RAM?

Also, how much difference does Q4/Q5/Q6 make in this situation?

I'm mostly interested in understanding what my laptop is capable of and what kind of local AI stuff I could reasonably use it for. Any tips from people with similar hardware would be appreciated.

reddit.com
u/Zestyclose-Bison-340 — 9 days ago