Local AI doesn’t replace Claude—but my 24 GB Mac mini became a much better complement than expected

Local AI doesn’t replace Claude—but my 24 GB Mac mini became a much better complement than expected

I rely heavily on Claude, but I recently tested whether the M4 Pro Mac mini I already own could handle useful local models without becoming a dedicated AI appliance. I saw what Network Chuck did with the $50K Mac Ultra 4 node cluster and was hoping to not need the same.

It could. A GPT-OSS 20B MoE model ran at roughly 63.9 tok/s on 24 GB unified memory, while a smaller 9B dense model ran around 44.8 tok/s. The result is a useful reminder that model architecture matters: MoE models may have large total parameter counts while activating substantially fewer parameters for each generated token.

My takeaway is not “cancel your cloud AI subscription.” It is that local models can be a compelling companion for private experiments, quick generation, offline work, and workloads where you want direct control over the model runtime.

I captured the full test, including MLX vs. GGUF performance and the impact of my running containers: https://www.youtube.com/watch?v=9_-bT62YWAI

How are you dividing work between Claude and local models today?

u/silent_lurker_69 — 11 days ago
▲ 134 r/ClaudeHomies+3 crossposts

M4 Pro Mac mini, 24 GB unified memory: local LLM benchmarks changed my expectations

I wanted to see what my M4 Pro Mac mini with 24 GB unified memory could do with local LLMs, expecting it to be a small-model-only machine.

My result: GPT-OSS 20B in MLX FP4 ran at about 63.9 tok/s while my 16-container OrbStack homelab was still running. A 9B dense model was slower at about 44.8 tok/s, which was a good reminder that total parameter count alone does not predict speed—active parameters and architecture matter.

Other results:

  • 4B MLX model: roughly 78 tok/s
  • 9B MLX model: roughly 44.8 tok/s
  • GPT-OSS 20B MLX FP4 MoE: roughly 63.9 tok/s
  • MLX was about 19% faster than GGUF in my controlled back-to-back test
  • Shutting down the homelab changed GPT-OSS 20B throughput by only about 1.6%

This is not a replacement for a high-memory Mac Studio cluster if you need to load enormous models. But for interactive local AI on a $1,600-ish Mac mini, I found the capability much better than I expected. Compared to the $50K 4 Mac Ultra cluster with 2TB of unified memory Network Chuck tested, I'm pleased.

Video with the testing and numbers: https://www.youtube.com/watch?v=9_-bT62YWAI

u/silent_lurker_69 — 11 days ago

Burning money for a problem I thought was fixed

Gotta love Claude. A while back I had Claude audit my setup. It flagged a plugin — 37 dev-workflow skills, none of them fit what I actually build — and I denied it. Skill(agentsystem-core:*) in the deny list. Done, I figured.

Ran a deeper audit a few days ago for a video. Turns out denying a skill only blocks Claude from calling it. The plugin itself was still enabled, so its full menu — all 37 names and descriptions — was still loading into every message. 8,800 tokens, every message, for over a month, for a tool I hadn't touched once.

Disabled in the config. Alive in the context.

Real fix is different: flip the plugin to false in enabledPlugins. That actually unloads it — denying it just gates permission to call it.

Had Claude patch its own audit skill with this so it doesn't make the same call on the next dead plugin. If you're running anything similar, worth checking — "denied" and "disabled" are not the same thing, and the gap between them is exactly where tokens go to die.

Free skill if you want to run this on your own setup:https://jimmygarciaiii.gumroad.com/l/ghost-token-audit

Full breakdown: https://youtu.be/1UtD3f44JME

Anyone else find something they thought was already fixed?

u/silent_lurker_69 — 21 days ago
▲ 718 r/macmini+1 crossposts

Still can't believe that something about the size of my mouse can be a full home server — Home Assistant, 16 Docker containers, zero open ports

Used windows for 20 of my 25 years in IT. Moved to an iMac Pro in 2017 and have been Apple ever since. To think this little M4 Pro could run circles around my iMac Pro is wild to me. I've got Orbstack running on it instead of Docker Desktop — 16 containers: Home Assistant, go2rtc, Nginx Proxy Manager, Pi-hole, Unbound, Mosquitto, govee2mqtt, ESPHome, n8n, Plex, Jellyfin, Audiobookshelf, Portainer, Arcane. Homebridge and cloudflared run native on macOS outside the containers.

No open ports — cloudflared tunnels in through Nginx Proxy Manager, DNS-level blocking through Pi-hole + Unbound.

Recorded a walkthrough of the whole build if anyone wants to see it running: https://www.youtube.com/watch?v=LjxY-orR0mw.

Would love to know what you are running and things I could optimize. Thanks, let me know if I can answer any questions.

u/silent_lurker_69 — 25 days ago

Ted rescued me back in 2022, thought he was a BMC not sure what breed he is

I love my boy. He literally got me back on my feet and out of bed. It’s taken a few years but he’s over his trauma too and is the sweet loving dog he always wanted to be. I thought Ted was a Black Mouth Cur but that community said he’s not. Curious for your thoughts.

u/silent_lurker_69 — 2 months ago

I may have niched down way too much

I did the whole “here is your differentiator advantage” path. I’m don’t know if it’s the packaging, my camera presence, presentation skills or what the issue is. All of the AI and VidIQ audits say “professional looking” and good but it is not translating. 23 subs in 6 months.

If anyone has any suggestions or constructive criticism, my ears are open.

The channel is about all the stuff leadership doesn’t tell you and what can keep you from moving up the ranks. The goal is to help others the way I was helped during parts of my career. I’ve sat in the rooms and heard what keeps the highest performers from getting the promotions they desire. I spend most of my day at work coaching. Trying to help others virtually.

My other goal is to be able to better provide form my special needs son who will always live with me.

Channel: https://youtube.com/@fromittoinfluence

u/silent_lurker_69 — 2 months ago
▲ 59 r/Boxer

Sassy Pants says hi!

Super smart, super feisty, super cuddled, and lives up to her name!

u/silent_lurker_69 — 3 months ago
▲ 30 r/Blackmouthcur+1 crossposts

He talks 24x7x365. Or he plays. Or he catches flies with his tongue. Or just does anything goofy. I love this dog!

u/silent_lurker_69 — 3 months ago