r/MacStudio

▲ 4 r/MacStudio+3 crossposts

Opinion on purchasing a used Mac mini M4 and Studio Display Standard Glass

Good morning everyone, I am about to buy a Mac mini M4 (256 GB, 10/10 CPU/GPU cores) and a Studio Display with standard glass for a total of €1,500. The display was purchased on Amazon in January 2026, and the Mac mini in 2025.
Do you think €1,500 is a good price?
Thanks for your answers!

reddit.com
u/AmazingMonarch — 3 hours ago
▲ 30 r/MacStudio+1 crossposts

Tried DFlash 2 on a Mac Studio M3 Ultra. Already ~63 tok/s with MTP — DFlash didn’t beat it.

You may have seen the clip: Qwen3.8-27B at 70 tok/s on a Mac, billed as 4.6× faster, day-one in oMLX.

I ran that on a Mac Studio M3 Ultra (60 GPU cores, 256 GB).

What I actually got, Qwen3.8-27B 4-bit, thinking off:

• No draft (plain AR): about 22 tokens/sec • Official MTP (already running here): about 63 • DFlash 2 (oMLX 0.6.3rc1): about 65

So speculation is real — roughly 3× vs a dumb decode. DFlash 2 vs the MTP that’s already on this Studio: a wash.

Where the launch numbers come from: • 70 tok/s is an M5 Max MacBook Pro demo, not a Studio • 4.6× is Muse Glimmer on an NVIDIA H200, not Qwen 27B on a Mac • Inco’s own Qwen3.8 table is 2.7–3.4× vs AR, and only about 1.2× vs MTP

Thinking-on didn’t change the story (DFlash ~62 vs MTP ~56 on a short prompt).

If you already run official MTP on a Studio, I would not rebuild the stack for DFlash 2. If you’re still decoding autoregressively, turn on MTP — that’s the jump.

Happy to share setup notes in the comments.

reddit.com
u/New_Guitar_9121 — 1 day ago

Which Mac for local LLM research?

Bear with me on this one. I currently own (personally) a 16" M5 Max 128gb/4tb machine. It's my daily driver. I run local models on this fairly regularly, and it works great -- but does get very warm when doing long duration inference.

I recently started a new job as an AI Architect and I've convinced my VP that I should have a company owned machine with enough power to do R&D using local LLMs. I can pick either another 16" M5 Max 128gb/2Tb machine, or I can instead go with an M3 Ultra Studio 32/80 96gb/2Tb for the same price.

I know the cooling of the Studio is going to be hugely beneficial, and memory bandwidth is slightly better (as is the GPU count).. and I can continue using my *personal* Macbook to ssh/screen share into the Studio, making it seem, on paper at least, that the Studio is a no-brainer.

My concern is memory. 96gb is likely fine for most of the upcoming LLM models. It seems either a model requires a 512gb machine, or it can run on something more pedestrian (ie 32/64gb) with little overlap.

I need to act while I have this opportunity -- I don't know if the budget will exist for an M5 Ultra if/when one is released. Even IF an M5 Ultra is released with capacity for 128gb/256gb, it'll likely be too expensive for my department budget anyway.

Thoughts?

Update - decided to request the 96gb M3U 80 core Studio. Thanks everyone for your input!

reddit.com
u/CatchInternational43 — 3 days ago
▲ 19 r/MacStudio+1 crossposts

How I made DeepSeek V4 Flash 12x faster on an M3 Ultra

I work with a Mac Studio M3 Ultra (512GB) serving DeepSeek V4 Flash on antirez/ds4 ("DwarfStar"). A chat turn took between 6 and 20 seconds. Now it takes 1.6s.

Kernels (+21% cold prefill at 64k, bit-exact). DeepSeek V4's sparse attention runs a "lightning indexer" that dominates long-context prefill. Three stacked PRs: threadgroup-tiled scorer (#830), register-blocked K-resident scorer (#831), and a streaming top-512 replacing the bitonic sort + merge cascade (#832). 392 → 475 t/s at 64k. Logits byte-identical at every context frontier, everything behind rollback envs.

Cache, the 10x (this part is useful way beyond DeepSeek or ds4). If you serve any model behind a chat API, check whether your client can actually hit the engine's KV cache, because a stateless client usually can't:

  • The live session ends in the exact reply the engine sampled. If your client doesn't resend that reply byte for byte (exact text, or the tool call by id), the prefix never matches and you re-prefill every turn. Replay it verbatim and cached_tokens ≈ everything.
  • Prewarming: max_tokens: 0. Send the conversation with zero tokens requested and the engine prefills it and stops exactly at the prompt, so the next real request extends the cache. max_tokens: 1 doesn't work: the one sampled token becomes part of the session and every later request misses. Great for warming a room/session before anyone asks anything, or re-warming after your slot got evicted.

Recipe with measurements: ds4#816.

Also you can read the things I tried that didn't work (single-stream decode is a wall, and I learned two Metal scheduling laws killing it) here:

https://adriangalilea.com/deepseek-on-a-mac-studio

EDIT: Regarding cache, I failed to mention there was an engine bug that was part of the 10x: the disk cache's eviction policy scored the only checkpoint a chat client can reuse as the first victim, so once the disk filled, every request prefilled from zero. Fixed in ds4#814.

u/Adrian_Galilea — 3 days ago

Sealed Mac Studio M1 Max 10-core CPU 32-core GPU 64GB/1TB - keep or not?

I’ve retrieved a sealed, unused Mac Studio M1 Max 10-core CPU 32-core GPU 64GB/1TB from storage. AppleCare+ has been kept up-to-date. Should I break this open or put towards an upgrade? Thanks in advance.

reddit.com
u/agentm74 — 2 days ago

Which UPS to buy for Mac Studio M3 Ultra in india

I recently got my mac studio m3 ultra, Im looking for a good ups for it. While searching I got to know that these expensive hardware should be used with pure sine wave ups. Looks like APC and Cyberpower are two big names in this but APC doesn't offer pure sine wave in india and Cyberpower is way out of my budget.

Which UPS you guys are using, any recommendations?

reddit.com
u/SureTrouble8022 — 3 days ago
▲ 381 r/MacStudio

My m3 ultra 512gb RAM setup

Mac m3 ultra 512gb RAM 4tb ssd— best purchase I ever made.
all I heard when I bought it for 11K was that it was overkill and that I wouldnt ever use all that RAM

Now over a year later and I am getting SOTA models completely locally , and consistently have 1000 chrome tabs open (not shown)

I use it for work with heavy media editing and CAD / CAM software , along with Local LLM workflows , as well as heavy claude code usage (3 claude max20 x subs) as well as my inveterate browser tab maxxing .

What do you guys use your mac studios for , and do you regret or still love them ?

u/tokentrillionaire — 5 days ago

Is an M1 Ultra Mac Studio with 128GB of RAM 1TB HD a good buy for $2,500?

I’m considering a used M1 Ultra Mac Studio with 128GB of unified memory for $2,500. My primary use would be running large local LLMs, coding agents, and other homelab/AI workloads not model training.
The 128GB unified memory is the main attraction, but I’m wondering whether the M1 Ultra is now too dated to justify that price. How is real world performance with MLX, Ollama, and larger quantized models? I’m particularly interested in generation speed, prompt processing, and which model sizes remain reasonably usable.
Would you buy this at $2,500, or put the money toward a newer Mac, Strix Halo system, or multi GPU PC instead? I’d especially appreciate feedback from anyone still running local models on an M1 Ultra.

reddit.com
u/Rabbit-09 — 4 days ago

Apple’s desktop Mac refresh strategy feels really unfair to customers

The M5‑Max chip is already mass‑produced and shipping in MacBook Pro.
Yet Mac Studio and Mac mini are still selling M4‑series hardware at full original price, no price cuts.

It’s frustrating for buyers: newer silicon already exists, but desktop Mac customers pay the same money for last‑gen chips.

I pre‑ordered a Mac Studio M4‑Max with estimated Sep 23 shipping. If Mac Studio M5‑Max launches in October before my order ships, I could end up receiving superseded hardware right out of the box.

Are there any real‑world cases where unshipped Mac Studio orders got automatically upgraded like what happened with some US MacBook M4→M5 orders?

reddit.com
u/Mobile-Chocolate985 — 4 days ago

Mac Studio M3 Ultra — 3+ Month Wait With No ETA

Specs: Apple Mac Studio M3 Ultra, 28-core CPU, 96 GB RAM, 2 TB storage.

I’m based in the EU and ordered a Mac Studio from my local authorised Apple retailer on May 9th. There’s no official Apple Store in my country, so I ordered through the retailer.

When I placed the order, I was told the expected delivery time would be around 3–4 weeks. After that period passed, I was told the delay was due to RAM shortages and that Apple apparently “does not provide an ETA.”
It has now been more than 3 months, and I still haven’t received any meaningful update from either the retailer or Apple. I don’t even have an estimated delivery date at this point.

I understand that supply shortages can happen, especially with custom configurations, but waiting more than three months with zero ETA or clear information is getting pretty frustrating. At this point, I’m starting to wonder whether I’m being given the runaround.

Has anyone else in the EU ordered a similar Mac Studio configuration around May and experienced the same thing? If so, how long did you end up waiting, and did you eventually get any kind of ETA?

reddit.com
u/Ilovetix — 5 days ago
▲ 1.7k r/MacStudio+2 crossposts

Unified Memory Architecture still unbeatable (when LLM size matters)

u/AntLife255 — 6 days ago
▲ 291 r/MacStudio+1 crossposts

Local Micro Center grab. 256 / 2TB

I have a 96GB on order with Apple and have Trackalaker scanning B&H for any inventory. I was pleasantly surprised to see one in stock at Micro Center online when I was browsing for something else.

This is replacing a M4 mini pro that has been running local AI models for projects we are not exposing online.

I cannot wait to get everything set up !

u/Magnum3k — 6 days ago

Mac Studio M4 Max vs Custom Windows PC for dual Samsung Odyssey OLED G9 49” trading setup — need advice”

Mac Studio M4 Max vs Custom Windows PC for dual Samsung Odyssey OLED G9 49” trading setup — need advice”

Hey everyone, need some real world advice
before I make a big purchase decision.

MY MONITOR SETUP:
- 2x Samsung Odyssey OLED G9 49"
(5120x1440 240Hz) stacked vertically
- Running both monitors simultaneously
- Used purely for stock trading
(web based platforms like TradingView)

OPTION 1 — Mac Studio M4 Max
- 16-core CPU
- 40-core GPU
- 64GB Unified Memory

Questions:

  1. Does Mac Studio M4 Max support full
  2. 5120x1440 resolution on BOTH G9
  3. monitors simultaneously?
  4. What is the actual maximum refresh
  5. rate achievable on these monitors
  6. with Mac Studio — is it really
  7. capped at 120Hz?
  8. Does the 240Hz cap actually matter
  9. for trading use specifically or
  10. is 120Hz perfectly fine?
  11. Any web based trading platform
  12. compatibility issues on macOS?
  13. Anyone running dual 49" ultrawides
  14. on Mac Studio M4 Max — how is
  15. the real world experience?

OPTION 2 — Custom Windows PC
- Intel Core Ultra 7
- 32GB RAM
- 1TB Gen4 NVMe SSD
- NVIDIA RTX 5060 8GB

Questions:

  1. Will RTX 5060 drive both G9 49"
  2. monitors at full 5120x1440 240Hz
  3. simultaneously without issues?
  4. Is this config overkill, just right,
  5. or underpowered for a dual 49"
  6. ultrawide trading workstation?
  7. Any GPU or driver issues with dual
  8. Samsung G9 OLED monitors on Windows?

MY CONSTRAINTS:
- No space for a full tower PC on desk
- Need compact, silent, low power
consumption machine
- Already have a Windows laptop for
other work
- This machine is purely dedicated
for trading on dual monitors

Which would you choose and why?
Especially looking for people who
are actually running dual 49"
ultrawides for trading.

Thanks!

reddit.com
u/kennykorani — 5 days ago
▲ 17 r/MacStudio+1 crossposts

My little Apple favorite.

I already have 2 headless Macs. M4 Mac Mini, and M2 Max Studio. My old 2013 Intel i7 Macbook Pro was starting to run so hot and noisy that I finally had to get rid of it. I was trying to make audio microphone recordings. But now I has no laptop except a Dell i5 Win 11 box. When started reading and watching about the Neo, the more I thought about it, it started making sense. I could not be happier. It has even replaced my TV. Every possible app is available either with a web version a Mac version or an iOS version. DOES NOT RUN HOT or even warm, even without a fan. I was a initially concerned about the low 8GB RAM. It has been no problem. I use it to remotely log into my headless Macs. For example, when my Mac Studio is doing a heavy process, I can control it completely from the Neo!

reddit.com
u/gretschplayer11 — 5 days ago

256GB Studio - local llm stack update

Been running local LLMs on a Mac Studio 256GB for a while. I used to keep the big 122B model loaded all day because it felt like the “real” one.

Then I ran the same 30 hard tests on Qwen3.8-27B and that 122B. Both scored 24/30. Same on the coding ones. 122 is faster on long answers. 38 is smaller, can see images, and leaves enough RAM to run Flash next to it.

So I simplified:

• Daily / coding: Qwen3.8-27B

• Planning / review: Heretic 35B

• Second opinion (not another Qwen): Gemma 4 31B

• Long documents only: 122B — I start it on purpose now

• One heavy model: DeepSeek V4 Flash. Deleted Hy3.

Also stopped the chat app from auto-loading 122. One wrong click used to eat most of the machine.

Not saying this beats ChatGPT. 38 is just a really good local coder on Apple silicon. Still one request at a time unless I give it two slots.

reddit.com
u/New_Guitar_9121 — 6 days ago