Published: September 26, 2026
- For local AI, buy memory first and chip second. The M6 Mac mini tops out at 32GB of unified memory with 170GB/s of bandwidth. The M5 Pro goes to 64GB at 307GB/s. Memory size decides which models fit; bandwidth decides how fast they run.
- The M6 with 24GB or 32GB is the right Mac for most people running 7B-13B models. It starts at $899 with 16GB, and the 24GB configuration is the sweet spot for daily use with Gemma 4, Llama 3 8B, and gpt-oss 20B.
- The M5 Pro earns its $1,699 starting price when you want 30B-class models at conversational speed, or a 70B model at all. Its 48GB and 64GB configurations are Apple build-to-order only, and at that point a 128GB Strix Halo mini PC is the cheaper way to hold big models.
What Apple Changed in August 2026
Apple announced the new Mac mini on August 25, 2026, and it started shipping September 22. The lineup is unusual: the entry model gets the brand-new M6, Apple's first 2nm chip, while the step-up model gets the M5 Pro that debuted in the MacBook Pro earlier this year. Apple reportedly has no plans for an M6 Pro, so this pairing stands until the M7 family arrives in 2027. The upshot is that the more expensive Mac mini is not simply a faster version of the cheaper one; the two chips trade wins.
Both models gain Wi-Fi 7, Bluetooth 6, and standard 2.5Gb Ethernet (10Gb is an option on the M5 Pro), which removes the old base model's single Gigabit port. Both keep the same five-inch-square chassis as the 2024 models. Pricing rose: the M6 starts at $899, $300 more than the M4 Mac mini launched at, and the M5 Pro starts at $1,699.
The Two Chips Side by Side
| Spec | Mac mini M6 | Mac mini M5 Pro |
|---|---|---|
| Starting price | $899 (16GB / 256GB) | $1,699 (24GB / 512GB) |
| CPU | 12-core | 15-core, or 18-core option |
| GPU | 12-core with Neural Accelerators | 16-core, or 20-core option |
| Unified memory options | 16GB, 24GB, 32GB | 24GB, 48GB, 64GB |
| Memory bandwidth | 170GB/s | 307GB/s |
| Storage | 256GB to 2TB | 512GB to 8TB |
| Ports | Thunderbolt 4 | Thunderbolt 5 |
| Ethernet | 2.5Gb | 2.5Gb, 10Gb optional |
| Single-core (Geekbench, Macworld test) | 4,055 | 3,720 |
| Multi-core (Geekbench, Macworld test) | 22,695 | 34,281 |
The benchmark rows show the trade. The M6 is the faster chip per core; the M5 Pro has more cores and, for local AI, the number that matters more: nearly twice the memory bandwidth.
Why Bandwidth Is the Number for Local AI
When a language model generates a token, the hardware reads every active weight in the model from memory. That makes generation speed a bandwidth problem: divide memory bandwidth by model size and you have the ceiling. A 7B model at Q4 is about 4.5GB, so the M6 can in theory produce up to about 38 tokens per second on it and the M5 Pro about 68. In practice you get somewhat less, but the ratio holds: the M5 Pro is roughly 1.8x faster on any given model.
The second number is how much of the unified memory macOS lets the GPU use. By default it is about 75%, so a 16GB Mac gives models roughly 12GB, a 24GB Mac about 18GB, 32GB about 24GB, 48GB about 36GB, and 64GB about 48GB. You can raise that limit with a terminal command, but the defaults are a fair guide to what fits comfortably.
One caveat on speed: independent tokens-per-second tests of the M6 and M5 Pro running local models were still thin at publication. The figures below are derived from bandwidth and from the M4 generation's measured results, so treat them as estimates rather than benchmarks.
Which Mac mini for Which Model
| Configuration | Memory for Models | Runs Well | Estimated 7B Speed | Where to Buy |
|---|---|---|---|---|
| M6, 16GB | ~12GB | Gemma 4 E4B and 12B, Llama 3 8B, Mistral 7B | 25-35 tok/s | M6 Mac mini on Amazon |
| M6, 24GB | ~18GB | Everything above plus gpt-oss 20B and Gemma 4 26B MoE | 25-35 tok/s | M6 Mac mini on Amazon |
| M6, 32GB | ~24GB | Everything above plus Gemma 4 31B and Qwen3-Coder 30B, at modest speed | 25-35 tok/s | Apple (build-to-order) |
| M5 Pro, 24GB | ~18GB | Same models as the 24GB M6, nearly twice as fast | 45-60 tok/s | M5 Pro Mac mini on Amazon |
| M5 Pro, 48GB | ~36GB | 30B-class dense models with long context, gpt-oss 20B at full context | 45-60 tok/s | Apple (build-to-order) |
| M5 Pro, 64GB | ~48GB | Llama 3.3 70B at Q4 with short context, roughly 5-7 tok/s | 45-60 tok/s | Apple (build-to-order) |
The Case for the M6
Most people running local AI at home want a 7B-13B model that answers quickly, and the M6 delivers that. Twenty-five to thirty-five tokens per second is faster than most people read, and the 24GB configuration holds Gemma 4's 26B Mixture of Experts model and OpenAI's gpt-oss 20B with room for context. The M6 is also the faster chip for everything else you do on the machine, since single-core speed is what most apps feel.
The M6 is the better buy when your models stay under 20GB, when the Mac will also be your everyday computer, and when the difference between 30 and 55 tokens per second is not worth $800. Get the 24GB configuration unless the 16GB price is the only one that fits; 16GB is workable but you will bump into the ceiling the first time you try a 20B model with a long document.
Apple 2026 Mac mini M6 (24GB, 512GB) on Amazon. This is the configuration we recommend, and it has been running about $30 under Apple's $1,299 list price. The 16GB base model and the 32GB build are selectable from Amazon's listing or from Apple.
The Case for the M5 Pro
The M5 Pro is the Mac for anyone who has already outgrown a 7B model. Its 307GB/s bandwidth makes 30B-class dense models feel conversational instead of sluggish, and it is the only Mac mini that can be configured to 64GB, which is the minimum for holding a 70B model at Q4. It also adds Thunderbolt 5 and the 10Gb Ethernet option, which matter if the Mac is serving models to other machines on your network.
The M5 Pro is the better buy when you run 30B models daily, when several people or apps hit the model at once, or when you want the single fastest quiet box that runs macOS. Note that memory is soldered: the configuration you order is the configuration you keep, so buy for the models you expect to run in two years.
Apple 2026 Mac mini M5 Pro (24GB, 512GB) on Amazon. Amazon carries the 24GB configuration; the 48GB and 64GB builds are ordered from Apple directly.
When Neither Is the Answer
If the model you want is bigger than 48GB, or the 64GB M5 Pro's price makes you wince, look at the AMD side. A Ryzen AI Max+ 395 mini PC with 128GB of unified memory holds a 70B model and 120B-class Mixture of Experts models that no Mac mini can load, runs Linux, and slots into a VLAN-isolated home lab the way a server should. Bandwidth is lower than the M5 Pro's, so mid-size models run a little slower, but capacity is the reason to buy it. We compare the options in Can a Mini PC Run a 70B Model?
And if you are not sure you need a Mac at all, our mini PC tier guide for local AI covers every option from an $80 Raspberry Pi up.
Setting Up Local AI on the Mac mini
The Mac is the easiest platform to get started on. Download the Ollama app from ollama.com, or install it from a terminal with the one-line script. Recent Ollama versions run compatible models on Apple's MLX framework by default on Apple Silicon, which is the fastest path available. When the first-run screen offers to sign in or continue locally, choose local; a signed-in install can route requests to cloud-hosted models, which defeats the point.
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
# Good first models for a 16-24GB Mac mini
ollama pull gemma4:e4b
ollama pull llama3.1:8b
ollama pull gpt-oss:20b # 24GB and up
# 32GB M6 or any M5 Pro
ollama pull gemma4:31b
ollama pull qwen3-coder:30b
# 64GB M5 Pro only
ollama pull llama3.3:70b
ollama run gemma4:e4b
Two settings worth knowing. Ollama defaults to a short context window; raise num_ctx in a Modelfile or per request if you feed it long documents. And if you want the GPU to use more than the default 75% of memory, sudo sysctl iogpu.wired_limit_mb=<megabytes> raises the ceiling until the next reboot; leave at least 6-8GB for macOS itself.
For a browser chat interface, LM Studio is a native Mac app that also supports MLX, or run Open WebUI in Docker Desktop and point it at Ollama.
Keep It Off the Flat Network
A Mac mini running Ollama and a web interface is a server, and the security advice from our tier guide applies here too. Put it on its own network segment at the router or switch, keep Ollama bound to localhost unless another machine needs to reach it, point its DNS at Pi-hole so you can see what it talks to, and read our MCP security guide before adding tool integrations.
One more monthly fee to kill
Renting your modem? That's $120+ a year.
Own it instead. It pays for itself in months. 90-day warranty and 30-day returns.
Pick your provider:
Frequently Asked Questions
Is the M6 or the M5 Pro Mac mini better for local AI?
The M6 with 24GB is the better value for 7B-13B models, which covers most home use. The M5 Pro is better for 30B-class models, for a 70B model (64GB configuration only), and for serving several users at once, because its memory bandwidth is nearly double.
How much memory do I need on a Mac mini to run a 7B model?
16GB is enough for a 7B or 8B model at Q4 with a normal context window. 24GB is the comfortable choice, since it adds room for 20B-class Mixture of Experts models and longer documents.
Can the Mac mini M6 run Gemma 4?
Yes. The 16GB M6 runs Gemma 4 E2B, E4B, and 12B. The 24GB M6 adds the 26B Mixture of Experts model. The 32GB M6 or any M5 Pro runs the 31B dense model, and the M5 Pro runs it noticeably faster.
Can a Mac mini run a 70B model?
Only the M5 Pro with 64GB, and only at Q4 with a short context window, at roughly 5-7 tokens per second. If 70B models are the goal, a 128GB AMD Strix Halo mini PC holds them with room to spare and costs less than the 64GB M5 Pro.
Is the older M4 Mac mini still worth buying for local AI?
At a closeout price, yes, for small models. The M4 has 120GB/s of bandwidth and tops out at 32GB, so it runs 7B-13B models at readable speed. It lacks Wi-Fi 7 and 2.5Gb Ethernet, and the M6 is meaningfully faster at AI tasks, so the M4 only makes sense at a steep discount.
Does Ollama run natively on the 2026 Mac mini?
Yes. Ollama supports Apple Silicon through Metal and, in its newest builds, runs compatible models on Apple's MLX framework by default. LM Studio and llama.cpp also run natively. No Docker or Linux is required, though Docker Desktop is available if you want the sandboxing.

