Last updated: August 2026
Key Takeaways
- Apple announced the new Mac Studio on August 25 with M5 Max and M5 Ultra chips. Preorders are open now, availability begins September 22, and the ceiling is 512GB of unified memory, not the rumored 768GB. The 512GB tier itself does not arrive until late October, unpriced.
- A 512GB M5 Ultra runs everything through the 340GB class, including Kimi K2.7, in a single quiet box for the first time. The largest open model, Kimi K3, still needs the new Thunderbolt 5 cluster path.
- Most local AI workloads still never needed an Ultra. The honest sweet spots are a Mac Mini M4, a 64GB unified-memory mini PC, a used RTX 3090, and the new 128GB M5 Max tier.
The new Mac Studio is official. Apple announced M5 Max and M5 Ultra models on the morning of August 25, with preorders open the same day and availability beginning September 22. The M5 Ultra tops out at 512GB of unified memory, not the 768GB the rumor cycle promised, and that 512GB configuration does not ship until late October, at a price Apple has not yet named.
This page has tracked the wait-or-buy question since July, when the honest read was that the memory market, not Apple's calendar, would decide this launch. That is how it played out. Below: what actually shipped, what each memory tier genuinely runs, the cluster math that changes the Kimi K3 answer, and whether you should buy, wait, or skip.
What Apple Actually Announced
Two machines, one chassis, and a memory ladder with a missing rung. The M5 Max model pairs an 18-core CPU with an up-to-40-core GPU, Neural Accelerators in every GPU core, and up to 128GB of unified memory at 614GB/s. The M5 Ultra scales to a 36-core CPU and an 80-core GPU, the first Ultra chip to get Neural Accelerators, with up to 512GB of unified memory at 1.2TB/s. Apple rates it at up to 4.3x the peak AI compute of the M3 Ultra and up to 4x faster LLM prompt processing in LM Studio, per its August 25 press release. The chassis adds Thunderbolt 5, Wi-Fi 7, Bluetooth 6, and a PCIe Gen 6 SSD architecture rated at twice the storage speed, a detail that matters when loading a model means a quarter-terabyte read.
Pricing is where the story lives. Per Apple's announcement, the M5 Max model starts at $2,499 and the M5 Ultra at $5,499. The fine print: that $5,499 buys a binned 30-core CPU, 64-core GPU part, while the full 36-core, 80-core chip in Apple's benchmark slides starts at $6,799, per launch-day configurator pricing compiled by Daring Fireball. Ultra memory jumps straight from 96GB to 256GB for $4,000 with nothing in between, and the 512GB option is listed only as coming in late October. AppleInsider notes the current maximum build already reaches $18,299 before that tier exists.
| Spec | M5 Max | M5 Ultra | M3 Ultra (outgoing) |
|---|---|---|---|
| CPU | 18 cores (6 super, 12 performance) | 30 cores base, up to 36 (12 super, 24 performance) | Up to 32 cores |
| GPU | Up to 40 cores, Neural Accelerators | 64 cores base, up to 80, Neural Accelerators | Up to 80 cores, no Neural Accelerators |
| Max unified memory | 128GB | 512GB (late October); 256GB until then | 512GB at launch; sold only in 96GB by mid-2026 |
| Memory bandwidth | 614GB/s | 1.2TB/s | 819GB/s |
| Availability | September 22 | September 22; 512GB tier late October | Replaced August 25, 2026 |
Specifications from Apple's August 25, 2026 announcement and launch-day configurator listings. Performance multipliers are Apple's own figures from July 2026 internal testing, not independent benchmarks.
The 768GB Rumor, Resolved
The 768GB configuration did not ship. The M5 Ultra's ceiling is 512GB, matching the outgoing M3 Ultra's original maximum, and even that tier arrives two months after the rest of the line. When the 768GB figure circulated this summer, this page called it a tested ceiling, not a shipping promise, and today Apple confirmed the distinction mattered. The same DRAM squeeze that stripped the M3 Ultra down to a single 96GB configuration over the course of 2026 shaped this launch in public view: the top tier exists on the spec sheet, ships late, and carries no price. The background on why is in our memory shortage analysis.
One more monthly fee to kill
Renting your modem? That's $120+ a year.
Own it instead. It pays for itself in months. 90-day warranty, free returns.
Pick your provider:
What 512GB Actually Runs
The rule of local inference has not changed: the model's weights must fit entirely in memory or generation speed collapses. What changed today is which weights fit in an Apple box. The floors below are roughly 4-bit quantization figures drawn from our local AI models by VRAM guide and our Kimi K3 hardware reality check, mapped against the two new memory ceilings.
| Model class | Memory floor (Q4 weights) | Fits in M5 Max 128GB? | Fits in M5 Ultra 512GB? |
|---|---|---|---|
| 70B-class dense (Llama tier) | 35 to 40GB | Yes, with room to spare | Yes, several at once |
| Step 3.7 Flash, 198B sparse MoE | 110 to 120GB | Borderline: needs wired-memory tuning, thin context headroom | Yes, comfortably |
| Frontier-adjacent 280B-class MoE | Roughly 130GB and up | No | Yes |
| Kimi K2.7-class open agent | About 340GB, needing 350GB-plus total | No | Yes, the first single box that holds it |
| Kimi K3, 2.8T parameters | 1-bit build is 594 to 620GB; 650GB-plus combined floor | No | Not in one box; yes in a two-box cluster |
Weights-only floors at roughly 4-bit quantization, per our VRAM-tier guide and K3 reality check. The operating system, inference runtime, and context window add overhead on top.
Read the table honestly and the transformative row is the fourth one. The 340GB-class open agents, which required multi-GPU server hardware a year ago, now fit in one quiet box you can order today for September delivery, at 256GB, or wait for the 512GB tier and hold them with real context headroom. The 128GB column deserves equal honesty in the other direction: the 110-to-120GB sparse models fit an M5 Max on paper, but macOS reserves a slice of unified memory for itself, so loading them means raising the GPU wired-memory limit by hand and accepting thin context margins.
The Cluster Asterisk: Kimi K3 and Thunderbolt 5
One 512GB box still cannot hold the largest open model, but two can, and Apple just made that path official. The M5 Ultra's headline feature for local AI is built-in Thunderbolt 5 clustering with RDMA, which pools memory across multiple Mac Studio systems; Apple says a four-box cluster delivers up to 3x the inference of a single machine. Kimi K3's smallest usable build, Unsloth's 1-bit dynamic GGUF, weighs 594 to 620GB with a 650GB-plus combined memory floor, per our K3 reality check, which named a 512GB Mac Studio linked to a second node as the narrowest realistic path back in July. Apple has now productized that exact architecture.
The honest ledger before anyone budgets for it: two high-memory Studios price in used-car territory, with AppleInsider expecting the maxed 512GB build alone to push well past $20,000; the open-source inference stacks have no track record on Apple's RDMA clustering yet, and K3's llama.cpp support was still living in a development branch weeks after the weights shipped; and a cluster is a rack project, not a quiet desk machine. The door is open. It is not cheap, and it is not turnkey.
Bandwidth Decides Your Tokens per Second
Memory capacity decides what loads; memory bandwidth decides how fast it talks back. The M5 Ultra's 1.2TB/s is roughly 50 percent above the M3 Ultra's 819GB/s and nearly double the M5 Max's 614GB/s, which makes it the fastest way to converse with a large model outside a GPU server. For calibration: a used RTX 3090 moves 936GB/s across its 24GB, and the 128GB Ryzen AI Max machines measure around 210 to 220GB/s in the real world, as our Ryzen AI Max 395 reality check found. That is the trade in one sentence: the PC side sells capacity cheaply and pays in speed, Apple sells both and charges accordingly.
Buy, Wait, or Skip: The Decision Tree
If you need maximum unified memory now
For the first time since spring, Apple will sell you one. The 256GB M5 Ultra is orderable today for September 22 arrival, which ends the stretch when the new-Apple path simply did not exist above 96GB. The used and refurbished market for 256GB and 512GB M3 Ultra units remains the value bridge, and those machines stay a capacity purchase rather than a speed one: the new generation beats them on CPU and AI compute, but their memory pools still hold what a 96GB machine cannot. If the 512GB M5 Ultra is the point, plan to order within hours of the late-October window opening, because the outgoing model's history says the top tier sells constrained first and vanishes second.
If your models fit in 24 to 128GB
This is most readers, and the honest answer survives its second product cycle: you never needed an Ultra. The best models per gigabyte live in this range, as our local AI hardware guide maps in detail. A Mac Mini M4 covers the 12B-to-30B class that handles everyday assistant and coding work. A 64GB unified-memory mini PC runs the strongest small MoE models with full working context for a fraction of Studio money, and our mini PC guide covers the whole tier. A used RTX 3090 remains the best dollar-per-VRAM value of 2026 for GPU inference.
Check Price on Amazon: Apple Mac Mini M4
Check Price on Amazon: MINISFORUM X1 Pro 370 (64GB)
Check Price on Amazon: NVIDIA RTX 3090 24GB (Renewed)
At the top of this range, the 40-core M5 Max configured to 128GB becomes the first-party Apple option for the one-128GB-computer class that our DeepSeek V4-Flash reality check maps, with AMD's Ryzen AI Max machines as the far cheaper, far slower PC counterweight. Match the box to the models you actually run, and none of these purchases gets better by waiting for a machine you would not fill.
If only the M5 Ultra will do
Then order early and order precisely. Treat the 512GB tier as a bonus if it ships on schedule rather than the plan you depend on, and mind the binned-base trap: the advertised starting price buys the 30-core, 64-core version of the chip, not the one in Apple's benchmark slides. Budget for the possibility that the unpriced 512GB tier lands above what the outgoing model's equivalents cost, because everything the shortage has touched so far has moved in exactly one direction. Our guide to when RAM prices will come down explains why relief is not close.
The Part Apple Cannot Control
The most vertically integrated company in consumer computing just published its supply chain's terms in its own price list. The M5 Max opens at $2,499, five hundred dollars above the M4 Max's $1,999 debut, a bump the line quietly absorbed mid-cycle; 9to5Mac notes the new price merely matches where the old model already sat. The M5 Ultra opens at $5,499, fifteen hundred above the M3 Ultra's $3,999 launch, though only two hundred above where the shortage had already pushed the outgoing machine. Memory climbs from 96GB to 256GB in a single $4,000 step, and the flagship tier ships late and unpriced. Three DRAM makers set those terms, not Cupertino. Whoever controls the infrastructure controls the experience, and this year the infrastructure is memory. The lesson lands at every scale, including yours: hardware you own outlasts product cycles you do not control, which is why the machines in our guides are chosen for what they run today, not what a supply chain might permit in October.
Frequently Asked Questions
When does the new Mac Studio come out?
Preorders opened August 25, 2026, and availability begins September 22 in stores and for delivery. The exception is the 512GB M5 Ultra configuration, which Apple lists as coming in late October without a stated price. Every other configuration, including the 256GB M5 Ultra, ships in the September wave.
Did the M5 Ultra Mac Studio ship with 768GB of RAM?
No. The rumored 768GB configuration never appeared; the M5 Ultra's ceiling is 512GB of unified memory, the same maximum the M3 Ultra launched with in 2025. The supply-chain reporting behind the 768GB figure described tested silicon, and the DRAM shortage appears to have decided what actually shipped.
How much does the Mac Studio M5 Ultra cost?
It starts at $5,499 with a 30-core CPU, 64-core GPU, and 96GB of memory, per Apple's August 25 announcement. The full 36-core, 80-core chip starts at $6,799, the 256GB memory upgrade adds $4,000, and launch-day coverage puts the current maximum build at $18,299. The 512GB tier is not yet priced.
Can the M5 Ultra run Kimi K3?
Not in a single box. K3's smallest usable build needs a 650GB-plus combined memory floor, above the 512GB ceiling. It becomes possible across two Mac Studios using the new Thunderbolt 5 clustering with RDMA, which pools memory between machines, though the open-source inference stacks have yet to prove themselves on that path.
Should I get the M5 Max or M5 Ultra for local AI?
Match the chip to your model sizes. The 128GB M5 Max covers everything through the 110-to-120GB sparse class, with tuning, at 614GB/s. The M5 Ultra exists for the 130GB-and-up tier, topping out with 340GB-class agents in one box, and its 1.2TB/s bandwidth roughly doubles generation speed on identical models.
Is a used M3 Ultra Mac Studio still worth buying?
For capacity, yes. Used and refurbished 256GB and 512GB M3 Ultra units remain the cheapest large unified-memory pools on the market until the 512GB M5 Ultra ships, and they hold models a 96GB machine cannot. Buy one for memory, not speed: the M5 generation clearly outruns it on compute.

