Last updated: August 2026
Key Takeaways
- The verified picks by tier, August 2026: Gemma 4 12B for 16GB, the new Qwen3.8-27B for 24-32GB, DeepSeek V4-Flash 0731 and Step 3.7 Flash sharing the 128-256GB tier, and GLM-5.2 at 384GB and up.
- Total parameters set your memory bill; active parameters only set speed. A model's quantized file size is your floor, and roughly 0.6GB per billion parameters is the Q4 rule of thumb.
- This page replaces two earlier guides at one permanent address. Every pick was re-verified on August 26 against its repository, license file, and runtime support — and gets re-verified on a dated schedule.
The best verified local AI models right now: Gemma 4 12B on 16GB machines, Qwen3.8-27B on a 24GB card, DeepSeek V4-Flash 0731 or Step 3.7 Flash on 128GB unified memory, and GLM-5.2 if you own server-class capacity. Three of the six tier picks changed since June — the field moved, and this page moved with it. Find your memory number below; every claim carries its source and date. If you have not picked hardware yet, start with our local AI hardware guide, then come back for the software side.
Jump to your tier: 8-12GB, 16GB, 24-32GB, 48-96GB, 128-256GB, 384GB and up, or the vetting checklist.
How to Read the Tiers
Three kinds of memory can hold a model — dedicated GPU VRAM, unified memory on Apple Silicon and AMD's Ryzen AI Max platforms, and plain system RAM — and the same two rules govern all of them. First: total parameters set your memory requirement, because every expert in a mixture-of-experts model must be resident for routing to work; the active count only sets generation speed. When a spec sheet leads with the active number, find the total before you plan a purchase. Second: a model's quantized file size is your floor, before the operating system, the runtime, and your conversation context take their share. Quantization to roughly 4-bit (Q4) cuts memory to about a quarter of trained precision with minor quality loss for most work — about 0.6GB per billion parameters — and the mechanics behind modern compression live in our TurboQuant explainer.
| Tier | Typical hardware | What fits at Q4 |
|---|---|---|
| 8-12GB | RTX 3060 12GB, older GPUs, 8-16GB laptops | Up to ~14B dense; small MoE |
| 16GB | RTX 4060 Ti 16GB, most modern laptops | Up to ~20B dense with context room |
| 24-32GB | RTX 3090/4090/5090, 32GB mini PCs | ~27-35B dense; mid MoE |
| 48-96GB | 64GB mini PCs, dual GPUs, 64-96GB Macs | ~70B dense; 35B-class MoE at high quality |
| 128-256GB | Ryzen AI Max+ 395 boxes, M5 Max, Mac Studio, DGX Spark | ~200B-class MoE; 284B-class at 3-bit |
| 384-512GB | M5 Ultra Mac Studio, multi-GPU rigs, servers | 400-750B frontier-class MoE |
8-12GB: Entry GPUs and Older Cards
The pick is LFM2.5-8B-A1B from Liquid AI: a sparse MoE with 8.3 billion total parameters and 1.5 billion active, unusually fast on modest hardware, with a 128K context and built-in step reasoning. At Q4 it occupies roughly 5GB, leaving real context headroom even on an 8GB card. One flag before you build on it: the LFM Open License v1.0 is Apache-based but limits free commercial use to companies under $10 million in annual revenue. For personal use and homelabs it changes nothing; above that line, you need a paid license. Google's Gemma 4 E4B is the clean Apache 2.0 alternative at this size, at 4.5GB by Google's own Q4 loading table. The hardware anchor remains the 12GB RTX 3060 class, still the cheapest reliable on-ramp to GPU-accelerated local AI.
Check Price on Amazon: RTX 3060 12GB
16GB: The Laptop-Class Sweet Spot
The pick is Gemma 4 12B Unified: a dense model handling text, images, audio, and video in one encoder-free architecture, with a 256K context window and day-one support in Ollama, LM Studio, and llama.cpp. At Q4 it loads in 6.7GB per Google's published memory table — a figure that already includes about 20 percent loading overhead — so long contexts fit comfortably on a 16GB machine. The license is the other half of the story: Apache 2.0, a meaningful shift from the restrictive terms that governed earlier Gemma generations. For coding-focused work at this size, JetBrains' Mellum2 posts the strongest self-reported LiveCodeBench score in class, with the standing caveat that its custom MoE architecture has run rough in Ollama and vLLM remains its supported path — verify current status before standardizing on it.
Check Price on Amazon: RTX 4060 Ti 16GB
24-32GB: New Pick — Qwen3.8-27B
The pick changed on August 14. Qwen3.8-27B is Alibaba's new dense 27.78B model under Apache 2.0: natively multimodal across text, images, and video, with a 262K context window and vendor-reported jumps over its predecessor — Terminal-Bench 2.1 rising from 63.4 to 73.0 and SWE-bench Pro from 53.5 to 61.7. Those are Alibaba's numbers, not independent reproductions; the early third-party datapoint is BenchLM's aggregate board, which ranks it second among all open-weight models with supported evidence as of August 25. At the house Q4 rule it lands around 16-17GB, fitting a 24GB card with comfortable context. Until independent benchmarks accumulate, Qwen3.6-27B remains the proven fallback with GGUF builds that run everywhere. If your 32GB is system RAM on an iGPU mini PC, Liquid AI's LFM2-24B-A2B was designed for exactly that fit, with the same revenue-clause flag as its smaller sibling. The hardware anchors: the renewed RTX 3090 for GPU builds, still the value king of local AI, and the 32GB mini PC class for an always-on box — our mini PC guide for local AI covers setup and network isolation.
Check Price on Amazon: RTX 3090 24GB (Renewed)
Check Price on Amazon: Beelink SER9 PRO+ (32GB)
48-96GB: Agents With Room to Breathe
The pick is Nex-N2-mini, a post-train of Qwen3.5-35B-A3B under Apache 2.0 from a lab with a published research paper behind its method. It technically squeezes onto a 24GB card at a tight Q4 near 20GB, but it belongs here, where higher-precision quants and long agentic contexts have room. Two standing flags: its benchmark numbers are self-reported on the lab's own suite, and Nex recommends serving through its customized sglang stack — budget an evening, not ten minutes. The smoother-running alternatives are Gemma 4's 26B A4B mixture (14.4GB at Q4 per Google's table) and the 31B dense at 17.5GB, or a 70B-class dense model when knowledge depth matters more than speed. The hardware anchor is the 64GB unified-memory mini PC, quietly the best dollars-per-capability play for local agents — and with DRAM at multi-year highs, machines shipping with memory installed beat build-it-later plans.
Check Price on Amazon: MINISFORUM X1 Pro 370 (64GB)
128-256GB: Two Picks for Two Jobs
Unified memory's tier now has two verified picks, split by job rather than ranked. For text and agent work, DeepSeek V4-Flash 0731: 284B total with 13B active, agent-tuned official weights since July 31, a measured quality ladder from 78 percent at 2-bit to lossless at 8-bit, and the widest runtime support in class — mainline llama.cpp, LM Studio, and Unsloth Studio all load its 103GB 3-bit build on a 128GB machine. Our V4-Flash reality check carries the full tier map. For multimodal work, Step 3.7 Flash from StepFun: a 196B-backbone MoE activating about 11B per token, native image and video understanding, Apache 2.0, and official GGUF builds with measured sizes — 75.8GB at IQ3_XXS, 103GB at Q3_K_L, 111GB at Q4_K_S — with llama.cpp in StepFun's own supported list and 128GB devices named as deployment targets.
The watch entry: GLM-5.3-Flash, released August 26 under MIT with native multimodality and day-one math that lands on these same machines — but fork-only tooling and estimated sizes at launch. Our GLM-5.3-Flash reality check tracks its runtime ledger; it earns a pick slot here the day mainline support lands. The hardware this tier runs on got a fresh entrant on August 25: Apple's M5-generation Mac Studio puts 128GB on the M5 Max from $2,499 and 256GB on the M5 Ultra shipping September 22 — our M5 wait-or-buy breakdown runs that math against the Ryzen AI Max+ 395 class, where the GMKtec EVO-X2's 128GB configuration remains the value entry.
Check Price on Amazon: GMKtec EVO-X2 (128GB)
384-512GB: The Open Flagship Tier
The pick updated: GLM-5.2 replaces GLM-5.1 as the open-weight flagship worth housing — 744B total with 40B active per Z.ai's documentation, a solid 1M-token context, and the strongest results of any plain-MIT model on the aggregate boards, with official weights and a community GGUF ladder. Plan around 410-420GB at Q4, family-scaled. The dated watch line: GLM-5.3, a post-train of this same base, launched August 14 with weights promised roughly two weeks later — as of August 26 its Hugging Face repository exists in gated form but has not opened. This section updates the day it does; the base being identical means the hardware math already holds. We covered the family's trajectory in our GLM-5.1 analysis. The Apache-licensed alternative is Nex-N2-Pro at 397B with 17B active, roughly 220GB at Q4, with the same self-reported-scores flag as its smaller sibling. Above this tier sit the models you can license but not house: Kimi K3 at 2.8 trillion parameters under its revenue-gated custom license — our K3 reality check priced the cluster — and Qwen3.8-Max at 2.4 trillion, open weights the size of a datacenter.
How to Vet a Model Before You Run It
Hugging Face hosts hundreds of new releases and fine-tunes weekly. Model weights are passive data files, so the risk is rarely malware; the risk is trusting capability claims that were never real. Before you download:
- Confirm the base model and its license. A fine-tune inherits its base's terms. If the card does not state the base, stop.
- Check who ran the benchmarks and on how many samples. A score from a 350-question slice is a signal, not a result. Vendor and author numbers are marketing until independently reproduced — the rule this page applies to Qwen3.8-27B above.
- Look for published evaluation logs and training-data disclosure. Authors who show their work earn more trust than a leaderboard screenshot.
- Check tooling support. GGUF builds that load in Ollama or LM Studio mean low friction. A custom serving fork means you are the QA department.
- Weight independent results over self-reported ones. If a model claims frontier wins and no third party confirms after weeks, that silence is data.
- If you cannot verify it exists as described, skip it. The same discipline you would apply to firmware from an unknown source.
Once a model passes, our zero-cost local agent stack guide covers turning it into something useful.
What Still Does Not Run at Home
Honesty requires the other half of the map. DeepSeek V4-Pro at 1.6 trillion parameters, Kimi K3 at 2.8 trillion, Qwen3.8-Max at 2.4 trillion, and the GLM flagship family all publish downloadable weights, and none runs on anything a household owns. Two license flags travel with that list: Kimi K3's bespoke license is revenue-gated, and Llama 4's terms exclude EU-domiciled users and add a separate license above 700 million monthly active users — exactly the conditions the Apache 2.0 and MIT class avoids. Downloading is free; the silicon and the electricity are not.
Where the Closed Frontier Fits
Nothing on this page matches Claude Fable 5 or its closed-frontier peers on the hardest long-horizon work, and pretending otherwise would insult your intelligence. But that capability is rented — our Fable 5 coverage lays out the access and retention terms. Every model on this page is the opposite arrangement: weights on your disk, terms that cannot change underneath you, no usage meter, nothing leaving your network. For a growing share of real work, that trade is no longer a sacrifice. It is just a choice.
All Picks at a Glance
| Model | Params (total / active) | License | Local size | Best tier |
|---|---|---|---|---|
| LFM2.5-8B-A1B | 8.3B / 1.5B | LFM v1.0 (revenue clause) | ~5GB Q4 | 8-12GB |
| Gemma 4 12B Unified | 12B dense | Apache 2.0 | 6.7GB Q4 (Google) | 16GB |
| Qwen3.8-27B | 27.78B dense | Apache 2.0 | ~16-17GB Q4 (est.) | 24-32GB |
| Nex-N2-mini | 35B / 3B | Apache 2.0 | ~20GB Q4 | 48-96GB |
| DeepSeek V4-Flash 0731 | 284B / 13B | MIT | 103GB 3-bit (measured) | 128-256GB |
| Step 3.7 Flash | 196B / ~11B | Apache 2.0 | 103GB Q3_K_L (measured) | 128-256GB |
| GLM-5.2 | 744B / 40B | MIT | ~410-420GB Q4 | 384-512GB |
| Nex-N2-Pro | 397B / 17B | Apache 2.0 | ~220GB Q4 | 384-512GB |
Sizes marked "measured" come from official or Unsloth repositories; others use the 0.6GB-per-billion Q4 estimate. Add several GB for context. Verified August 26, 2026.
Frequently Asked Questions
What is the best local AI model for 16GB of RAM or VRAM?
Gemma 4 12B Unified: dense, multimodal across text, image, audio, and video, 256K context, Apache 2.0, and 6.7GB at Q4 by Google's own loading table, leaving real context room on a 16GB machine. The stronger Qwen3.8-27B needs a 24GB card; on 16GB, Gemma remains the pick.
What is the best model for a 24GB GPU in 2026?
Qwen3.8-27B, released August 14 under Apache 2.0: dense 27.78B, natively multimodal, 262K context, roughly 16-17GB at Q4. Its large benchmark gains are vendor-reported and await independent reproduction, so if you need a proven quantity today, Qwen3.6-27B remains the reliable fallback with GGUF builds in every runtime.
Can I run DeepSeek V4 at home?
V4-Flash, yes — since July 2026 it is the recommended pick for 128GB machines, with the 103GB 3-bit build loading in mainline llama.cpp, LM Studio, and Unsloth Studio, and the agent-tuned 0731 weights downloadable since July 31. V4-Pro at 1.6 trillion parameters remains datacenter hardware by any definition.
What is the difference between total and active parameters?
Total parameters are everything the model stores; active parameters are the subset used per token. Active count determines speed. Total count determines memory, because the entire model must be resident for expert routing to work. When evaluating any MoE model, find the total parameter count first — it is the number your hardware pays for.
Are community fine-tunes from Hugging Face safe to download?
Model weights are passive data files, not executable code, so malware risk from the weights themselves is low. The real risk is unverified capability claims. Check the base model, the license, who ran the benchmarks and on how many samples, and whether independent results exist before relying on one — the checklist above walks each step.
Should I wait for cheaper RAM before trying local AI?
No — start with the memory you have. Models are improving faster than memory is getting cheaper, which means your current machine runs better AI every quarter without a hardware change. Three of this page's six tier picks changed in one summer. Upgrade when a specific model you want demands it, not in anticipation.

