Can You Run Qwen3.8-27B Locally? Hardware Requirements, Speed, and License

Qwen3.8-27B now loads on 8GB of RAM, but the useful floor is 12GB and a 24GB GPU or Mac runs the full 4-bit build. Re-verified September 1, 2026, with the independent benchmark score, the corrected 1-bit story, and the Apache 2.0 versus custom-license split.

Updated on
Can You Run Qwen3.8-27B Locally? Hardware Requirements, Speed, and License

Last updated: September 2026

Key Takeaways

  • Qwen3.8-27B loads on as little as 8GB of RAM, but the useful floor is 12GB. Unsloth's own follow-up testing shows the 1-bit builds break down on tool use and multi-step work, and Unsloth now points to its 9.8GB 2-bit build as the smallest one worth running. The 4-bit default is 16.5GB on a 24GB GPU or a 24GB Mac.
  • Independent numbers now exist. Artificial Analysis scores the 27B at 52 on its Intelligence Index at maximum reasoning effort, tying a proprietary frontier model, with the caveat that it burns about three times the median token count to get there.
  • Only the 27B is open in the way that matters: Apache 2.0 with vision intact. The 2.4T-A95B ships text-only under a custom license. The community's "uncensored" re-uploads still have no independent benchmark showing added capability, and their own model cards claim parity, not gains. Treat modified third-party weights as untrusted software.

Yes. Qwen3.8-27B runs on anything from an 8GB laptop at 1-bit precision to full-quality 4-bit on a 24GB GPU or 24GB Mac, with official Ollama, LM Studio, and llama.cpp support live. The useful floor is closer to 12GB, and the 2.4-trillion-parameter sibling is still not a home model.

Two things moved since our August 21 pass, and both are corrections rather than additions. Artificial Analysis had published independent scores on August 17 that we missed, and Unsloth followed its Dynamic 3.0 launch with a fuller analysis that narrows what the 1-bit and 2-bit builds are good for. DFlash 2, the speed headline from launch week, is still not in a shipped llama.cpp release. Every figure below was re-verified on September 1, 2026, against the primary source named next to it.

Qwen3.8-27B Hardware Requirements: 8GB Loads It, 12GB Uses It

A 24GB GPU or a 24GB-and-up unified-memory machine runs Qwen3.8-27B at full quality. Below that, the map has two lines instead of one: the 8GB line where the model loads, and the 12GB line where Unsloth's own testing says it starts behaving like the model on the benchmark chart. Unsloth's Dynamic 3.0 release drew the first line on August 19; its follow-up analysis drew the second. Here is the corrected map, sized by the memory you already own.

Memory You Have Build to Run Download Size What You Get
8GB RAM UD-IQ1_S (1-bit) 6.2GB It loads and answers short factual questions. Unsloth's multi-token test drops below 10 percent agreement with full precision at this size; no tool use, no agent work.
12GB RAM or 8GB VRAM UD-Q2_K_XL (2-bit) 9.8GB The smallest build Unsloth itself recommends for everyday use. Chat, drafts, and summaries hold up; expect visible misses on hard reasoning and code.
16GB RAM or 12GB VRAM UD-Q3_K_XL (3-bit) 13.1GB The budget lane. Daily-drivable for chat, documents, and light coding.
24GB GPU or 24GB Mac UD-Q4_K_M (4-bit) 16.5GB The practical default and the build most numbers in this article assume. Full capability for most work. UD-Q4_K_XL (17.6GB) is what Unsloth's own launch command uses if you have the headroom.
32GB unified memory or better UD-Q6_K (6-bit) 22GB Near-reference quality with the least quantization risk, plus room for long sessions.
RTX 50-series (Blackwell) Unsloth NVFP4 Fits 24GB VRAM About 1.5x BF16 speed on Blackwell cards with 92 to 97 percent top-1 retention; vLLM or a from-source SGLang build only.

File sizes are from Unsloth's Hugging Face repository as of September 1, 2026. Builds at 8.4GB and below have the multi-token prediction module stripped out to save about 500MB; a separate 1.4GB MTP file restores it. Download size covers weights only; long contexts add KV-cache memory on top, so a 16.5GB file on a 24GB card does not mean 262K context on a 24GB card.

The 24GB tier is still the sweet spot, and the value play there has not changed all year: a renewed RTX 3090 holds the 16.5GB build with context room to spare and runs it at full GPU speed. It remains the card our local AI hardware guide calls the answer for exactly this model class.

Check Price on Amazon: RTX 3090 24GB (Renewed)

If you want an appliance instead of a build, a 32GB unified-memory mini PC clears the 4-bit tier with the operating system's share included, and the 64GB class buys what this model rewards: long contexts, the 6-bit build, and a second loaded model. Our mini PC guide for local AI walks the tiers; the 64GB pick below is its top-shelf option.

Check Price on Amazon: MINISFORUM X1 Pro (64GB)

And the plain version of the small tiers: an 8GB or 16GB machine now runs Qwen3.8-27B, which is different from running the model the benchmark chart describes. If your memory is fixed and your workload is serious, the right-sized picks in our models-by-VRAM guide may still beat a heavily compressed 27B at the same footprint.

What You Lose at 1-Bit and 2-Bit

You lose the tasks that made the model famous, and Unsloth now says so in its own documentation. The headline number for the 1-bit build was 77 percent top-1 accuracy retention (the documentation page puts it at about 72 percent), and by Unsloth's own account that number flatters the build. Top-1 accuracy checks one predicted token at a time. Unsloth's stricter test, 300 held-out prompts from coding and math benchmarks decoded 32 tokens deep, tells a different story: the 2-bit UD-Q2_K_XL build matches full precision on roughly a quarter of those runs, and every build smaller than it falls below 10 percent.

Unsloth's guidance for anyone running below UD-Q2_K_XL is blunt. Do not use the model for tool calling; it will fail to call tools, call them in loops, or skip them entirely. Keep thinking enabled at least on low, because non-thinking mode can return empty responses. Set presence_penalty to 1.5 or higher to stop repetition loops. Only general knowledge survives the compression.

The practical read: the 2-bit build is a usable assistant for summaries, drafts, translation, and question-answering, the 1-bit build is a curiosity that answers trivia, and neither is the model that scored 61.7 on SWE-bench Pro. Those launch benchmarks were measured at full precision. "Runs on 8GB" and "scores like the model card" are two different sentences, and no quantization makes them the same sentence.

What 262K Context Costs

The 262,144-token native window is a memory bill that competes with the weights for the same RAM or VRAM. The bill is smaller than a conventional 27B would charge: only 16 of the model's 64 layers use full attention, so only those layers keep a KV cache, and with four KV heads at 256 dimensions that works out to about 64KB per token at 16-bit precision, roughly a quarter of the usual rate.

A quarter-million tokens of loaded repository is still gigabytes of cache on top of the file size, which is the argument for the 32GB and 64GB tiers above. And the advertised path to 1M tokens is YaRN scaling, a configuration you apply yourself, not a default you receive.

The Independent Benchmarks Landed

Our August 21 version said no independent evaluation had been published. That was wrong: Artificial Analysis had scored the model on August 17, and the result is the strongest independent number a local model has posted. At its default xhigh reasoning effort, Qwen3.8-27B scores 52 on the Intelligence Index, a composite of nine evaluations covering coding, science, reasoning, and professional tasks. That ties GPT-5.6 Luna at maximum effort and sits one point behind GLM-5.2 and DeepSeek V4 Pro, both mixture-of-experts models many times its size. On the separate Agentic Index it scores 51.

The score has a price, and Artificial Analysis itemizes it. The model generated about 160 million output tokens across the evaluation against a median near 48 million, which is why the same site rates it as notably slow and very verbose. Turn the reasoning dial down and the score follows: 44 at medium effort, 35 with thinking off. Both of those still lead the 27B class; they just are not the headline.

Two things this does not settle. The Artificial Analysis run used hosted full-precision serving, so a 4-bit local build sits somewhere below 52 that nobody has measured. And Qwen's own agentic-coding claims, including the SWE-bench Pro figure, remain vendor-run; the independent composite points the same direction without confirming any single number.

Speed: DFlash 2 Is Still Waiting, MTP Works Today

The launch-week speed story has cooled, and the shipping path is the boring one. DFlash 2, released August 18 by Inco AI from work seeded at Z Lab, pairs the 27B with a small drafter that proposes whole blocks of tokens for the big model to verify in one pass. The output is provably unchanged, and the launch demo reported 70 tokens per second on an M5 Max MacBook Pro. As of September 1 the llama.cpp integration is still an open pull request that you have to build from source, and the vLLM and SGLang paths are likewise development branches.

The independent results are mixed enough to wait on. One M5 Max test of the pull-request build measured 11 to 35 tokens per second across reasoning levels, well under the advertised figure, while a plain Ollama MLX setup on the same machine reached 24 to 56 without any drafter at all. On data-center hardware the technique holds up better: an SGLang demo on a single A100 went from about 29 to 59 tokens per second. Single-GPU tinkerers are posting around 76 tokens per second on an RTX 4090 with a custom 2-bit drafter, which is an enthusiast result on an unmerged branch, not a default.

What works today is the model's own multi-token prediction head, which mainline llama.cpp already supports. Serving the model with llama-server's draft-mtp speculative mode, one DGX Spark comparison came in about 72 percent faster than LM Studio's default build. That needs the MTP module present, which is why Unsloth ships it separately for the sub-9GB builds.

If your install feels slow, check two dials before blaming the hardware. Thinking defaults to xhigh reasoning effort, which spends heavily on simple prompts; medium keeps most of the quality at a fraction of the tokens. Qwen's recommended sampling is temperature 1.0 with top_p 0.95 and top_k 20 in thinking mode, or temperature 0.7 with top_p 0.80 and presence_penalty 1.5 with thinking off. And if generation crawls at single digits, the build is spilling out of VRAM into system memory; pick the next file size down.

Ollama, LM Studio, and llama.cpp Status (Verified September 1)

All three consumer lanes are live and settled. The official Ollama library serves qwen3.8:27b as an 18GB package with the vision projector bundled in the manifest and has passed 1.3 million pulls, plus a qwen3.8:27b-mlx variant that runs MLX-native on Apple Silicon. LM Studio lists the model in its catalog with GGUF and MLX builds; its 17GB Q4_K_M is the one most published hands-on tests used. Mainline llama.cpp support has been live since day one, including MTP speculative decoding.

One warning from the Kimi K3 era survives, upgraded: name-alike repositories squatted the Qwen3.8-27B name on Hugging Face before the weights existed, and the model tree now lists more than 900 quantized derivatives. Download from the official Qwen organization or a quantizer you already trust, confirm the model card exists, and sanity-check file sizes. A 27B at 4-bit cannot weigh 400MB.

One Family, Two Licenses: The 27B Is the Open One

The same release used "open weights" for two very different deals, and the small model got the better one. The 27B carries Apache 2.0, the industry's cleanest standard grant: use it, modify it, ship it commercially, no revocation lever. The 2.4T-A95B repository ships a custom Qwen3.8-Max License, arrives text-only with thinking permanently on, and needs a 450GB-plus memory floor that puts it in the retired-server class our Kimi K3 reality check mapped for its 2.8T cousin.

Spec Qwen3.8-27B Qwen3.8-2.4T-A95B
License Apache 2.0 Custom Qwen3.8-Max License
Vision Yes, images and video No, text only
Context 262K native, YaRN to 1M Below the hosted tier's default
Weights on disk 6.2GB (1-bit) to 54.7GB (BF16) 4.9TB full; 397GB at 1-bit
Independent score 52 on Artificial Analysis Intelligence Index (xhigh) 58 on the same index
Realistic home hardware Yes, 12GB RAM and up for daily use No, 450GB-plus memory floor

This is the pattern we flag every time it appears, whoever ships it: MiniMax M2.7's modified MIT restricts commercial use, Llama 4 carries an EU restriction and a monthly-active-user clause, and Kimi K3 arrived under a bespoke license a widely shared guide misidentified as Apache 2.0. The LICENSE file in the repository is the governing document. Read it before anything commercial touches the weights.

Are the Uncensored Qwen3.8-27B Builds More Capable?

No independent benchmark shows that, and the modified builds' own model cards now agree. The refusal-stripped re-uploads that appeared within days of launch report capability within about a point of the base model on standard benchmarks, with one claiming a single-point gain on a saturated knowledge test. Those are the publishers grading their own work, and even taken at face value they describe parity, not improvement. The capability everyone is reacting to lives in the base weights Alibaba shipped and in the independent 52 that Artificial Analysis measured on those weights.

What the modified builds do prove is structural, and it cuts both ways. Apache 2.0 weights mean the model's behavior, refusals included, ships inside a file you own, and within days the community published versions with that behavior removed. The property that makes the model yours to keep is the same property that makes it anyone's to edit. That is what "whoever controls the infrastructure controls the experience" looks like when the infrastructure is a 17GB file, and it is why open weights carry policy debates alongside ownership.

The consumer-safety read is simpler, and this month supplied a clean example. One modification team disclosed that its merge process had silently dropped the model's multi-token prediction head, fifteen tensors, and that it had to graft the original weights back in by hand. That was caught by a careful publisher; a careless one would not have noticed. A third-party re-upload with modified behavior is untrusted software that will hold your documents and, in agent setups, your network access. Verify the publisher, prefer the official weights, and if you experiment beyond them, do it on an isolated box using the practices in our local AI security guide.

One more monthly fee to kill

Renting your modem? That's $120+ a year.

Own it instead. It pays for itself in months. 90-day warranty, free returns.

Pick your provider:

Who Should Download It, and Who Should Skip It

Three lanes, plainly drawn. If you own a 24GB GPU or a 24GB-and-up unified-memory machine, download it today: this is the most capable Apache 2.0 model that fits your hardware, it now has an independent score to back the claim, and the license means it stays yours. If you own 12GB to 16GB, take the 2-bit or 3-bit Dynamic 3.0 build from the table above and treat it as a strong chat assistant rather than a coding agent. If you own 8GB, the 1-bit build is a demo of what is coming, not a tool you will keep open. And if you need 1M context by default or zero setup, Alibaba's hosted tier is the answer; that ceiling is precisely what the company kept for itself.

Every consumer build is a data-cap rounding error, 6.2GB to 22GB and minutes on a gigabit line, while the 4.9TB A95B would burn four months of a 1.2TB cap on one file; our download-time breakdown by internet plan has the tier-by-tier math. However you run it, the reason to run it locally has not changed: weights you hold, on hardware you own, under terms nobody can revise on their schedule.

Frequently Asked Questions

What hardware does Qwen3.8-27B need?

About 16.5GB of RAM or VRAM for the practical 4-bit build, which makes a 24GB GPU such as a used RTX 3090 the sweet spot and a 24GB-and-up Mac or mini PC the appliance path. Smaller Dynamic 3.0 builds run on 16GB and 12GB machines with quality tradeoffs, and a 1-bit build loads on 8GB for basic question-answering only.

Can Qwen3.8-27B run on 8GB of RAM?

It loads, but Unsloth itself does not recommend it for anything beyond short factual questions. The 6.2GB 1-bit Dynamic 3.0 build fits in 8GB of RAM and retains roughly 72 to 77 percent top-1 accuracy, but Unsloth's deeper 32-token test shows it agreeing with full precision under 10 percent of the time, and tool calling breaks. The 9.8GB 2-bit build on a 12GB machine is the smallest configuration worth keeping.

Is there an independent benchmark for Qwen3.8-27B?

Yes. Artificial Analysis scored it 52 on its Intelligence Index at xhigh reasoning effort on August 17, 2026, tying GPT-5.6 Luna at maximum effort and trailing GLM-5.2 and DeepSeek V4 Pro by one point. It scores 44 at medium effort and 35 with thinking off, and it generated about three times the median token count to reach the top score. Qwen's own agentic-coding figures remain vendor-run.

Does Qwen3.8-27B work in Ollama and LM Studio?

Yes, both, since release week. The official Ollama library serves qwen3.8:27b as an 18GB download with vision bundled, plus a qwen3.8:27b-mlx variant for Apple Silicon. LM Studio lists the model in its catalog with GGUF and MLX builds, and mainline llama.cpp support has been live since day one.

Is Qwen3.8-27B free for commercial use? (Qwen3.8 license explained)

Yes. The 27B ships under Apache 2.0: commercial use, modification, and redistribution are permitted with no usage agreement. Open-weight is the precise term; the training data and pipeline are not published. The larger Qwen3.8-2.4T-A95B is different: it carries a custom license, so read the LICENSE file in whichever repository you build on.

How big is the Qwen3.8-27B download?

Between 6.2GB and 22GB depending on the build: 6.2GB at 1-bit, 9.8GB at 2-bit, 13.1GB at 3-bit, 16.5GB at the 4-bit default, and 22GB at 6-bit, per Unsloth's Dynamic 3.0 file sizes as of September 1, 2026. Ollama's default package is 18GB. Any of them downloads in minutes on a gigabit line.

Is the uncensored version of Qwen3.8-27B more capable?

No independent benchmark shows that as of September 1, 2026, and the modified builds' own model cards report capability within about a point of the base model, which is parity, not a gain. Capability lives in the base weights, which are the ones Artificial Analysis scored. Treat third-party modified weights as untrusted software first, and as a research topic second.

Does the local Qwen3.8-27B have vision?

Yes, natively. Qwen3.8-27B accepts images and video in the open weights, the Ollama manifest bundles the vision projector, and LM Studio's builds carry it too. Note the family split: the open 2.4T-A95B is text-only, with vision reserved for Alibaba's hosted Max tier. Confirm your runtime's vision path before relying on it.

USA-Based Modem & Router Technical Support Expert

Our entirely USA-based team of technicians each have over a decade of experience in assisting with installing modems and routers. We are so excited that you chose us to help you stop paying equipment rental fees to the mega-corporations that supply us with internet service.

Updated on

Leave a comment

Please note, comments need to be approved before they are published.