Can You Run Qwen 4 27B Locally? What the Announcement Changes for Home Hardware

Qwen 4 27B was named on stage September 22, not shipped. The Qwen3.8-27B hardware baseline, what the new architecture could change, and what to wait for.

Updated on
Can You Run Qwen 4 27B Locally? What the Announcement Changes for Home Hardware

Last updated: September 2026

Key Takeaways

  • Qwen 4 27B is a name from a September 22 conference stage, not a download. Alibaba's press release says Qwen 4 is "currently in training"; the four-tier lineup comes from keynote coverage. No Qwen 4 tier has a model card, weights, license, or date.
  • The only home-scale Qwen you can run today is Qwen3.8-27B under Apache 2.0: 16.5 GB at 4-bit on a 24 GB GPU or 24 GB Mac, with a useful floor of 12 GB. Nothing announced this week changes that math.
  • Do not buy a card on a press-conference name. The checklist below names the artifacts that turn the name into a purchase decision, and this page updates at each one.

No. There is nothing to run yet. Qwen 4 27B has no weights, no model card, no license, and no date as of September 22, 2026. What exists is a lineup name spoken at Alibaba's Apsara Conference, an architecture preview you can download today as Qwen3.8-Flash-Next, and a shipping dense 27B, Qwen3.8-27B, that sets the hardware baseline any Qwen 4 27B will be measured against.

Every figure below was checked on September 22, 2026 against the source named beside it, and the page will be updated in place as each Qwen 4 artifact lands.

What Alibaba Announced on September 22, and What It Did Not

Two sources exist, and they say different things. Alibaba's press release of September 22 says the company "revealed that its next-generation model, Qwen 4, is currently in training," and gives a roadmap for Qwen 4.5 and Qwen 5 "projected to scale up to 5 to 10 trillion parameters." It names no tiers, no sizes, and no dates. The four-model lineup (Max, Flash, Plus, and 27B) comes from keynote coverage: attendees posted the slide, and vendor write-ups repeated it within hours. This article treats the name as unofficial until a Qwen model card or blog post uses it. Here is the scorecard.

Claim Source Status
Qwen 4 is in training Alibaba press release Official
Four tiers: Max, Flash, Plus, 27B Keynote coverage Reported; in no official document
5 to 10 trillion parameters Alibaba press release Official, but for Qwen 4.5 and Qwen 5
Qwen 4 architecture Qwen3.8-Flash-Next model card Official preview, downloadable since August 26
Qwen 4 27B size, license, dense or sparse None Unknown

The 5-to-10-trillion figure is a roadmap target for the generations after Qwen 4, not a Qwen 4 specification. No parameter count has been published for any Qwen 4 tier.

Why the 27B Is the Only Tier That Matters at Home

The Max tier will not be a home model regardless of what Qwen 4 changes. The current flagship, Qwen3.8-Max, shipped its open weights on August 12 as Qwen3.8-2.4T-A95B: 2.4 trillion parameters, about 95 billion active per token, a memory floor above 450 GB even at 1-bit, and a custom license. Our Qwen3.8-27B guide put it in the retired-server class alongside the 2.8T Kimi K3 of our K3 reality check. A successor in that class moves the floor, not the category.

The 27B tier is different in three ways. It has shipped dense and under Apache 2.0 for two generations, Qwen3.6-27B in April 2026 and Qwen3.8-27B on August 14, 2026. It is the tier people run: the Qwen3.8-27B repository showed 7.15 million downloads in the past month on September 22 and more than 1,190 quantized derivatives. And the ecosystem forms within days: GGUF, MLX, and NVFP4 builds all appeared in launch week. If a Qwen 4 27B ships under the same terms, it is the release most readers of this page will use.

The Baseline: What Qwen3.8-27B Runs On Today

Any Qwen 4 27B will be judged against these numbers. The table is the map our Qwen3.8-27B guide keeps current; file sizes are Unsloth's Dynamic 3.0 builds as listed on September 22, 2026.

Memory You Have Qwen3.8-27B Build Download Size What You Get
8 GB RAM UD-IQ1_S (1-bit) 6.2 GB Loads and answers short questions; no tool use.
12 GB RAM or 8 GB VRAM UD-Q2_K_XL (2-bit) 9.8 GB Smallest build Unsloth recommends; chat, drafts, and summaries.
16 GB RAM or 12 GB VRAM UD-Q3_K_XL (3-bit) 13.1 GB Daily use for chat, documents, and light coding.
24 GB GPU or 24 GB Mac UD-Q4_K_M (4-bit) 16.5 GB The practical default; full capability for most work.
32 GB unified memory or more UD-Q6_K (6-bit) 22 GB Near-reference quality with room for long sessions.
RTX 50-series (Blackwell) NVFP4 Fits 24 GB VRAM Blackwell-native format; about 1.5x BF16 speed on those cards.

Download size covers weights only. Only 16 of 64 layers use full attention, so context costs about a quarter of a conventional 27B's KV cache, but a 200,000-token session still adds gigabytes.

Speed on the hardware people own

Measured decode figures, sourced and dated: about 49 tokens per second for the Q4_K_M build on an RTX 4090-class card in llama.cpp with no draft model, the baseline on DimInfer's DSpark drafter card (August 2026); and 24 to 56 tokens per second from a plain Ollama MLX setup on an M5 Max, per the tests our Qwen3.8-27B guide tracks. All are short-context figures; speed at 40,000-plus tokens of loaded context runs well under them, and thinking mode defaults to the highest reasoning effort, which multiplies output tokens.

What the independent score says

Artificial Analysis lists Qwen3.8-27B at 52 on its Intelligence Index at the default xhigh reasoning effort, 44 at medium, and 43 at low, as of September 22, 2026. The 52 was measured on hosted full-precision serving; a 4-bit local build sits somewhere below it that nobody has measured. That is the bar a Qwen 4 27B has to clear to be worth a download.

Dense 27B or Flash-Next Design: What the Architecture Does to Memory

The announcement leaves this open, and it decides whether "27B" means what it meant in August. Qwen's Flash-Next release text calls the model "an early preview of the architecture used in Qwen4," released early so the community can examine the changes before the full family is built on them. The model card gives the accounting: 125 billion main parameters in a sparse mixture of experts (512 experts, 10 routed plus 1 shared, about 6 billion active per token), a separate 51-billion-parameter n-gram embedding table indexed by bigrams and trigrams, and a 4-billion-parameter multi-token-prediction head. The package is about 180 billion parameters, roughly 360 GB on disk in BF16, and its configuration declares a model type of qwen4_exp.

The n-gram table is the part that changes hardware budgets. It is designed to sit in host memory and be fetched asynchronously rather than living in VRAM; a SGLang pull request opened September 18 to stage that file-backed table from host memory cites a 47.7 GiB table at full precision. In llama.cpp, initial qwen4exp support merged in late August through pull request 27742 with optimization still pending; the table is kept in CPU memory by override, and the upstream converter drops the MTP head. None of that is finished, and all of it is newer than the dense path Qwen3.8-27B runs on.

Two possibilities follow, and the model card will settle it in one field.

Scenario What "27B" Would Mean Home Hardware Consequence
Dense, on the Qwen3.8-27B lineage (qwen3_5-family model type) 27 billion weights, all active on every token Same budget as today: about 16.5 GB at 4-bit, a 24 GB GPU or Mac, VRAM is the constraint, day-one runtime support likely.
Flash-Next design (qwen4_exp model type) About 27 billion main weights with a subset active per token, plus a separate n-gram table and an MTP head Faster decode; more bytes on disk than the name suggests; system RAM joins VRAM as a budget; runtime support on the newer qwen4exp path.

Until the card exists, plan for the dense case and treat the sparse case as a reason to wait, not a reason to buy RAM.

License and Sovereignty: Apache 2.0, Community License, or Hosted Max

The 3.8 generation shipped three different deals under one name, and the small model got the best one. Qwen3.8-27B is Apache 2.0: use, modify, redistribute, and ship commercially, with vision intact. Qwen3.8-Flash-Next is under the Qwen Community License 1.0, with clauses that apply above a monthly-active-user and revenue threshold. Qwen3.8-2.4T-A95B ships under a custom Qwen3.8-Max license, text-only. The hosted tier on Qwen Cloud keeps the 1-million-token default context and built-in tools; the open 27B gets 262,144 tokens natively and reaches 1 million only through YaRN scaling you configure. Expect the same mix for Qwen 4, and read the LICENSE file per tier rather than the headline.

The reason to care about the small open tier is not benchmark position. Weights you hold, on hardware you own, run under terms nobody can revise after the fact, and a repository you feed to a local model never leaves your machine. A hosted Max endpoint is the opposite deal on every one of those points.

The same rule, one box closer to home

The gateway your ISP rents you is the same arrangement as a hosted model: a monthly fee for something you never own and cannot change.

Most rented gateways run remote-management firmware (TR-069) that gives the provider administrative access to the device at the edge of your network. An owned modem and router ends the fee and the remote access. Our guide to owning your modem covers what changes.

Shop Modems and Routers You Own

What to Wait For Before Buying Another Card

Eight artifacts turn a slide into a hardware decision. This page updates at each one.

  1. A model card under the official Qwen organization on Hugging Face. Name-alike repositories squatted the Qwen3.8-27B name before its weights existed.
  2. The LICENSE file in that repository. Apache 2.0, Community License 1.0, or custom decides what you may do with the weights.
  3. The model type field in config.json. A qwen3_5-family value means dense on the 3.8 lineage; qwen4_exp means the sparse design with the n-gram table.
  4. Official or trusted GGUF and MLX conversions with plausible file sizes. A 27B at 4-bit cannot weigh 400 MB.
  5. An Ollama library tag and an LM Studio catalog entry. Day-one mainline support is likely if the model is dense and not guaranteed if it is not.
  6. Whether the MTP head survives conversion. For Flash-Next the upstream converter dropped it; it is the difference between headline speed and shipped speed.
  7. Whether vision ships in the small checkpoint. Qwen3.8-27B included image and video input; the open 2.4T did not.
  8. An independent score. Artificial Analysis took three days to score Qwen3.8-27B; vendor charts before that point are claims.

The Verdict: Do Not Upgrade on a Press-Conference Name

Three lanes. A 24 GB GPU or a 24 GB-and-up Mac or mini PC is already the machine a dense Qwen 4 27B and today's Qwen3.8-27B target; nothing announced gives you a reason to spend. If you own 12 GB to 16 GB, run the 2-bit or 3-bit Qwen3.8-27B build from the table above now; a new generation does not shrink a 27B's bytes. If you own 8 GB, the 1-bit build is a demo, and the right move is to wait.

If a 24 GB card was already the plan, the 27B class justifies it on what ships today. The renewed RTX 3090 remains the value pick our local AI hardware guide names for this class, and our models-by-VRAM guide has the right-sized pick at every tier. What none of this justifies is buying host RAM against a 51 GB table for a model that may not use one. Apple's M5 Max and M5 Ultra Mac Studio went on sale today, and our Mac Studio guide makes the same argument: a machine bought today is bought for what runs today.

One purchase here does remove a recurring fee. If you rent a router or gateway, replacing it with one you own ends the equipment charge and puts DNS, firewall rules, and firmware under your control, the same reason to run a model locally.

Shop WiFi Routers

Frequently Asked Questions

Is Qwen 4 released?

No. As of September 22, 2026, Alibaba has said only that Qwen 4 is in training, and no Qwen 4 tier has a model card, open weights, a license, an API identifier, a price, or a date. Its architecture is public as Qwen3.8-Flash-Next, an open-weight preview released August 26, 2026, which is a different model under a different license.

What is Qwen 4 27B?

Qwen 4 27B is the reported open-weight small tier of the Qwen 4 family, named in coverage of Alibaba's Apsara Conference keynote on September 22, 2026, alongside Qwen 4 Max, Flash, and Plus. The name appears in attendee posts and vendor write-ups, not in Alibaba's press release, and no specification, license, or date has been published.

When will Qwen 4 27B be released?

No date has been given, and two precedents disagree: the 3.8 generation went from Max launch (August 3) to open weights (August 12 to 14) in under two weeks, while the last architecture preview, Qwen3-Next in September 2025, preceded the Qwen3.5 family by about five months. Treat "soon" as unquantified until a model card exists.

What hardware will Qwen 4 27B need?

Unknown until the model card exists. If it is dense like Qwen3.8-27B, expect about 16.5 GB at 4-bit, a 24 GB GPU or Mac for the full build, and a useful floor around 12 GB with smaller quantizations. If it uses the Flash-Next design, expect fewer active parameters per token, a separate n-gram table in system RAM, and newer, less optimized runtime support.

Will Qwen 4 27B be Apache 2.0?

Not announced. The 27B tier has shipped under Apache 2.0 for two generations, Qwen3.6-27B and Qwen3.8-27B, while Qwen3.8-Flash-Next uses the Qwen Community License 1.0 and the 2.4T model uses a custom license. The LICENSE file is the governing document; read it before anything commercial touches the weights.

Will Qwen 4 27B run in Ollama or LM Studio?

Unknown. A dense 27B on the Qwen3.5 through 3.8 lineage would run in Ollama, LM Studio, and llama.cpp on day one, as Qwen3.8-27B did. A 27B on the Flash-Next design depends on the qwen4exp path, which merged into llama.cpp in late August 2026 with optimization pending.

USA-Based Modem & Router Technical Support Expert

Our entirely USA-based team of technicians each have over a decade of experience in assisting with installing modems and routers. We are so excited that you chose us to help you stop paying equipment rental fees to the mega-corporations that supply us with internet service.

Updated on

Leave a comment

Please note, comments need to be approved before they are published.