Last updated: October 2026
Key Takeaways
- Yes, at 48GB and up: community 3-bit and 4-bit builds of Aleph Alpha's Kolibri-1 run on 48GB+ Apple Silicon Macs and 64GB+ PCs, at reported speeds from about 12 to 116 tokens per second.
- The catch: as of October 4, 2026, Ollama, LM Studio, mainline llama.cpp, and released mlx-lm cannot load Kolibri-1. Every consumer path today runs on a community patch or a custom architecture file.
- Kolibri-1 has 78B parameters but uses only 3.46B per token, so memory capacity decides whether it runs, not GPU compute. The weights are Apache 2.0.
Yes, you can run Kolibri-1 locally, but today only through community builds. A 3-bit MLX conversion needs about 35GB and runs on a 48GB Mac. A 4-bit GGUF needs 47.5GB and runs on a 64GB PC with no GPU at all. Aleph Alpha's own stated minimum is datacenter hardware: two A100 80GB cards or one H200.
Aleph Alpha released Kolibri-1 on October 3, 2026. It is a 78-billion-parameter German-English mixture-of-experts model with a context window of up to 1,048,576 tokens, published under Apache 2.0. Within about 36 hours, independent developers had uploaded more than 20 community builds to Hugging Face.
What Runs Kolibri-1 Today, by Hardware
Any machine with about 48GB of fast memory can run a community build of Kolibri-1; which build depends on how much memory you have. The table collects every build that published a size and a hardware result, as of October 4, 2026.
| Hardware | Build | Size | Runtime | Reported speed |
|---|---|---|---|---|
| 48GB Apple Silicon Mac | MLX 3-bit | 34.9GB | mlx-lm 0.32+ with a custom architecture file | 52–56 tok/s on M1 Max |
| 64GB Apple Silicon Mac | MLX 4-bit | About 41–44GB | mlx-lm with a custom architecture file | No figure published yet |
| 96GB–128GB Apple Silicon Mac | MLX 6-bit | 63.5GB | mlx-lm fork | 91–116 tok/s on M5 Max 128GB |
| Desktop PC, 64GB+ DDR5, no GPU | GGUF Q4_K_M | 47.5GB | Patched llama.cpp | 11.9–14.9 tok/s, Ryzen 7 7800X3D, CPU only |
| Strix Halo mini PC, 128GB | GGUF Q4_K_M or Q6_K | 47.5GB or 64.2GB | Patched llama.cpp | No Kolibri result yet; Vulkan and ROCm untested |
| NVIDIA DGX Spark, 128GB | Official FP8 or NVFP4 | About 78GB or 42.6GiB | vLLM with Aleph Alpha's plugin | 52.1 tok/s single stream (NVFP4), about 8% faster than the official FP8 weights |
| Two 24GB GPUs | W4A16 GPTQ | 41GiB | vLLM with Aleph Alpha's plugin | Fits on paper; tested only on one A100 80GB |
| Datacenter: 2x A100 80GB, 2x H100 SXM5, or 1x H200, B200, or B300 | Official FP8 | About 78GB | vLLM 0.29 with Aleph Alpha's plugin | No speed published; Aleph Alpha's stated minimum |
Speeds are generation figures reported by each build's own uploader on October 3–4, 2026, mostly from short tests at 4K–8K context. None are independently verified. Sources: the MLX 3-bit, MLX 6-bit, Q4_K_M GGUF, Q6_K GGUF, NVFP4, and GPTQ model cards, plus the official model card.
A second llama.cpp port, the independent kolibri-llama-cpp project, reports 64 tokens per second on a 48GB Mac through Metal, and a KV cache validated to 262,144 tokens.
Why a 78B Model Runs Like a 3.5B One
Kolibri-1 activates only 3.46 billion of its 78 billion parameters for each token, so the work per token is small even though the whole model has to sit in memory.
Each of its 50 layers holds 384 small expert networks. For every token, a router picks 6 of them plus one shared expert that always runs. At 4-bit precision, that means the machine reads roughly 1.7GB of weights per generated token instead of the roughly 39GB a dense 78B model would need. That is why a desktop CPU on ordinary dual-channel DDR5 manages about 14 tokens per second, and why Apple Silicon's high-bandwidth unified memory reaches 50 to 116.
The trade-off is capacity. The router can pick any expert at any time, so all 78 billion parameters must be loaded. A 4-bit build occupies about 44–48GB before you add the operating system and the context cache. This is the reverse of a dense 70B model, where memory bandwidth caps the speed; our explainer on unified memory and running 70B models on a mini PC covers that side.
Long context is cheap on paper
Kolibri-1's long context costs far less memory than its 1M-token headline suggests. Forty of its 50 layers only look back 512 tokens; just 10 keep a full-length cache, each with 4 key-value heads.
Our calculation from the published architecture: the cache grows by about 10KB per token at FP8, or 20KB at 16-bit. A 262,144-token context therefore adds roughly 2.7GB at FP8 (5.4GB at 16-bit), and the full 1M tokens adds about 10.7GB at FP8. Most community builds have only been tested at 4K–8K context, and Aleph Alpha recommends staying at or below 262,144 tokens for serving efficiency.
The Catch: No One-Click Runtime Yet
As of October 4, 2026, none of the popular local AI apps can load Kolibri-1, because its kolibri1 architecture is not in mainline llama.cpp or in a released version of mlx-lm.
Ollama's library returns no Kolibri model, and LM Studio's catalog does not list it. Both inherit model support from llama.cpp and MLX, so every build in the table works around the gap in one of two ways:
- Patched llama.cpp. GGUF builds need llama.cpp compiled at a pinned commit with a community patch; stock binaries refuse them.
-
MLX with a custom architecture file. MLX conversions ship a Python file, usually
kolibri1.py, that mlx-lm loads with the weights, or require a forked mlx-lm.
Read the code before you run it
Running a model locally keeps your prompts off someone else's server. It does not protect you from code you have not read. A custom architecture file is executable Python, and a llama.cpp patch is code you compile and run with your own permissions. Both come from individual uploaders, not Aleph Alpha. Use builds that publish that code openly, skim it first, and move to the mainline runtime once support lands.
The official path
Aleph Alpha supports exactly one runtime: vLLM 0.29 with its aleph-alpha-inference plugin, which is also Apache 2.0. It serves the official FP8 weights, about 78GB, and needs datacenter GPUs or a 128GB-class box such as the DGX Spark. Aleph Alpha has not announced a public hosted API for Kolibri-1, so a third-party "Kolibri API" means sending your prompts to whoever runs that server.
Which Hardware Fits Kolibri-1
For Kolibri-1, buy memory capacity first: 64GB is the practical floor for a 4-bit build, and 128GB leaves room for higher-quality builds and long context.
48GB to 64GB Apple Silicon Mac
A 48GB Mac runs the 3-bit MLX build, but only after raising macOS's GPU memory limit; the build's uploader recommends a 40GB allocation. In that uploader's perplexity measurements, 3-bit lands close to 4-bit, while 2-bit drops off more sharply. A 64GB Mac runs the 4-bit build with room for context. The Mac Studio with M5 Max offers up to 128GB, enough for the 6-bit build that posted the fastest result in the table.
128GB unified memory: Strix Halo and DGX Spark
A 128GB Ryzen AI Max+ 395 (Strix Halo) mini PC holds the Q6_K build at 64.2GB with plenty left for context. No Kolibri-1 result on Strix Halo has been published, and Vulkan and ROCm are untested, so today the patched llama.cpp runs on its CPU cores. The same chassis also ships with 64GB; confirm 128GB in the listing title, as our Ryzen AI Max+ 395 reality check explains.
Check Price on Amazon: GMKtec EVO-X2 (128GB)
The 128GB NVIDIA DGX Spark is the only desk-sized machine with a published result running the official FP8 weights on Aleph Alpha's own vLLM stack. A 64GB DGX Spark configuration also exists through partners, and it cannot hold those weights.
Check Price on Amazon: NVIDIA DGX Spark (128GB)
PCs and mini PCs with 64GB or more
A desktop with 64GB of DDR5 runs the Q4_K_M build on the CPU alone, the cheapest working path to Kolibri-1. 64GB is tight; 96GB to 128GB is comfortable, though filling four DIMM slots often forces lower memory speeds. RAM prices remain elevated, as our 2026 RAM shortage explainer covers.
Check Price on Amazon: DDR5 64GB Desktop Kit (2x32GB)
A 64GB mini PC fits the 33.9GB Q3_K_S build comfortably and the Q4_K_M build only tightly.
Check Price on Amazon: MINISFORUM X1 Pro 370 (64GB)
Single-GPU owners
A single 24GB or 32GB graphics card cannot hold even the 4-bit builds, which start at 41GB. Splitting layers between GPU and system RAM is untested. Do not buy a GPU for this model. If you already own one, our guide to the best local AI models by VRAM and memory lists models that fit your card today.
Is Kolibri-1 Worth Running? The Benchmarks, With a Caveat
On Aleph Alpha's own benchmarks, Kolibri-1 posts the top overall score in its comparison group but trails on tool calling and coding, and no independent score exists yet.
The model card reports 84.3 on GPQA Diamond (Qwen3.5 scores 83.8 in the same table), 96.9 on AIME 2025 (second to Apertus at 97.9), and 96.0 on AIME 2026, plus an English overall score of 75.5 (Nemotron 3 Super: 73.0) and a German overall score of 70.8. The weaker results are in the same table: 61.4 on the BFCL v4 tool-calling benchmark against 70.5 for Qwen3.5, and 85.9 on LiveCodeBench v6 against 93.8 for Qwen3.8.
Every one of those numbers comes from Aleph Alpha, and Artificial Analysis had not scored Kolibri-1 as of October 4. Results this strong from 3.46B active parameters deserve independent confirmation.
Kolibri-1 makes the most sense for German-English work, long documents, and commercial projects that need a clean license. If you depend on Ollama or LM Studio, or lean on tool calling, waiting costs you little. Our GLM-5.3-Flash 128GB reality check covers another model in the same hardware class.
What Apache 2.0 Gives You
Apache 2.0 lets you download, modify, and use Kolibri-1 commercially, with no user-count cap or regional carve-out in the license itself.
The model card adds a responsible-use statement that rules out unlawful uses and uses that violate the EU AI Act. That is a statement of intended use, not a license limit on who may run the model. Kolibri-1 is text-only, covers German and English, and has a stated knowledge cutoff of June 18, 2026.
Ownership is also in motion. On September 16, 2026, Cohere and Aleph Alpha agreed to merge under the Cohere name, pending regulatory approval. That does not change the weights you download today, because Apache 2.0's grant is perpetual and irrevocable. That is the difference between owning a model and renting access to one.
How to Try Kolibri-1 Today
The quickest route is an MLX build on a Mac or a patched llama.cpp build on a PC; each build's README carries the exact commands.
- Match the build to your memory: 3-bit MLX for a 48GB Mac, 4-bit for 64GB, 6-bit for 96GB to 128GB, or Q4_K_M GGUF for a PC with 64GB or more of RAM.
- On a Mac, install the mlx-lm version or fork named on the model card, raise the GPU memory limit as instructed, and start the model with the included launcher.
- On a PC, check out llama.cpp at the pinned commit, apply the patch, build
llama-server, and load the GGUF file. - Use Aleph Alpha's recommended sampling: temperature 1.0, top_p 0.97, top_k 128.
If you would rather wait, watch for mainline llama.cpp support; Ollama and LM Studio follow it. We will update this page when it lands.
Frequently Asked Questions
Can I run Kolibri-1 in Ollama or LM Studio?
Not as of October 4, 2026. Ollama's library has no Kolibri-1 model and LM Studio's catalog does not list it. Both apps depend on llama.cpp or MLX support for the kolibri1 architecture, which has not reached mainline llama.cpp or a released mlx-lm. Today you need a patched llama.cpp build or an MLX conversion that ships its own architecture file.
How much RAM do I need to run Kolibri-1?
About 48GB is the floor. The smallest practical build, 3-bit MLX at 34.9GB, runs on a 48GB Apple Silicon Mac. On a PC, the 47.5GB Q4_K_M GGUF needs at least 64GB of system RAM, and 96GB to 128GB is comfortable. Aleph Alpha's official FP8 weights need about 78GB plus cache, which means datacenter GPUs or a 128GB machine.
How fast is Kolibri-1 on a Mac?
Community builds report 52 to 56 tokens per second for the 3-bit MLX build on an M1 Max, and 91 to 116 tokens per second for the 6-bit MLX build on a 128GB M5 Max. A separate llama.cpp port reports 64 tokens per second on a 48GB Mac using Metal. These are uploader-reported, short-context results, not independent benchmarks.
Is Kolibri-1 free for commercial use?
Yes. Kolibri-1 is released under Apache 2.0, which permits commercial use, modification, and redistribution with no user-count cap. The model card also includes a responsible-use statement excluding unlawful uses and uses that violate the EU AI Act. Apache 2.0's grant is perpetual and irrevocable, so ownership changes at Aleph Alpha do not affect weights you already have.
Does Kolibri-1 really have a 1 million token context?
Yes, with limits. Kolibri-1's native context is 262,144 tokens, and Aleph Alpha says it extends to 1,048,576. Aleph Alpha recommends 262,144 or fewer for serving efficiency. Because only 10 of its 50 layers keep a full-length cache, long context costs little memory on paper, but most community builds have only been tested at 4K to 8K tokens.
Will Kolibri-1 run on a single RTX 4090 or RTX 5090?
Not fully. The smallest 4-bit builds start at 41GB, more than a 24GB or 32GB card holds. Splitting the model between GPU and system RAM is possible in principle with the patched llama.cpp, but no results have been published. A PC with 64GB of RAM runs it on the CPU alone at 12 to 15 tokens per second.
Is Aleph Alpha's Kolibri-1 the same as the Kolibri learning platform?
No. Kolibri-1 is a large language model released by the German AI company Aleph Alpha in October 2026. The Kolibri learning platform is an unrelated open-source offline education app from the nonprofit Learning Equality. They share a name, the German word for hummingbird, and nothing else. Search for "Kolibri-1" or "Aleph Alpha Kolibri" to find the model.

