Last updated: July 2026
Key Takeaways
- On capability, open-weight models genuinely sit inside the frontier pack in July 2026, but the parity arrived at 744-billion-to-2.8-trillion-parameter sizes that run in datacenters, not homes.
- The viral six-model scoreboards circulating this week mix two shipped, verifiable models with four promises, at least one of which is inflated and one of which has no primary source at all.
- The place open models have most clearly caught up is price: top open-weight API rates now run roughly 20x to 57x cheaper than the closed flagships, and that gap is spendable today.
A scoreboard has been making the rounds on X this week: six open-source AI models, each tagged "almost Opus-level," "almost Fable-level," or "expected to beat Opus," ending with a single word: "Holy." Versions of it have pulled hundreds of thousands of views, and the underlying feeling is real. Three Chinese labs now hold the top of the open-weight leaderboard, and the week of July 16 alone saw a 2.8-trillion-parameter open-weight launch and a 2.4-trillion-parameter open-weight tease land three days apart.
So has open-source AI actually caught up to closed? The honest answer splits into three questions, and they get three different answers. Measured on capability, the gap has narrowed to something like a photo finish, with asterisks. Measured on hardware, the gap did not close; it moved somewhere you cannot follow. Measured on price, open models did not just catch up. They lapped the field, and that is the part you can actually spend. But before any of that, the scoreboard itself deserves an audit, because half of it has not shipped.
The Viral Scoreboard vs. the Verified Ledger
Here is each entry from the six-model lists circulating this week, checked against what actually exists as of July 19, 2026. The pattern to notice is the sourcing column.
| Model | The Viral Claim | Status Today | What the Record Shows | Source Quality |
|---|---|---|---|---|
| GLM-5.2 (Z.ai) | "Almost Opus-level" | Shipped, weights downloadable | 744B parameters, MIT license, leads several coding benchmarks on Z.ai's own table | Official release; partial third-party data |
| Kimi K3 (Moonshot) | "Almost Fable-level" | Shipped hosted; weights committed by Jul 27 | 2.8T parameters, #1 on a major frontend coding arena, fourth on an independent intelligence index | Official release; partial third-party data |
| Qwen3.8 (Alibaba) | "2.4T, expected to beat Opus" | Hosted preview only | No license, date, model card, or benchmarks published; open weights promised "soon" | Official teaser, no artifacts |
| DeepSeek V4 | "GA soon, $0.0028/M, expected to beat Opus" | Already shipped, April 2026 | Not "soon" — V4-Pro is out under MIT; $0.0028/M is its live cache-hit input rate, not a future price | Official release; claim mislabels a shipped product |
| MiniMax M3-Pro | "3T, expected to be Fable-level" | Reported plan, not announced | The Information reports 2.7T (not 3T) targeting Q3, citing two people familiar; nothing official from MiniMax | Single-source press report |
| GLM-5.5 (Z.ai) | "Expected to beat Opus" | Nothing to verify | No announcement, report, or primary source we could locate as of July 19, 2026 | None located |
Status as of July 19, 2026. "Expected to beat Opus" is not a measurement for any unreleased model; it is a mood.
Two shipped models, one official teaser, one already-released model mislabeled as upcoming, one single-sourced report with an inflated number, and one entry with no source at all. That is not a scoreboard; it is a ledger with the debts and assets mixed together. The two assets, though, are real, and they are enough to answer the question seriously.
Caught Up on Capability? Closer Than It Has Ever Been, With Asterisks
Take only what shipped. Kimi K3 debuted at #1 on a major blind-preference frontend coding arena, ahead of Claude Fable 5 and GPT-5.6 Sol, and it ranks fourth of 189 models on Artificial Analysis's independent intelligence index, on par with Claude Opus 4.8 and behind only Fable 5 and GPT-5.6 Sol. That is an open-weight model inside the frontier's top five on third-party measurement, something that had never happened before this month. GLM-5.2, at a comparatively svelte 744 billion parameters under a clean MIT license, posts leading scores on long-horizon software engineering benchmarks, though those particular leads come from Z.ai's own published table rather than independent testing.
The asterisks matter. Preference arenas measure taste, not verified correctness. Most benchmark breadth beyond the headline results is still vendor-reported, K3's technical report is not out, and the closed flagships still lead on the hardest reasoning suites. We went through K3's numbers line by line in our K3 versus Fable 5 versus GPT-5.6 Sol breakdown. The fair summary: on capability, "almost caught up" is now defensible with third-party evidence, and "caught up, full stop" is not yet.
Caught Up on Hardware? No — and This Is the Fine Print Every List Skips
Here is the part the scoreboards never mention. The models that closed the gap are 744 billion to 2.8 trillion parameters. GLM-5.2, the smallest of the three shipped leaders, needs roughly a terabyte of VRAM at full precision, on the order of eight datacenter-class GPUs. DeepSeek V4-Pro at 1.6 trillion needs more. Kimi K3 is the heaviest: Moonshot recommends serving it on 64 or more accelerators, and even at 4-bit quantization its weights alone run to roughly 1.4 terabytes — we did the full hardware math on K3 the week it launched. No consumer machine, including a maxed-out Mac Studio, comes within a factor of two of holding any of them.
So the honest phrasing is this: open weights reached the frontier, but the frontier moved to a size class where "you can download it" and "you can run it" stopped meaning the same thing. Parity at 2.8 trillion parameters is parity you rent. That does not make the weights worthless — it changes who competes to serve them, which is exactly where the next answer comes from.
Caught Up on Price? Yes. This Is the Actual Win, and It Is Spendable Today
Compare list rates as of July 2026. Claude Fable 5 costs $10 per million input tokens and $50 per million output. GPT-5.6 Sol runs $5 and $30. Claude Opus 4.8, the yardstick every viral list measures against, is $5 and $25. Now the open side: DeepSeek V4-Pro's first-party rate is $0.435 input and $0.87 output per million — a 75 percent launch discount that DeepSeek made permanent on May 31, per its pricing page — with cached input at $0.0028 per million, which is the number the viral lists quote without explaining. Kimi K3 launched at $3 and $15. GLM-5.2's median across hosts is about $1.40 and $4.40.
Run the division: on output tokens, V4-Pro is roughly 57 times cheaper than Fable 5 and about 29 times cheaper than Opus 4.8, while sitting a rung or two below them on capability rather than a different league. And because the weights are public, no single vendor sets the price. GLM-5.2 is served by roughly two dozen hosts on OpenRouter at input rates spanning $0.95 to $3.00 — that spread is what open weights structurally buy you: providers competing to serve the same model, margins compressing, and no deprecation risk, because a model whose weights are public cannot be taken away from the people using it. This is the mechanism we flagged when Moonshot priced K2.7-Code at a tenth of closed rates, now running at full scale. One tradeoff stays constant: every hosted API, open or closed, routes your prompts through someone else's servers under someone else's terms, a tradeoff we documented clause by clause in our read of Kimi's own policies.
What "Caught Up" Means on Hardware You Actually Own
None of the frontier trio runs in your house, so what does this race deliver to a machine you own? Two things, both compounding. First, the runnable tier keeps absorbing frontier techniques: the strongest 8-to-70-billion-parameter open models available today exist substantially because trillion-scale open models serve as teachers for distillation, and each frontier open release refreshes that pipeline. Our guide to the best local AI models by VRAM tier maps what runs at every memory level from an 8GB laptop up, and the quality floor at each tier has risen every quarter. Second, the falling API prices and the improving local models bracket the closed vendors from both sides — which is the quiet consumer story of 2026: whichever way you go, your leverage improved.
If you want the owned-hardware route, a dedicated box does not require datacenter money. Our mini PC guide for local AI covers capable setups from Raspberry Pi to 64GB-class machines, including the network isolation an always-on AI box should have, and our full local AI hardware guide covers the GPU route. A model running on hardware you control has no pricing page, no usage credits, and no preview window in which it can "continuously evolve" out from under you.
The Ledger to Watch
The promises column of the table resolves on checkable dates and artifacts, so here is what turns the maybes into facts. July 27 is the committed date for Kimi K3's public weight release; if it lands, the largest open-weight model in history becomes downloadable, and every self-reported K3 number becomes independently testable. Qwen3.8 owes four artifacts before its row changes: a license text, an active-parameter disclosure, a Hugging Face model card, and third-party benchmarks. MiniMax M3-Pro needs anything official at all — a 2.7-trillion-parameter Q3 open release would be enormous, and it currently rests on one outlet's reporting. GLM-5.5 needs to exist in public before it belongs on any list. We will keep this page updated as each of those lands or slips.
Frequently Asked Questions
Has open-source AI caught up to ChatGPT and Claude?
On measured capability, nearly: Kimi K3 ranks fourth on an independent intelligence index, on par with Claude Opus 4.8 and behind only Claude Fable 5 and GPT-5.6 Sol. On price, open models are far ahead, at rates 20x to 57x cheaper. On consumer hardware, no: the open models that reached the frontier are far too large to run at home.
What is the best open-source AI model right now?
Among shipped models as of July 2026: Kimi K3 has the strongest overall third-party standing, GLM-5.2 leads on vendor-reported software-engineering benchmarks with a clean MIT license, and DeepSeek V4-Pro offers the best price-to-capability ratio with downloadable weights. Which is "best" depends on whether you weight capability, license, or cost.
Can I run these frontier open models at home?
No. The smallest of the three, GLM-5.2 at 744 billion parameters, needs roughly a terabyte of memory at full precision, and Kimi K3 needs a 64-plus accelerator cluster. What you can run at home is the 8-to-70-billion-parameter tier, which improves partly by distilling from these larger models.
Is open-source AI free to use?
The weights are free to download where released, but serving a trillion-parameter model requires datacenter hardware, so most people pay for hosted API access. Those rates are the real bargain: DeepSeek V4-Pro at $0.435 input and $0.87 output per million tokens versus $10 and $50 for Claude Fable 5, as of July 2026.
What is the difference between open weight and open source?
Open weight means the trained model parameters are published for download and self-hosting. Open source, strictly, also releases training code and data under a permissive license. Most of the models discussed here are open weight; the MIT-licensed ones like DeepSeek V4 and GLM-5.2 come closest to full open source. License terms still vary, so read them before building commercially.
Which open-source models are releasing next?
Committed: Kimi K3's full weights by July 27, 2026. Promised without a date: Alibaba's Qwen3.8 at a claimed 2.4 trillion parameters. Reported but unconfirmed: MiniMax's M3-Pro at 2.7 trillion parameters targeting Q3, per The Information. Rumored with no primary source: GLM-5.5. Treat each according to its evidence.

