I Scored Every "Open Source" AI Model Released in 2026 Against the Actual Definition. None of Them Pass.

Abhishek Dash12 min read
A classroom of sad AI-model chips at exam desks, every answer sheet stamped FAIL, an examiner holding an OSAID clipboard, and a spectrum board from closed to open source AI with a G7 declaration circled as open washing

In short

The Open Source AI Definition 1.0 requires data information sufficient to rebuild a substantially equivalent system, complete training and inference code, and weights under an OSI-approved license. RedMonk found zero of 40 non-closed models qualify. Moonshot requires commercial agreements above $20M in sales with up to 30% revenue sharing, and Alibaba now requires revenue sharing from major Qwen users.

Key takeaways

  • RedMonk's May 2026 survey of 68 models found zero of the 40 non-closed ones qualify as Open Source AI; half fell into Open Weights AI, half into Weights Available AI.
  • The Kimi K3 License requires businesses with over $20M in annual sales to enter a commercial agreement with Moonshot, which may seek up to 30% revenue sharing.
  • Alibaba now requires major commercial users of Qwen3.8-Max, a 2.4-trillion-parameter model, to share a portion of revenue, with the rate still undecided.
  • The G7 Digital and Technology Ministers framework of 2 June 2026 replaced the binary open/closed label with a spectrum and named the practice "open washing."
  • The OSI has no enforcement mechanism; its plan is community objection. Google and Microsoft dropped the term, Meta has not and publicly disputes the definition.

Let me rule out the obvious misreading before it starts: this post is for the term "open source," not against the models. Several of the releases below are permissively licensed and genuinely excellent, and MIT is a real freedom. The target is one word being used for four different legal situations at once.

The site already ran the one-model version of this. In the K2 Horizon post I verified the five datasets on Hugging Face under Apache 2.0 and CC-BY-4.0, with sizes and row counts. I found the training-code repos were still placeholders as of 4 September 2026. I found no working GGUF existed as of 20 September 2026, with the reliable path being transformers 5.17+ in BF16. I flagged the 32B as a stage-1 checkpoint that early benchmarks put behind the 7B. I noted IFM's own reward-hacking audit: TerminalBench at 70.2% dropping to 66.9% after removing flagged trials. That was one model scored on two of three criteria, with dates and citations. This post generalises it, and then goes one level down to the part with almost no coverage anywhere: "open" has stopped meaning open and started meaning revenue-share licensing.

The standard, exactly

The Open Source AI Definition 1.0 was published by the Open Source Initiative on 28 October 2024, after a two-year consultation with machine learning and NLP academics, philosophers, and Creative Commons people. Three requirements:

Data Information: "sufficiently detailed information about the data used to train the system so that a skilled person can build a substantially equivalent system." That includes the complete description of all training data, including, where it cannot be shared, its provenance, scope, acquisition and selection, labeling procedures, and processing and filtering methodologies. Plus a listing of all training data, public or obtainable from third parties, including for fee, and where to get it.

Code: "the complete source code used to train and run the system," the full specification of how data was processed, filtered, and trained. Preprocessing, training arguments and settings, validation, tokenizers, hyperparameter search, inference code, architecture. All of it.

Parameters: weights and configuration, potentially including checkpoints from key intermediate training stages and the final optimizer state.

One nuance in the final version: the standard initially pushed for training data availability but settled for detailed documentation about the data where the data itself cannot be shared. So the bar is lower than "publish the corpus" and higher than "publish nothing." Remember that when someone tells you the definition is impossible to meet, because the definition itself already flinched.

The number that should end the marketing

RedMonk, May 2026, Stephen O'Grady, 68 models surveyed. Of the 40 non-closed models, none qualified as Open Source AI or Open Source AI with Open Data. Half fell into Open Weights AI and half into Weights Available AI.

That is the cleanest single number available on this subject and it is devastating. Zero for forty. And the models that do actually meet the standard, Pythia, OLMo, MIT-licensed Boltz-2 with code, weights, and training data all public, are not the ones topping leaderboards. The best-documented models and the best-performing models are two different lists, and every marketing page in the industry is exploiting the gap between them.

The OSI's own explanation of why open weights is not open source, from August 2025: open weights exclude the training code, the training dataset when legally possible, and comprehensive data transparency. "By withholding these critical elements, developers only provide a glimpse into the final state of the model, making it difficult for others to replicate, audit, or deeply understand the training process."

Three bodies, one funeral for the binary label

The G7 Digital and Technology Ministers published a framework on 2 June 2026 that replaces the binary open/closed distinction with a spectrum and names the practice: "open washing," defined as "publishers claiming the Apache 2.0 label while keeping weights, training data or training code private." It landed in parallel with OpenMDW 1.1, a Linux Foundation license framework released 28 May 2026, and the OSI published its first official draft on Open-Weight AI Model Licensing on 1 September 2026.

Three bodies, three directions, one conclusion: the binary label is being retired, and the G7 is the one that said the quiet part with a name.

Meanwhile the enforcement picture is a joke, and the OSI knows it: it has no enforcement mechanism. Stefano Maffulli's plan is community pressure: "Our hope is that when someone tries to abuse the term, the AI community will say, 'we don't recognize this as open source,' and it gets corrected." Maffulli on the history: "In 26 years of OSI, it has contended with numerous organizations claiming varying degrees of openness as 'open source.' This issue is now mirrored in AI. Open Source is binary: either users have full rights or they don't."

Who folded, who did not: after discussion with the OSI, Google and Microsoft agreed to drop the term for models that are not fully open. Meta has not, and publicly disputes the definition through spokesperson Faith Eischen: "no single open source AI definition, and previous open source definitions do not encompass the complexities of today's rapidly advancing AI models." Note the asymmetry: Meta disputes an existing standard and has offered no replacement. Hugging Face CEO Clement Delangue called the OSAID "a huge help in shaping the conversation around openness in AI."

The category shift: "open" as a commercial tier

The actual headline: Moonshot AI released Kimi K3 as an "open model," 2.8 trillion parameters, performance approaching Claude Fable 5 and GPT-5.6 Sol. Read the Kimi K3 License: businesses with annual sales exceeding $20M must enter a commercial agreement with Moonshot, which may seek up to 30% revenue sharing. Details undisclosed.

Alibaba shipped Qwen3.5 and Qwen3.6 as open models, Qwen3.7 closed, Qwen3.8 expected open again, and now requires major commercial users of Qwen3.8-Max, a 2.4-trillion-parameter model described as "performance second only to Claude Fable 5," to share revenue. Reuters reported it on 7 August 2026. The rate is undecided.

Read those two together. A company with $20M or more in annual revenue does not get unrestricted rights to these weights. It gets a negotiated contract. That is a licensing tier, not an open release, and the G7's spectrum framework is, in effect, the formalisation of exactly that tiering. And note: the labs are not hiding any of this. Moonshot and Alibaba both published the terms. It is normal commercial licensing, and it is normal precisely because these are not open releases. A permissive license has no revenue condition by definition.

Richard Fontana, principal commercial counsel at Red Hat, on what the vocabulary costs: "That combination of lack of familiarity, an emphasis on restriction and unclear legal drafting introduces a degree of legal risk that does not exist when a model is released under a real open source license." His baseline: weights must be released under a genuine open source license, and "there continue to be strong arguments that training data is the AI analogue of source code."

Here is the 2026 licensing table, tiered by what the license actually does:

Tier License Examples 2026 Restrictions
Fully permissive MIT GLM 5.2 (744B, 40B active), DeepSeek R1, Phi-4, Boltz-2 (code + weights + training data) None
Permissive Apache 2.0 Qwen3.5, Mistral Large 3, Gemma 4 (switched from Gemma ToU) None, but watch trademark and patent terms
Commercially gated Kimi K3 License Moonshot Kimi K3 Above $20M annual sales requires commercial agreement; up to 30% revenue share
Revenue-share Qwen3.8 terms Alibaba Mandatory revenue share for major commercial users
Conditionally permissive Llama Community License Llama 4, Llama 3.x 700M MAU threshold; no training competing models; geographic and use-case restrictions; attribution
Remotely revocable Gemma Terms of Use Gemma 3 (predecessor) Google can remotely restrict usage
Closed Proprietary Claude, GPT, Gemini Everything

Fontana's point lands hardest on the Llama and Gemma rows. A company building on Llama 4 faces different legal exposure than one building on Qwen, even though both are universally described as "open source." If licensing risk matters, and for anyone shipping a product it obviously does, MIT and Apache 2.0 are the only two rows that behave the way the word implies. There is also the Linux Foundation's Model Openness Framework, where Jim Zemlin pitches it as "a way to evaluate if a model is open or not open. It allows you to grade models" across three tiers, fully recreatable down to data-descriptions-only.

And the decay precedent: Pythia-7B was once compliant with a stricter draft and stopped being defensible when its foundational dataset, the Pile, became unavailable. Openness degraded through non-action. Under the final OSAID's documentation-based bar that is less fatal, which is arguably the point of the softening.

GLM 5.2, released at 5:21pm on June 13

The geopolitical subplot is also the best available hook. On 13 June 2026, hours after the US government suspended access to Claude Fable 5 and Mythos 5, Zhipu AI, operating as Z.ai, released GLM 5.2: 744B mixture-of-experts, 40B active, 1M context, MIT license, with DeepSeek Sparse Attention integration, claiming state-of-the-art on SWE-Bench Pro and leading open-model scores on NL2Repo and Terminal-Bench 2.0.

Jie Tang, founder of Z.ai and Tsinghua professor, posted at 5:21pm on June 13, the same clock time the Anthropic directive had arrived the day before: "Today, the sudden restriction of certain frontier models is deeply regrettable." And: "Science should be global. The path to AGI must never be enclosed by high walls." The weights went out under MIT the following week.

The honest reading holds both facts at once. A genuinely permissive license released in direct response to an export control, by a national-lab-adjacent lab, is worth taking seriously on its merits. MIT is MIT. The strategic motive is also plainly visible: Zhipu is a Chinese state-linked lab and a counter-move is obviously part of the picture. Both are true and neither cancels the other. And the structural consequence is the important one: if a license can be withdrawn or withheld by jurisdiction, then no open-weights release is a full substitute for domestic capability. That is the ownership-versus-access argument, now visible in a license file.

My scorecard work, with receipts

The scorecard is ten models, three criteria each, citations and dates, and it is cheap because it needs no GPU time. The set: top ten by Hugging Face downloads and trending among models marketed as open in 2026, including GLM 5.2, Qwen3.8, DeepSeek V4, Gemma 4, Kimi K3, Tencent Hy3, and Llama 4 Scout, plus one genuinely compliant model as a control so the table is not just a wall of failures.

The scoring, per model. Parameters: is a permissive OSI-approved license attached, and does it carry an MAU threshold, geographic restriction, use-case restriction, competing-model prohibition, or remote-revocation clause? Quote it verbatim. Code: is training code present or a placeholder, and are hyperparameters, training arguments, tokenizer config, and inference code published? Record the commit hash and the date checked. Data: is there a data card with provenance, scope, acquisition method, selection criteria, labeling procedure, and filtering methodology? Could a skilled person build a substantially equivalent system from the documentation alone?

Then the empirical checks, which make this first-party instead of a literature review. Does the claimed architecture actually run in llama.cpp, vLLM, or transformers, with the stated parameter count and context length matching the config? Is a working quantized build available? The K2 Horizon finding, no working GGUF as of 20 September 2026, is the standing test case: "open weights" you cannot actually run is a narrower claim than it sounds. Does the model card's benchmark set include anything the lab did not run? Can dataset claims be verified independently? I did that for K2 Horizon, five datasets checked. Doing it again, model by model, publishing the rubric, dates, and commit hashes.

The optional extension: attempt a reproduction on the highest-scoring, most permissively licensed model. The gap between "weights are open" and "the system is reproducible" is where criteria one and two get separately falsified, and a reproduction failure is the most concrete possible evidence.

Where this could be wrong

Three real counter-arguments, taken seriously. First: the OSI definition is impossible to meet for frontier models, and that is a fair objection. Requiring full reproducibility of a frontier model is not a realistic bar, and the final standard's softening to documentation was a response to exactly that. Grant it. The point survives: if the bar is aspirational, the honest response is to stop using a word that implies compliance. Second: weights alone are enormously valuable. Control, privacy, local inference, fine-tuning without a vendor relationship, and not having your roadmap depend on a counterparty are all real, and that is what open weights buy. It is not the same as the ability to audit how a model was built or reproduce it. Third: Meta is right that there is no single definition, and the G7 agreed by abandoning the binary. But "no agreed definition" and "the existing definitions are fine" are different positions, and Meta has disputed one without offering another.

One factual caution, stated because my credibility is on it: RedMonk's exact 68/40 split and tier definitions should be read from the primary report, not a summary, and Moonshot's up-to-30% revenue-share figure is reported rather than confirmed in a filing. Also, I am not your lawyer. The licensing-risk view belongs to named counsel, Fontana at Red Hat, and if you are choosing a base model for a product, ask your own.

Three terms that need to stay distinct

Open source AI means three criteria met, permissive terms, reproducible. Open weights means parameters only. Commercially gated means parameters with a revenue condition attached. Almost every argument in this space is two people using the third term to mean the first, and the scorecard exists because nobody selling you a model has any incentive to fix that.

On this page

Sources

  1. Open Source AI Definition 1.0 (OSI, 28 Oct 2024)Open Source Initiative, 2024
  2. Open Weights: not quite what you've been told (OSI, 22 Aug 2025)Open Source Initiative, 2025
  3. A legal minefield: Open source licensing for AI models (TechTarget, 9 Jun 2026)TechTarget, 2026
  4. Alibaba to charge large users of its next open-source model (Reuters, 7 Aug 2026)Reuters, 2026
  5. Open Source AI Definition weekly update (OSI, 10 Jun 2024)Open Source Initiative, 2024

Frequently asked questions

What does the Open Source AI Definition actually require?

Three things. Data Information, meaning sufficient detail about the training data for a skilled person to build a substantially equivalent system, including provenance, scope, acquisition and selection, labeling procedures, and data processing and filtering methodology, plus a listing of all training data and where to obtain it. Code, meaning the complete source used to train and run the system, including preprocessing, filtering, training arguments and settings, validation, tokenizers, hyperparameter search, and inference code. And Parameters, meaning weights and configuration, potentially including intermediate checkpoints. All three, not one.

How many 2026 models actually qualify?

By the evidence available, effectively none of the widely-marketed ones. The RedMonk survey of 68 models published in May 2026, conducted by Stephen O'Grady, found that of 40 non-closed models, none qualified as Open Source AI or Open Source AI with Open Data, with half falling into Open Weights AI and half into Weights Available AI. The models that do meet the standard, such as Pythia, OLMo, and MIT-licensed Boltz-2, are not the ones topping leaderboards. That inverse relationship is the finding.

Is the label just being abused?

No, and this is the more interesting story. A genuine open-weight ecosystem exists and the licenses have improved substantially, with Qwen, GLM, and Gemma 4 all moving to Apache 2.0 or MIT in 2026. What has changed is that the commercial terms attached to open releases have shifted toward revenue sharing, with Moonshot requiring commercial agreements from businesses above $20M in annual sales and Alibaba announcing mandatory revenue sharing for major commercial users of Qwen3.8-Max. "Open" has become a tiered commercial category rather than a legal one.