Meta's Llama 4 blog on 5 April 2025 called Scout and Maverick open-weight in the first line, then talked about a "commitment to open source." The file that actually ships is the Llama 4 Community License. Apache and MIT are other files, not this one. If your product, or your affiliates', crossed 700 million monthly active users last calendar month, you ask Meta for a second licence. Meta can say no.
The blog and the licence file do not match. On 11 August 2026 I sat down with the licence cards for Llama 4, Qwen 3.x, Gemma 4, DeepSeek from R1 through V4, Mistral's open line, OpenAI's gpt-oss, and Microsoft Phi-4. A lab can swap a card later. These rows are from that day.
I scored four cells. Weights: can you download the parameters and run them with no API key. Data: is the training corpus public so someone else can rebuild. Code: is the training or inference stack under an OSI licence, or just a README screenshot. Use: can you sell a product on top of it, and does that right die at a user count, a country, or a banned-use list.
All 7 families let you download the weights. None publish the training corpus. The code cell is mixed. Use rights usually hold until a cap, a country bar, or a prohibited-use list. If "open source" still means source, the source is missing. What shipped is a checkpoint you can host and a training set you cannot inspect.
The Open Source Initiative published OSAID 1.0 on 28 October 2024. It asks for the 4 software freedoms on the system: use, study, modify, share. It wants enough "data information" to rebuild, not necessarily every raw token. Meta objected that week. The Software Freedom Conservancy said OSAID was already too soft, because it does not demand the actual dump. Meta wants the phrase without the data, so Llama stays the default download. OSI wants a definition labs will accept. People who wrote the original open-source definition want the books, which would empty most of the hub's "open" filter.
Almost nobody runs the model that publishes the data. Allen Institute's OLMo line puts weights, training code, and a documented corpus in public. The models people host skip the data cell. A 70B checkpoint you can put behind an API sells. Teams took that and kept calling it open.
Llama. The 700 million clause
The 700 million clause sat in Llama 2 and 3. Counsel notes have walked through the text: if products made available by you or your affiliates were used by more than 700 million monthly active users in the preceding calendar month, you request a licence, which Meta may grant in its sole discretion, before any such use. Google, Microsoft, Amazon, and a handful of consumer apps sit above that line. A startup does not.
The clause is a legal hook on the 3 companies that could make Llama the default inside a competing cloud. Meta does not need to charge the startup. A veto you never use still shapes who builds their own model and who comes to Menlo Park. The startup spreads the file for free.
The same file bans using Llama outputs to train a competing model and attaches an acceptable-use policy. Llama 4 added a limit the Llama 3 text did not have. The acceptable use policy says the rights in the community licence are not granted, for the multimodal models, to a person domiciled in the EU or a company with its principal place of business there. Text-only cards are a different sentence. EU end users can still hit a service that wraps Llama 4 if the service sits outside the Union. The person who wants to fine-tune Scout in Berlin cannot. Meta's wording moved from "open source" on Llama 2 and 3 to "open-weight" on Llama 4. The 5 April blog used both. The licence file used neither.
The EU multimodal bar is the AI Act written into a licence. Meta did not wait for an Office inspection. It stopped granting those rights. That costs less than documentation, copyright summaries, and a systemic-risk file. It also shows who Meta thinks the Act will touch: people who host the multimodal weights, not people who chat with them.
For most builders the file is free. For Meta's peers it is a commercial negotiation. That is what 700 million is for.
Qwen and Gemma switched to Apache
Alibaba's early Qwen models used a custom Tongyi Qianwen licence. From Qwen3 the team put Apache 2.0 on the open-weight cards. Qwen-Max stays behind an API. The 2026 cards for Qwen3.5 and Qwen3.6, including the 27B dense and 35B-A3B sparse builds, stay on Apache. The Qwen3.6 repo says so in one line. Qwen3.5-Plus, Qwen3.6-Plus, and the 3.7 Max and Plus variants sit behind Alibaba Cloud. The same company now sells a downloadable line and a closed line under one logo.
Google did the same walk later. Gemma 1 and 2 sat under Google's own terms. Gemma 4 cards list Apache 2.0. Hugging Face's 2 April 2026 write-up called that out: "truly open with Apache 2 licenses." Google's own Gemma 4 post put a number on the funnel: more than 400 million downloads of the family, more than 100,000 variants. Apache on Gemma 4 keeps that funnel from moving to Qwen.
Neither family publishes the corpus. Apache on the weights is not Apache on the books, the crawl, or the synthetic traces. You can fork the file. You cannot check what it was trained on. Labs found they could win developers with a licence change and keep the training set. A competitor cannot copy that set by downloading the card.
Alibaba open-weights the 27B and closes the Plus. Google puts Apache on Gemma and keeps Gemini behind the API. Mistral prints Open and Premier on the same homepage. The open card brings users. The closed card is the paid product. Calling both "the Qwen family" or "Google's open models" mixes the download with the company.
DeepSeek switched the file
DeepSeek V3 sat under the DeepSeek Model License: commercial use allowed, not MIT. From DeepSeek-R1 the weights themselves are MIT. V4-Pro and the July 2026 V4-Flash cards stay on MIT. The code repos were already MIT. For a Western company that wants to host the checkpoint, that is a shorter memo than Llama.
MIT still hides the training data. It will not stop a government from blocking the download or make a trillion-parameter mixture cheap to run. It does make the licence memo short enough to vendor. The file got cleaner while the model got close enough to the frontier that CAISI spent April 2026 measuring it.
CAISI put V4 Pro about 8 months behind the US frontier on an aggregate, and much further on cyber: 32% against GPT-5.5's 71% on their unpublished suite. The UK AI Security Institute put GLM-5.2 and DeepSeek V4-Pro a few months behind closed frontier models on cyber, not a year. A MIT file on a model 8 months off the frontier is something a cloud might host. A 7B MIT file is not. Western labs that spent 2024 calling open weights a safety problem now face a MIT file 8 months back and a Llama file with a 700 million veto.
Mistral, gpt-oss, Phi
Mistral's docs split the catalogue into Open and Premier. Open, on the current cards, is Apache 2.0 for 7B, Mixtral, NeMo, Pixtral 12B, Small, Magistral Small, and the Mistral 3 family, including Large 3. Premier stays behind the API. Some older or mid-tier cards still show a Research License or a modified MIT. Read the card, not the homepage. The logo does not change when the file does. Mistral cannot match Meta's 700 million clause, so the open shelf is actually Apache.
OpenAI's gpt-oss 120B and 20B weights are Apache 2.0 plus a usage policy. The policy is a prohibited-use list, not a 700 million gate. Lawyers treat that as lighter than Llama and heavier than raw Apache. OpenAI spent years saying weights would not be released. gpt-oss puts a card into the same download funnel Gemma and Qwen already own. The usage policy lets OpenAI cut off marketplace use without a user-count veto.
Microsoft Phi-4 variants are MIT. Small enough to run locally. Clean enough to put in a product without a Meta letter. The limit is quality. Nobody writes a 3,000-word memo about Phi's terms. They write it about Llama's, because Llama is the one a cloud would ship.
What the hub actually sorts
Hugging Face lets you filter by licence tag. The tag is whatever the uploader typed. A Llama derivative often inherits "llama3" or "other." A Qwen fine-tune sometimes keeps Apache and sometimes does not. Fine-tunes inherit the base licence unless the tuner has the right to relicense, which they usually do not. A "Qwen-Llama merge" with an Apache tag is a filing error.
Builders search Apache, pull a merge, and ship. The tag can be wrong. The base file is the licence. Counsel who treat the hub facet as a legal opinion will find out at the first diligence. The facet shows what people want to believe. It does not show what they may do.
Cloud marketplaces add a second contract. AWS, Azure, and Together can cut you off for acceptable-use breaches even when the weight file is MIT. A MIT card plus a marketplace AUP can block you even if the file does not. You can leave Together and go to a bare-metal box. Most teams will not. For anyone too small to hit 700 million users and too lazy to own the GPU, that second contract is the restriction that actually binds.
The Act already ignored the blog
The EU AI Act does not treat "open" as a free pass. Article 2(12) exempts some free and open-source systems unless they are high-risk, prohibited, or under Article 50 transparency rules. GPAI models with systemic risk stay in scope even if the weights are downloadable. The Commission's GPAI guidelines say an Apache card does not cancel documentation or copyright duties. Provider obligations applied from 2 August 2025. Enforcement powers, including fines of up to the higher of 3% of worldwide turnover or €15 million, switched on 2 August 2026.
A Llama Community file is even less likely to count as a "free and open licence" under that text. The 700 million veto is a field-of-use limit. Field-of-use limits fail the exemption. Meta's EU multimodal withdrawal fits that reading. The company treated the Act as a reason to stop granting rights, not as a reason to publish the training data.
Llama, Gemma, gpt-oss, and several Mistral cards also attach a use policy: no bioweapons, no unsolicited mass persuasion, no child sexual abuse material, sometimes no military. Apache and MIT do not contain those sentences. The policy file does. Breaking it can get you banned from the hub or a marketplace. A policy file is not a law. The AI Office can still fine a GPAI provider. A US company above 700 million MAU still needs Meta's letter.
On 11 August 2026, none of the 7 families had published the training corpus. The cards still said Apache, MIT, Community, or Research.