Open-weight licences.
The thing that will stop you shipping an open-weight model is almost never copyleft — it is a clause that attaches to who you are rather than to what you built, which means two companies can run byte-identical weights and only one of them is compliant. That is the distinction worth carrying: a standard licence (Apache 2.0, MIT) can be cleared once by reading a file, while a bespoke model agreement has to be re-cleared against your own user counts, your own jurisdiction and a usage policy that lives at a URL and changes without a version bump. The good news is that the industry has spent 2026 moving toward the first kind; the bad news is that your already-deployed models are the ones under the second.
Three separate documents get called "the licence".
"Open weights" says something about distribution — you can download the parameters — and nothing at all about permission. The permissions come from up to three documents, written by different people, changing on different schedules, and a compliance answer that cites only one of them is not an answer.
- The weight licence. Governs copying, modifying, redistributing and serving the parameter files. This is the document people mean, and it is the only one of the three that is usually versioned alongside the release.
- The acceptable-use policy. A list of prohibited uses, almost always incorporated by reference to a live URL rather than pinned into the licence text. Apache 2.0 weights can still arrive with one attached — the permissive licence covers the copyright grant, and the AUP rides alongside it.
- The output terms. Whether you may train another model on this model's outputs, and whether anything is claimed over what your deployment generates. This is where the "distillation" question lives, and it is separate from your right to run the thing at all — see distillation and quantization and IP and copyright for agent output.
The middle document is the one that moves. A licence file in a repository has a commit history you can diff; a policy at a vendor URL does not necessarily, and your entitlement to keep running last year's checkpoint can change without anything in your lockfile changing. If you rely on open weights for a reason — air-gapping, residency, unit cost — snapshot all three at download time, next to the digest, for the reasons in pinning and verification.
The axis is standard versus bespoke, not permissive versus restrictive.
Engineers sort licences by how much they permit. For an agent team the more useful sort is by how much work it takes to know where you stand, because that cost recurs on every model swap and every new market. A standard licence has been read by thousands of lawyers and has no conditions that reference you; a bespoke agreement is a contract drafted by one counterparty, with conditions that may need legal review each time your business changes shape.
# Two questions, in the order that decides the answer Q1 Is this a standard, unmodified OSI licence (Apache-2.0, MIT, BSD) with no rider? yes -> clear once, per licence, forever no -> continue Q2 Do any conditions reference facts about the LICENSEE rather than facts about the USE? (user counts, headquarters, sector, affiliates) yes -> re-clear on every business change no -> clear once, per model # Landscape as of late 2026 Apache-2.0 Qwen3 and later, Gemma 4 (Apr 2026), gpt-oss MIT DeepSeek R1 and V3/V4 checkpoints bespoke Llama 4 Community License
Two changes in 2026 made this a live question rather than a settled one. Google moved Gemma 4 to Apache 2.0 in April, retiring the source-available Gemma Terms of Use that governed Gemma 3 and earlier; Alibaba has released Qwen under Apache 2.0 from Qwen3 onward; OpenAI's gpt-oss shipped Apache 2.0 with a usage policy attached. Meta is the conspicuous holdout, which matters less for what you deploy next than for what you deployed in 2024 and 2025 and never revisited.
Apache 2.0 and MIT are not interchangeable, and the difference is the one thing in this whole subject that is purely legal. Apache 2.0 contains an express patent grant from the contributors and a termination clause if you sue them over it; MIT says nothing about patents. If you are shipping a product rather than running an experiment, that grant is the reason to prefer Apache 2.0 weights where both exist.
What bespoke terms attach to, with the Llama 4 agreement as the worked example.
The Llama 4 Community License is worth reading in full once, not because Llama is uniquely restrictive but because it is the clearest specimen of conditions that are about the licensee rather than the use. Three of its clauses illustrate the whole category.
- A user-count threshold. If your products had more than 700 million monthly active users in the calendar month before the release date, you must request a licence from Meta, granted at Meta's sole discretion. Almost nobody trips this — but note the shape: it is a condition your growth can violate, on terms the counterparty decides, and no amount of care in your code affects it.
- A jurisdiction carve-out. Companies and individuals domiciled in the EU are not granted rights to the multimodal models, while end users of products built with them are unaffected. So the same weights are licensed differently depending on where the deploying entity sits, which turns a legal question into an architecture question the moment you open an EU subsidiary. This is the licence half of data residency and sovereignty.
- Attribution and naming obligations. Displaying "Built with Llama", carrying the copyright notice, and including "Llama" in the name of derivative models. These are cheap to comply with and easy to forget, and they are the clauses most likely to be violated silently by a fine-tuned checkpoint that someone renamed.
Notice what none of these are: a restriction on commercial use. Bespoke model licences are mostly commercially generous and conditionally personal, which is exactly backwards from the mental model most engineers import from open-source software, where the conditions attach to what you distribute.
Fine-tuning is where the obligations get inherited and then lost. A LoRA adapter trained on a bespoke-licensed base is a derivative under most of these agreements, so the naming rule, the AUP and the jurisdiction carve-out all follow it — including into the internal model registry where nobody recorded the base. If your registry stores a checkpoint without its lineage, you have no way to answer the question later; that record is part of agent supply-chain security, not a separate chore.
Record four facts per model and the question becomes answerable.
The practical control is small. For every set of weights you run, store four things alongside the digest, and the compliance question stops requiring a lawyer each time it is asked.
# One record per deployed checkpoint licence SPDX id, or "bespoke" + local copy of the text aup URL + retrieval date + local copy outputs may we train on this model's outputs? y/n/unclear lineage base model + digest, recursively # Derived once, re-derived only on business change exposure which bespoke conditions could our growth, our jurisdictions or our sectors violate?
- Prefer standard licences when the capability is close. The gap between the best Apache 2.0 or MIT open-weight model and the best bespoke one has narrowed to the point where, for most agent workloads, the licence is a legitimate tiebreaker rather than a sacrifice. Score it explicitly in choosing a model instead of discovering it at launch review.
- Audit the models you already shipped, not the ones you are choosing. Selection gets legal attention; the checkpoint quietly serving production since 2024 does not. That is where the unrecorded lineage and the unmet naming obligation live.
- Do not let "open" do argumentative work. Open weights buy you air-gapping, residency control and a unit cost you own — the reasons in small and local models and self-hosted inference for agents. They do not buy you freedom from a counterparty, and a bespoke licence keeps one in the picture indefinitely, which is the same risk register entry as third-party model and vendor risk.
- Treat the AUP as a moving dependency. Diff it on a schedule. It is the only part of this that can change under a model you have already deployed and pinned by digest.
Do this today: list every set of open weights running in production, and for each one write down the SPDX identifier from memory. The ones you cannot name are the finding — not because they are non-compliant, but because nobody has checked, and the check costs ten minutes per model while the answer at launch review costs a week. Then read open vs closed models for what the weights actually buy you, and model deprecation and migration for how to leave one behind when the terms change.