The Open-Weights Families
Lesson 2 of 3 in Model Families.
Seven families anchor the Open weights side of the map as of the fact table’s date. For each one, the walk extracts the same four fields — who develops it, how the weights are actually distributed, what the license says (with the honesty hedge about which model that was verified on), and one trait the family’s own documentation claims. Names and numbers below are quoted from the official pages recorded in the fact table; everything else about these families you should look up, not remember.
Llama is Meta’s family — the Hugging Face org page calls itself the home of “The Llama Family — From Meta”, covering Llama, Llama Guard, and Prompt Guard models. The weights are open but not simply downloadable: you get them after you “accept the license terms and acceptable use policy” — Gated distribution in its canonical form. The current license title, verbatim from Meta’s license page, is the Llama 4 Community License Agreement (earlier Llama versions’ license names were not re-read in this verification run) — a Community license that requires “Built with Llama” attribution, derived model names starting with “Llama”, and a separate license for entities exceeding 700 million monthly active users. Documented trait: Llama 4 models are “natively multimodal” and use “a mixture-of-experts architecture”. (Sources: huggingface.co/meta-llama; developer.meta.com/ai/llama4/license/.)
Mistral / Mixtral is Mistral AI’s catalog (“Frontier AI. In Your Hands.”) — and it is mixed on purpose: open-weights models tagged Apache 2.0, “Modified MIT”, or CC BY-NC 4.0 on the official models overview, alongside “Premier” commercial models. The current flagship is described as “our first flagship merged model” with configurable reasoning effort per request; the family’s namesake trait is historical — Mixtral-8x7B was “a pretrained generative Sparse Mixture of Experts”, an early open Mixture of experts (MoE) — and the Mixtral, Magistral, Devstral, and Pixtral lines now appear only in the docs’ deprecated/retired table. One family: three license classes, a commercial tier, and a live demonstration that even sub-family names deprecate. (Sources: huggingface.co/mistralai; docs.mistral.ai models overview.)
Qwen is “the large language model family built by Alibaba Cloud”, per its org card, with model cards crediting the Qwen Team; the org says it releases “large language models (LLM), large multimodal models (LMM), and other AGI-related projects”. The license column is a lesson in hedging done right: the Qwen3.8-27B card carries apache-2.0 in its metadata, but that is verified for that model only — the org page names no licenses at all, so each card gets its own read. Documented trait, from that same card: “Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos”, a 262,144-token native context “extensible up to 1,000,000 tokens”, and a thinking mode that is on by default. (Sources: huggingface.co/Qwen; huggingface.co/Qwen/Qwen3.8-27B.)
DeepSeek introduces itself on its org card without ceremony: “DeepSeek (深度求索), founded in 2023, is a Chinese company dedicated to making AGI a reality.” Weights are published on Hugging Face; the flagship card shows “License: mit” — again, verified for that model only. Documented trait: the DeepSeek-V4.1-Flash card describes “a multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters” built on a “Causal Encoder-Decoder (CED)” design that can “activate only 8B parameters per token during prefill”, with a “continuously controllable reasoning effort” dial that takes an integer from 1 to 100. (Sources: huggingface.co/deepseek-ai; huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash.)
| Family | Developer | License (as documented) | Documented trait |
|---|---|---|---|
Llama | Meta | Llama 4 Community License Agreement — “Built with Llama” attribution, “Llama”-prefixed derivative names, separate license above 700M monthly active users; weights gated behind license acceptance | Llama 4: “natively multimodal”, uses “a mixture-of-experts architecture” |
Mistral / Mixtral | Mistral AI | Mixed per model: Apache 2.0, “Modified MIT”, CC BY-NC 4.0 — alongside “Premier” commercial models | Flagship: “our first flagship merged model” with per-request reasoning effort; Mixtral-8x7B was “a pretrained generative Sparse Mixture of Experts” |
Qwen | Alibaba Cloud (Qwen Team) | apache-2.0 on the Qwen3.8-27B card — verified for that model only; the org page names no licenses | “Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos”; thinking mode on by default |
DeepSeek | DeepSeek | MIT — “License: mit” on the DeepSeek-V4.1-Flash card, verified for that model only | “Causal Encoder-Decoder (CED)” MoE; “continuously controllable reasoning effort” (integer 1–100) |
Gemma | Google DeepMind | Gemma Terms of Use — the terms themselves note Gemma 4 has a separate Apache 2 license | “Our most capable open models”; terms include a Prohibited Use Policy and a right to “remotely restrict violating usage” |
Phi | Microsoft | MIT — “License: mit” on the Phi-4-reasoning-vision-15B card, verified for that model only | Hybrid reasoning in one model: chain-of-thought for math/science vs direct answers for perception; SigLIP-2 vision encoder |
gpt-oss | OpenAI | “Permissive Apache 2.0 license” | Two variants (gpt-oss-120b, gpt-oss-20b) with adjustable reasoning effort and “Full chain-of-thought”; MoE weights in MXFP4 quantization |
Gemma is Google DeepMind’s open family — “Our most capable open models”, distributed via Kaggle, Hugging Face, Ollama, LM Studio, and Google Cloud. Its license is the case study in why “open-weights” and “permissive” are different claims: the Gemma Terms of Use (last modified April 1, 2026) apply to the models listed in the terms’ Appendix, include a Prohibited Use Policy, reserve Google’s right to “remotely restrict violating usage”, and state that Google claims no rights in generated outputs — while noting that Gemma 4 carries a separate Apache 2 license. The pitch for the current generation: “Frontier-level capabilities and mobile-first AI”. A family can hand you the weights and still keep a contractual kill switch — the license row is where you find that out. (Sources: deepmind.google/models/gemma; ai.google.dev/gemma/terms.)
Phi is Microsoft’s compact family — the open-weights end of the Small language model (SLM) story told later in this domain. The one card verified in this run describes “a compact open-weight multimodal reasoning model” under “License: mit” — the hedge applies once more: verified for that model only, on an org page listing hundreds of models. Documented trait: hybrid reasoning in a single model — Chain-of-thought (CoT) for math and science versus direct answers for perception tasks — with a SigLIP-2 Vision encoder in a mid-fusion architecture, and a card that leans hard on data and compute efficiency rather than scale. (Sources: huggingface.co/microsoft; huggingface.co/microsoft/Phi-4-reasoning-vision-15B.)
gpt-oss is the reminder that the open/closed line runs through vendors, not between them: the card calls these “OpenAI’s open-weight models”, under a “Permissive Apache 2.0 license”. Two variants — gpt-oss-120b (“117B parameters with 5.1B active parameters”, sized to fit a single 80GB GPU) and gpt-oss-20b (“21B parameters with 3.6B active parameters”) — with adjustable reasoning effort, “Full chain-of-thought” access flagged as not intended for end users, a required harmony response format, and MoE weights shipped in MXFP4 Quantization. (Source: huggingface.co/openai/gpt-oss-120b.)
Interactive sorting exercise: Classify each documented license situation from the 2026-09-16 fact table. The skill: telling permissive OSS licenses from custom vendor terms — and knowing when the honest answer is “check each model”.
Interactive checkpoint quiz (1 questions) — open this page in a browser to take it.