Small Language Models
The counter-trend that matters: why small models keep winning real workloads, who builds them, and how small gets made good.
Content last verified 2026-09.
Lessons
Sources
- Gemma 3 model card (gemma-3-4b-it) — sizes, training tokens, context, license
- Qwen3-0.6B model card — parameters, context length, thinking modes
- SmolLM3-3B model card — a fully open 3B model with published training details
- Phi-4-reasoning-vision-15B model card — a compact open-weight multimodal reasoning model
- Mistral models overview — Ministral 3 sizes and license tags