Back to glossary

Term

Llama 3.1

Llama 3.1 (July 2024) is Metas open-weight family in 8B, 70B and 405B with a 128K context and multilinguality — the 405B was the first open model on par with GPT-4o and Claude 3.5 Sonnet.

Llama 3.1 — explained in more detail

Llama 3.1 is a family of open-weight language models from Meta, released on 23 July 2024. It came in three sizes — 8 billion, 70 billion and 405 billion parameters — each as a pre-trained base and an instruction-tuned (Instruct) variant. The 405B model was at that point the largest openly available language model, and also the first open model to be competitive on benchmarks with the leading API models of the time, GPT-4o and Claude 3.5 Sonnet.

The core additions over Llama 3 were a context window extended to 128,000 tokens and official multilinguality in eight languages: English, German, French, Italian, Portuguese, Hindi, Spanish and Thai. In benchmarks Llama 3.1 405B Instruct reached, among others, 87.3 percent on MMLU (5-shot), 96.8 percent on GSM-8K (maths, CoT), 89.0 percent on HumanEval (code) and 91.6 on MGSM (multilingual maths) — the last on par with Claude 3.5 Sonnet. The models are under the Llama community license, which permits commercial use but carries some conditions (such as an acceptable-use policy and a threshold for very large providers) and therefore does not count as a fully free open-source license like Apache 2.0.

Example / Practical context

The size tiering makes Llama 3.1 flexible in practice: the 8B model runs on a single consumer GPU or, quantized, on a laptop and suits local chat, extraction or classification services; the 70B is a strong all-rounder for server deployments; the 405B targets demanding tasks and serves as an open alternative to closed frontier models. A typical use: a company runs Llama 3.1 70B self-hosted to process customer data without sending it to an external API — the open weights are exactly what make that possible. Llama 3.1 also served as a popular base for countless fine-tunes, including German-language variants.

Within the Llama line, 3.1 follows Llama 3 (same generation but without long context and multilinguality) and was later superseded by Llama 3.2 (adding small and multimodal models) and the Llama 4 generation. The central difference from API-only models (GPT, Claude, Gemini) is the open weights. Note the license: the Llama community license is open but not as unrestricted as Apache 2.0 (as with Mistral) — so Llama is more precisely described as open weight than open source. All 3.1 models are dense transformers, not Mixture-of-Experts.

See everything in one place:Llama