Term
DeepSeek R1
Open reasoning model from the Chinese lab DeepSeek, released on 20 January 2025 under the MIT license. A mixture-of-experts with 671B parameters (37B active per token), heavily trained via reinforcement learning; the first open model to match OpenAI's o1 level.
DeepSeek R1 — explained in more detail
DeepSeek R1 is a reasoning model from the Chinese AI lab DeepSeek, released on 20 January 2025 under the permissive MIT license. It drew wide attention because it was the first openly available model to match the reasoning performance of OpenAI’s o1 — at significantly lower reported training costs. The open weights can be freely downloaded, reused and fine-tuned.
Technically, R1 is a mixture-of-experts (MoE) model with 671 billion parameters, of which only around 37 billion are active per token. It builds on the DeepSeek-V3 architecture. Its training approach is distinctive: DeepSeek used large-scale reinforcement learning (including the GRPO algorithm) so that the model develops reasoning strategies such as self-verification and error correction on its own. The model “thinks” via a visible chain-of-thought before answering.
Alongside the large R1, DeepSeek released distilled, smaller variants (such as 1.5B, 7B, 8B, 14B, 32B and 70B) that transfer the reasoning capabilities into more manageable, locally runnable models.
Example / practical relevance
DeepSeek R1 suits tasks with a high logic and mathematics component: multi-step problem solving, code reasoning or scientific questions. Because the weights are open and MIT-licensed, companies can self-host the model — relevant for data protection, cost control and independence from a single API provider.
The distilled variants are especially practical when the full 671B infrastructure is not available: a 32B or 70B model runs on far smaller hardware while retaining part of the reasoning strength.
Distinction from similar terms
R1 is DeepSeek’s reasoning line and differs from the base model DeepSeek-V3, which is geared more toward general language tasks. Compared with proprietary reasoning models (OpenAI o1, Claude with Extended Thinking), R1 is openly licensed and self-hostable. “Distilled” R1 models are smaller models trained on the large R1’s outputs — not the original itself.