Comparing AI models — who builds what and how to choose
The major model families in 2026 at a glance. Who builds Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen — and which model to pick when.
in AI Models
DeepSeek's open MoE and reasoning models (China) — the V series as general-purpose MoE, the R series for reasoning, and the multimodal Janus line; known for strong performance at low cost.
Open reasoning model from the Chinese lab DeepSeek, released on 20 January 2025 under the MIT license. A mixture-of-experts with 671B parameters (37B active per token), heavily trained via reinforcement learning; the first open model to match OpenAI's o1 level.
DeepSeek V4.1-Flash (September 10, 2026) is the smallest model in the new DeepSeek V4.1 architecture family — the first DeepSeek generation with native multimodal image processing. A speed-/cost-optimized "Flash" model with open weights (MIT), available via boostN.
DeepSeek V4-Pro (2026) is DeepSeeks open-weight flagship — a 1.6-trillion-parameter MoE with about 49B active parameters per token, a 1M-token context and an MIT license.
DeepSeek V4-Flash (July 2026) is an open MoE model with 284B parameters (13B active), a 1M-token context and an MIT license — freely downloadable on Hugging Face.
Janus-Pro-7B is DeepSeek's open-source multimodal model from January 2025 — unifying image understanding and image generation in a single 7B architecture, MIT-licensed.
DeepSeek V3 is an open model from the Chinese provider DeepSeek, released in December 2024. It uses a Mixture-of-Experts architecture with 671 billion parameters (37 billion active per token) and a context window of around 128,000 tokens.
DeepSeek V3.1 is DeepSeeks open-weights hybrid model from August 2025 — it unifies a think and a non-think mode in a single model, with a 128K context and strong agentic skills.
DeepSeek V3.2 (December 1, 2025) is an open MoE model with 685B parameters under an MIT license — especially efficient on long contexts thanks to DeepSeek Sparse Attention.
DeepSeek V4 (April 2026) is the fourth generation of the Chinese open-weight MoE model — released as Pro (1.6T parameters) and Flash (284B) under MIT license, with a 1M-token context window.
The major model families in 2026 at a glance. Who builds Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen — and which model to pick when.
Claude, GPT, Gemini, Llama & co. — who builds what, where each family shines, and how to pick the right model for your own use case.
DeepSeek V4.1-Flash: rank 1 on coding under 1h, sharp drop on agentic work over 1h. Prices, benchmarks and an honest read in the lexicon.
DeepSeek V as an open model line: positioning, use profile and how to choose. Flagship DeepSeek V4-Pro — cheap and self-hostable.
DeepSeek makes its 75% V4-Pro discount permanent: $0.435 input, $0.87 output per million tokens. What the price pressure means for model choice.
July to August 2026 brought a wave of open models: GLM-5.3, DeepSeek V4-Pro, Qwen3.8. What the permissive licenses mean for self-hosting.
DeepSeek ships V4.1-Flash: tops its own coding benchmarks, falls far behind on long-horizon agentic tasks — at a fraction of the competition's price.