Image & Video
in AI Models
Generative models for image and video creation from text or image prompts — from text-to-image and inpainting to text-to-video. A cross-vendor category, e.g. FLUX.1, Stable Diffusion, Sora, and Veo.
Glossary
FLUX.2 is the text-to-image model family from Black Forest Labs. It spans the proprietary Pro and Flex variants plus the open 32-billion-parameter Dev model and the compact Klein series, unifying image generation and image editing in a single model.
GPT Image 2 is OpenAIs native image model (April 2026) that reasons before drawing, renders text very reliably and produces high-resolution photorealistic images.
Veo 3.1 is Googles video model (DeepMind) that turns text or an image into 8-second clips with natively synchronized audio. Since the January 2026 update it delivers true 4K (3840x2160) and native vertical formats.
Google Imagen is the text-to-image model family from Google DeepMind. The current generation, Imagen 4, launched in 2025 in Fast, Generate and Ultra tiers, is available via the Gemini API and Vertex AI, and is known for strong typography and prompt adherence.
Kling 3.0 is Kuaishous video model, released on February 5, 2026. It generates photorealistic clips of up to 15 seconds with native audio across multiple languages and is regarded as the strongest price-to-performance video generator.
Midjourney v8 is the eighth generation of Midjourneys image generator (alpha March 2026) with roughly five times faster generation, native 2K resolution and improved text rendering.
Runway Gen-4.5 is Runways video model that turns text or an image into 5- and 10-second clips with strong prompt adherence and cinematic motion. It was built with NVIDIA and uses an Autoregressive-to-Diffusion technique.
Seedream 4.0 is ByteDances multimodal image model (September 2025) that processes text and multiple images as input, renders up to 4K and runs more than ten times faster than Seedream 3.0.