Term
Kokoro
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
Kokoro — explained in more detail
Kokoro (Kokoro-82M) is a text-to-speech (TTS) model whose version 1.0 was released on January 27, 2025 under the permissive Apache-2.0 license. What stands out most is its small size: with only 82 million parameters it is a fraction of many competing speech-synthesis models, yet it delivers comparable quality — at significantly higher speed and lower cost. The weights occupy roughly 1 GB of VRAM, so Kokoro runs even on modest hardware or in the browser.
The v1.0 release ships 54 voices across 8 languages. The model was deliberately trained only on permissive, non-copyrighted audio data with IPA phoneme labels — which is what makes the permissive licensing possible in the first place. Kokoro is therefore the TTS counterpart to speech-to-text models like Whisper: instead of turning audio into text, it turns text into audible speech.
Example / In practice
Kokoro is well suited for voiceovers, screen readers, in-app speech output or narrating articles and podcasts — locally and without API costs. Because it is small and Apache-licensed, it can be embedded into your own products without being tied to a cloud provider.
Distinction from similar terms
Kokoro generates speech (TTS), Whisper transcribes it (STT) — both belong to the Speech & Voice use case. Compared to commercial services such as ElevenLabs, Kokoro wins on open weights and self-hosting, while those offer broader voice variety and voice cloning.
Discover more
AI in Review — August 2026: A Price U-Turn, the Open-Weight Wave and an Enforceable EU AI Act
August 2026 in AI: Sonnet 5 price freeze, Qwen3.8-Max, the open-weight wave, Google AI Mode on Gemini Flash and the enforceable EU AI Act — sorted out.
GlossaryChatterbox
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
EncyclopediaFLUX.2
FLUX.2 by Black Forest Labs (Freiburg): open-weight image model, variants from 4B to 32B, licenses, hardware needs and how it stacks up in the market.