Term
Deepgram Nova-3
Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.
Deepgram Nova-3 — explained in more detail
Nova-3 is the current speech-to-text model from the US provider Deepgram, released in February 2025. It is a closed model delivered through Deepgrams cloud API and via marketplaces; the weights are not public. Its focus is real-time transcription: Nova-3 streams audio over a WebSocket endpoint with end-to-end latency of around 200 to 300 milliseconds under good conditions, returning interim transcripts that may be revised, then finalized transcripts that are locked as more audio arrives.
On accuracy, Deepgram reports a median streaming WER of about 6.8 percent and roughly 5.3 percent in batch mode. The model supports multilingual real-time transcription across more than 30 languages, offers self-serve keyterm prompting (up to 100 terms) for domain vocabulary, and real-time redaction of sensitive data. This combination of low latency and customization makes Nova-3 a leader for real-time and streaming among commercial ASR services.
Example / In practice
A typical use case is voice assistants and live captions: in a voice agent or call-center system, microphone audio is streamed continuously to Nova-3, which returns the spoken text with minimal delay. This lets an agent react while the speaker is still talking, and domain-specific terms can be recognized reliably via keyterm prompting.
Distinction from similar terms
While AssemblyAI Universal-2 targets maximum accuracy and ready-made formatting in batch mode, Nova-3 is optimized for low latency in real-time streaming — serving the use-case segment of voice agents and live transcription. Unlike open models such as NVIDIA Canary, Nova-3 is a closed managed service without self-hosting of the weights.
Discover more
Better speech recognition in boostN-CLI: Audio normalization + a flexible Whisper model
Why dictated text got swallowed, how SoX normalization and a model switcher fixed it — and what is actually happening under the hood.
GlossaryAssemblyAI Universal-2
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
EncyclopediaComparing AI models — who builds what and how to choose
The major model families in 2026 at a glance. Who builds Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen — and which model to pick when.