
Deepgram
Voice AI infrastructure for speech-to-text, text-to-speech, and voice agents
Gallery
3 items



About Deepgram
Deepgram is a foundational voice AI company providing developer-first APIs for speech-to-text transcription, text-to-speech synthesis, and real-time conversational voice agents. The company was founded on August 18, 2015 by Scott Stephenson, Adam Sypniewski, and Noah Shutty, three University of Michigan physicists who had previously worked on dark matter detection research at a lab in China. The origin story is genuinely unusual: Stephenson and Shutty had built a wearable device to record audio around the clock as a personal life-logging experiment, but found no good way to search or index the hundreds of hours of audio they accumulated. Realizing the waveform-analysis techniques from their physics research could apply to speech, they built what would become Deepgram, going through Y Combinator's Winter 2016 batch shortly after.
Deepgram's core product is Nova-3, its third-generation automatic speech recognition model, which delivers a 47-54% reduction in word error rate compared to competitors and was the first voice AI model to offer self-serve customization, letting developers adapt vocabulary instantly without retraining. Flux, a newer model built specifically for voice agent use cases, adds model-integrated end-of-turn detection and configurable turn-taking dynamics on top of Nova-3-level accuracy, directly addressing one of the hardest problems in conversational AI: knowing when a user has actually finished speaking. Aura-2 handles text-to-speech with 90ms latency for real-time applications, and the Voice Agent API bundles speech-to-text, LLM orchestration, and text-to-speech into a single unified endpoint, removing the need for developers to stitch together three separate services with their own latency and reliability tradeoffs.
Beyond core transcription and synthesis, Deepgram's Audio Intelligence layer adds built-in summarization, sentiment analysis, topic detection, and intent recognition directly on top of transcribed audio, while speaker diarization, punctuation, and custom keyterm prompting round out the platform for production call center and meeting transcription use cases. The company has processed over 50,000 years of audio and transcribed more than 1 trillion words, serving over 1,300 organizations including NASA (transcribing ISS-to-Mission-Control communications), Spotify, Twilio, Citi, and Vapi. In January 2026, Deepgram raised $130 million in a Series C round at a $1.3 billion valuation, becoming the newest unicorn in the voice AI space while remaining cash-flow positive since 2025.
Deepgram's pricing is genuinely usage-based and billed by the exact second with no rounding: new accounts get $200 in free credit with no credit card required, covering tens of thousands of minutes of pre-recorded transcription, before moving to transparent Pay-As-You-Go rates starting around $0.0043/minute for Nova-3 pre-recorded transcription. This makes Deepgram best suited for developers and engineering teams building production voice pipelines, call center analytics, or real-time voice agents who specifically need low latency and high accuracy at scale, rather than non-technical users looking for a simple, no-code transcription dashboard.
Key Features
- Nova-3 speech-to-text with 47-54% lower word error rate than competitors
- Flux model with built-in end-of-turn detection for natural voice agent conversations
- Voice Agent API unifying STT, LLM orchestration, and TTS in one real-time endpoint
- Aura-2 text-to-speech with 90ms latency for real-time conversational applications
- Self-serve model customization for instant vocabulary adaptation without retraining
- Audio Intelligence: built-in summarization, sentiment, topic detection, and intent recognition
- Speaker diarization, punctuation, and custom keyterm prompting
- On-premise and self-hosted deployment for security, compliance, and latency-sensitive use cases
Pros
- Nova-3 delivers industry-leading accuracy with a 47-54% lower word error rate than competing models
- Voice Agent API eliminates the need to stitch together separate STT, LLM, and TTS services
- $200 free credit with no credit card required is one of the more generous evaluation tiers in the category
- Billing by the exact second with transparent, published per-minute rates avoids hidden pricing surprises
- Proven at massive scale: 50,000+ years of audio processed, 1 trillion+ words transcribed, used by NASA, Spotify, and Twilio
Cons
- Growth tier discounted pricing requires a $4,000+ annual prepayment commitment
- On-premise/self-hosted deployment requires a custom enterprise agreement rather than self-serve setup
- Steeper learning curve than simpler transcription apps, built for developers rather than non-technical dashboard users
- Multi-channel (stereo) audio transcription costs meaningfully more than mono per-minute rates
- Flux TTS pricing shifts from free to standard rates after September 12, 2026, which teams building now should plan around
Pricing
Deepgram offers $200 in free credit on signup with no credit card required, usable across any service (Speech-to-Text, Text-to-Speech, Voice Agent API, and Audio Intelligence) and credits do not expire. Pay-As-You-Go pricing for Nova-3 starts at approximately $0.0043/minute for mono pre-recorded transcription (roughly $0.0052/minute for multi-channel/stereo audio) and $0.0077/minute for real-time streaming transcription, with rates dropping to around $0.003/minute at high volume. The Voice Agent API, which bundles speech-to-text, LLM orchestration, and text-to-speech, is priced at $4.50/hour, or roughly $0.08/minute (with a lower bring-your-own-LLM rate around $0.07/minute). Flux text-to-speech is free to use through September 12, 2026 (up to 45 concurrent streaming connections globally, 5 in EU/AU), with standard pricing applying starting September 13, 2026. A Growth tier offers roughly 15-20% discounted rates in exchange for a $4,000+ annual prepayment commitment. On-premise and self-hosted deployment for security, compliance, or latency-sensitive workloads requires a custom Enterprise agreement.
Claim Verified Creator Badge
Are you the founder of Deepgram? Display this listing's verified badge on your website to show your customers that your product has been vetted and listed on AI Central Resources.
<a href="https://www.aicentralresources.com/tool/deepgram" target="_blank" rel="noopener"> <img src="https://www.aicentralresources.com/badges/featured-badge-dark.svg" alt="Featured on AICentralResources" width="200" height="54" style="border: none;" /> </a>
* Place this HTML snippet in your website's footer, landing page, or press section. This creates a search-friendly backlink directly to your verification page.
Connect with Deepgram
Frequently Asked Questions
Deepgram is a voice AI infrastructure platform providing APIs for speech-to-text transcription, text-to-speech synthesis, and real-time conversational voice agents, used by developers building call center analytics, meeting transcription, voice assistants, and other audio-based applications.
Deepgram gives new accounts $200 in free credit with no credit card required, covering roughly 46,000+ minutes of pre-recorded transcription. After that, Pay-As-You-Go pricing starts around $0.0043/minute for pre-recorded Nova-3 transcription and $0.0077/minute for streaming, with a Growth tier offering discounted rates for a $4,000+ annual prepayment, and custom Enterprise pricing for on-premise deployment.
Deepgram's Voice Agent API is a single, unified API that combines speech-to-text, LLM orchestration, and text-to-speech into one real-time conversational pipeline, eliminating the need to stitch together three separate services. It's priced at $4.50/hour and includes built-in barge-in detection and turn-taking prediction.
Nova-3 is Deepgram's third-generation speech-to-text model, delivering a 47-54% reduction in word error rate compared to competitors, with self-serve customization for instant vocabulary adaptation. Flux is Deepgram's newer model built specifically for voice agents, adding model-integrated end-of-turn detection and configurable turn-taking dynamics on top of Nova-3-level accuracy.
Deepgram was founded in 2015 by Scott Stephenson, Adam Sypniewski, and Noah Shutty, three University of Michigan physicists who had worked on dark matter detection research. After struggling to search their own recorded audio archives with existing tools, they applied their waveform analysis background to build a deep learning speech recognition company, going through Y Combinator in Winter 2016.
Similar AI Tools to Deepgram
View all alternatives of Deepgram
AssemblyAI
Voice AI infrastructure for speech-to-text, voice agents, and conversation intelligence

Voice.ai
Real-time AI voice changing, cloning, text-to-speech, and voice agents in one platform

Murf AI
Studio-quality AI voiceovers and voice agents in 200+ voices across 35+ languages


