Home / Directory / AI SaaS tooling / Voice, speech & multimodal pipelines / Deepgram

Deepgram

Real-time voice APIs for speech-to-text, text-to-speech, and voice agents, available as cloud APIs or self-hosted.

7.6/10
Overall Score
Recommend

Deepgram is a specialist speech stack for production voice agents and transcription, with cloud and self-hosted options and unpublished revenue.

Best for

Product teams building production voice agents or transcription that need a specialist speech API.

Not ideal for

Buyers that want studio avatar video or a consumer voice-cloning app.

Verdict

Hire Deepgram when the product needs production speech: transcription, text-to-speech, or a voice agent, as a cloud API or self-hosted. The fetched product set includes Nova-3, Aura-2, Flux, and a Voice Agent API.

A Series C of $130 million at a $1.3 billion valuation was announced January 13, 2026. The company said it had raised over $215 million.

Score Breakdown

How Deepgram scores on the jobs buyers hire it for, and on the company and commercial side of the deal.

Buyer outcomes

Speech-to-text and text-to-speech
7.8
Voice agent API
7.7
Cloud and self-hosted options
7.6
Production speech fit
7.4

Company & commercial

Innovation & product leadership
7.7
Project management & communication
7.4
Pricing
7.5
Contract fairness
7.7

Use-case matrix

Use caseFitNotes
Production speech-to-textStrongA core API.
Voice agentsStrongA published Voice Agent API.
Self-hosted speechStrongSold beside the cloud API.
Avatar video productionWeakA different creative tool.

Review Excerpts

Below are excerpts from public reviews. Our team scoured public reviews, forums, and chat rooms to get a balanced view of customers' experience with this company. Paid reviews and pay-for-play sites such as Clutch were excluded.

What people like

“The fast transcription rates significantly improve time management by transcribing hours of audio in seconds. Deepgram's API is straightforward to integrate, and the quality of transcriptions is greatly enhanced.”

Muhammad A. · G2 · source
What people don't like

“Sometimes the multi-language support doesn't work perfectly. When calling customers who speak Indonesian or Dubai dialects, it doesn't detect their language well.”

Aman S. · G2 · source

Methodology

This page is an independent evaluation of Deepgram for buyers comparing options in voice, speech & multimodal pipelines. AI Industry Reviews accepts no sponsorships, advertising, or pay-for-placement fees. Deepgram did not pay for this review.

What we scored

The headline number is an Overall Score on a 0-10 scale. Eight criteria sit under it in two groups.

  • Buyer outcomes (for voice, speech & multimodal pipelines)
    • Speech-to-text and text-to-speech
    • Voice agent API
    • Cloud and self-hosted options
    • Production speech fit
  • Company & commercial
    • Innovation & product leadership
    • Project management & communication
    • Pricing
    • Contract fairness

Pricing measures whether the price looks fair for the value delivered, including packaging and renewal friction that show up in real buying cycles.

Score Composition

InputWeightWhat it covers
Reviews40%A proprietary read of what practitioners say about likes, complaints, and day-to-day use, including public review sites, forums, and private chat rooms. Paid reviews and pay-for-play sites such as Clutch are out of scope.
Product35%Hands-on look at screens and workflows.
Pricing15%Whether the price looks fair for what you get.
Docs & training10%Docs, tutorials, and training material.

How we balanced the evidence

The Overall Score is the simple average of the eight criteria. Recommendation language follows that score and the fit pattern described above.

Scope

Deepgram is graded here as voice, speech & multimodal pipelines. Criteria scores can move as more review volume and product checks are added.

← Back to Voice, speech & multimodal pipelines