Home / Founders / Scott Stephenson
Scott Stephenson
Co-founder and CEO
Deepgram
Scott Stephenson is co-founder and CEO of Deepgram, the speech company he started after leaving particle-physics research. He builds developer-facing speech systems that other teams meter into products, and he still talks about the work in the language of waveform analysis and first principles he learned underground in dark-matter labs.
Early days in particle physics
He grew up in academic physics, earning a bachelor's degree at the University of Missouri-St. Louis and then a master's and PhD in particle physics at the University of Michigan. At Michigan his research centered on dark-matter detection: building sensitive liquid-xenon experiments deep underground, including a laboratory roughly two miles below the surface, and writing analysis software that hunted rare event signatures inside huge waveform datasets.
In a Madrona Featured Leader conversation, he described that lab as "basically in a James Bond lair" and recalled the stretch after the detector started working: months of data-taking with little to do but wait. With Noah Shutty, a collaborator from the same physics world, he built wearable recorders that captured voice notes, conversations, and ambient audio around the clock. The devices produced "over 1,000 hours of audio," he told Madrona, and nobody wanted to listen back. In a RedMonk Conversation he tied the physics habit to speech directly: terabytes of mostly uninteresting signal, a few moments that matter, and models that can find "a needle in a haystack of recorded audio." The speech tools of the mid-2010s still looked like older pipelines. Stephenson and his collaborators were already applying end-to-end deep learning to detector waveforms, and they wanted the same approach for speech.
Founding Deepgram
In 2015 he left a postdoctoral post at UC Davis and, with Shutty and fellow Michigan physicist Adam Sypniewski, founded Deepgram in the Bay Area. On Madrona he has said they "didn't set out to build a company," but once they saw even cutting-edge firms still using "antiquated technology" for open-ended audio, "it became obvious that we were the ones to build it." He framed the career switch in first-principles language: "what we were working on in physics was discovering the fundamental laws of nature. But working on Deepgram would be discovering the fundamental laws of intelligence." The early company tried to be a search engine for sound, then found commercial footing in business audio such as support-call corpora before concentrating on speech recognition APIs other teams could meter into their own products. Deepgram joined Y Combinator's Winter 2016 batch.
That first embedding-and-search product, he later admitted, arrived years before buyers were ready. In the Madrona interview he distilled the timing lesson: "Build for a year or two into the future." On Madrona's Founded & Funded episode he put the pivot more bluntly: early researchers found speech-to-text "boring" next to embeddings and speech-to-speech dreams, but the company had to "earn our license to learn in this market" by becoming undeniable in one domain first.
Building speech as infrastructure
As CEO he has kept the company on foundational speech models that product teams can buy as infrastructure: transcription first, then speech generation and tooling for conversational voice agents. Talking to Voicebot about the developer motion, he said: "Think of Twilio, think of Stripe, think of an API integration you use, which is a piece of a greater system you build yourself internally, rather than buying a behemoth one-size-fits-all instead." And: "One-size-fits-all just won't do it." Early on he sold the product himself while the technical team chased accuracy and latency. On Madrona he recalled an almost year-and-a-half stretch when the model was already strong and the company barely trained: "all we did for that year and a half was smile and dial," because distribution and reliability, not another accuracy point, were what blocked growth.
On first principles and voice interfaces
In interviews he carries a physics habit of first principles into AI. On Founded & Funded he compared foundational AI builders to Tesla versus Ford or SpaceX versus Lockheed: the visible product may look similar, but "the difference is not so much what is the chemical makeup of the metal" as "the methodology by how they arrive at their product." Internally, he said, Deepgram asks "Should we do it at all?" then "Can we automate this?" before defaulting to hiring. Academia had taught him that superior methods spread in months after a paper; the business world taught the harder lesson. "Many folks say distribution matters even more than a product," he told Madrona. "That's probably true in the long run."
His public case for Deepgram is that spoken interfaces fail when recognition, reasoning, and speech generation do not share context, and when latency or cost makes the loop feel broken. On RedMonk he argued that voice is finally having its moment because speech-to-text, language models, and text-to-speech can be stitched into agents people will talk to, and that B2B teams need controllable speech infrastructure rather than consumer demos built for someone else's product. He hosts The Scott Stephenson AI Show, where the same first-principles framing shows up when he talks with other builders about voice and intelligence systems.
Index score breakdown
Overall 73.5 · Rank 19 on the AI Founders to Watch Index
| Factor | Score |
|---|---|
| Innovation | 8 |
| Impact | 8 |
| Company success | 7 |
| Vision clarity | 8 |
| Credibility | 9 |
| Momentum | 7 |
| Independence of signal | 3 |
