Home / Directory / AI infrastructure / Inference & model hosting / Cerebrium
Cerebrium
Serverless GPU platform, founded in Cape Town and now based in New York, for running voice agents, video models, and LLMs with per-second billing.
Teams running real-time voice agents, video models, or custom LLMs that want serverless GPUs, multi-region deploys, and data residency without an infrastructure team.
Teams that want a hosted model API with no containers to manage, or workloads that run GPUs flat out all month where reserved capacity is cheaper.
Verdict
Cerebrium is a serverless AI infrastructure company founded in Cape Town, South Africa, by Michael Louis and Jonathan Irwin and now headquartered in New York City. Engineers bring their own code or Dockerfile, and Cerebrium runs it on GPUs across several clouds and regions with 2 to 4 second cold starts, autoscaling, and per-second billing, for customers such as Tavus, Deepgram, and Vapi. It matters because real-time voice and video AI needs low latency and sudden scale, which is hard for small teams to run themselves.
Score Breakdown
How Cerebrium scores in the categories that matter to its buyers.
Pricing
Cerebrium bills compute by the second on top of a platform plan. Sample GPU rates are shown for the Standard plan.
| Plan | Price | What stands out |
|---|---|---|
| Hobby | Free + compute | 3 seats, up to 3 deployed apps, 5 concurrent GPUs |
| Standard | $100 / month + compute | Unlimited seats and apps, 30 concurrent GPUs, custom domains |
| Enterprise | Custom | Unlimited GPU concurrency, volume discounts, dedicated Slack |
| H100 compute | $0.000944 / second | About $3.40 an hour |
| A100 80GB compute | $0.000583 / second | About $2.10 an hour |
The Field at a Glance
How Cerebrium compares with other Inference & model hosting vendors we reviewed, by Overall Score and relative typical engagement cost.
Cerebrium scores 7.8 in Inference & model hosting, under Modal (8.4) and Replicate (8.0), and ahead of Together AI (7.3). Relative cost lands mid-band for this subcategory.
Use-case matrix
| Use case | Fit | Notes |
|---|---|---|
| Real-time voice agents | Strong | Templates cover Pipecat, LiveKit, and Twilio agents. |
| Bursty traffic | Strong | Containers launch in seconds and scale automatically. |
| Data residency | Strong | Workloads can be pinned to specific regions. |
| Bring your own code | Strong | No custom SDK or decorators required. |
| Ready-made model APIs | Weak | You deploy your own models rather than calling a catalog. |
| Constant full-time GPU use | Mixed | Per-second billing favors variable workloads. |
Who it’s for
Good fit
- Voice AI startups
- Video and avatar model teams
- Companies needing regional data residency
Poor fit
- Teams that want a model catalog API
- Always-on training clusters
- Teams with no container experience
Review Excerpts
Below are excerpts from public reviews. Our team scoured public reviews, forums, and chat rooms to get a balanced view of customers' experience with this company. Paid reviews and pay-for-play sites such as Clutch were excluded.
“We run a range of real-time audio and video models, and performance is everything. We tried a number of solutions, but Cerebrium consistently delivered the speed and reliability we needed without the overhead.”
“Even as we’ve scaled rapidly and gone viral, they’ve kept up with our compute demands and delivered the stability we rely on. It has become a core part of our infrastructure!”
The Standard plan costs $100 a month before any compute and caps GPU concurrency at 30; unlimited concurrent GPUs require a custom Enterprise contract.
Methodology
This page is an independent evaluation of Cerebrium for buyers comparing options in inference & model hosting. AI Industry Reviews accepts no sponsorships, advertising, or pay-for-placement fees. Cerebrium did not pay for this review.
What we scored
The headline number is an Overall Score on a 0-10 scale. Eight criteria sit under it in two groups.
- Buyer outcomes (for inference & model hosting)
- Cold start and latency
- Real-time voice and video fit
- Regions and data residency
- Breadth of managed model APIs
- Company & commercial
- Innovation & product leadership
- Project management & communication
- Pricing
- Contract fairness
Pricing measures whether the price looks fair for the value delivered, including packaging and renewal friction that show up in real buying cycles.
Score Composition
| Input | Weight | What it covers |
|---|---|---|
| Reviews | 40% | A proprietary read of what practitioners say about likes, complaints, and day-to-day use, including public review sites, forums, and private chat rooms. Paid reviews and pay-for-play sites such as Clutch are out of scope. |
| Product | 35% | Hands-on look at screens and workflows. |
| Pricing | 15% | Whether the price looks fair for what you get. |
| Docs & training | 10% | Docs, tutorials, and training material. |
How we balanced the evidence
The Overall Score is the simple average of the eight criteria. Recommendation language follows that score and the fit pattern described above.
Scope
Cerebrium is graded here as inference & model hosting. Criteria scores can move as more review volume and product checks are added.
