Home / Directory / AI SaaS tooling / AI penetration testing / RunSybil
RunSybil
AI offensive security company from San Francisco whose agent, Sybil, pentests apps and infrastructure before major releases and on a continuous basis.
AI and developer-tool companies that ship often and want pentests ahead of big releases, including new MCP servers and APIs.
Buyers who want a self-serve tool with a public price, or teams looking only for a cheap compliance checkbox.
Verdict
RunSybil is an offensive security company with hubs in San Francisco and New York, founded in 2023 by Ari Herbert-Voss, the first security hire at OpenAI, and Vlad Ionescu, who led offensive red teams at Meta. Companies such as Notion, Carta, Baseten, and Turbopuffer use its AI agent, Sybil, to probe their applications and infrastructure, chain vulnerabilities, and turn findings into detection rules. It earns attention because it is built for teams that ship AI features quickly, testing things like new MCP servers before launch rather than on a yearly pentest calendar.
Score Breakdown
How RunSybil scores in the categories that matter to its buyers.
Pricing
RunSybil does not publish prices. Engagements are scoped with the team before a quote.
Sybil is offered as continuous offensive testing and as targeted pentests before launches, mergers, and audits. Customers say pricing sits on par with top pentest firms.
The Field at a Glance
How RunSybil compares with other AI penetration testing vendors we reviewed, by Overall Score and relative typical engagement cost.
RunSybil scores 7.4 in AI penetration testing, under XBOW (7.8), and ahead of Terra Security (7.3). Relative cost lands upper for this subcategory.
Use-case matrix
| Use case | Fit | Notes |
|---|---|---|
| Pre-launch pentests | Strong | Notion used it to pressure-test systems before a major release. |
| New MCP servers and APIs | Strong | Found two chained high-severity issues in a Carta MCP server before launch. |
| M&A security reviews | Strong | Carta also used it for a review during an acquisition. |
| Turning findings into detections | Mixed | Carta feeds findings into its detection patterns; this takes in-house work. |
| Self-serve testing | Weak | Engagements are scoped with the team. |
| Budget compliance pentests | Weak | Priced like a top pentest firm, not a scanner. |
Who it’s for
Good fit
- AI startups shipping new features weekly
- Fintech and SaaS companies before launches
- Security teams reviewing acquisitions
Poor fit
- Teams that want a self-serve scanner
- Buyers who need a public price
- Organizations with only network targets
Review Excerpts
Below are excerpts from public reviews. Our team scoured public reviews, forums, and chat rooms to get a balanced view of customers' experience with this company. Paid reviews and pay-for-play sites such as Clutch were excluded.
“RunSybil was an excellent partner for us. They pressure-tested our systems ahead of a major release and delivered fast, high-quality results at a competitive price on par with top pen-testing firms.”
“We wanted a world-class partner for Turbopuffer’s ongoing pentesting needs, and we couldn’t be happier with our relationship with RunSybil. Quick turnaround, attention to detail, and fun to work with…”
“RunSybil’s expertise was instrumental in enhancing our security posture, providing us with critical insights for a confident launch.”
“There's no price on the website, so we had to get on a call and scope the whole test before we knew if it fit our budget. For a small security team, that added a couple of weeks.”
Methodology
This page is an independent evaluation of RunSybil for buyers comparing options in ai penetration testing. AI Industry Reviews accepts no sponsorships, advertising, or pay-for-placement fees. RunSybil did not pay for this review.
What we scored
The headline number is an Overall Score on a 0-10 scale. Eight criteria sit under it in two groups.
- Buyer outcomes (for ai penetration testing)
- Depth on modern app stacks
- Speed before releases
- Findings into detections
- Pricing transparency
- Company & commercial
- Innovation & product leadership
- Project management & communication
- Pricing
- Contract fairness
Pricing measures whether the price looks fair for the value delivered, including packaging and renewal friction that show up in real buying cycles.
Score Composition
| Input | Weight | What it covers |
|---|---|---|
| Reviews | 40% | A proprietary read of what practitioners say about likes, complaints, and day-to-day use, including public review sites, forums, and private chat rooms. Paid reviews and pay-for-play sites such as Clutch are out of scope. |
| Product | 35% | Hands-on look at screens and workflows. |
| Pricing | 15% | Whether the price looks fair for what you get. |
| Docs & training | 10% | Docs, tutorials, and training material. |
How we balanced the evidence
The Overall Score is the simple average of the eight criteria. Recommendation language follows that score and the fit pattern described above.
Scope
RunSybil is graded here as ai penetration testing. Criteria scores can move as more review volume and product checks are added.
