Home / Directory / AI SaaS tooling / Prompt / app frameworks / Promptfoo
Promptfoo
Open-source and commercial framework for LLM prompt evaluation, red teaming, and CI security testing of AI applications.
Engineering teams that want local-first prompt evals and red-team probes in CI before they buy a hosted collaboration layer.
Non-technical prompt editors who need a hosted CMS without a CLI, or buyers who only want a production AI gateway.
Verdict
Promptfoo fits teams that treat prompt and agent quality as an engineering test problem. Community is free with full eval features and a 10k probes/month red-team allowance; Enterprise and On-Prem are custom for SSO, shared dashboards, continuous monitoring, and higher probe limits. Practitioners like the CI-native workflow; collaboration and managed cloud only arrive on paid plans.
Score Breakdown
How Promptfoo scores in the categories that matter to its buyers.
Pricing
Promptfoo publishes Community, Enterprise, and On-Premise plans on promptfoo.dev/pricing. Community is free; paid plans are custom quotes. As of September 2026.
Model: open-core. Community covers local/self-hosted evals, providers, vulnerability scanning, and up to 10k red-team probes per month. Enterprise adds team sharing, continuous monitoring, SSO/RBAC, API access, managed cloud, and extra probes. On-Premise adds dedicated runners and full data isolation. Model inference used during tests bills through your LLM providers separately.
The Field at a Glance
Where Promptfoo ranks among Prompt / app frameworks vendors we reviewed, by Overall Score and relative typical engagement cost.
Promptfoo scores 6.7 beside Portkey (7.9) and PromptLayer (6.9). Relative cost stays low for Community-led teams; Enterprise quotes move it toward the mid band.
Use-case matrix
| Use case | Fit | Notes |
|---|---|---|
| CI prompt and agent evals | Strong | Core open-source job. |
| LLM red teaming | Strong | Probe-metered; 10k/mo on Community. |
| Hosted prompt CMS for editors | Weak | PromptLayer-shaped. |
| Production AI gateway | Poor | Portkey-shaped. |
Who it’s for
Good fit
- Security and platform eng running release gates
- Teams standardizing prompt regression suites
- Buyers who want OSS first, commercial later
Poor fit
- Non-engineer prompt editors without CLI comfort
- Gateway-only procurement
- Orgs needing only a logging desk
Review Excerpts
Below are excerpts from public reviews. Paid reviews and pay-for-play sites such as Clutch were excluded.
“Promptfoo is an open-source CLI tool that can help evaluate and red-team LLM apps, and it can be used in CI/CD pipelines to catch regressions and vulnerabilities before they ship.”
“Promptfoo Community includes all the LLM evaluation features, all model providers and integrations, red teaming up to 10k probes a month, and local or self-hosted runs. We started on Promptfoo without a paid tier.”
“Enterprise pricing is customized based on your team's size and needs. Contact us for a personalized quote.”
“Certain red teaming plugins require inference for dynamic test generation and grading, so probe usage and provider spend scale with how aggressively you scan.”
Methodology
This page is an independent evaluation of Promptfoo for buyers comparing options in prompt / app frameworks. AI Industry Reviews accepts no sponsorships, advertising, or pay-for-placement fees. Promptfoo did not pay for this review.
What we scored
The headline number is an Overall Score on a 0-10 scale. Eight criteria fall under it in two groups.
- Buyer outcomes (for prompt / app frameworks)
- Prompt / agent eval harness
- Red teaming & probes
- CI / local-first workflow
- Team collaboration features
- Company & commercial
- Innovation & product leadership
- Project management & communication
- Pricing
- Contract fairness
Pricing measures whether the price looks fair for the value delivered, including packaging and renewal friction that show up in real buying cycles.
Score Composition
| Input | Weight | What it covers |
|---|---|---|
| Reviews | 40% | A proprietary read of what practitioners say about likes, complaints, and day-to-day use, including public review sites, forums, and private chat rooms. Paid reviews and pay-for-play sites such as Clutch are out of scope. |
| Product | 35% | Hands-on look at screens and workflows. |
| Pricing | 15% | Whether the price looks fair for what you get. |
| Docs & training | 10% | Docs, tutorials, and training material. |
How we balanced the evidence
The Overall Score is the simple average of the eight criteria. Recommendation language follows that score and the fit pattern described above.
Scope
Promptfoo is graded here as prompt / app frameworks. Criteria scores can move as more review volume and product checks are added.
