About Promptfoo
Promptfoo is an open-source command-line tool and library for testing applications built on large language models. Developers describe test cases and checks in declarative configuration files, run them against one or more models, and review the results, which replaces manual trial and error with repeatable evaluations. It is written in TypeScript and released under the MIT license.
Beyond quality testing, Promptfoo includes red teaming and vulnerability scanning for AI systems, probing prompts, agents, and retrieval-augmented generation (RAG) pipelines for security and safety weaknesses. It can compare models from providers such as OpenAI, Anthropic, Azure, Bedrock, and Ollama side by side, and it works with any LLM API or programming language. Its documentation also covers a code-scanning feature that reviews pull requests for LLM-related security and compliance issues.
The project is designed to run locally: the maintainers state that evaluations execute on the developer's own machine, so prompts only leave it when a chosen model provider receives them. It can be installed with Homebrew or pip, or run without installing via npx, and it plugs into CI/CD so checks run automatically on each change. The README notes that Promptfoo is now part of OpenAI while remaining open source and MIT licensed.
Key features
- Declarative test configs for prompts and models
- Side-by-side comparison of multiple LLM providers
- Red teaming and vulnerability scanning for AI apps
- Evaluations run locally on your own machine
- CI/CD integration for automated checks
- Pull request code scanning for LLM risks
- Live reload and result caching
- Shareable results and security reports
Good fit for
- →Regression testing prompts before release
- →Comparing models for a specific task
- →Security testing of chatbots and agents
- →Gating LLM changes in CI pipelines
- Built with
- TypeScript
- Tags
- llm
- evaluation
- red-teaming
- prompt-testing
- rag
- ci-cd
- ai-security
- cli
- open-source
Open-source alternatives to Promptfoo
See all
Langfuse
AI Infrastructure
Trace, evaluate, and improve AI agents with one open platform. Use production data to unde
OSSvs Weights & Biases★ 35k
Opik
AI Infrastructure
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows wit
Apache-2.0vs Weights & Biases★ 22k
Arize Phoenix
AI Infrastructure
AI Observability & Evaluation
OSSvs Weights & Biases★ 12k
LangWatch
AI Infrastructure
The platform for LLM evaluations and AI agent testing
Apache-2.0vs LangSmith★ 4.9k
Laminar
AI Infrastructure
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
Apache-2.0vs Weights & Biases★ 3.3k
UpTrain
AI Infrastructure
UpTrain is an open-source unified platform to evaluate and improve Generative AI applicati
Apache-2.0vs Weights & Biases★ 2.4k
SaaS alternatives to Promptfoo
See all
Braintrust
AI Infrastructure
Evaluation and observability platform for building and testing AI applications
SaaS
Galileo
AI Infrastructure
Evaluation and monitoring platform for generative AI applications
SaaS
Vellum
AI Infrastructure
Platform for building, testing and monitoring LLM prompts and workflows
SaaS
LangSmith
AI Infrastructure
Observability, tracing and evaluation platform for LLM applications
SaaS
Arize AI
AI Infrastructure
Observability and evaluation platform for machine learning models and LLM applications
SaaS
Portkey
AI Infrastructure
AI gateway and observability platform for managing LLM requests
SaaS

