6,598 open-source and SaaS tools, with GitHub stats refreshed every day.

Promptfoo

Open source

Open-source CLI and library for evaluating prompts, agents, and RAG apps, and for red-teaming LLM applications to find security weaknesses.

Open-source alternative to

promptfoo.dev
Promptfoo homepage screenshot
GitHub stars
26k
Last commit
today
Repository age
3 years
Version
0.123.1
Licence
MIT
Self-hosted
Yes

About Promptfoo

Promptfoo is an open-source command-line tool and library for testing applications built on large language models. Developers describe test cases and checks in declarative configuration files, run them against one or more models, and review the results, which replaces manual trial and error with repeatable evaluations. It is written in TypeScript and released under the MIT license.

Beyond quality testing, Promptfoo includes red teaming and vulnerability scanning for AI systems, probing prompts, agents, and retrieval-augmented generation (RAG) pipelines for security and safety weaknesses. It can compare models from providers such as OpenAI, Anthropic, Azure, Bedrock, and Ollama side by side, and it works with any LLM API or programming language. Its documentation also covers a code-scanning feature that reviews pull requests for LLM-related security and compliance issues.

The project is designed to run locally: the maintainers state that evaluations execute on the developer's own machine, so prompts only leave it when a chosen model provider receives them. It can be installed with Homebrew or pip, or run without installing via npx, and it plugs into CI/CD so checks run automatically on each change. The README notes that Promptfoo is now part of OpenAI while remaining open source and MIT licensed.

Key features

  • Declarative test configs for prompts and models
  • Side-by-side comparison of multiple LLM providers
  • Red teaming and vulnerability scanning for AI apps
  • Evaluations run locally on your own machine
  • CI/CD integration for automated checks
  • Pull request code scanning for LLM risks
  • Live reload and result caching
  • Shareable results and security reports

Good fit for

  • →Regression testing prompts before release
  • →Comparing models for a specific task
  • →Security testing of chatbots and agents
  • →Gating LLM changes in CI pipelines
Built with
TypeScript
Tags
llm
evaluation
red-teaming
prompt-testing
rag
ci-cd
ai-security
cli
open-source

Open-source alternatives to Promptfoo

See all

SaaS alternatives to Promptfoo

See all