RAGFlow
An open-source retrieval-augmented generation engine that combines document understanding with agent capabilities to give LLMs grounded context.
- GitHub stars
- 92k
- Last commit
- today
- Latest release
- v1.0.0-rc1
- Licence
- Apache-2.0
- Self-hosted
- Yes
- Hosted version
- Available

RAGFlow is an open-source engine for retrieval-augmented generation (RAG). It combines RAG with agent capabilities to give large language models a context layer built from your own documents, so applications can answer from real data. It is released under the Apache-2.0 license and, per the repository, is currently at a 1.0.0 release candidate.
It extracts knowledge from unstructured documents with complicated formats using deep document understanding, and offers pre-built agent templates. Recent updates mentioned in the README include website ingestion through sitemaps, Google BigQuery data sources with incremental sync, knowledge compilation that generates wikis, graphs, trees and mind maps, and agentic RAG with multiple thinking modes.
You can try the managed cloud service or deploy it locally by following the local deployment guide. The project provides documentation, a roadmap and a Discord community. It suits developers and enterprises building document question answering and knowledge-based assistants who want control over their data pipeline.
Key features
- Deep document understanding for unstructured files
- Retrieval-augmented generation workflows
- Pre-built agent templates
- Agentic RAG with adjustable thinking modes
- Website ingestion through sitemaps
- Knowledge compilation into wikis and graphs
- Google BigQuery data source sync
Pricing: Free plan with 500 credits per month. Starter shows $59 and Pro $259 per month (a lower unlabeled price of $29 and $129 likely applies with annual billing); Enterprise is quoted and can be deployed on-premises.



