About MyScaleDB
MyScaleDB is a database that adds high-performance vector search and full-text search to ClickHouse, so developers can build retrieval-augmented generation and other AI applications using ordinary SQL. It is a fork of ClickHouse rather than a separate service layered on top, and it keeps the OLAP architecture that makes ClickHouse suited to large datasets.
Because it speaks SQL, queries can combine vector similarity, metadata filters and joins across tables, which the project says improves RAG accuracy compared with filtering after retrieval. It manages structured data, text, vectors, JSON, geospatial and time-series data in one system, and the README stresses scalability as data grows.
The source code is released under Apache-2.0 and can be self-hosted. A managed offering, MyScale Cloud, runs MyScaleDB with extra premium features at billion-scale and is positioned against specialized vector databases that use custom APIs. It suits teams already comfortable with SQL who want one system for vectors and analytics.
Key features
- SQL interface with vector-related functions
- Vector search with metadata filtering
- Full-text search alongside vectors
- SQL-vector join queries
- Built on the ClickHouse OLAP architecture
- Supports JSON, geospatial and time-series data
Good fit for
- →Retrieval-augmented generation backends
- →Image and similarity search
- →Combining analytics with vector search in SQL
- Built with
- C++
- Tags
- vector-database
- sql
- clickhouse
- vector-search
- full-text-search
- rag
- ai
- similarity-search
MyScaleDB: questions and answers
- What is MyScaleDB used for?
- MyScaleDB is an open-source SQL vector database built on ClickHouse, combining vector search, full-text search and analytics for AI applications. It is a good fit for retrieval-augmented generation backends, image and similarity search, and combining analytics with vector search in SQL.
- Is MyScaleDB open source?
- Yes. MyScaleDB is open source under the Apache-2.0 licence. Its source code is on GitHub at myscale/MyScaleDB and is written mainly in C++.
- Is MyScaleDB free?
- Yes. MyScaleDB is open source, so the software itself is free to use. A managed cloud version is also available.
- Can I self-host MyScaleDB?
- Yes. MyScaleDB can be self-hosted on your own server or infrastructure.
- What is MyScaleDB an alternative to?
- MyScaleDB is an open-source alternative to Pinecone. Other open-source alternatives to Pinecone include Milvus, Qdrant and Chroma.
- Is MyScaleDB actively maintained?
- The most recent commit to MyScaleDB was on 5 February 2025. The project has 1k stars on GitHub.
Open-source alternatives to MyScaleDB
See all
Milvus
Databases
Milvus is a high-performance, cloud-native vector database built for scalable vector ANN s
Apache-2.0vs Pinecone★ 46k
Qdrant
Databases
Qdrant - High-performance, massive-scale Vector Database and Vector Search Engine for the
Apache-2.0vs Pinecone★ 35k
Chroma
Databases
Search infrastructure for AI
Apache-2.0vs Pinecone★ 29k
Weaviate
Databases
Weaviate is an open-source vector database that stores both objects and vectors, allowing
OSSvs Pinecone★ 17k
Infinity
Databases
The AI-native database built for LLM applications, providing incredibly fast hybrid search
Apache-2.0vs Pinecone★ 4.7k
Epsilla VectorDB
Databases
A high performance Vector Database leveraging parallel graph computing
GPL-3.0vs Pinecone★ 875
SaaS alternatives to MyScaleDB
See all
Pinecone
Databases
Managed vector database for semantic search and retrieval apps
SaaS
DataStax Astra DB
Databases
Cassandra-based serverless database with vector search for AI applications
SaaS
Upstash
Databases
Serverless Redis, Kafka and vector database with per-request pricing
SaaS
Exasol
Databases
In-memory analytics database for fast SQL queries on large datasets
SaaS
Tinybird
Databases
Real-time analytics database that turns SQL queries into API endpoints
SaaS
Aerospike
Databases
Real-time NoSQL database for low-latency, high-throughput applications
SaaS

