7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

10 alternatives ranked by real activity

Open-source Google BigQuery alternatives

A curated, ranked list of the 10 best open-source alternatives to Google BigQuery.

The best open-source alternative to Google BigQuery is ClickHouse. If that doesn't suit you, other good options are Apache Doris, StarRocks, Databend and Apache Hive.

Google BigQuery alternatives are mainly databases, but some are also data pipeline & ETL tools. 6 of them shipped code in the last 30 days, 10 can be self-hosted, and 8 use a permissive licence.

Last updated October 2, 2026 · ranked by GitHub stars, growth and recent commits

ClickHouse

An open-source column-oriented database management system built for real-time analytical reports, available to self-host or as the managed ClickHouse Cloud service.

GitHub stars
50k
Last commit
today
Latest release
v26.9.8.3-stable
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
clickhouse.comClickHouse homepage screenshot

ClickHouse is an open-source database management system built to generate analytical reports in real time. Its column-oriented storage makes scans and aggregations over very large tables efficient, which is why it is widely used for event data, logs, metrics and product analytics. It is written in C++ and queried with SQL.

The project's topics describe it as an OLAP, massively parallel and distributed system, and it is used for big data workloads where query latency matters. You can install it on Linux, macOS or FreeBSD, follow a tutorial that sets up a small cluster, or skip installation by trying ClickHouse Cloud, a managed service built by its creators. Releases arrive monthly, with a community call for each one, and there are Slack and Telegram channels for help.

ClickHouse is Apache-2.0 licensed. It is a good fit for analytics engineers and platform teams who need fast aggregations over billions of rows and are comfortable operating a database, or who prefer to hand that operation to the vendor's cloud.

Key features

  • Column-oriented storage for analytics
  • SQL queries over large datasets
  • Distributed, massively parallel processing
  • Real-time analytical reporting
  • Install on Linux, macOS or FreeBSD
  • Managed ClickHouse Cloud option

Pricing: The open-source distribution is free. ClickHouse Cloud is usage-based, with plans from $53 (Basic), $437 (Scale) and $571 (Enterprise) per month and a 30-day trial with $300 in credits.

Read more about ClickHouseWebsite GitHub

Apache Doris

Apache Doris is an open-source, MPP real-time analytics and hybrid search database for fast SQL, lakehouse query acceleration and vector and text search.

GitHub stars
16k
Last commit
today
Latest release
4.1.4.1
Licence
Apache-2.0
Self-hosted
Yes
doris.apache.orgApache Doris homepage screenshot

Apache Doris is an open-source analytics database built on a massively parallel processing architecture. It offers fast SQL analytics, acceleration of queries over lakehouse data, and hybrid search spanning structured, text and vector data, and it presents itself as a real-time analytics and search engine suitable for AI agents. It is written in Java.

The README lists use cases that include customer-facing analytics with sub-second interactive queries for external users, data warehousing across business domains, observability for high-throughput logs, events and metrics analyzed with SQL, and AI use where vector, text, JSON and structured search run in one SQL engine. Topics mention lakehouse formats Hudi, Iceberg and Delta Lake and comparisons with warehouses such as BigQuery, Redshift and Snowflake.

Doris is a project of the Apache Software Foundation and Apache-2.0 licensed. You can deploy it on your own clusters, and the website provides release notes, use cases and user stories, with documentation in a large number of languages. It suits data engineering teams that want an open-source, high-concurrency analytical database.

Key features

  • MPP architecture for fast SQL analytics
  • Lakehouse query acceleration
  • Hybrid search across structured, text and vector data
  • Real-time ingestion and analysis
  • Support for Iceberg, Hudi and Delta Lake
  • Apache Software Foundation project

Pricing: Free and open source under the Apache-2.0 license.

Read more about Apache DorisWebsite GitHub

StarRocks

A Linux Foundation project: an analytical SQL engine for real-time and ad-hoc queries, running on its own storage or directly over data lakehouse tables.

GitHub stars
12k
Last commit
today
Latest release
4.1.3
Licence
Apache-2.0
Self-hosted
Yes
starrocks.ioStarRocks homepage screenshot

StarRocks is a query engine for analytics that returns answers quickly to multi-dimensional, real-time and ad-hoc queries. It is a Linux Foundation project written in Java, and can be used both on its own tables and over data that sits in a data lake or lakehouse without first moving it.

It uses a vectorized SQL engine that takes advantage of CPU parallelism, supports standard ANSI SQL and the MySQL protocol so existing clients and BI tools can connect, and applies a cost-based optimizer to complex queries. Primary-key tables support upserts and deletes with efficient querying during concurrent updates, and materialized views refresh during data import and are chosen automatically at query time.

Data in Hive, Iceberg, Delta Lake and Hudi tables can be queried in place, and its topics point to star-schema, MPP and distributed-database designs. The software is licensed under Apache-2.0 and can be downloaded and run yourself, with documentation, benchmarks and a demo linked from the README. It suits data teams that want fast dashboards and lakehouse analytics without heavy denormalization.

Key features

  • Vectorized SQL engine for fast analytics
  • ANSI SQL with MySQL protocol compatibility
  • Cost-based query optimizer
  • Real-time upserts and deletes by primary key
  • Automatically maintained materialized views
  • Direct queries over Hive, Iceberg, Delta Lake and Hudi

Pricing: Free and open source under the Apache-2.0 licence.

Read more about StarRocksWebsite GitHub

Databend

An open-source cloud data warehouse written in Rust that unifies analytics, vector search and full-text search on object storage, with sandboxed UDFs for AI agents.

GitHub stars
9.5k
Last commit
today
Latest release
v1.2.881
Self-hosted
Yes
Hosted version
Available
docs.databend.comDatabend homepage screenshot

Databend is an enterprise-grade, open-source data warehouse written in Rust. It combines large-scale analytics, vector search, full-text search and automatic schema evolution in one engine, and keeps data in object storage such as S3, Azure Blob Storage or Google Cloud Storage with elastic, cloud-native compute.

The project now positions itself as a warehouse that is ready for AI agents. Sandboxed Python user-defined functions run agent logic, SQL handles orchestration, transactions provide reliability, and Git-like branching lets agents experiment safely on production snapshots. A three-layer architecture separates the control plane, the execution plane and the sandbox workers. Typical use cases listed include AI agents, analytics and BI, and search and RAG.

You can use the managed Databend Cloud, run it locally from Python for development and testing, or start the full warehouse in Docker. The repository lists the licence as Other, so review the licence file before commercial use. Topics reference Snowflake and Elasticsearch as comparable products.

Key features

  • Large-scale SQL analytics
  • Vector and full-text search
  • Automatic schema evolution
  • Sandboxed Python UDFs for agents
  • Git-like data branching
  • Object storage on S3, Azure or GCS
  • Cloud, Docker and local Python options

Pricing: Databend Cloud offers a free start; the repository lists the license as Other.

Read more about DatabendWebsite GitHub

Apache Hive

Apache Hive is data warehouse software that lets you read, write and manage large datasets in distributed storage using SQL on top of Hadoop.

GitHub stars
6k
Last commit
yesterday
Licence
Apache-2.0
Self-hosted
Yes
hive.apache.orgApache Hive homepage screenshot

Apache Hive is data warehouse software built on Apache Hadoop. It makes it possible to read, write and manage very large datasets in distributed storage using SQL, so analysts can run extract-transform-load jobs, reports and data analysis without writing low-level processing code.

Hive provides tools for accessing data through SQL, a way of imposing structure on a variety of data formats, and access to files stored in Apache HDFS or other systems such as Apache HBase. Queries can run on the Apache Tez framework, which is designed for interactive queries and has far lower overhead than MapReduce. It supports much standard SQL functionality, including many analytics features from the 2003 and 2011 SQL standards.

Hive is an Apache Software Foundation project written in Java and licensed under Apache-2.0. You deploy it on your own Hadoop-based clusters, though hosted big data platforms from cloud providers also bundle it. It suits data engineering teams maintaining batch analytics on large distributed data.

Key features

  • SQL access to distributed datasets
  • Built on Apache Hadoop
  • Reads data from HDFS and HBase
  • Query execution on Apache Tez
  • ETL, reporting and analysis workloads
  • Standard SQL analytics functions

Pricing: Free and open source under the Apache-2.0 license.

Read more about Apache HiveWebsite GitHub

Apache Cloudberry

Apache Cloudberry is an open-source MPP database based on PostgreSQL for data warehousing and large-scale analytics, positioned as a Greenplum alternative.

GitHub stars
1.4k
Last commit
today
Latest release
2.1.0-incubating
Licence
Apache-2.0
Self-hosted
Yes
cloudberry.apache.orgApache Cloudberry homepage screenshot

Apache Cloudberry, currently in the Apache incubator, is an open-source massively parallel processing database. It was created by original developers of Greenplum Database and evolves from the open-source version of Pivotal Greenplum, but uses a newer PostgreSQL kernel and adds enterprise capabilities. It is described as an open-source alternative to Greenplum Database.

It can serve as a data warehouse and is also intended for large-scale analytics and AI or machine learning workloads, distributing queries across many nodes. Because it is based on PostgreSQL, it uses familiar SQL. The main repository sits alongside ecosystem repositories for the website, a backup utility, Go libraries and the Platform Extension Framework (PXF) for connecting to external data.

Cloudberry is written in C and licensed under Apache-2.0. You can build it from source on Linux, including RHEL, Rocky Linux and Ubuntu, and macOS, or try it quickly through a Docker-based sandbox. Community help is available through Slack and GitHub Discussions. It suits data engineering teams running large analytical workloads that want an open-source MPP engine they can operate themselves.

Key features

  • Massively parallel processing SQL engine
  • PostgreSQL-based kernel
  • Data warehouse and analytics workloads
  • Backup utility and PXF connectors
  • Docker sandbox for trying it out
  • Builds on Linux and macOS

Pricing: Free and open source under the Apache-2.0 licence.

GlareDB

Lightweight SQL database written in Rust for running analytics quickly across a variety of data sources.

GitHub stars
1k
Last commit
10 mo ago
Latest release
v25.6.3
Licence
MIT
Self-hosted
Yes
glaredb.comGlareDB homepage screenshot

GlareDB is a light and fast SQL database aimed at analytics. The project describes it as a lightweight, ergonomic database for quickly running analytics across a variety of sources, and it is written in Rust. It is meant for developers and analysts who want to query data with SQL without setting up a heavy data warehouse.

Installation is a single command from the project's site, and documentation, a blog and a changelog are available there. GlareDB uses calendar versioning in a year, month and patch format, so version 25.5.0 was the first release in May 2025. The code is licensed under MIT, and the company behind it is GlareDB, Inc. The README itself is brief, so see the documentation for supported data sources and query features.

Key features

  • SQL database focused on analytics
  • Queries across a variety of data sources
  • Written in Rust
  • Single-command installation
  • Calendar-versioned releases

Pricing: Free and open source under the MIT license.

Read more about GlareDBWebsite GitHub

Dremio OSS

Dremio OSS is a free, Apache-licensed, open-source data lakehouse platform in Java that provides fast SQL analytics directly on data stored in warehouses and lakes.

GitHub stars
1.5k
Last commit
1 yr ago
Licence
Apache-2.0
Self-hosted
Yes
dremio.comDremio OSS homepage screenshot

Dremio is a data analytics platform aimed at organizations that want to query and analyze large volumes of data without first moving or copying it into a separate, dedicated analytics database. It positions itself as filling a gap between raw data storage and the tools people need to get value out of that data.

The open-source edition can be built and run locally, exposing a web UI at localhost:9047 once started, or installed as a production tarball for server deployment, with an embedded mode also available. Building it requires a specific combination of JDK versions (21 as default, with 17 and 11 configured in the Maven toolchain for certain tests) plus Maven, reflecting a substantial, enterprise-grade Java codebase.

The README notes that to provide the best possible experience, the build includes some dependencies distributed under non-open-source licenses, so the fully open-source experience may differ slightly depending on build configuration. Dremio is written in Java and the open-source edition is released under the Apache-2.0 license, with full documentation hosted separately.

Key features

  • SQL analytics directly on lake and warehouse data
  • Web-based query interface
  • Local and production deployment modes
  • Built on a Java codebase with Maven tooling
  • No data duplication required for analytics

Pricing: Dremio Cloud is pay-as-you-go at $0.20 per compute unit, with a 30-day trial that includes $400 in credit. Dremio Enterprise, which can be self-hosted, is priced through sales.

Read more about Dremio OSSWebsite GitHub

TensorBase

TensorBase was a free, Apache-licensed, open-source big-data warehouse written in Rust aiming for ClickHouse compatibility with faster write and query performance; active development has paused.

GitHub stars
1.5k
Last commit
4 yr ago
Latest release
v2021.07.05
Licence
Apache-2.0
Self-hosted
Yes
tensorbase.ioTensorBase homepage screenshot

TensorBase was a modern big-data warehouse project written in Rust, aimed at engineers who wanted a high-performance analytical database inspired by, and aiming to be compatible with, ClickHouse. It explored several firsts in database engineering, including a no-LSM, write-and-read optimized storage layer and running on real-world RISC-V hardware.

The project reported faster write throughput and faster simple aggregation query speed than ClickHouse in its own benchmarks, alongside experiments with a 'copy-free, lock-free, async-free, dyn-free' design in its critical path and an early prototype of a whole-lifecycle JIT SQL query engine, some of which was shared only through blog posts, presentations and videos rather than released in the open-source repository.

The README states plainly that the project has paused in the general data-warehousing space, explaining that the maintainers did not want open source to become a copy-and-fork game and chose to step back rather than compete on GitHub star counts. They recommend ClickHouse for anyone seeking a production-ready data warehouse today, while TensorBase remains available for people who want to study database internals or high-performance Rust techniques. It is released under the Apache-2.0 license.

Key features

  • ClickHouse-compatible query interface
  • No-LSM, write-and-read optimized storage layer
  • RISC-V hardware support
  • Experimental JIT SQL query engine prototype
  • High-performance Rust implementation

Pricing: Free and open source under the Apache-2.0 license. General development has paused; the maintainers recommend ClickHouse for production use.

Read more about TensorBaseWebsite GitHub

EventQL

EventQL is a free, open-source, distributed columnar database and massively parallel SQL query engine written in C++, built for large-scale event collection and analytics.

GitHub stars
1.2k
Last commit
9 yr ago
Latest release
v0.4.1
Self-hosted
Yes

EventQL is a distributed database built specifically for large-scale data collection and analytics workloads, aimed at teams that need to ingest very high volumes of streaming data while still running fast SQL and MapReduce-style queries over it. It is designed as a columnar, massively parallel (MPP) system rather than a general-purpose transactional database.

Tables are automatically partitioned by primary key and distributed across machines without manual shard configuration, and the system supports idempotent primary-key-based INSERT, UPSERT and DELETE operations, with UPSERT suited to exactly-once ingestion from streaming sources. Its compact columnar storage engine drastically reduces I/O for analytical workloads compared with row-oriented systems, and it supports nearly complete SQL 2009, including joins, with queries automatically parallelized across many machines. It claims to scale to petabytes across equally privileged servers and supports streaming, low-latency writes, with mutations immediately visible and minimal query latency as low as 0.1 milliseconds.

EventQL natively supports both time-series and relational data models through the same automatic partitioning mechanism. It is written in C++11, and the repository metadata lists its license as Other, so check the license file for exact terms before adopting it, particularly given the project's age and apparently limited recent activity.

Key features

  • Automatic table partitioning across machines
  • Columnar storage for fast analytical queries
  • Idempotent UPSERT for streaming ingestion
  • Near-complete SQL 2009 support including joins
  • Petabyte-scale distributed architecture
  • Sub-millisecond query latency on fresh writes

Pricing: Free and open source. The repository lists its license as Other, so check the license file for exact terms.

Read more about EventQLWebsite GitHub

Google BigQuery alternatives: questions

What is the best open-source alternative to Google BigQuery?
ClickHouse is the top-ranked open-source alternative to Google BigQuery on Enlisted: An open-source column-oriented database management system built for real-time analytical reports, available to self-host or as the managed ClickHouse Cloud service. Other strong options are Apache Doris, StarRocks, Databend and Apache Hive.
Are these Google BigQuery alternatives free?
All 10 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer. 2 also offer a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Google BigQuery alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 6 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all