7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

6 alternatives ranked by real activity

Open-source Databricks alternatives

A curated, ranked list of the 6 best open-source alternatives to Databricks.

The best open-source alternative to Databricks is Apache Spark. If that doesn't suit you, other good options are Apache Zeppelin, Apache Cloudberry, Pachyderm and Dremio OSS.

Databricks alternatives are mainly data pipeline & ETL tools, but some are also databases and BI & dashboard tools. 3 of them shipped code in the last 30 days, 6 can be self-hosted, and 5 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Apache Spark

An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming.

GitHub stars
44k
Last commit
today
Licence
Apache-2.0
Self-hosted
Yes
spark.apache.orgApache Spark homepage screenshot

Apache Spark is an analytics engine that brings batch, SQL, machine learning, graph and streaming workloads together for processing data at scale. It is maintained by the Apache Software Foundation under the Apache-2.0 license. It is written mainly in Scala and offers high-level APIs in Scala, Java and Python, with an R API that is now deprecated, on top of an engine optimized to execute general computation graphs.

A set of higher-level tools is included: Spark SQL for queries and DataFrames, MLlib for machine learning, GraphX for graphs, Structured Streaming for streams, and a pandas-style API for existing pandas workloads. Together they let teams run batch analytics, machine learning and streaming jobs on a single engine and share code between them.

The README points to the project website for full documentation and programming guides, and it contains only basic setup and build instructions. Spark suits data engineers and data scientists who process datasets too large for a single machine, and organizations that want one framework for ETL, analytics and machine learning on clusters.

Key features

  • High-level APIs in Scala, Java and Python
  • Optimized engine for computation graphs
  • Spark SQL and DataFrames
  • MLlib for machine learning
  • GraphX for graph analytics
  • Structured Streaming for stream processing
  • Pandas API on Spark

Pricing: Free and open source under the Apache-2.0 license.

Apache Zeppelin

Apache Zeppelin is a web-based notebook for interactive data analytics, with real-time collaboration, many language interpreters, visualizations and scheduling.

GitHub stars
6.7k
Last commit
yesterday
Licence
Apache-2.0
Self-hosted
Yes
zeppelin.apache.orgApache Zeppelin homepage screenshot

Apache Zeppelin is a web-based notebook for interactive data analytics. You build data-driven, collaborative documents using SQL, Scala and other languages, and the notebook-style editor supports real-time collaboration between users.

Core features include support for many languages through a pluggable interpreter architecture with process isolation, including Spark, Flink, Python, SQL and Shell and more than 20 interpreters, built-in visualization and dynamic forms, and notebook scheduling with cron. Flexible deployment options are local, Docker, Kubernetes and YARN. Topics place it in the big data ecosystem alongside Spark, Flink, NoSQL and databases.

Zeppelin is an Apache Software Foundation project written in Java and JavaScript and licensed under Apache-2.0. You install it from a binary package or build from source, with user guides, mailing lists and a Jira issue tracker. It suits data engineers and analysts who work on cluster-based data platforms and want shareable, scheduled notebooks.

Key features

  • Web notebook with real-time collaboration
  • 20+ interpreters including Spark and Flink
  • Process isolation per interpreter
  • Built-in visualizations and dynamic forms
  • Cron scheduling of notebooks
  • Local, Docker, Kubernetes and YARN deployment

Pricing: Free and open source under the Apache-2.0 licence.

Apache Cloudberry

Apache Cloudberry is an open-source MPP database based on PostgreSQL for data warehousing and large-scale analytics, positioned as a Greenplum alternative.

GitHub stars
1.4k
Last commit
today
Latest release
2.1.0-incubating
Licence
Apache-2.0
Self-hosted
Yes
cloudberry.apache.orgApache Cloudberry homepage screenshot

Apache Cloudberry, currently in the Apache incubator, is an open-source massively parallel processing database. It was created by original developers of Greenplum Database and evolves from the open-source version of Pivotal Greenplum, but uses a newer PostgreSQL kernel and adds enterprise capabilities. It is described as an open-source alternative to Greenplum Database.

It can serve as a data warehouse and is also intended for large-scale analytics and AI or machine learning workloads, distributing queries across many nodes. Because it is based on PostgreSQL, it uses familiar SQL. The main repository sits alongside ecosystem repositories for the website, a backup utility, Go libraries and the Platform Extension Framework (PXF) for connecting to external data.

Cloudberry is written in C and licensed under Apache-2.0. You can build it from source on Linux, including RHEL, Rocky Linux and Ubuntu, and macOS, or try it quickly through a Docker-based sandbox. Community help is available through Slack and GitHub Discussions. It suits data engineering teams running large analytical workloads that want an open-source MPP engine they can operate themselves.

Key features

  • Massively parallel processing SQL engine
  • PostgreSQL-based kernel
  • Data warehouse and analytics workloads
  • Backup utility and PXF connectors
  • Docker sandbox for trying it out
  • Builds on Linux and macOS

Pricing: Free and open source under the Apache-2.0 licence.

Pachyderm

Pachyderm is an Apache-licensed, Kubernetes-based platform for data pipelines that adds automatic versioning and lineage tracking to language-agnostic data transformations.

GitHub stars
6.3k
Last commit
1 yr ago
Latest release
v2.12.2
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available

Pachyderm is a data pipeline platform built for data engineering teams that need to automate complex transformations across large and varied datasets. It positions itself as a CI/CD engine for data, combining pipeline automation with built-in version control and lineage so that every change to data and every pipeline run can be traced.

Pipelines trigger automatically when Pachyderm detects a change in input data, and each stage can be written in any language since Pachyderm treats containers as the unit of processing. It tracks immutable data lineage and versioning for any data type, scales and parallelizes processing automatically on top of Kubernetes, and stores data in standard object stores with automatic deduplication. It runs on all major cloud providers as well as on-premises.

Pachyderm is written in Go and released under the Apache-2.0 license. It can be run locally for evaluation or deployed on AWS, GCE or Azure in a short setup, and the project provides tutorials, example projects and case studies through its documentation.

Key features

  • Automatic pipeline triggers on data changes
  • Immutable data versioning and lineage tracking
  • Language-agnostic, container-based pipeline stages
  • Kubernetes-based autoscaling and parallel processing
  • Works with standard object storage and deduplication
  • Runs on major clouds or on-premises

Pricing: Free and open source under the Apache-2.0 license; it can also be deployed on supported cloud providers.

Dremio OSS

Dremio OSS is a free, Apache-licensed, open-source data lakehouse platform in Java that provides fast SQL analytics directly on data stored in warehouses and lakes.

GitHub stars
1.5k
Last commit
1 yr ago
Licence
Apache-2.0
Self-hosted
Yes
dremio.comDremio OSS homepage screenshot

Dremio is a data analytics platform aimed at organizations that want to query and analyze large volumes of data without first moving or copying it into a separate, dedicated analytics database. It positions itself as filling a gap between raw data storage and the tools people need to get value out of that data.

The open-source edition can be built and run locally, exposing a web UI at localhost:9047 once started, or installed as a production tarball for server deployment, with an embedded mode also available. Building it requires a specific combination of JDK versions (21 as default, with 17 and 11 configured in the Maven toolchain for certain tests) plus Maven, reflecting a substantial, enterprise-grade Java codebase.

The README notes that to provide the best possible experience, the build includes some dependencies distributed under non-open-source licenses, so the fully open-source experience may differ slightly depending on build configuration. Dremio is written in Java and the open-source edition is released under the Apache-2.0 license, with full documentation hosted separately.

Key features

  • SQL analytics directly on lake and warehouse data
  • Web-based query interface
  • Local and production deployment modes
  • Built on a Java codebase with Maven tooling
  • No data duplication required for analytics

Pricing: Dremio Cloud is pay-as-you-go at $0.20 per compute unit, with a 30-day trial that includes $400 in credit. Dremio Enterprise, which can be self-hosted, is priced through sales.

SnappyData

SnappyData, also called TIBCO ComputeDB, is an in-memory analytics database built on Apache Spark and Apache Geode; the repository is now legacy.

GitHub stars
1k
Last commit
3 yr ago
Latest release
v1.3.1
Self-hosted
Yes

SnappyData is a distributed, memory-optimized analytics database that combines an in-memory hybrid database with Apache Spark. In one cluster it provides analytic query processing, mutable data with transactions, access to many big data sources and stream processing. TIBCO also marketed it as ComputeDB.

A typical use is interactive analytics over large datasets with little pre-processing, avoiding the need to pre-aggregate data or build cubes for ad hoc visual analytics. It manages data in memory, generates code with vectorization optimizations and makes use of multi-core CPUs to keep response times low. The README stresses it is not a data warehouse replacement but a compute and caching cluster that augments warehouses and data lakes.

The repository is provided for legacy users and informational purposes only: TIBCO states it provides no updates, including security updates, and that the code may contain vulnerabilities. The TIBCO code is under the Apache License 2.0, although the repository metadata lists the license as Other. Evaluate it carefully before any new deployment.

Key features

  • In-memory distributed analytics database
  • Built on Apache Spark and Apache Geode
  • Transactions and mutable data
  • Stream processing in the same cluster
  • Access to many big data sources
  • Ad hoc analytics without pre-aggregation

Pricing: Source code under the Apache License 2.0; TIBCO provides no updates to this repository.

Databricks alternatives: questions

What is the best open-source alternative to Databricks?
Apache Spark is the top-ranked open-source alternative to Databricks on Enlisted: An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming. Other strong options are Apache Zeppelin, Apache Cloudberry, Pachyderm and Dremio OSS.
Are these Databricks alternatives free?
All 6 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer. 1 also offers a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Databricks alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 3 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all