7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

2 alternatives ranked by real activity

Open-source Google Cloud Dataflow alternatives

A curated, ranked list of the 2 best open-source alternatives to Google Cloud Dataflow.

The best open-source alternative to Google Cloud Dataflow is Apache Spark. If that doesn't suit you, another good option is RisingWave.

Google Cloud Dataflow alternatives are mainly data pipeline & ETL tools. 2 of them shipped code in the last 30 days, 2 can be self-hosted, and 2 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Apache Spark

An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming.

GitHub stars
44k
Last commit
today
Licence
Apache-2.0
Self-hosted
Yes
spark.apache.orgApache Spark homepage screenshot

Apache Spark is an analytics engine that brings batch, SQL, machine learning, graph and streaming workloads together for processing data at scale. It is maintained by the Apache Software Foundation under the Apache-2.0 license. It is written mainly in Scala and offers high-level APIs in Scala, Java and Python, with an R API that is now deprecated, on top of an engine optimized to execute general computation graphs.

A set of higher-level tools is included: Spark SQL for queries and DataFrames, MLlib for machine learning, GraphX for graphs, Structured Streaming for streams, and a pandas-style API for existing pandas workloads. Together they let teams run batch analytics, machine learning and streaming jobs on a single engine and share code between them.

The README points to the project website for full documentation and programming guides, and it contains only basic setup and build instructions. Spark suits data engineers and data scientists who process datasets too large for a single machine, and organizations that want one framework for ETL, analytics and machine learning on clusters.

Key features

  • High-level APIs in Scala, Java and Python
  • Optimized engine for computation graphs
  • Spark SQL and DataFrames
  • MLlib for machine learning
  • GraphX for graph analytics
  • Structured Streaming for stream processing
  • Pandas API on Spark

Pricing: Free and open source under the Apache-2.0 license.

Read more about Apache SparkWebsite GitHub

RisingWave

An open-source event streaming platform that ingests data from databases, streams and webhooks, processes it incrementally with SQL and serves fresh results at low latency.

GitHub stars
9.4k
Last commit
today
Latest release
v3.1.0
Licence
Apache-2.0
Self-hosted
Yes
go.risingwave.comRisingWave homepage screenshot

RisingWave is an event streaming platform, positioned for agentic AI and real-time applications, that continuously takes in data, transforms it and serves it. It is meant to replace the familiar chain of Debezium (change data capture), Kafka (transport), Flink (processing) and a serving database with one system, removing the latency and operational work of each hop.

Sources include webhooks from SaaS applications, native change data capture from PostgreSQL and MySQL, event streams such as Kafka, Pulsar and Kinesis, and batch data from S3 and warehouses, all unified under a SQL interface where streams and tables can be joined. Processing is incremental: if upstream data changes, just the affected results get recomputed, which keeps materialized views up to date without full recomputation. Results are queryable directly, and the topics mention PostgreSQL compatibility and Apache Iceberg.

RisingWave is written in Rust and licensed under Apache-2.0. It can be tried with Docker or Kubernetes using the quick start guide, and it suits data engineers who want streaming pipelines defined in SQL instead of several separate services.

Key features

  • Ingestion from CDC, Kafka, webhooks and batch sources
  • Incremental stream processing in SQL
  • Always up-to-date materialized views
  • Low-latency serving of query results
  • Interface compatible with PostgreSQL clients
  • Docker and Kubernetes deployment

Pricing: Free and open source under the Apache-2.0 licence.

Read more about RisingWaveWebsite GitHub

Google Cloud Dataflow alternatives: questions

What is the best open-source alternative to Google Cloud Dataflow?
Apache Spark is the top-ranked open-source alternative to Google Cloud Dataflow on Enlisted: An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming. Another strong option is RisingWave.
Are these Google Cloud Dataflow alternatives free?
Both are open source, so the code is free to use under its licence, and both can be self-hosted on your own server or computer.
How is this list of Google Cloud Dataflow alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 2 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all