7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

Apache Spark

Open source

An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming.

spark.apache.org
Apache Spark homepage screenshot
GitHub stars
44k
Last commit
today
Repository age
12 years
Licence
Apache-2.0
Self-hosted
Yes

About Apache Spark

Apache Spark is an analytics engine that brings batch, SQL, machine learning, graph and streaming workloads together for processing data at scale. It is maintained by the Apache Software Foundation under the Apache-2.0 license. It is written mainly in Scala and offers high-level APIs in Scala, Java and Python, with an R API that is now deprecated, on top of an engine optimized to execute general computation graphs.

A set of higher-level tools is included: Spark SQL for queries and DataFrames, MLlib for machine learning, GraphX for graphs, Structured Streaming for streams, and a pandas-style API for existing pandas workloads. Together they let teams run batch analytics, machine learning and streaming jobs on a single engine and share code between them.

The README points to the project website for full documentation and programming guides, and it contains only basic setup and build instructions. Spark suits data engineers and data scientists who process datasets too large for a single machine, and organizations that want one framework for ETL, analytics and machine learning on clusters.

Key features

  • High-level APIs in Scala, Java and Python
  • Optimized engine for computation graphs
  • Spark SQL and DataFrames
  • MLlib for machine learning
  • GraphX for graph analytics
  • Structured Streaming for stream processing
  • Pandas API on Spark

Good fit for

  • →Large-scale ETL pipelines
  • →Distributed machine learning
  • →Stream processing
Built with
Scala
Python
Java
Tags
big-data
analytics
data-processing
scala
spark-sql
machine-learning
streaming
apache

Apache Spark: questions and answers

What is Apache Spark used for?
Apache Spark is an analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming. It is a good fit for large-scale ETL pipelines, distributed machine learning and stream processing.
Is Apache Spark open source?
Yes. Apache Spark is open source under the Apache-2.0 licence. Its source code is on GitHub at apache/spark and is written mainly in Scala.
Is Apache Spark free?
Yes. Apache Spark is open source, so the software itself is free to use.
Can I self-host Apache Spark?
Yes. Apache Spark can be self-hosted on your own server or infrastructure.
What is Apache Spark an alternative to?
Apache Spark is an open-source alternative to Databricks, AWS Glue, Google Cloud Dataflow and Amazon EMR. Other open-source alternatives to Databricks include Dremio OSS and Pachyderm.
Is Apache Spark actively maintained?
Yes. The most recent commit to Apache Spark was on 2 October 2026. The project has 44k stars on GitHub.

Open-source alternatives to Apache Spark

See all

SaaS alternatives to Apache Spark

See all