Apache Spark
An analytics engine for processing large datasets, with APIs in Scala, Java, Python and R and libraries for SQL, machine learning, graphs and streaming.
- GitHub stars
- 44k
- Last commit
- today
- Licence
- Apache-2.0
- Self-hosted
- Yes

Apache Spark is an analytics engine that brings batch, SQL, machine learning, graph and streaming workloads together for processing data at scale. It is maintained by the Apache Software Foundation under the Apache-2.0 license. It is written mainly in Scala and offers high-level APIs in Scala, Java and Python, with an R API that is now deprecated, on top of an engine optimized to execute general computation graphs.
A set of higher-level tools is included: Spark SQL for queries and DataFrames, MLlib for machine learning, GraphX for graphs, Structured Streaming for streams, and a pandas-style API for existing pandas workloads. Together they let teams run batch analytics, machine learning and streaming jobs on a single engine and share code between them.
The README points to the project website for full documentation and programming guides, and it contains only basic setup and build instructions. Spark suits data engineers and data scientists who process datasets too large for a single machine, and organizations that want one framework for ETL, analytics and machine learning on clusters.
Key features
- High-level APIs in Scala, Java and Python
- Optimized engine for computation graphs
- Spark SQL and DataFrames
- MLlib for machine learning
- GraphX for graph analytics
- Structured Streaming for stream processing
- Pandas API on Spark
Pricing: Free and open source under the Apache-2.0 license.



