About Delta Lake
Delta Lake is an open-source storage framework for building a lakehouse architecture on top of data lake storage. It works with compute engines including Apache Spark, PrestoDB, Flink, Trino and Hive, and offers APIs for Scala, Java, Rust, Ruby and Python. Its topics reference ACID transactions, big data and analytics.
The repository documents a transaction protocol, requirements for the underlying storage systems and concurrency control, and ships a quick start for Scala, Java and Python. A set of integrations connects the format to other tools: Spark can read and write Delta tables, Flink can write to them, PrestoDB and Hive can read them, Trino can read and write, Delta Standalone serves JVM projects and the Delta Rust API gives low-level access to tables with Python and Ruby bindings.
This repository is one of several in the Delta Lake organization, alongside delta-rs, delta-sharing and kafka-delta-ingest. The code is written in Scala and licensed under Apache-2.0. It is used by data engineering teams that want reliable tables on object storage, and it is a library and format rather than a hosted product.
Key features
- ACID transactions on data lake storage
- Works with Spark, Flink, Trino and Hive
- APIs for Scala, Java, Rust, Ruby and Python
- Open transaction protocol specification
- Delta Standalone for JVM projects
- Delta Rust API with Python bindings
Good fit for
- →Building a lakehouse on object storage
- →Reliable batch and streaming tables
- →Sharing table access across multiple engines
- Built with
- Scala
- Tags
- lakehouse
- data-lake
- acid
- spark
- big-data
- analytics
- delta-lake
- scala
- storage-format
Delta Lake: questions and answers
- What is Delta Lake used for?
- Delta Lake is an open-source storage layer that adds ACID transactions to data lakes and supports lakehouse architectures with Spark, Flink, Trino and other engines. It is a good fit for building a lakehouse on object storage, reliable batch and streaming tables, and sharing table access across multiple engines.
- Is Delta Lake open source?
- Yes. Delta Lake is open source under the Apache-2.0 licence. Its source code is on GitHub at delta-io/delta and is written mainly in Scala.
- Is Delta Lake free?
- Yes. Delta Lake is open source, so the software itself is free to use.
- Can I self-host Delta Lake?
- Yes. Delta Lake can be self-hosted on your own server or infrastructure.
- What are some alternatives to Delta Lake?
- Similar open-source tools in the Data Pipelines & ETL category include Apache Spark, Dremio OSS and Kafka. SaaS products in the same category include Databricks, Amazon EMR and Cloudera.
- Is Delta Lake actively maintained?
- Yes. The most recent commit to Delta Lake was on 2 October 2026, and the latest release is v4.4.0, published on 20 August 2026. The project has 9k stars on GitHub.
Open-source alternatives to Delta Lake
See all
Apache Spark
Data Pipelines & ETL
Apache Spark - A unified analytics engine for large-scale data processing
Apache-2.0vs Databricks★ 44k
Dremio OSS
Data Pipelines & ETL
Dremio - the missing link in modern data
Apache-2.0vs Snowflake★ 1.5k
Kafka
Data Pipelines & ETL
Apache Kafka - A distributed event streaming platform
Apache-2.0vs Striim★ 34k
Jitsu
Data Pipelines & ETL
Jitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for mo
MITvs Segment★ 5.1k
GraphScope
Data Pipelines & ETL
🔨 🍇 💻 🚀 GraphScope: A One-Stop Large-Scale Graph Computing System from Alibaba | 一站式图计
Apache-2.0vs TigerGraph★ 3.6k
Apache StreamPipes
Data Pipelines & ETL
Apache StreamPipes - A self-service (Industrial) IoT toolbox to enable non-technical users
Apache-2.0★ 751
SaaS alternatives to Delta Lake
See all
Databricks
Data Pipelines & ETL
Lakehouse platform for data engineering, analytics and machine learning
SaaS
Amazon EMR
Data Pipelines & ETL
Managed big data platform on AWS for running Spark, Hive, Presto and other frameworks
SaaS
Cloudera
Data Pipelines & ETL
Hybrid data platform for data engineering, warehousing and machine learning on Hadoop and Spark
SaaS
Microsoft Fabric
Data Pipelines & ETL
Unified analytics platform combining data engineering, warehousing and Power BI
SaaS
Alteryx
Data Pipelines & ETL
Analytics automation platform for data prep, blending and predictive workflows
SaaS
AWS Glue
Data Pipelines & ETL
Serverless data integration service on AWS for ETL jobs and data catalogs
SaaS

