7,363 open-source and SaaS tools, with GitHub stats refreshed every day.

8 alternatives ranked by real activity

Open-source Azure Data Factory alternatives

A curated, ranked list of the 8 best open-source alternatives to Azure Data Factory.

The best open-source alternative to Azure Data Factory is Airflow. If that doesn't suit you, other good options are Kestra, Prefect, Airbyte and Dagster.

Azure Data Factory alternatives are mainly data pipeline & ETL tools. 8 of them shipped code in the last 30 days, 8 can be self-hosted, and 7 use a permissive licence.

Last updated October 3, 2026 · ranked by GitHub stars, growth and recent commits

Airflow

An Apache workflow orchestration platform where data teams author, schedule and monitor pipelines as Python code, widely used for data engineering and ELT jobs.

GitHub stars
47k
Last commit
today
Latest release
3.3.2
Licence
Apache-2.0
Self-hosted
Yes
airflow.apache.orgAirflow homepage screenshot

Apache Airflow is a workflow platform where you write, schedule and watch over pipelines in code. Workflows are defined as Python code in the form of directed acyclic graphs, or DAGs, where each task and its dependencies are explicit, so pipelines can be versioned, tested and reviewed like any other software.

A scheduler runs tasks on a defined cadence and workers execute them, while a web interface shows the status of each run, lets you inspect logs and retry failures. Airflow is used heavily in data engineering, including data integration, ELT and ETL pipelines, data orchestration and machine learning workflows, and its topics cover data science and automation. It has a large ecosystem of providers and operators for connecting to databases, cloud services and other systems, and the 3.x line is the current series.

Airflow is an Apache Software Foundation project under the Apache-2.0 license and runs on your own infrastructure, from a single machine to Kubernetes. Several vendors offer managed Airflow services, but the project itself is self-hosted software. It suits data engineers who prefer code-first orchestration over drag-and-drop tools.

Key features

  • Workflows defined as Python DAGs
  • Scheduler for recurring and dependent tasks
  • Web UI for monitoring runs and logs
  • Retries and task dependency management
  • Large provider ecosystem of integrations
  • Runs on single machines or Kubernetes

Pricing: Free and open source under the Apache-2.0 license.

Read more about AirflowWebsite GitHub

Kestra

An open-source, event-driven orchestration and scheduling platform where data, AI and infrastructure workflows are defined declaratively in YAML and managed from a UI.

GitHub stars
29k
Last commit
today
Latest release
v2.0.4
Licence
Apache-2.0
Self-hosted
Yes
go.kestra.ioKestra homepage screenshot

Kestra is an orchestration platform, open source, for data pipelines, AI workflows and infrastructure automation. It brings scheduled and event-triggered automation together under one declarative interface that does not depend on a programming language, applying infrastructure-as-code practices to pipelines so that reliable workflows can be defined in a few lines of YAML.

Workflows can be built in the UI, written by a built-in AI Copilot, or generated from coding agents such as Claude Code and Cursor using agent skills, and everything can be kept as code with Git integration even when it was authored visually. A large plugin ecosystem connects it to external systems, and key concepts cover tasks, triggers and flows. The project is written in Java, with topics covering orchestration, high availability, pipeline-as-code and DevOps, and the 2.0 release adds new capabilities.

Kestra is Apache-2.0 licensed and can be self-hosted, with a quick start that gets a first workflow running in minutes. It targets data engineers, platform teams and automation developers who want an alternative to schedulers such as Airflow with a more declarative style.

Key features

  • Declarative YAML workflow definitions
  • Scheduled and event-driven triggers
  • Visual editor and AI Copilot
  • Git version control integration
  • Large plugin ecosystem
  • Agent skills for coding agents

Pricing: Free and open source under the Apache-2.0 license.

Read more about KestraWebsite GitHub

Prefect

A Python workflow orchestration framework for turning scripts into resilient, observable data pipelines with scheduling, retries, caching and event-driven automation.

GitHub stars
24k
Last commit
today
Latest release
3.8.7
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
prefect.ioPrefect homepage screenshot

Prefect is a Python framework for orchestrating workflows and building data pipelines. It aims to be the simplest way to lift an ordinary script into a production workflow: you add flow and task decorators, and Prefect adds the machinery that makes pipelines dependable and visible.

That machinery includes scheduling, caching, retries and event-based automations, and flows can express dependencies and complex branching logic. The goal is dynamic pipelines that react to changes in the world and recover from unexpected failures. Run activity is tracked, and you can monitor it from a self-hosted Prefect server or from the managed Prefect Cloud dashboard. Prefect requires Python 3.10 or newer, and the documentation covers installation, quickstart, building and deploying workflows and agent setup.

Prefect is licensed under Apache-2.0. Topics list data engineering, ML ops and observability among its uses. It is aimed at data teams who prefer to write workflows as plain Python code, and who want a choice between running the orchestrator themselves or using a hosted control plane.

Key features

  • Flows and tasks defined with Python decorators
  • Scheduling and event-based automations
  • Automatic retries and caching
  • Complex branching and dependencies
  • Self-hosted server or Prefect Cloud monitoring
  • Deployment of workflows to infrastructure

Pricing: Prefect Cloud's Hobby tier is free. Starter is $100 per month and Team $100 per user per month, both billed monthly; Enterprise is custom and billed annually.

Read more about PrefectWebsite GitHub

Airbyte

An open-source data movement platform with hundreds of connectors for ELT pipelines from APIs, databases and files into warehouses, lakes and AI applications.

GitHub stars
22k
Last commit
today
Latest release
v2.0.0
Self-hosted
Yes
Hosted version
Available
airbyte.comAirbyte homepage screenshot

Airbyte is an open-source data movement platform for moving data from APIs, databases and files into data warehouses, data lakes and AI applications. Its premise is that only an open-source project can reach the long tail of data sources while letting engineers tailor existing connectors, with the aim of moving data from any source to any destination.

It offers a catalog of 600 or more connectors for APIs, databases, warehouses, lakes and AI applications, with warehouse destinations such as BigQuery, Redshift and Snowflake, and topics also mention change data capture. The README distinguishes products by job: Airbyte Open Source or Airbyte Cloud for ELT and ETL into warehouses, lakes or databases, and Airbyte Agents, a managed data and context layer, plus an open-source Agent SDK for giving AI agents and MCP clients real-time access to business data.

Airbyte is written in Python and Java, and its license is listed as 'Other' on GitHub because the project uses a mix of licenses, so review the terms for your use. It can be self-hosted or run as the managed Airbyte Cloud, and suits data teams that need many integrations without writing each pipeline by hand.

Key features

  • 600+ source and destination connectors
  • ELT pipelines into warehouses and lakes
  • Change data capture support
  • Customizable open-source connectors
  • Agent SDK for AI agent data access
  • Self-hosted or Airbyte Cloud

Pricing: The open-source Core edition is free to self-host. Managed cloud starts at $20 a month (Standard) or $189 a month for 40 credits (Plus); Pro and Enterprise Flex are custom. 30-day trial.

Read more about AirbyteWebsite GitHub

Dagster

A Python data orchestration platform for building, scheduling and observing data pipelines around the data assets they produce; Dagster is now part of Prefect.

GitHub stars
16k
Last commit
yesterday
Latest release
1.13.25
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
dagster.ioDagster homepage screenshot

Dagster is an orchestration platform for developing, running and observing data assets. Instead of thinking only in terms of tasks, teams define the tables, files and models their pipelines produce, and Dagster tracks how they are built, scheduled and monitored. It is written in Python and aimed at data engineering, analytics and machine learning workloads.

The product surface covers data orchestration, a data catalog, data quality checks, cost insights and many integrations, with an enterprise offering for larger organizations. Comparison pages on the website position it against Airflow, dbt Cloud, Azure Data Factory and AWS Step Functions, and topics include ETL, scheduling and MLOps. The homepage now announces that Dagster is part of Prefect, and points visitors looking for agentic orchestration or MCP support to Prefect.

The open-source core is Apache-2.0 licensed and can be self-hosted, while a managed cloud and enterprise plans are offered with pricing on the company site. Dagster University and documentation help with learning. It suits data teams that want a software-engineering-style approach to pipelines.

Key features

  • Asset-centric data orchestration
  • Scheduling and monitoring of pipelines
  • Data catalog and data quality features
  • Cost insights for pipelines
  • Integrations with common data tools
  • Pipelines developed in Python

Pricing: Dagster+ Solo is $10 and Starter $100 per month, each plus per-credit usage charges, with a 30-day free trial. Pro is quoted through sales.

Read more about DagsterWebsite GitHub

Mage AI

Mage is an open-source data pipeline tool for building, scheduling and debugging ETL and transformation jobs in Python, SQL or R through a notebook-style UI.

GitHub stars
8.8k
Last commit
21 days ago
Latest release
0.9.79
Licence
Apache-2.0
Self-hosted
Yes
Hosted version
Available
mage.aiMage AI homepage screenshot

Mage, called Mage OSS in its README, is a self-hosted development environment for building data pipelines. It is meant for teams that automate ETL tasks, design data flows or orchestrate transformations, and it presents the work in a notebook-style interface made of modular blocks of code.

Pipelines are written block by block in Python, SQL or R. Jobs can be run manually or on a schedule, including cron expressions, and prebuilt connectors reach databases, APIs and cloud storage. Debugging is visual, with logs, live data previews and execution you can follow step by step, and dbt models can be built and run inside Mage. Topics on the repository also reference Spark, reverse ETL and machine learning workloads.

Mage OSS is licensed under Apache-2.0 and installs with Docker, pip or conda, with no cloud account required. For larger deployments the vendor offers Mage Pro, a commercial platform that adds enterprise orchestration, collaboration and AI-assisted workflows.

Key features

  • Modular pipelines in Python, SQL or R
  • Notebook-style interactive editor
  • Prebuilt connectors for databases, APIs and storage
  • Manual and cron-based scheduling
  • Visual debugging with logs and previews
  • dbt model support inside pipelines

Pricing: Managed Mage starts at $29 per month (Starter) and $100 per month (Team), plus $0.50 per CPU-hour, with a 7-day free trial. Enterprise is by contract.

Read more about Mage AIWebsite GitHub

Bruin

An open-source data pipeline tool that combines ingestion, SQL and Python transformations, and quality checks in one framework you can run locally or in CI.

GitHub stars
1.8k
Last commit
today
Latest release
v0.11.767
Licence
Apache-2.0
Self-hosted
Yes
getbruin.comBruin homepage screenshot

Bruin is a data pipeline tool that brings data ingestion, transformation and data quality into a single framework. Instead of stitching together separate tools for loading, modeling and testing, you describe a pipeline with SQL, Python or R assets and run it against the major data platforms, such as BigQuery and Snowflake.

Pipelines can ingest data from different sources, materialize tables and views, and build incremental tables. Python assets run in isolated environments using uv, Jinja templating cuts down on repetition, and built-in quality checks validate results. A dry-run mode validates a pipeline end to end, secrets are injected through environment variables, and a VS Code extension supports day-to-day development.

Bruin is written in Go, licensed under Apache-2.0 and distributed as a command-line tool that is easy to install. It runs on your local machine, an EC2 instance or GitHub Actions, so no dedicated orchestration server is required. It suits data engineers and analytics teams who want pipelines defined as code in a repository.

Key features

  • SQL, Python, and R transformations
  • Data ingestion from multiple sources
  • Table and view materializations with incremental loads
  • Built-in data quality checks
  • Dry-run validation of whole pipelines
  • VS Code extension for development

Pricing: Free and open source under the Apache-2.0 licence.

Read more about BruinWebsite GitHub

Duckle

An open-source ETL and ELT platform built on DuckDB that you deploy yourself, with visual or SQL pipelines, dbt, CDC, data quality and an MCP server.

GitHub stars
1.3k
Last commit
yesterday
Latest release
v0.7.4
Licence
Apache-2.0
Self-hosted
Yes
duckle.orgDuckle homepage screenshot

Duckle is an open-source ETL and ELT platform for teams that want their data pipelines running on their own infrastructure. You build pipelines on a visual canvas, in Python or in SQL, then ship the same file to your own server or cloud account, where a headless runner executes it on a schedule.

Pipelines compile to SQL on DuckDB and use all the cores available, so a larger machine runs them faster. The platform lists support for a large catalog of components, dbt, change data capture, data quality checks, reverse ETL and lineage, along with a web console, roles, an audit trail and an MCP server so AI agents such as Claude or Cursor can work with it. Each pipeline is a single file that can live in git.

Duckle is written in Rust and licensed under Apache-2.0, and it runs headless in Docker or on a plain server, with Kubernetes among its topics. It states there is no vendor cloud and no per-row billing. It is an independent project by SlothFlowLabs and is not affiliated with DuckDB Labs or MotherDuck. It suits data engineers who want a self-hosted alternative to managed pipeline services.

Key features

  • Visual canvas, Python, or SQL pipeline authoring
  • Runs on DuckDB across all CPU cores
  • dbt, CDC, and reverse ETL support
  • Data quality checks and lineage
  • Web console with roles and audit trail
  • MCP server for AI agents

Pricing: Free and open source under the Apache-2.0 licence, with no vendor cloud or per-row billing according to the README.

Read more about DuckleWebsite GitHub

Azure Data Factory alternatives: questions

What is the best open-source alternative to Azure Data Factory?
Airflow is the top-ranked open-source alternative to Azure Data Factory on Enlisted: An Apache workflow orchestration platform where data teams author, schedule and monitor pipelines as Python code, widely used for data engineering and ELT jobs. Other strong options are Kestra, Prefect, Airbyte and Dagster.
Are these Azure Data Factory alternatives free?
All 8 are open source, so the code is free to use under its licence, and all of them can be self-hosted on your own server or computer. 4 also offer a paid or managed cloud version if you'd rather not host it yourself.
How is this list of Azure Data Factory alternatives ranked?
By a score built from GitHub stars, star growth over the last 30 days and how recently the code changed. 8 of these projects shipped code in the last 30 days. Data is refreshed daily, and nobody can pay to move up.

People also look for alternatives to…

View all