About Pachyderm
Pachyderm is a data pipeline platform built for data engineering teams that need to automate complex transformations across large and varied datasets. It positions itself as a CI/CD engine for data, combining pipeline automation with built-in version control and lineage so that every change to data and every pipeline run can be traced.
Pipelines trigger automatically when Pachyderm detects a change in input data, and each stage can be written in any language since Pachyderm treats containers as the unit of processing. It tracks immutable data lineage and versioning for any data type, scales and parallelizes processing automatically on top of Kubernetes, and stores data in standard object stores with automatic deduplication. It runs on all major cloud providers as well as on-premises.
Pachyderm is written in Go and released under the Apache-2.0 license. It can be run locally for evaluation or deployed on AWS, GCE or Azure in a short setup, and the project provides tutorials, example projects and case studies through its documentation.
Key features
- Automatic pipeline triggers on data changes
- Immutable data versioning and lineage tracking
- Language-agnostic, container-based pipeline stages
- Kubernetes-based autoscaling and parallel processing
- Works with standard object storage and deduplication
- Runs on major clouds or on-premises
Good fit for
- →Versioned data pipelines for machine learning
- →Auditable data transformations across teams
- Tags
- data-pipelines
- data-versioning
- kubernetes
- big-data
- data-lineage
- golang
- docker
- data-engineering
Open-source alternatives to Pachyderm
See all
Apache Spark
Data Pipelines & ETL
Apache Spark - A unified analytics engine for large-scale data processing
Apache-2.0vs Databricks★ 44k
Dremio OSS
Data Pipelines & ETL
Dremio - the missing link in modern data
Apache-2.0vs Snowflake★ 1.5k
Apache Zeppelin
BI & Dashboards
Web-based notebook that enables data-driven, interactive data analytics and collaborative
Apache-2.0vs Hex★ 6.7k
SnappyData
Databases
Project SnappyData - memory optimized analytics database, based on Apache Spark™ and Apach
OSSvs Databricks★ 1k
Apache Cloudberry
Databases
One advanced and mature open-source MPP (Massively Parallel Processing) database. Open sou
Apache-2.0vs Snowflake★ 1.4k
Airflow
Data Pipelines & ETL
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Apache-2.0vs Astronomer★ 47k
SaaS alternatives to Pachyderm
See all
Databricks
Data Pipelines & ETL
Lakehouse platform for data engineering, analytics and machine learning
SaaS
Astronomer
Data Pipelines & ETL
Managed Apache Airflow platform for orchestrating data pipelines
SaaS
Amazon EMR
Data Pipelines & ETL
Managed big data platform on AWS for running Spark, Hive, Presto and other frameworks
SaaS
AWS Glue
Data Pipelines & ETL
Serverless data integration service on AWS for ETL jobs and data catalogs
SaaS
Cloudera
Data Pipelines & ETL
Hybrid data platform for data engineering, warehousing and machine learning on Hadoop and Spark
SaaS
Google Cloud Dataflow
Data Pipelines & ETL
Managed stream and batch data processing service on Google Cloud based on Apache Beam
SaaS
