About AWS Glue
AWS Glue is a serverless data integration service from Amazon Web Services. It is used to find data across a company's AWS storage and databases, prepare it and run extract, transform and load jobs, without provisioning or managing servers for those jobs.
A central part of the service is its data catalog, which records the structure of datasets so that other AWS analytics tools can query them. Glue also provides job authoring for Spark-based transformations, scheduling and workflow features. AWS positions it as a way to discover, prepare, integrate and modernize ETL processes, and it works alongside services such as Amazon EMR, Athena and Redshift.
AWS Glue is proprietary and available only as a managed service inside AWS, so there is no self-hosted edition. It is billed on usage under AWS pricing, and the vendor maintains a dedicated pricing page.
Key features
- Serverless ETL job execution
- Central data catalog for datasets
- Automated data discovery
- Spark-based data transformation
- Job scheduling and workflows
- Integration with AWS analytics services
Good fit for
- →Building ETL pipelines on AWS
- →Cataloging data lake contents
- →Preparing data for analytics and machine learning
- Tags
- aws
- etl
- serverless
- data-catalog
- data-integration
- spark
- data-pipelines
- cloud-service
AWS Glue: questions and answers
- What is AWS Glue used for?
- AWS Glue is a serverless data integration service from AWS for discovering, preparing and moving data with ETL jobs and a central data catalog. It is a good fit for building ETL pipelines on AWS, cataloging data lake contents, and preparing data for analytics and machine learning.
- How much does AWS Glue cost?
- AWS Glue is a paid product with no free plan. Usage-based. The Data Catalog is free for the first million objects and requests each month, then $1.00 per 100,000 objects or per million requests. Table optimization, statistics and crawlers cost $0.44 per DPU-hour.
- Is AWS Glue open source?
- No. AWS Glue is proprietary (closed-source) software and can't be self-hosted. Open-source alternatives to AWS Glue include Airbyte, Mage AI and Duckle.
- What are some alternatives to AWS Glue?
- AWS Glue competes with Azure Data Factory, Fivetran and Informatica. For open-source options, see Enlisted's ranked list of open-source AWS Glue alternatives.
Open-source alternatives to AWS Glue
See all
Airbyte
Data Pipelines & ETL
Airbyte is the open-source data movement platform. Run ELT pipelines across 700+ connector
OSSvs Fivetran★ 22k
Mage AI
Data Pipelines & ETL
🧙 Build, run, and manage data pipelines for integrating and transforming data.
Apache-2.0vs Fivetran★ 8.8k
Duckle
Data Pipelines & ETL
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-
Apache-2.0vs SnapLogic★ 1.3k
Bruin
Data Pipelines & ETL
Build data pipelines with SQL and Python, ingest data from different sources, add quality
Apache-2.0vs Fivetran★ 1.8k
Apache Spark
Data Pipelines & ETL
Apache Spark - A unified analytics engine for large-scale data processing
Apache-2.0vs Databricks★ 44k
Airflow
Data Pipelines & ETL
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
Apache-2.0vs Astronomer★ 47k
SaaS alternatives to AWS Glue
See all
Azure Data Factory
Data Pipelines & ETL
Managed Azure service for building data integration and ETL pipelines
SaaS
Fivetran
Data Pipelines & ETL
Managed data pipeline service that syncs SaaS and database sources to warehouses
SaaS
Informatica
Data Pipelines & ETL
Enterprise data integration, quality and governance platform
SaaS
Google Cloud Dataflow
Data Pipelines & ETL
Managed stream and batch data processing service on Google Cloud based on Apache Beam
SaaS
Talend
Data Pipelines & ETL
Data integration, quality and governance platform now part of Qlik
SaaS
Matillion
Data Pipelines & ETL
Cloud data integration and transformation platform for loading data into warehouses
SaaS

