About SnappyData
SnappyData is a distributed, memory-optimized analytics database that combines an in-memory hybrid database with Apache Spark. In one cluster it provides analytic query processing, mutable data with transactions, access to many big data sources and stream processing. TIBCO also marketed it as ComputeDB.
A typical use is interactive analytics over large datasets with little pre-processing, avoiding the need to pre-aggregate data or build cubes for ad hoc visual analytics. It manages data in memory, generates code with vectorization optimizations and makes use of multi-core CPUs to keep response times low. The README stresses it is not a data warehouse replacement but a compute and caching cluster that augments warehouses and data lakes.
The repository is provided for legacy users and informational purposes only: TIBCO states it provides no updates, including security updates, and that the code may contain vulnerabilities. The TIBCO code is under the Apache License 2.0, although the repository metadata lists the license as Other. Evaluate it carefully before any new deployment.
Key features
- In-memory distributed analytics database
- Built on Apache Spark and Apache Geode
- Transactions and mutable data
- Stream processing in the same cluster
- Access to many big data sources
- Ad hoc analytics without pre-aggregation
Good fit for
- →Interactive analytics over large datasets
- →Maintaining legacy Spark-based deployments
- Built with
- Scala
- Tags
- analytics
- in-memory-database
- apache-spark
- apache-geode
- streaming
- big-data
- scala
- legacy
SnappyData: questions and answers
- What is SnappyData used for?
- SnappyData, also called TIBCO ComputeDB, is an in-memory analytics database built on Apache Spark and Apache Geode; the repository is now legacy. It is a good fit for interactive analytics over large datasets and maintaining legacy Spark-based deployments.
- Is SnappyData open source?
- Yes. SnappyData is open source under a custom licence. Its source code is on GitHub at TIBCOSoftware/snappydata and is written mainly in Scala.
- Is SnappyData free?
- Yes. SnappyData is open source, so the software itself is free to use under the terms of its own licence.
- Can I self-host SnappyData?
- Yes. SnappyData can be self-hosted on your own server or infrastructure.
- What is SnappyData an alternative to?
- SnappyData is an open-source alternative to Databricks and Snowflake. Other open-source alternatives to Databricks include Apache Cloudberry and Dremio OSS.
- Is SnappyData actively maintained?
- The most recent commit to SnappyData was on 21 November 2022, and the latest release is v1.3.1, published on 12 June 2022. The project has 1k stars on GitHub.
Open-source alternatives to SnappyData
See all
Apache Cloudberry
Databases
One advanced and mature open-source MPP (Massively Parallel Processing) database. Open sou
Apache-2.0vs Snowflake★ 1.4k
Dremio OSS
Data Pipelines & ETL
Dremio - the missing link in modern data
Apache-2.0vs Snowflake★ 1.5k
ClickHouse
Databases
ClickHouse is a fast open-source column-oriented database management system that allows ge
Apache-2.0vs Snowflake★ 50k
TensorBase
Databases
TensorBase is a new big data warehousing with modern efforts.
Apache-2.0vs Snowflake★ 1.5k
Apache Doris
Databases
Apache Doris is a real-time analytics and hybrid search database for AI agents.
Apache-2.0vs Snowflake★ 16k
StarRocks
Databases
The world's fastest open query engine for sub-second analytics both on and off the data la
Apache-2.0vs Snowflake★ 12k
SaaS alternatives to SnappyData
See all
Databricks
Data Pipelines & ETL
Lakehouse platform for data engineering, analytics and machine learning
SaaS
Snowflake
Databases
Cloud data platform for warehousing, data sharing and analytics
SaaS
Azure Synapse Analytics
Databases
Microsoft analytics service combining data warehousing, big data and data integration
SaaS
Google BigQuery
Databases
Serverless data warehouse for SQL analytics on large datasets
SaaS
Amazon Athena
Databases
Serverless interactive query service for analyzing data in Amazon S3 with SQL
SaaS
Amazon Redshift
Databases
Managed cloud data warehouse on AWS for SQL analytics at scale
SaaS
