DuckDB
An analytical, in-process SQL database that runs inside your application or from a CLI, with a rich SQL dialect and easy CSV and Parquet querying.
- GitHub stars
- 42k
- Last commit
- yesterday
- Latest release
- v1.5.6
- Licence
- MIT
- Self-hosted
- Yes

DuckDB is an analytical SQL database management system that runs in-process, meaning it is embedded in your application or notebook rather than running as a separate server. The goals are speed, reliability, portability and ease of use for analytical queries, and it is written in C++.
Its SQL dialect goes well beyond the basics: it handles arbitrary nested correlated subqueries, window functions, collations and complex types such as arrays, structs and maps, plus extensions that make SQL more convenient. CSV and Parquet files can be queried by naming them in a FROM clause. DuckDB is available as a standalone command line tool and has clients for Python, R, Java and WebAssembly, with deep integration into packages such as pandas and dplyr.
The project is MIT licensed. Building from source needs CMake, Python 3 and a C++17 compiler. Because it needs no server, DuckDB is popular with data analysts and engineers for exploring files and running local analytics, and it complements larger warehouses rather than replacing them.
Key features
- In-process analytical SQL engine
- Direct querying of CSV and Parquet files
- Window functions and correlated subqueries
- Arrays, structs and maps data types
- CLI plus Python, R, Java and Wasm clients
- Integrations with pandas and dplyr
Pricing: Free and open source under the MIT license.
