Apache Iceberg: Data Lakehouse Engineering
Tuesday, June 3, 2025
Add Comment
Design scalable, versioned, and ACID-compliant data lakehouse solutions using Apache Iceberg from the ground up.
Preview this Course
What you'll learn
- Gain a deep understanding of Apache Iceberg’s architecture, its role in the modern data lakehouse ecosystem, and why it outperforms traditional table formats li
- Learn how to create, manage, and query Iceberg tables using Python (PyIceberg), SQL interfaces, and metadata catalogs — with practical examples from real-world
- Build high-performance batch and streaming data pipelines by integrating Iceberg with leading engines like Apache Spark, Polars Trino, and DuckDB.
- Explore how to use cloud-native storage with AWS S3, and design scalable Iceberg tables that support large-scale, distributed analytics.
- Apply performance tuning techniques such as file compaction, partition pruning, and metadata caching to optimize query speed and reduce compute costs.
- Work with modern Python analytics tools like Polars and DuckDB for fast in-memory processing, enabling rapid exploration, testing, and data validation workflows

0 Response to "Apache Iceberg: Data Lakehouse Engineering"
Post a Comment