
Opis
Building a modern data platform is no longer just about moving data. Organizations must scale reliably, control costs, enforce governance, and accelerate analytics. This book shows you how to design and operate production-grade data platforms using AWS Glue and related AWS analytics services.You will begin with core data management concepts before moving into ingestion from diverse sources, data preparation strategies, metadata management, security controls, and cross-account data sharing. Learn how to design efficient data layouts, orchestrate pipelines, implement CI CD practices, and manage the full lifecycle of data integration workloads.
This updated edition expands coverage of open table formats such as Apache Hudi, Delta Lake, and Apache Iceberg, along with performance tuning, observability, cost optimization, and real-world troubleshooting. You will also explore integrations with machine learning and generative AI workflows powered by Glue and SageMaker.
Written by AWS engineers and architects with deep hands-on experience in large-scale enterprise data lakes, this guide blends architecture principles with real-world implementation insight.
By the end of this book, you will be able to design, deploy, monitor, and optimize scalable serverless ETL pipelines and governed data platforms on AWS.