Running Databricks on AWS
(1 day)
Course Description
This course is ideal for data engineers looking to master the Databricks Lakehouse platform on AWS, as well as cloud engineers and architects specializing in AWS who need to effectively integrate and manage Databricks within their cloud infrastructure. It bridges the gap between data processing and cloud deployment, providing practical skills for building and maintaining modern data solutions.
Prerequisites
- Databricks Data Engineering Associate Level of understanding
- Some experience administering and/or deploying workloads on AWS
Course Outline
Module 1: Foundations of Databricks and AWS
- Introduction to the Databricks Lakehouse Platform
- AWS Essentials for Databricks
Module 2: Deploying Databricks on AWS
- Workspace Deployment within a VPC
- Integrating with S3 for Storage
- Establishing Trust with IAM Roles
- Exercise: Deploy a workspace to a customer managed VPC
Module 3: Managing Databricks Clusters and AWS Resources
- Databricks Cluster Types and EC2 Instances
- Autoscaling and Spot Instances
- Exercise: Deploying a cluster with Spot Instances
- Resource Organization with Tagging
- Secure Access with Instance Profiles
- Exercise: Connecting to AWS Textract service from within a cluster
Module 4: Working with S3 and Delta Lake
- Databricks Integration with S3
- How Unity Catalog gives clusters permissions to S3
- Exercise: Creating a new S3 bucket and setting up an External Location
- Deep Dive into Delta Lake
- Understanding CDC Storage Costs
Module 5: Monitoring and Logging across AWS and Databricks
- Monitoring with Amazon CloudWatch
- Analyzing Databricks Logs
- Leveraging the Spark UI
- Using all three for debugging bottlenecks
- Exercise: Creating a performance analysis Dashboard