Course Outline
Introduction, Goals, and Migration Strategy
- Course objectives, alignment with participant profiles, and key success metrics
- Overview of high-level migration strategies and associated risks
- Configuration of workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Core Lakehouse concepts, an overview of Delta Lake, and Databricks architecture
- Contrasting SMP vs MPP architectures and their impact on migration
- Designing the Medallion (Bronze→Silver→Gold) architecture and an introduction to Unity Catalog
Day 1 Lab — Converting a Stored Procedure
- Practical migration of a sample stored procedure to a notebook
- Translating temp tables and cursors into DataFrame transformations
- Validating results by comparing them against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- ACID transactions, commit logs, versioning, and time travel features
- Using Auto Loader, MERGE INTO patterns, upserts, and handling schema evolution
- Techniques for OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization
Day 2 Lab — Incremental Ingestion & Optimization
- Setting up Auto Loader ingestion and MERGE workflows
- Applying OPTIMIZE, Z-ORDER, and VACUUM, followed by result validation
- Evaluating improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Advanced SQL features: window functions, higher-order functions, and JSON/array manipulation
- Interpreting the Spark UI: DAGs, shuffles, stages, tasks, and identifying bottlenecks
- Query optimization techniques: broadcast joins, hints, caching, and reducing spill
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Utilizing Spark UI traces to pinpoint and resolve skew and shuffle issues
- Conducting before/after benchmarks and documenting tuning steps
Day 4 — Practical PySpark: Replacing Procedural Logic
- Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Modularization techniques, using UDFs/pandas UDFs, widgets, and building reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Converting procedural ETL scripts into modular PySpark notebooks
- Incorporating parametrization, unit-style testing, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Databricks Workflows: job design, task dependencies, triggers, and error management
- Designing incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI pipelines, and testing strategies for PySpark logic
Day 5 Lab — Building a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated by Workflows
- Implementing logging, auditing, retry mechanisms, and automated validations
- Executing the full pipeline, validating outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage, and access control
- Managing cost, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and operational runbooks
Final Review, Knowledge Transfer, and Next Steps
- Presentations by participants detailing their migration work and key takeaways
- Gap analysis, recommended follow-up activities, and handover of training materials
- Providing references, further learning pathways, and support options
Requirements
- A solid grasp of core data engineering concepts
- Practical experience with SQL and stored procedures (e.g., Synapse \/ SQL Server)
- Familiarity with ETL orchestration principles (ADF or comparable tools)
Target Audience
- Technology managers with a data engineering background
- Data engineers transitioning from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing Databricks adoption
Testimonials (1)
All the topics covered, although many were very quick, give us an idea of what we will need to delve into further. Additionally, I liked that we got to do some hands-on practice, although I still believe the course deserves more.
Sandra Mariela Lopez Bernal - Kueski
Course - Databricks
Machine Translated