Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction, Learning Objectives, and Migration Strategy
- Defining course goals, aligning participant profiles, and establishing success metrics
- Discussing high-level migration methodologies and associated risk factors
- Configuring workspaces, repositories, and lab datasets
Day 1 — Migration Fundamentals and Architecture
- Overview of Lakehouse concepts, Delta Lake, and the Databricks architecture
- Analyzing the differences between SMP and MPP systems and their impact on migration
- Exploring Medallion (Bronze→Silver→Gold) design principles and the Unity Catalog
Day 1 Lab — Translating a Stored Procedure
- Practical migration of a sample stored procedure into a notebook format
- Converting temporary tables and cursors into DataFrame transformations
- Validating the new solution against the original output
Day 2 — Advanced Delta Lake & Incremental Loading
- Understanding ACID transactions, commit logs, versioning, and time travel features
- Working with Auto Loader, MERGE INTO patterns, upserts, and schema evolution
- Applying OPTIMIZE, VACUUM, Z-ORDER, partitioning, and storage optimization techniques
Day 2 Lab — Incremental Ingestion & Optimization
- Setting up Auto Loader ingestion and MERGE workflows
- Executing OPTIMIZE, Z-ORDER, and VACUUM operations; verifying results
- Evaluating improvements in read/write performance
Day 3 — SQL in Databricks, Performance & Debugging
- Utilizing advanced SQL features: window functions, higher-order functions, and JSON/array processing
- Interpreting the Spark UI to analyze DAGs, shuffles, stages, tasks, and diagnose bottlenecks
- Implementing query tuning strategies: broadcast joins, hints, caching, and spill mitigation
Day 3 Lab — SQL Refactoring & Performance Tuning
- Refactoring complex SQL processes into optimized Spark SQL
- Using Spark UI traces to detect and resolve data skew and shuffle issues
- Benchmarking performance before and after tuning; documenting the process
Day 4 — Tactical PySpark: Replacing Procedural Logic
- Understanding the Spark execution model: driver, executors, lazy evaluation, and partitioning strategies
- Converting loops and cursors into vectorized DataFrame operations
- Implementing modularization, UDFs/pandas UDFs, widgets, and reusable libraries
Day 4 Lab — Refactoring Procedural Scripts
- Transforming procedural ETL scripts into modular PySpark notebooks
- Incorporating parametrization, unit-style tests, and reusable functions
- Conducting code reviews and applying best-practice checklists
Day 5 — Orchestration, End-to-End Pipeline & Best Practices
- Designing Databricks Workflows: job structure, task dependencies, triggers, and error handling
- Architecting incremental Medallion pipelines with quality rules and schema validation
- Integrating with Git (GitHub/Azure DevOps), CI, and testing strategies for PySpark logic
Day 5 Lab — Build a Complete End-to-End Pipeline
- Assembling a Bronze→Silver→Gold pipeline orchestrated by Workflows
- Implementing logging, auditing, retries, and automated validations
- Executing the full pipeline, verifying outputs, and preparing deployment documentation
Operationalization, Governance, and Production Readiness
- Best practices for Unity Catalog governance, lineage tracking, and access controls
- Managing costs, cluster sizing, autoscaling, and job concurrency patterns
- Creating deployment checklists, rollback strategies, and runbooks
Final Review, Knowledge Transfer, and Next Steps
- Participant presentations on migration work and key takeaways
- Gap analysis, recommended follow-up activities, and handover of training materials
- Providing references, further learning paths, and support options
Requirements
- A solid grasp of data engineering concepts
- Practical experience with SQL and stored procedures (Synapse / SQL Server)
- Knowledge of ETL orchestration concepts (ADF or similar tools)
Target Audience
- Technology managers with a background in data engineering
- Data engineers moving from procedural OLAP logic to Lakehouse patterns
- Platform engineers overseeing the adoption of Databricks
35 Hours