Databricks Certified Data Engineer Professional -Preparation

Udemy Certificate USD 99.99
Enroll now →
Databricks Certified Data Engineer Professional -Preparation

About this course

If you are interested in becoming a Certified Data Engineer Professional from Databricks, you have come to the right place! This study guide will help you with preparing for this certification exam.By the end of this course, you should be able to:1- Develop Code for Data Processing using Python and SQLUsing Python and Tools for developmentDesign and implement a scalable Python project structure optimized for Databricks Asset Bundles (DABs), enabling modular development, deployment automation, and CI/CD integration.Manage and troubleshoot external third-party library installations and dependencies in Databricks, including PyPI packages, local wheels, and source archives.Develop User-Defined Functions (UDFs) using Pandas/Python UDFBuilding and Testing an ETL pipeline with Lakeflow Declarative Pipelines, SQL, and Apache Spark on the Databricks platformBuild and manage reliable, production-ready data pipelines, for batch and streaming data using Lakeflow Declarative Pipelines and Autoloader.Create and Automate ETL workloads using Jobs via UI/APIs/CLI.Explain the advantages and disadvantages of streaming tables compared to materialized views.Use APPLY CHANGES APIs to simplify CDC in Lakeflow Declarative Pipelines.Compare Spark Structured Streaming and Lakeflow Declarative Pipelines to determine the optimal approach for building scalable ETL pipelines. ● Create a pipeline component that uses control flow operators (e.g. if/else, foreach, etc.)Choose the appropriate configs for environments and dependencies, high memory for notebook tasks, and auto-optimization to disallow retries.Develop unit and integration tests using assertDataFrameEqual, assertSchemaEqual, DataFrame.transform, and testing frameworks, to ensure code correctness, including a built-in debugger.

What you'll learn

  • Develop code for data processing using Python and SQL
  • Create and manage reliable data pipelines for batch and streaming data
  • Implement User-Defined Functions (UDFs) in PySpark
  • Use Lakeflow Declarative Pipelines for building ETL workloads
  • Automate job management through APIs and CLI

Course objectives

  • Prepare for the Databricks Certified Data Engineer Professional exam
  • Understand and implement scalable Python project structures
  • Learn to manage external dependencies in Databricks
  • Compare Spark Structured Streaming and Lakeflow for ETL
  • Develop unit and integration tests for code correctness

Skills you'll gain

Related courses

Course details are provided by the platform and may change — always confirm on the provider's site. Links may be affiliate links.