Gain the skills for building efficient and scalable data pipelines. Explore essential data engineering platforms (Hadoop, Spark, and Snowflake) and learn how to optimize them using Python, PySpark, and MLflow.
What you'll learn
understand the architecture and functionalities of Spark, Hadoop, and Snowflake
learn to build and optimize data pipelines using Python
apply PySpark and MLflow for data processing and model management
Course objectives
equip learners with the skills to construct scalable data pipelines
provide knowledge on optimizing data engineering tools