Hadoop and Spark Fundamentals: Unit 2

Coursera MOOC / Non-credit USD 49
Enroll now →
Hadoop and Spark Fundamentals: Unit 2

About this course

This course introduces the fundamentals of modern data processing for data engineers, analysts, and IT professionals. You will learn the basics of Hadoop MapReduce, including how it works, how to compile and run Java MapReduce programs, and how to debug and extend them using other languages. The course includes practical exercises such as word counts across multiple files, log file analysis, and large-scale text processing with datasets like Wikipedia. You will also cover advanced MapReduce features and use tools like Yarn and the Job Browser. The course then covers higher-level tools such as Apache Pig and Hive QL for managing data workflows and running SQL-like queries. Finally, you will work with Apache Spark and PySpark to gain experience with modern data analytics platforms. By the end of the course, you will have practical skills to work with big data in various environments.

What you'll learn

  • Understanding the fundamentals of Hadoop and MapReduce
  • Compiling and running Java MapReduce programs
  • Debugging MapReduce jobs and using advanced MapReduce features
  • Utilizing tools like Yarn and the Job Browser
  • Working with Apache Spark and PySpark for data analytics

Skills you'll gain

Related courses

Course details are provided by the platform and may change — always confirm on the provider's site. Links may be affiliate links.