Get in Touch
 Duration 21 hours

Course Outline

Introduction:

  • Apache Spark within the Hadoop Ecosystem
  • Brief Overview of Python and Scala

Foundational Concepts (Theory):

  • Spark Architecture
  • Resilient Distributed Datasets (RDD)
  • Transformations and Actions
  • Stages, Tasks, and Dependencies

Exploring Fundamentals in a Databricks Environment (Hands-on Workshop):

  • Practical exercises using the RDD API
  • Core action and transformation functions
  • PairRDD operations
  • Join operations
  • Caching strategies
  • Practical exercises using the DataFrame API
  • SparkSQL
  • DataFrame operations: select, filter, group, sort
  • User Defined Functions (UDF)
  • Exploring the Dataset API
  • Structured Streaming

Deployment in an AWS Environment (Hands-on Workshop):

  • Fundamentals of AWS Glue
  • Distinguishing between AWS EMR and AWS Glue
  • Implementing sample jobs in both environments
  • Analyzing advantages and limitations

Additional Topics:

  • Introduction to Apache Airflow for orchestration

Requirements

Programming experience (ideally with Python and Scala)

Fundamentals of SQL

Number of participants


Price per participant

Testimonials (3)

Upcoming Courses

Related Categories