Interactive course · 30 lessons · ~10 hours

Understand PySpark properly, in a weekend.

Not a function list. This course builds the mental model first — what the driver does, what the executors do, why a plan splits into stages — and only then shows the API. Every concept is explained twice: once in plain English with an analogy, once in engine-level technical detail, with a diagram in between.

0 of 30 lessons complete0%

🛠 Want to build rather than read? Start with the project.

Three tables (500 rows each, 20 columns total) in three formats — CSV, JSON Lines and Parquet — deliberately full of real defects. Clean them, join them, window them, and write the results. Runs on your laptop or on Databricks with Delta tables. The data generator is included.

Open the project →

The whole stack on one page

Before the first lesson, here is everything this course covers and how the pieces sit together. Come back to this diagram whenever a lesson feels disconnected — every topic below is one labelled box here.

1 · Your PySpark code df.filter(...).groupBy(...).agg(...) · spark.sql("...") · MLlib pipelines · readStream Lessons 05, 07–17, 25–27 — the API surface you actually type 2 · Driver process Python interpreter JVM (SparkContext) Py4J socket bridge — lesson 03 Catalyst optimizer → DAG scheduler → task scheduler (18, 06) 3 · Cluster manager Hands out containers. Spark does not own the machines. Standalone YARN Kubernetes Databricks Deploy modes: client vs cluster — lesson 02 4 · Executors — where your data and your work actually live Executor 1 task task task idle cache + shuffle + Python workers Executor 2 task task task task memory model + spill — lesson 21 Executor 3 task task skewed task skew + AQE — lesson 23 Shuffle exchange all-to-all network transfer — lesson 19 5 · Storage — Spark has none of its own S3 / ADLS / GCS · HDFS · Parquet, ORC, Avro, Delta · Kafka · JDBC databases  —  lessons 09, 26
Read it top-down. Your Python code (1) never touches the data. It builds a plan on the driver (2), which asks a cluster manager (3) for machines, then ships tasks to executors (4) that read and write real storage (5). Every performance problem you will ever have is a problem in layer 4 or the arrow into it.

Pick a path

All 30 lessons are worth reading, but you may be on a deadline. These are the honest shortcuts:

~2 hours

I need to write PySpark tomorrow

Lessons 02, 05, 06, 08, 10, 11, 13, 14. You will be productive and you will not embarrass yourself in code review.

~3 hours

I have an interview this week

Lessons 01–03, 06, 14, 18, 19, 20, 23, then 29. Architecture, shuffle and skew are 80% of the questions.

~2.5 hours

My job is slow and I must fix it

Lessons 06, 18, 19, 21, 22, 23, 24. Learn to read the plan and the UI before you change a single config.

~9 hours

I want to actually know this

Start at lesson 01 and go in order. Each lesson assumes the previous one. Mark them complete as you go.

How each lesson is built

🧠 Plain-English first

Every concept opens with an analogy from a kitchen, a library or a warehouse — no jargon until the idea has landed.

⚙️ Then the engine truth

What the JVM, the scheduler and the optimizer are really doing, at the level a senior engineer would explain it.

📊 A diagram per idea

Hand-drawn SVG, not screenshots. They adapt to light and dark mode and stay readable on a phone.

💻 Code that runs

Complete, copy-pasteable snippets with the real output printed underneath so you can check yourself.

🚨 Gotchas called out

The mistakes that cost real teams real money, marked clearly, with the fix next to them.

✅ Checkpoint quizzes

Two or three questions per lesson with explanations — if you can answer them, you understood it.

The curriculum

Keyboard

/ search · ← → move between lessons · progress is saved in your browser only, nothing is uploaded anywhere.