Interactive course · 18 lessons · ~6 hours

Databricks: Spark, made into a platform.

The PySpark course teaches the engine. This one teaches everything built around it — governance with Unity Catalog, incremental ingestion, declarative pipelines, orchestration, SQL warehouses, MLflow, and how to ship it all like software without a surprise bill.

0 of 18 lessons complete0%

The one picture this course is built on

Databricks is split in two. Understanding what runs where explains its security model, its networking, its billing and most of its error messages.

CONTROL PLANE — runs in Databricks' cloud account Workspace web app notebooks, UI, REST API (lessons 03, 05) Jobs scheduler orchestration, triggers (lesson 12) Cluster manager starts / stops compute (lesson 04) Unity Catalog metadata, permissions, lineage (lesson 06) commands + credentials flow down; your data does NOT flow up COMPUTE PLANE — where Spark actually runs Classic: VMs in YOUR cloud account (your VPC/VNet) Serverless: VMs in Databricks' account, isolated per workspace Clusters driver + executors SQL warehouses Photon SQL engine Model serving endpoints (lesson 14) YOUR CLOUD STORAGE S3 · ADLS · GCS Delta tables = Parquet files + a transaction log Your data stays in your account, in open formats. Why this matters in practice A cluster that "won't start" is usually a compute-plane cloud problem (quota, network); a "permission denied" on a table is Unity Catalog.
Databricks orchestrates; your cloud (or its serverless pool) computes; your storage holds the data. Lesson 02 covers each boundary and what crosses it.
Prerequisites, and a note on names

Do PySpark lessons 02, 06, 08 and 09 first — Databricks runs Spark, and the mental model transfers directly. Databricks renames products often (Delta Live Tables became Lakeflow Spark Declarative Pipelines; Workflows became Lakeflow Jobs; Asset Bundles became Declarative Automation Bundles). Examples were checked against the documentation in September 2026 but could not be executed while writing, because they need a workspace. This course uses the current names and mentions the old ones, because you will see both in documentation and job adverts.

Pick a path

~2 hours

Starting a Databricks job next week

Lessons 01, 02, 04, 05, 06, 12. Architecture, compute, notebooks, Unity Catalog and jobs.

~2.5 hours

Building data pipelines

Lessons 06, 07, 08, 09, 10, 13. Ingest, medallion layers, declarative pipelines and streaming.

~2 hours

Certification prep (Data Engineer Associate)

Lessons 02, 04, 06, 07, 08, 10, 12, 16 — then the project end to end.

~1.5 hours

It's slow and expensive

Lessons 04, 15, 16. Right-size compute, use Photon and clustering, and read the billing system tables.

Ships with a hands-on project

🛠 An end-to-end lakehouse on Free Edition

Land raw files in a Unity Catalog volume, ingest them with Auto Loader, build bronze → silver → gold in a declarative pipeline with data-quality expectations, schedule it as a job, and put a Databricks SQL dashboard on the gold layer. It reuses the retail data from the PySpark project, so you can compare the two approaches directly.

Open the project →

The curriculum