Databricks: Spark, made into a platform.
The PySpark course teaches the engine. This one teaches everything built around it — governance with Unity Catalog, incremental ingestion, declarative pipelines, orchestration, SQL warehouses, MLflow, and how to ship it all like software without a surprise bill.
The one picture this course is built on
Databricks is split in two. Understanding what runs where explains its security model, its networking, its billing and most of its error messages.
Do PySpark lessons 02, 06, 08 and 09 first — Databricks runs Spark, and the mental model transfers directly. Databricks renames products often (Delta Live Tables became Lakeflow Spark Declarative Pipelines; Workflows became Lakeflow Jobs; Asset Bundles became Declarative Automation Bundles). Examples were checked against the documentation in September 2026 but could not be executed while writing, because they need a workspace. This course uses the current names and mentions the old ones, because you will see both in documentation and job adverts.
Pick a path
Starting a Databricks job next week
Lessons 01, 02, 04, 05, 06, 12. Architecture, compute, notebooks, Unity Catalog and jobs.
Building data pipelines
Lessons 06, 07, 08, 09, 10, 13. Ingest, medallion layers, declarative pipelines and streaming.
Certification prep (Data Engineer Associate)
Lessons 02, 04, 06, 07, 08, 10, 12, 16 — then the project end to end.
It's slow and expensive
Lessons 04, 15, 16. Right-size compute, use Photon and clustering, and read the billing system tables.
Ships with a hands-on project
🛠 An end-to-end lakehouse on Free Edition
Land raw files in a Unity Catalog volume, ingest them with Auto Loader, build bronze → silver → gold in a declarative pipeline with data-quality expectations, schedule it as a job, and put a Databricks SQL dashboard on the gold layer. It reuses the retail data from the PySpark project, so you can compare the two approaches directly.
Open the project →