What Databricks is
Apache Spark is an open-source engine: give it data and code and it runs them across many machines. A company needs much more than an engine: somewhere to store data reliably, rules about who can see what, scheduling, notebooks, SQL for analysts, machine-learning tooling, monitoring and cost control. Databricks is a managed platform that wraps Spark with all of that, built around an idea it calls the lakehouse.
An engine versus a car
Spark is a superb engine. You could bolt it into a frame yourself, add wheels, brakes, dashboard, locks and insurance, and maintain it all forever. Databricks sells the whole car: the same engine (plus a faster one, Photon), with the steering, safety systems and service plan already fitted. You pay for the convenience, and you can still open the bonnet.
1. From warehouses and lakes to the lakehouse
2. What Databricks adds beyond open-source Spark
| Need | Open-source Spark alone | Databricks | Lesson |
|---|---|---|---|
| Reliable tables | Delta Lake (open source) you set up yourself | Delta by default, plus liquid clustering, predictive optimisation | 07 |
| Who can see what | whatever your storage permissions allow | Unity Catalog: grants, lineage, audit, one catalog for everything | 06 |
| Compute | you run and patch clusters (YARN, Kubernetes) | managed clusters, serverless, SQL warehouses | 04 |
| Speed | the JVM Spark engine | Photon, a native C++ engine for SQL/DataFrames | 04, 15 |
| Ingestion | write your own file tracking | Auto Loader, Lakeflow Connect | 08 |
| Pipelines | hand-written jobs | Lakeflow Spark Declarative Pipelines with data-quality expectations | 10 |
| Orchestration | Airflow or cron | Lakeflow Jobs | 12 |
| Analysts | Spark Thrift server | Databricks SQL, AI/BI dashboards, Genie | 11 |
| Machine learning | install MLflow yourself | managed MLflow, feature tables, model serving | 14 |
| Shipping code | your own tooling | Git folders, CLI, Declarative Automation Bundles | 17 |
Databricks renames products often. Delta Live Tables became Lakeflow Declarative Pipelines and then Lakeflow Spark Declarative Pipelines once the framework was contributed to Apache Spark 4.1. Workflows became Lakeflow Jobs. Databricks Asset Bundles became Declarative Automation Bundles. This course uses the names in the documentation at the time of writing (2026) and mentions the older ones, because blog posts, certification material and job adverts still use them.
3. What stays exactly the same
The code you write is ordinary PySpark and SQL. Everything from the PySpark course applies unchanged:
the driver/executor model, lazy evaluation, the Catalyst optimiser, shuffles, partitions and the Spark UI.
A Databricks notebook simply gives you a ready spark session.
# `spark` already exists in every Databricks notebook; no SparkSession.builder needed
trips = spark.read.table("samples.nyctaxi.trips") # a sample dataset every workspace has
(trips
.groupBy("pickup_zip")
.agg({"fare_amount": "avg", "*": "count"})
.orderBy("count(1)", ascending=False)
.limit(5)
.display()) # display(): Databricks' rich table/chart output
Databricks code runs on Databricks, not on a laptop, so these examples were not executed while the course was written. They follow the current documentation, and any output shown is labelled as illustrative. Free Edition (lesson 03) lets you run every example yourself at no cost.
4. Who uses what
Data engineers
Ingestion, pipelines, jobs, Unity Catalog, performance and cost. Most of this course.
Analysts & BI
SQL warehouses, the SQL editor, dashboards, alerts; connecting Power BI or Tableau.
Data scientists & ML engineers
Notebooks, MLflow experiments, feature tables, model serving, vector search for AI apps.
Recap
- Databricks = managed Spark + the platform around it: storage format, governance, compute, orchestration, SQL, ML, DevOps.
- Lakehouse = open files on cheap storage + a transaction log + governance, serving every workload from one copy.
- Your PySpark and SQL skills transfer directly; Databricks adds services around them.
- Expect renames; learn the concept, then map the current name.