Module 1 Β· The platform

What Databricks is

Beginner 14 min read Spark, plus everything a company needs around it

Apache Spark is an open-source engine: give it data and code and it runs them across many machines. A company needs much more than an engine: somewhere to store data reliably, rules about who can see what, scheduling, notebooks, SQL for analysts, machine-learning tooling, monitoring and cost control. Databricks is a managed platform that wraps Spark with all of that, built around an idea it calls the lakehouse.

πŸš—

An engine versus a car

Spark is a superb engine. You could bolt it into a frame yourself, add wheels, brakes, dashboard, locks and insurance, and maintain it all forever. Databricks sells the whole car: the same engine (plus a faster one, Photon), with the steering, safety systems and service plan already fitted. You pay for the convenience, and you can still open the bonnet.

1. From warehouses and lakes to the lakehouse

Data warehouse βœ“ SQL, ACID transactions βœ“ governance, fast BI βœ— expensive per TB βœ— proprietary storage format βœ— poor for ML, images, logs 1990s β†’ Teradata, Oracle Data lake βœ“ cheap object storage βœ“ any format, open files βœ— no transactions: half- written files, no updates βœ— weak governance β†’ "swamp" 2010s β†’ Hadoop, S3 + Spark Lakehouse βœ“ open files (Parquet) on cheap object storage βœ“ + a transaction log (Delta) βœ“ + governance (Unity Catalog) βœ“ SQL, BI, streaming and ML on the SAME copy of data 2020s β†’ Databricks, others The lakehouse keeps the lake's cheap, open storage and adds the warehouse's reliability on top, so you stop copying data between the two.

2. What Databricks adds beyond open-source Spark

NeedOpen-source Spark aloneDatabricksLesson
Reliable tablesDelta Lake (open source) you set up yourselfDelta by default, plus liquid clustering, predictive optimisation07
Who can see whatwhatever your storage permissions allowUnity Catalog: grants, lineage, audit, one catalog for everything06
Computeyou run and patch clusters (YARN, Kubernetes)managed clusters, serverless, SQL warehouses04
Speedthe JVM Spark enginePhoton, a native C++ engine for SQL/DataFrames04, 15
Ingestionwrite your own file trackingAuto Loader, Lakeflow Connect08
Pipelineshand-written jobsLakeflow Spark Declarative Pipelines with data-quality expectations10
OrchestrationAirflow or cronLakeflow Jobs12
AnalystsSpark Thrift serverDatabricks SQL, AI/BI dashboards, Genie11
Machine learninginstall MLflow yourselfmanaged MLflow, feature tables, model serving14
Shipping codeyour own toolingGit folders, CLI, Declarative Automation Bundles17
Names change; concepts do not

Databricks renames products often. Delta Live Tables became Lakeflow Declarative Pipelines and then Lakeflow Spark Declarative Pipelines once the framework was contributed to Apache Spark 4.1. Workflows became Lakeflow Jobs. Databricks Asset Bundles became Declarative Automation Bundles. This course uses the names in the documentation at the time of writing (2026) and mentions the older ones, because blog posts, certification material and job adverts still use them.

3. What stays exactly the same

The code you write is ordinary PySpark and SQL. Everything from the PySpark course applies unchanged: the driver/executor model, lazy evaluation, the Catalyst optimiser, shuffles, partitions and the Spark UI. A Databricks notebook simply gives you a ready spark session.

# `spark` already exists in every Databricks notebook; no SparkSession.builder needed
trips = spark.read.table("samples.nyctaxi.trips")          # a sample dataset every workspace has
(trips
   .groupBy("pickup_zip")
   .agg({"fare_amount": "avg", "*": "count"})
   .orderBy("count(1)", ascending=False)
   .limit(5)
   .display())                                               # display(): Databricks' rich table/chart output
About the code in this course

Databricks code runs on Databricks, not on a laptop, so these examples were not executed while the course was written. They follow the current documentation, and any output shown is labelled as illustrative. Free Edition (lesson 03) lets you run every example yourself at no cost.

4. Who uses what

Data engineers

Ingestion, pipelines, jobs, Unity Catalog, performance and cost. Most of this course.

Analysts & BI

SQL warehouses, the SQL editor, dashboards, alerts; connecting Power BI or Tableau.

Data scientists & ML engineers

Notebooks, MLflow experiments, feature tables, model serving, vector search for AI apps.

Recap

  • Databricks = managed Spark + the platform around it: storage format, governance, compute, orchestration, SQL, ML, DevOps.
  • Lakehouse = open files on cheap storage + a transaction log + governance, serving every workload from one copy.
  • Your PySpark and SQL skills transfer directly; Databricks adds services around them.
  • Expect renames; learn the concept, then map the current name.

Checkpoint

1 Β· What does the lakehouse add to a plain data lake?
The files stay open and cheap. The transaction log provides ACID updates, time travel and schema enforcement, and the catalog provides governance, which is what lakes lacked.
2 Β· You know PySpark. What changes when you move to Databricks?
Databricks runs Apache Spark (and the Photon engine, which accepts the same APIs). The execution model you learned still explains performance there.