Module 1 · The platform

Getting started & the workspace

Beginner 14 min read From sign-up to your first query

Databricks Free Edition is a free, serverless workspace for learning, with no cloud account and no credit card. It has limits (below), but it runs everything in this course: notebooks, Unity Catalog, volumes, pipelines, jobs, SQL and dashboards. This lesson gets you from sign-up to a query over your own uploaded file.

1. Sign up and look around

  1. Go to the Databricks website and choose Free Edition ("Try Databricks Free Edition"). Sign up with an email, Google or Microsoft account.
  2. You land in a workspace. The left sidebar is your map: Workspace (files and notebooks), Catalog (data), Jobs & Pipelines, Compute, SQL Editor, Dashboards, Genie and more. Exact labels move occasionally, but the sections are the same.
  3. Open Catalog. You will see a workspace catalog you can write to, and the read-only samples catalog with ready-made datasets.
Free Edition limit (per the documentation)What it means for this course
Serverless compute only; no custom clustersYou will not configure classic clusters (lesson 04 explains them anyway)
One SQL warehouse, size 2X-SmallFine for the project's dashboards
Up to 5 concurrent job tasks per accountKeep jobs small; the project's job has 3 tasks
One active pipeline per pipeline typeRun one declarative pipeline at a time
One workspace, one metastore; Python and SQL (no R or Scala)Everything here is Python and SQL
Restricted outbound internet; no commercial useUpload files instead of downloading from the internet

2. Your first notebook

Workspace → Create → Notebook. Attach it to Serverless compute (the default in Free Edition) and run a cell with Shift+Enter.

df = spark.read.table("samples.tpch.orders")
print(df.count())
display(df.limit(5))            # an interactive table with sorting, charts and download
%sql
SELECT o_orderpriority, COUNT(*) AS orders, ROUND(AVG(o_totalprice), 2) AS avg_price
FROM samples.tpch.orders
GROUP BY o_orderpriority
ORDER BY orders DESC;

3. Bring your own data: a volume

customers.csv on your laptop upload Volume (governed files) /Volumes/workspace/retail/raw/ catalog / schema / volume / path read + save Delta table workspace.retail.customers queryable from SQL, BI, jobs Files go in volumes; tables go in schemas. Both are governed by Unity Catalog (lesson 06).
CREATE SCHEMA IF NOT EXISTS workspace.retail;
CREATE VOLUME IF NOT EXISTS workspace.retail.raw;

Then in Catalog → workspace → retail → raw, choose Upload to this volume and upload customers.csv from the PySpark course's project/data folder.

path = "/Volumes/workspace/retail/raw/customers.csv"
customers = (spark.read
             .option("header", True)
             .option("inferSchema", True)
             .csv(path))
customers.printSchema()
customers.write.mode("overwrite").saveAsTable("workspace.retail.customers")

display(spark.sql("SELECT city, COUNT(*) AS n FROM workspace.retail.customers GROUP BY city ORDER BY n DESC"))
No paths to storage buckets

You never typed s3:// or an access key. Volumes and managed tables are governed locations: Unity Catalog decides where the files physically live and who may read them. On a company workspace that is exactly how it should stay (lesson 06).

4. The workspace, section by section

SectionUse it forLesson
WorkspaceNotebooks, files, Git folders (your repositories)05, 17
CatalogBrowse catalogs, schemas, tables, volumes; permissions; lineage06
Jobs & PipelinesLakeflow Jobs and declarative pipelines: schedules, runs, repairs10, 12
ComputeClusters, SQL warehouses, policies (limited in Free Edition)04
SQL Editor, Queries, Dashboards, Alerts, GenieAnalytics and BI11
Experiments, Models, ServingMLflow and model endpoints14

Recap

  • Free Edition: free, serverless, enough for the whole course; know its limits.
  • Notebook cells can be Python or SQL (%sql); spark and display() are ready.
  • Files → volumes (/Volumes/catalog/schema/volume/…); tables → schemas.
  • saveAsTable creates a managed Delta table that every tool can query.

Checkpoint

1 · Where should an uploaded CSV live in a Unity Catalog workspace?
Volumes are Unity Catalog's governed home for files: permissions, lineage and auditing apply, with no credentials in code.
2 · Why can't you create a cluster with 16 GPU nodes in Free Edition?
Free Edition is serverless-only by design. Classic clusters with chosen instance types exist in paid workspaces (lesson 04).