Getting started & the workspace
Databricks Free Edition is a free, serverless workspace for learning, with no cloud account and no credit card. It has limits (below), but it runs everything in this course: notebooks, Unity Catalog, volumes, pipelines, jobs, SQL and dashboards. This lesson gets you from sign-up to a query over your own uploaded file.
1. Sign up and look around
- Go to the Databricks website and choose Free Edition ("Try Databricks Free Edition"). Sign up with an email, Google or Microsoft account.
- You land in a workspace. The left sidebar is your map: Workspace (files and notebooks), Catalog (data), Jobs & Pipelines, Compute, SQL Editor, Dashboards, Genie and more. Exact labels move occasionally, but the sections are the same.
- Open Catalog. You will see a
workspacecatalog you can write to, and the read-onlysamplescatalog with ready-made datasets.
| Free Edition limit (per the documentation) | What it means for this course |
|---|---|
| Serverless compute only; no custom clusters | You will not configure classic clusters (lesson 04 explains them anyway) |
| One SQL warehouse, size 2X-Small | Fine for the project's dashboards |
| Up to 5 concurrent job tasks per account | Keep jobs small; the project's job has 3 tasks |
| One active pipeline per pipeline type | Run one declarative pipeline at a time |
| One workspace, one metastore; Python and SQL (no R or Scala) | Everything here is Python and SQL |
| Restricted outbound internet; no commercial use | Upload files instead of downloading from the internet |
2. Your first notebook
Workspace → Create → Notebook. Attach it to Serverless compute (the default in Free Edition) and run a cell with Shift+Enter.
df = spark.read.table("samples.tpch.orders")
print(df.count())
display(df.limit(5)) # an interactive table with sorting, charts and download
%sql SELECT o_orderpriority, COUNT(*) AS orders, ROUND(AVG(o_totalprice), 2) AS avg_price FROM samples.tpch.orders GROUP BY o_orderpriority ORDER BY orders DESC;
3. Bring your own data: a volume
CREATE SCHEMA IF NOT EXISTS workspace.retail; CREATE VOLUME IF NOT EXISTS workspace.retail.raw;
Then in Catalog → workspace → retail → raw, choose Upload to this volume and upload
customers.csv from the PySpark course's project/data folder.
path = "/Volumes/workspace/retail/raw/customers.csv"
customers = (spark.read
.option("header", True)
.option("inferSchema", True)
.csv(path))
customers.printSchema()
customers.write.mode("overwrite").saveAsTable("workspace.retail.customers")
display(spark.sql("SELECT city, COUNT(*) AS n FROM workspace.retail.customers GROUP BY city ORDER BY n DESC"))
You never typed s3:// or an access key. Volumes and managed tables are governed locations:
Unity Catalog decides where the files physically live and who may read them. On a company workspace
that is exactly how it should stay (lesson 06).
4. The workspace, section by section
| Section | Use it for | Lesson |
|---|---|---|
| Workspace | Notebooks, files, Git folders (your repositories) | 05, 17 |
| Catalog | Browse catalogs, schemas, tables, volumes; permissions; lineage | 06 |
| Jobs & Pipelines | Lakeflow Jobs and declarative pipelines: schedules, runs, repairs | 10, 12 |
| Compute | Clusters, SQL warehouses, policies (limited in Free Edition) | 04 |
| SQL Editor, Queries, Dashboards, Alerts, Genie | Analytics and BI | 11 |
| Experiments, Models, Serving | MLflow and model endpoints | 14 |
Recap
- Free Edition: free, serverless, enough for the whole course; know its limits.
- Notebook cells can be Python or SQL (
%sql);sparkanddisplay()are ready. - Files → volumes (
/Volumes/catalog/schema/volume/…); tables → schemas. saveAsTablecreates a managed Delta table that every tool can query.