Module 1 · Foundations & Architecture

Why Spark exists

Beginner 12 min read No setup needed

You cannot understand Spark's design until you feel the problem it was built for. This lesson has no code. It gives you the three constraints — size, coordination and failure — that explain every single design decision in the rest of the course.

After this lesson you can…

  • Explain why adding RAM to one machine eventually stops working.
  • Say what Spark is and, more importantly, what it is not.
  • Describe why Spark replaced MapReduce in one sentence that is actually correct.
  • Decide, honestly, whether a given job should use Spark at all.

1. The problem: data outgrew the machine

🍳

The one-chef kitchen

One excellent chef can serve 40 covers a night. For 80, you buy a bigger stove and a faster knife — that works for a while. For 4,000 covers you cannot buy a chef who is 100× faster, because that chef does not exist. You have to hire 100 chefs, split the menu between them, and accept a completely new problem: keeping 100 chefs coordinated. That trade — from "a faster one" to "many cooperating ones" — is the entire subject of this course.

Put numbers on it

Hardware has a ceiling and it is closer than people think. A fast NVMe SSD reads roughly 2 GB/s. A good cloud instance has maybe 256 GB of RAM. So on one machine:

~8 min

To read 1 TB from one local SSD at 2 GB/s — before you compute anything at all.

~13 hrs

To read 100 TB the same way. Your daily job now takes longer than a day.

0 rows

What pandas.read_csv() returns on a 400 GB file with 256 GB of RAM. It raises MemoryError.

Now split that 100 TB across 200 machines. Each one reads 500 GB locally, in parallel, in about four minutes. The work did not get smaller — it got divided. That is the only trick. Everything else in Spark is bookkeeping to make the trick safe.

Scale UP (vertical) buy a bigger machine — simple, and it runs out 8 GB 64 GB 256 GB 4 TB? exists, costs a fortune Hard ceiling: one motherboard, one memory bus, one failure = total loss Scale OUT (horizontal) buy more ordinary machines — cheap, and hard 16 GB 16 GB 16 GB 16 GB 16 GB 16 GB 16 GB DEAD 16 GB 16 GB No ceiling: add machines for more capacity… …but now you must handle splitting, coordinating and dying nodes Spark's job in one line Give you the programming comfort of the left-hand picture, with the capacity of the right-hand one. You write code as if there is one dataset. Spark splits it, schedules it, and recovers it when a box dies.
Scaling up is easy and finite. Scaling out is unbounded and difficult. Spark exists to hide that difficulty behind an API that looks like it is operating on a single table.

2. The three hard problems of going distributed

The moment your data lives on more than one machine, three problems appear that never existed before. Every Spark concept in this course is an answer to one of them. Learn the three names and you have a filing cabinet for everything else.

PROBLEM 1 — Splitting How do you cut the data up? one 100 TB dataset Spark's answer: partitions lessons 07, 19 PROBLEM 2 — Coordination Rows that must meet are apart. A:1 B:2 A:3 must move over the network Spark's answer: the shuffle lessons 06, 19, 23 PROBLEM 3 — Failure 200 machines × 4 hours = one dies. ok ok lost recompute it Spark's answer: lineage + retries lessons 07, 20
Splitting, coordination, failure. When you meet a new Spark concept later, ask which of these three it is solving. Partitions and formats solve (1). Shuffles, joins and windows solve (2). Lineage, retries and checkpoints solve (3).
The one that will bite you

Problem 1 is easy and Problem 3 is handled for you. Problem 2 is where all your pain will come from. Moving data between machines is thousands of times slower than reading it locally, and almost every slow Spark job is slow because it moves too much data. Keep that sentence.

3. What came before: MapReduce and its tax

Hadoop MapReduce (2004) solved all three problems and made large-scale processing genuinely possible. Its design choice was to be maximally paranoid: after every single step, write the complete intermediate result to disk (in HDFS, replicated three times). If a machine dies, the next step reads that file and continues. Extremely robust. Extremely slow.

The problem is that real analytics is never one step. A join followed by a group-by followed by a filter is three MapReduce jobs, and each one pays the full disk-write-and-replicate tax. Iterative work — machine learning, graph algorithms — is dozens of steps, so it pays that tax dozens of times on the same data.

MapReduce — disk between every step input step 1 step 2 step 3 result 4 × (write to disk + replicate 3× + read back) = the tax Spark — memory between steps, disk only when it must input step 1 step 2 step 3 result intermediates stay in executor memory — no round trip 10–100× faster on multi-step and iterative workloads
This is the whole "Spark is faster than MapReduce" story. Not magic — Spark simply keeps intermediate results in memory and only spills to disk when memory runs out. The fault tolerance that MapReduce bought with disk writes, Spark gets instead from lineage: it remembers the recipe and recomputes a lost partition (lesson 07).
Interview trap

"Spark is faster because it is in-memory" is the answer that gets you marked down. Say instead: "Spark avoids materialising intermediate results to replicated disk between stages, and it optimises the whole multi-step plan before running it. In-memory caching helps most when the same data is reused, as in iterative ML." That is the correct, complete answer.

4. What Spark is — and what it is not

Spark is a work-splitter. You hand it a description of a calculation over a large pile of data. It figures out how to cut the pile into pieces, sends a copy of your instructions to a fleet of computers, makes them do their piece, moves data between them when two pieces need to meet, and hands you back one answer. It does not own the data and it does not remember anything after the job ends.

Spark is a distributed, general-purpose, in-memory data processing engine. It exposes a functional, lazily-evaluated API over a partitioned collection abstraction; a query optimiser (Catalyst) compiles that API into a physical plan of stages; a DAG scheduler splits stages at shuffle boundaries; and a task scheduler dispatches one task per partition to executor JVMs, with automatic retry and recomputation from lineage on failure.

The four things Spark is not

Spark is NOT…Because…What you actually use
A storage systemSpark has no filesystem and no persistent tables of its own. Kill the cluster and nothing remains.S3, ADLS, GCS, HDFS + Parquet/Delta
A databaseNo indexes, no primary keys, no row-level updates, no millisecond point lookups. A scan engine, not a lookup engine.Postgres, Cassandra, DynamoDB
Always fasterStarting a cluster costs 10–60s and every shuffle costs network. Under a few GB, pandas or DuckDB usually wins outright.pandas, Polars, DuckDB
A schedulerSpark runs a job. It does not decide that the job runs at 02:00 daily, retry it tomorrow, or alert you.Airflow, Dagster, Databricks Jobs

5. When you should not use Spark

Reaching for Spark when you do not need it is the most common and most expensive mistake in data engineering. Be honest with this checklist.

🚫 Don't use Spark when…

Your data fits comfortably in one machine's RAM (roughly < 20–50 GB) · You need sub-second responses to single-row lookups · You are doing a lot of small, frequent updates · The job is a one-off you will run twice · Your team has no one who can debug a cluster at 3am.

✅ Do use Spark when…

Data genuinely exceeds one machine, or grows toward it · You need the same code to run over 1 GB today and 10 TB next year · You are joining several large datasets · You need batch and streaming with one codebase · You need ML over data too big to collect locally.

A useful rule of thumb

If a single modern laptop can hold your data in memory, a single laptop will beat a Spark cluster — every time. Spark's break-even point is roughly where data stops fitting in RAM comfortably. Below that you are paying coordination cost for no parallelism benefit.

6. One engine, four libraries

Spark's other winning idea was unification. Before Spark you needed separate systems for SQL, streaming and ML, each with its own API, cluster and operational burden. Spark put them all on one core so they share one optimiser, one memory model and one deployment — and so that a DataFrame produced by streaming can be joined to one produced by batch, with no conversion.

Spark SQL DataFrames · SQL · Catalog Structured Streaming same API, unbounded input MLlib pipelines · distributed fit GraphX / GraphFrames graph algorithms (niche) Catalyst optimiser + Tungsten execution engine one query planner and one memory/codegen layer serving all four libraries — lesson 18 Spark Core — RDDs, DAG scheduler, task scheduler, shuffle, memory manager everything above compiles down to this — lessons 02, 06, 07, 19, 21
You will spend 90% of your time in the top-left box (Spark SQL / DataFrames) and 90% of your debugging time in the bottom box (Spark Core). That is exactly why this course covers both, and why the next lesson goes straight to the architecture underneath.
A short, honest history — useful context for interviews
  • 2004 — Google publishes the MapReduce paper; Hadoop implements it open-source.
  • 2009–2010 — Matei Zaharia starts Spark at UC Berkeley's AMPLab, targeting the iterative workloads MapReduce handled badly.
  • 2012 — The RDD paper formalises lineage-based fault tolerance: remember the recipe rather than replicating the result.
  • 2014 — Spark 1.0; Spark SQL and the DataFrame API arrive shortly after, moving users off raw RDDs.
  • 2016 — Spark 2.0: Datasets, Structured Streaming, and whole-stage code generation (Tungsten).
  • 2020 — Spark 3.0: Adaptive Query Execution and dynamic partition pruning — the engine starts re-planning itself at runtime (lesson 23).
  • 2022+ — Spark Connect decouples the client from the driver; the lakehouse (Delta, Iceberg) becomes the default storage layer (lesson 26).

The trend across all of it: push more decisions from the user into the engine. Early Spark made you hand-optimise RDDs. Modern Spark re-plans your query while it runs.

Recap

  • One machine has a ceiling. Past it, the only option is dividing work across many machines.
  • Distribution creates three problems: splitting (partitions), coordination (shuffle), failure (lineage).
  • Coordination is the expensive one. Almost every slow job moves too much data across the network.
  • Spark beat MapReduce by not materialising every intermediate step to replicated disk, and by optimising the whole plan at once — not merely by "being in-memory".
  • Spark is compute only. No storage, no database, no scheduler — and not faster on small data.
  • One core engine, four libraries. SQL, streaming, ML and graph share Catalyst and Spark Core.

Checkpoint

1 · Your team processes 12 GB of CSV once a day and it takes 4 minutes on one server. A colleague proposes moving it to a 20-node Spark cluster. What is the best response?
Spark's break-even point is roughly where data stops fitting comfortably in one machine's memory. At 12 GB you pay cluster startup, scheduling and network cost to parallelise work that was never the bottleneck. The right answer is to revisit this when the data grows, not before.
2 · Which statement about Spark vs MapReduce is the one an interviewer wants to hear?
"All data in RAM" is false — Spark streams through partitions and spills to disk when needed. The real wins are (a) no replicated disk materialisation between stages and (b) whole-plan optimisation. Fault tolerance is not removed, it is achieved differently, through lineage.
3 · Of the three hard problems of distribution, which one causes most real-world Spark performance problems?
Network transfer is orders of magnitude slower than local reads, and joins, group-bys and windows all trigger it. Splitting is mostly automatic, and failure is handled transparently by lineage and retries. Coordination — the shuffle — is where your time and money go, which is why lessons 19 and 23 exist.