Why Spark exists
You cannot understand Spark's design until you feel the problem it was built for. This lesson has no code. It gives you the three constraints — size, coordination and failure — that explain every single design decision in the rest of the course.
After this lesson you can…
- Explain why adding RAM to one machine eventually stops working.
- Say what Spark is and, more importantly, what it is not.
- Describe why Spark replaced MapReduce in one sentence that is actually correct.
- Decide, honestly, whether a given job should use Spark at all.
1. The problem: data outgrew the machine
The one-chef kitchen
One excellent chef can serve 40 covers a night. For 80, you buy a bigger stove and a faster knife — that works for a while. For 4,000 covers you cannot buy a chef who is 100× faster, because that chef does not exist. You have to hire 100 chefs, split the menu between them, and accept a completely new problem: keeping 100 chefs coordinated. That trade — from "a faster one" to "many cooperating ones" — is the entire subject of this course.
Put numbers on it
Hardware has a ceiling and it is closer than people think. A fast NVMe SSD reads roughly 2 GB/s. A good cloud instance has maybe 256 GB of RAM. So on one machine:
To read 1 TB from one local SSD at 2 GB/s — before you compute anything at all.
To read 100 TB the same way. Your daily job now takes longer than a day.
What pandas.read_csv() returns on a 400 GB file with 256 GB of RAM. It raises MemoryError.
Now split that 100 TB across 200 machines. Each one reads 500 GB locally, in parallel, in about four minutes. The work did not get smaller — it got divided. That is the only trick. Everything else in Spark is bookkeeping to make the trick safe.
2. The three hard problems of going distributed
The moment your data lives on more than one machine, three problems appear that never existed before. Every Spark concept in this course is an answer to one of them. Learn the three names and you have a filing cabinet for everything else.
Problem 1 is easy and Problem 3 is handled for you. Problem 2 is where all your pain will come from. Moving data between machines is thousands of times slower than reading it locally, and almost every slow Spark job is slow because it moves too much data. Keep that sentence.
3. What came before: MapReduce and its tax
Hadoop MapReduce (2004) solved all three problems and made large-scale processing genuinely possible. Its design choice was to be maximally paranoid: after every single step, write the complete intermediate result to disk (in HDFS, replicated three times). If a machine dies, the next step reads that file and continues. Extremely robust. Extremely slow.
The problem is that real analytics is never one step. A join followed by a group-by followed by a filter is three MapReduce jobs, and each one pays the full disk-write-and-replicate tax. Iterative work — machine learning, graph algorithms — is dozens of steps, so it pays that tax dozens of times on the same data.
"Spark is faster because it is in-memory" is the answer that gets you marked down. Say instead: "Spark avoids materialising intermediate results to replicated disk between stages, and it optimises the whole multi-step plan before running it. In-memory caching helps most when the same data is reused, as in iterative ML." That is the correct, complete answer.
4. What Spark is — and what it is not
Spark is a work-splitter. You hand it a description of a calculation over a large pile of data. It figures out how to cut the pile into pieces, sends a copy of your instructions to a fleet of computers, makes them do their piece, moves data between them when two pieces need to meet, and hands you back one answer. It does not own the data and it does not remember anything after the job ends.
Spark is a distributed, general-purpose, in-memory data processing engine. It exposes a functional, lazily-evaluated API over a partitioned collection abstraction; a query optimiser (Catalyst) compiles that API into a physical plan of stages; a DAG scheduler splits stages at shuffle boundaries; and a task scheduler dispatches one task per partition to executor JVMs, with automatic retry and recomputation from lineage on failure.
The four things Spark is not
| Spark is NOT… | Because… | What you actually use |
|---|---|---|
| A storage system | Spark has no filesystem and no persistent tables of its own. Kill the cluster and nothing remains. | S3, ADLS, GCS, HDFS + Parquet/Delta |
| A database | No indexes, no primary keys, no row-level updates, no millisecond point lookups. A scan engine, not a lookup engine. | Postgres, Cassandra, DynamoDB |
| Always faster | Starting a cluster costs 10–60s and every shuffle costs network. Under a few GB, pandas or DuckDB usually wins outright. | pandas, Polars, DuckDB |
| A scheduler | Spark runs a job. It does not decide that the job runs at 02:00 daily, retry it tomorrow, or alert you. | Airflow, Dagster, Databricks Jobs |
5. When you should not use Spark
Reaching for Spark when you do not need it is the most common and most expensive mistake in data engineering. Be honest with this checklist.
🚫 Don't use Spark when…
Your data fits comfortably in one machine's RAM (roughly < 20–50 GB) · You need sub-second responses to single-row lookups · You are doing a lot of small, frequent updates · The job is a one-off you will run twice · Your team has no one who can debug a cluster at 3am.
✅ Do use Spark when…
Data genuinely exceeds one machine, or grows toward it · You need the same code to run over 1 GB today and 10 TB next year · You are joining several large datasets · You need batch and streaming with one codebase · You need ML over data too big to collect locally.
If a single modern laptop can hold your data in memory, a single laptop will beat a Spark cluster — every time. Spark's break-even point is roughly where data stops fitting in RAM comfortably. Below that you are paying coordination cost for no parallelism benefit.
6. One engine, four libraries
Spark's other winning idea was unification. Before Spark you needed separate systems for SQL, streaming and ML, each with its own API, cluster and operational burden. Spark put them all on one core so they share one optimiser, one memory model and one deployment — and so that a DataFrame produced by streaming can be joined to one produced by batch, with no conversion.
A short, honest history — useful context for interviews
- 2004 — Google publishes the MapReduce paper; Hadoop implements it open-source.
- 2009–2010 — Matei Zaharia starts Spark at UC Berkeley's AMPLab, targeting the iterative workloads MapReduce handled badly.
- 2012 — The RDD paper formalises lineage-based fault tolerance: remember the recipe rather than replicating the result.
- 2014 — Spark 1.0; Spark SQL and the DataFrame API arrive shortly after, moving users off raw RDDs.
- 2016 — Spark 2.0: Datasets, Structured Streaming, and whole-stage code generation (Tungsten).
- 2020 — Spark 3.0: Adaptive Query Execution and dynamic partition pruning — the engine starts re-planning itself at runtime (lesson 23).
- 2022+ — Spark Connect decouples the client from the driver; the lakehouse (Delta, Iceberg) becomes the default storage layer (lesson 26).
The trend across all of it: push more decisions from the user into the engine. Early Spark made you hand-optimise RDDs. Modern Spark re-plans your query while it runs.
Recap
- One machine has a ceiling. Past it, the only option is dividing work across many machines.
- Distribution creates three problems: splitting (partitions), coordination (shuffle), failure (lineage).
- Coordination is the expensive one. Almost every slow job moves too much data across the network.
- Spark beat MapReduce by not materialising every intermediate step to replicated disk, and by optimising the whole plan at once — not merely by "being in-memory".
- Spark is compute only. No storage, no database, no scheduler — and not faster on small data.
- One core engine, four libraries. SQL, streaming, ML and graph share Catalyst and Spark Core.