Ollama: run open LLMs locally for development, privacy and offline work
Ollama makes running open models on your own machine a one-command job and exposes them through a local API. Here is the workflow, how to call it from code, and what to expect on real hardware.
Author's connection to this tool: No connection. An independent overview written by the RecallRun editors.
Calling a hosted LLM API is the fastest way to build an AI feature, but it isn't always the right one. Sometimes the data can't leave the machine, you want to develop offline, or you'd like to experiment without a usage bill. Ollama is an open-source tool that downloads, runs and serves open-weight models locally with very little setup.
Install and run a model
Install Ollama for macOS, Windows or Linux from the official site, then:
ollama pull llama3.2 # download a model
ollama run llama3.2 # chat in the terminal
ollama list # models on this machine
Ollama's model library includes many open-weight families in several sizes, usually already quantised so they fit in normal amounts of memory. Each model page lists its licence; check it before using a model in a product.
Call it from code
Ollama runs a local server (by default on port 11434) with a simple HTTP API:
curl http://localhost:11434/api/chat -d '{
"model": "llama3.2",
"messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}],
"stream": false
}'
There's an official Python client:
import ollama
resp = ollama.chat(
model="llama3.2",
messages=[{"role": "user", "content": "Write a regex that matches ISO dates."}],
)
print(resp["message"]["content"])
Ollama also offers an OpenAI-compatible endpoint, so many existing tools and SDKs can point at it by changing the base URL:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama") # key is ignored
r = client.chat.completions.create(model="llama3.2", messages=[{"role": "user", "content": "Hi"}])
That makes it easy to develop against a local model and switch to a hosted one in production with configuration alone.
Embeddings, too
Many embedding models run under Ollama as well, which is useful for building a local RAG prototype:
emb = ollama.embed(model="nomic-embed-text", input=["first chunk", "second chunk"])
vectors = emb["embeddings"]
Customise with a Modelfile
A Modelfile packages a base model with a system prompt and parameters:
FROM llama3.2
PARAMETER temperature 0.2
SYSTEM """You are a concise code reviewer. Point out bugs first, style last."""
ollama create reviewer -f Modelfile
ollama run reviewer
What to expect on real hardware
- Memory is the main limit. As a rough guide, small models (around 1 to 4 billion parameters) run comfortably on most laptops; 7 to 8 billion parameter models want around 8 GB of free memory or more; larger ones need serious GPUs or lots of unified memory.
- A GPU or Apple Silicon makes a big difference to speed. CPU-only works but is slow for long answers.
- Local models are smaller than frontier hosted models. Expect weaker reasoning and more mistakes on complex tasks. They shine at focused jobs: classification, extraction, summarising, drafting and code completion.
Good uses
- Developing and testing LLM features offline, or in CI without API keys.
- Processing sensitive documents that can't be sent to a third party.
- Cheap batch jobs where a smaller model is accurate enough.
- Trying many prompts and models quickly without cost.
Things to keep in mind
- The server listens on localhost by default. If you expose it on a network, put authentication in front of it; the API itself has none.
- Pin model versions (tags) in your code so behaviour doesn't change when a model is updated.
- Evaluate on your own task before trusting a model: write 20 to 50 examples and compare outputs.
Verdict
Ollama is the easiest way to get an open model running locally and callable from code. Use it for development, privacy-sensitive work and focused tasks, and measure quality honestly against a hosted model before deciding which one belongs in production.
Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.
More from the community
- Tech articles
Prompt injection: practical defences for LLM apps that read untrusted text
Any LLM feature that reads emails, web pages or user uploads can be steered by text hidden inside them. Here is a layered set of defences that work in production, and why no single one is enough.
- Tech articles
Vector embeddings explained for developers: similarity, normalisation and the mistakes that hurt search
What an embedding really is, how cosine similarity and dot product relate, why you should normalise, and the practical mistakes that quietly make semantic search worse.
- Tools
Pydantic v2: validate data at the edges of your Python application
Pydantic turns type hints into fast runtime validation and serialisation. Here are the core patterns for API payloads, settings and LLM outputs, plus the v2 changes that trip people up.
Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.