Community/Tools

Ollama: run open LLMs locally for development, privacy and offline work

Ollama makes running open models on your own machine a one-command job and exposes them through a local API. Here is the workflow, how to call it from code, and what to expect on real hardware.

Author's connection to this tool: No connection. An independent overview written by the RecallRun editors.

Calling a hosted LLM API is the fastest way to build an AI feature, but it isn't always the right one. Sometimes the data can't leave the machine, you want to develop offline, or you'd like to experiment without a usage bill. Ollama is an open-source tool that downloads, runs and serves open-weight models locally with very little setup.

Install and run a model

Install Ollama for macOS, Windows or Linux from the official site, then:

ollama pull llama3.2        # download a model
ollama run llama3.2         # chat in the terminal
ollama list                 # models on this machine

Ollama's model library includes many open-weight families in several sizes, usually already quantised so they fit in normal amounts of memory. Each model page lists its licence; check it before using a model in a product.

Call it from code

Ollama runs a local server (by default on port 11434) with a simple HTTP API:

curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2",
  "messages": [{"role": "user", "content": "Explain idempotency keys in two sentences."}],
  "stream": false
}'

There's an official Python client:

import ollama

resp = ollama.chat(
    model="llama3.2",
    messages=[{"role": "user", "content": "Write a regex that matches ISO dates."}],
)
print(resp["message"]["content"])

Ollama also offers an OpenAI-compatible endpoint, so many existing tools and SDKs can point at it by changing the base URL:

from openai import OpenAI

client = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")  # key is ignored
r = client.chat.completions.create(model="llama3.2", messages=[{"role": "user", "content": "Hi"}])

That makes it easy to develop against a local model and switch to a hosted one in production with configuration alone.

Embeddings, too

Many embedding models run under Ollama as well, which is useful for building a local RAG prototype:

emb = ollama.embed(model="nomic-embed-text", input=["first chunk", "second chunk"])
vectors = emb["embeddings"]

Customise with a Modelfile

A Modelfile packages a base model with a system prompt and parameters:

FROM llama3.2
PARAMETER temperature 0.2
SYSTEM """You are a concise code reviewer. Point out bugs first, style last."""
ollama create reviewer -f Modelfile
ollama run reviewer

What to expect on real hardware

  • Memory is the main limit. As a rough guide, small models (around 1 to 4 billion parameters) run comfortably on most laptops; 7 to 8 billion parameter models want around 8 GB of free memory or more; larger ones need serious GPUs or lots of unified memory.
  • A GPU or Apple Silicon makes a big difference to speed. CPU-only works but is slow for long answers.
  • Local models are smaller than frontier hosted models. Expect weaker reasoning and more mistakes on complex tasks. They shine at focused jobs: classification, extraction, summarising, drafting and code completion.

Good uses

  • Developing and testing LLM features offline, or in CI without API keys.
  • Processing sensitive documents that can't be sent to a third party.
  • Cheap batch jobs where a smaller model is accurate enough.
  • Trying many prompts and models quickly without cost.

Things to keep in mind

  • The server listens on localhost by default. If you expose it on a network, put authentication in front of it; the API itself has none.
  • Pin model versions (tags) in your code so behaviour doesn't change when a model is updated.
  • Evaluate on your own task before trusting a model: write 20 to 50 examples and compare outputs.

Verdict

Ollama is the easiest way to get an open model running locally and callable from code. Use it for development, privacy-sensitive work and focused tasks, and measure quality honestly against a hosted model before deciding which one belongs in production.

Written by

RecallRun Editors

Practical guides and independent tool overviews from the RecallRun team. Every post is written to be tested on your own machine.

Website

Written by RecallRun Editors for the RecallRun community. Community posts are checked for safety and reviewed by our editors before publishing, but the views and claims are the author's own. Links are the author's; open them with care. Report this post.

More from the community

Write for RecallRun

Share a tech article or a tool you built. Every post is checked and reviewed before it goes live.

Start writing