Welcome to HowToShipIt — practical how-to guides for developers: code, AI tools, and servers, explained step by step.

How to Run Ollama: Install, Models, and API in 30 Minutes

If you want to know how to run Ollama on your own machine, this guide gets you there in about 30 minutes: install, your first model, chat in the terminal, custom models with Modelfiles, and calling the API from Python. No API keys, no usage bills, no sending your code to someone else’s servers.

Ollama is a single tool that downloads, runs, and manages open-weight large language models on your laptop, workstation, or home server. Think of it as Docker for LLMs: you pull models by name, run them with one command, and get a local HTTP API your code can call. It is free, it works on macOS, Linux, and Windows, and the commands below are exactly what working developers use every day.

Last verified: 2 October 2026. All install commands, API endpoints, and configuration were checked against the official Ollama repository and documentation on the date of publication.

Why run Ollama instead of just calling an API

Cloud APIs are great, but local models win in four situations:

  • Privacy: the prompt and the answer never leave your machine. Client code, unreleased work, personal notes — none of it goes over the wire.
  • Cost: after the hardware you already own, every token is free. Experiment all day without watching a usage dashboard.
  • Offline work: planes, spotty hotel Wi-Fi, locked-down corporate networks. Your model still works.
  • Deterministic environments: pin an exact model version and get the same behaviour every time, unlike hosted APIs that change silently.

It is not a replacement for frontier models. A 3B-parameter model on your laptop will not out-reason a 2026 flagship cloud model. But for summarising, drafting, regex generation, code review, log parsing, and prototyping, it is shockingly good — and free.

What you need before you start

If your computer runs a modern browser, it can probably run a small model. You need:

  • macOS, Linux, or Windows (a recent CPU; GPUs speed things up but are optional)
  • A few GB of free disk space per model
  • Enough RAM for the model size you pick

The practical rule of thumb: a quantised model needs roughly its parameter count in gigabytes of RAM. A 3B model needs ~2–3 GB, an 8B model needs ~4–6 GB, a 32B model needs ~20 GB, and a 70B model needs ~40 GB. So an 8 GB laptop comfortably runs 3B models and small 7B ones; a 16 GB machine handles 8B models; 32B+ models need serious RAM or a GPU with lots of VRAM. Check the exact download size for each tag on ollama.com/library before pulling.

Step 1: Install Ollama

macOS: the fastest route is Homebrew.

brew install ollama

Alternatively, download the macOS app from ollama.com/download, or use the install script:

curl -fsSL https://ollama.com/install.sh | sh

Linux: same install script.

curl -fsSL https://ollama.com/install.sh | sh

Windows: open PowerShell and run the official installer:

irm https://ollama.com/install.ps1 | iex

Or grab the installer from the download page. Docker users: there is an official ollama/ollama image on Docker Hub if you would rather containerise it.

Verify the install:

ollama --version

On the desktop apps, Ollama starts a background server automatically. On Linux or headless servers, start it yourself with ollama serve — that is what you use whenever you want Ollama running without the desktop application.

Step 2: Pick a model that fits your machine

Models are named like Docker tags: name:tag. A few sensible starting points that all exist on the Ollama library:

  • llama3.2:3b (~2 GB download) — great first model for any laptop with 8 GB of RAM.
  • llama3.1:8b — the workhorse for 16 GB machines.
  • qwen3:8b — strong alternative in the 8B class, good for code and multilingual text.
  • mistral:7b — a classic fast all-rounder.
  • codellama:7b — Meta’s code-focused model, useful if you mostly want code help.
  • deepseek-r1:8b — a distilled reasoning model; slower but more deliberate on hard problems.
  • gemma4 — Google’s latest open model, featured in Ollama’s own quickstart.

Heavier hitters like qwen3:32b, gpt-oss:20b, or 70B-class models exist too — just make sure your RAM budget covers them before you pull.

Download the model first, or let ollama run fetch it automatically on first use:

ollama pull llama3.2:3b

Step 3: Chat with your model in the terminal

ollama run llama3.2:3b

This drops you into an interactive chat session. Type a prompt, hit enter, get an answer. Useful commands inside the session:

  • /? — show all commands and shortcuts
  • /bye — exit the session
  • /clear — clear the conversation context
  • Wrap text in """ for multi-line input

For one-off prompts in scripts, skip the interactive session and pass the prompt as an argument:

ollama run llama3.2:3b "Explain recursion in Python with a short example."

You can also pipe text in, which is perfect for feeding a file into a prompt:

cat README.md | ollama run llama3.2:3b "Summarise this file in three bullet points."

Step 4: Manage your models

Models pile up fast — the CLI keeps that painless:

ollama list            # everything downloaded, with sizes
ollama ps              # what is currently loaded in memory
ollama show llama3.2   # details: parameters, template, system prompt
ollama show llama3.2 --modelfile   # see the full Modelfile
ollama cp llama3.2 my-copy         # cheap rename/copy (blobs are shared)
ollama rm my-copy                  # delete
ollama stop llama3.2               # unload from memory right now

ollama ps deserves special attention: it shows which models are resident and how memory is split, which is the first thing to check when things feel slow or memory-hungry.

Step 5: Build a custom model with a Modelfile

This is where Ollama gets interesting for developers. A Modelfile is like a Dockerfile for a model: it pins a base model, sets parameters, and bakes in a system prompt. Say you want a code reviewer that is terse and strict. Create a file called Modelfile:

FROM qwen3:8b

PARAMETER temperature 0.2
PARAMETER num_ctx 16384

SYSTEM """
You are a strict senior code reviewer. Review only what you are given.
Be terse: file:line references, the problem, the fix. No praise, no preamble.
"""

Build it and use it:

ollama create code-reviewer -f Modelfile
git diff | ollama run code-reviewer "Review this diff."

The directives that matter most: FROM (base model, required), PARAMETER (temperature, top_p, num_ctx, num_predict, stop), SYSTEM (the baked-in system prompt), and TEMPLATE (the prompt template in Go template syntax). Low temperature (0.1–0.3) is what you want for deterministic dev tooling; raise it for creative writing.

Step 6: How to run Ollama as an API server

The terminal is nice, but your apps want HTTP. Ollama serves a REST API on http://localhost:11434 by default. Make sure the server is running (ollama serve on headless setups), then test it:

# Health check: lists installed models
curl http://localhost:11434/api/tags

# Chat completion
curl http://localhost:11434/api/chat -d '{
  "model": "llama3.2:3b",
  "messages": [{ "role": "user", "content": "Why is the sky blue?" }],
  "stream": false
}'

There is also /api/generate for single-prompt completions, /api/show for model info, and /api/embed for embeddings. Responses stream token-by-token by default — set "stream": false when you want the whole answer at once.

Use the OpenAI-compatible endpoint from Python

Here is the genuinely convenient part: Ollama exposes an OpenAI-compatible API at http://localhost:11434/v1, so every OpenAI SDK client, agent framework, and IDE plugin can point at it with no model-provider code changes. The API key is not validated — any placeholder works.

# pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",  # Ollama does not check this — any value works
)

response = client.chat.completions.create(
    model="qwen3:8b",
    messages=[{"role": "user", "content": "Write a Python function that retries with backoff."}],
)
print(response.choices[0].message.content)

Prefer the official client? There are libraries for Python and JavaScript:

pip install ollama

from ollama import chat

response = chat(
    model="qwen3:8b",
    messages=[{"role": "user", "content": "Why is the sky blue?"}],
)
print(response.message.content)

Agent tools integrate directly too — the official quickstart shows integrations with Claude Code, OpenCode, Codex, Copilot CLI, and more via ollama launch.

Three developer workflows that make Ollama worth it

1. Terminal code review. With the code-reviewer Modelfile from step 5, reviewing a branch is one line:

git diff main...HEAD | ollama run code-reviewer "Review this diff."

2. Cheap AI inside your scripts. Anything that calls the OpenAI SDK now has a free local backend. Prototyping a log summariser, a commit-message generator, or a test-data factory? Point base_url at Ollama, iterate for free, and only switch to a paid model when the local one genuinely cannot do the job.

3. A coding assistant in your editor. Extensions like Continue connect to Ollama for autocomplete and chat directly in VS Code — private, offline, no subscription.

Tune it: the environment variables that matter

Server behaviour is controlled by environment variables you set before starting ollama serve (or the desktop app):

  • OLLAMA_HOST (default 127.0.0.1:11434) — bind address. Set to 0.0.0.0:11434 to serve your LAN or a remote box.
  • OLLAMA_KEEP_ALIVE (default 5m) — how long an idle model stays loaded. Set to -1 to keep it loaded forever (fast, always-warm responses; costs RAM) or 0 to unload immediately.
  • OLLAMA_NUM_PARALLEL — parallel request slots per model; each needs its own slice of memory.
  • OLLAMA_MAX_LOADED_MODELS — how many different models can sit in memory at once. Lower it on tight machines.
  • OLLAMA_MODELS — where model blobs live (default ~/.ollama/models). Point it at a big disk if your home partition is small.
  • OLLAMA_FLASH_ATTENTION (set to 1) and OLLAMA_KV_CACHE_TYPE (q8_0 roughly halves KV-cache memory with negligible quality loss) — the memory-saver combo when running long contexts.

A security note you should not skip: the local Ollama API has no authentication. Keep it on loopback. If you set OLLAMA_HOST=0.0.0.0 to reach it from another machine, put it behind an authenticated HTTPS proxy — never expose an unauthenticated inference endpoint to an untrusted network.

Troubleshooting the common failures

“Connection refused” from the API. The server is not running. Start it with ollama serve, confirm with curl http://localhost:11434 (you should get an “Ollama is running” response), and make sure nothing else holds port 11434.

“Model not found.” The tag does not match what you pulled. Run ollama list and use the exact name:tag from that list.

Out of memory or painfully slow. Drop to a smaller model or a more aggressive quant. Watch ollama ps to see the memory split, and consider OLLAMA_MAX_LOADED_MODELS=1 so two models never fight over RAM. Enabling flash attention also cuts memory use as context grows.

Desktop app vs server behaving differently. On Linux installs from the script, Ollama often runs as a systemd service; use systemctl restart ollama after changing environment variables, with the variables set in a systemd override file rather than your shell.

When you should NOT use Ollama

Honesty section. Do not build on local models when you need: frontier-level reasoning (complex agents, hard maths), very long contexts handled reliably (small models degrade fast past their trained context), the newest model the day it launches (the library lags behind frontier releases), or guaranteed uptime for production users (you now own the ops). For those, a hosted API is the right tool. Ollama shines as your free, private, always-on workbench — not as a datacentre replacement.

Further Reading & References

Leave a Comment