30-SECOND SUMMARY
What to take away
- Use the official installer for your operating system.
- Check CPU/GPU loading with `ollama ps` before diagnosing performance.
- Local and `:cloud` models have different data paths; verify the model name.
When generation feels slow
Change one condition at a time.
- 01REPEAT
First run or every run?
- 02PROCESSOR
Check ollama ps
- 03MEMORY
Close heavy apps
- 04MODEL
Try one size smaller
Prepare the computer
Confirm the operating system version, CPU architecture, free storage, and installation permissions. Model downloads can consume many gigabytes even when the Ollama application itself is small, and supported acceleration depends on the operating system, hardware, and drivers.
Use the official download page, connect a laptop to power, and close memory-heavy applications before the first large download. On managed equipment, obtain installation and model-use approval before creating local caches.
Reserve space for more than one download: the operating system, model cache, logs, and a second test model all need headroom. Record the OS version, free storage, RAM, and CPU/GPU name before installation; this makes later compatibility and performance troubleshooting much faster.
Install and verify the service
On macOS and Windows, run the official installer and confirm that the Ollama application is active. On Linux, follow the official Linux guide; the install script, manual installation, and service setup are documented there, while service details can vary by distribution.
Open Terminal or PowerShell and run `ollama` or `ollama -v`. If a server is not already managed by the desktop app or system service, `ollama serve` can expose the startup error directly. Do not start a duplicate server when one is already listening.
Verify two independent signals: the version or help output proves the CLI is available, and a model request followed by `ollama ps` proves the local server responds. If either fails, preserve the message and consult the official OS-specific log location before reinstalling.
# Verify the CLI
ollama -v
# Start manually only when no app or service is already running
ollama serveRun the first model
The first `run` downloads the model, loads it, and enters chat mode, so several kinds of delay are mixed together. Start with a short public question that is easy to verify and exit with `/bye`.
The example below uses `gemma3:4b`, but model availability and requirements can change. Confirm the exact local tag in the official model library rather than guessing a name or accidentally choosing a `:cloud` model.
Run the same prompt twice. The first response may include download or load work; the second is a better indication of normal interactive use. Success means the answer completes, `ollama ls` includes the exact tag, and another run starts without downloading the model again.
ollama run gemma3:4b
>>> Explain local AI in three sentences.
/byeConfirm where the model is loaded
`ollama ls` lists downloaded models. `ollama ps` shows models currently in memory, their context allocation, and the processor split. A model may run partly on the GPU and partly on the CPU when VRAM is limited.
Read PROCESSOR as a placement report, not a speed score: `100% GPU` is fully GPU-resident, `100% CPU` uses system-memory inference, and a split means both are used. None is automatically an error if the task still meets its quality and latency target.
Run a short prompt and inspect `ollama ps` immediately, because an unloaded model may not appear. Record the exact model tag, PROCESSOR, CONTEXT, first-response wait, and repeated-run time; another model can produce a different placement on the same computer.
ollama ls
ollama psManage downloads, memory, and deletion
Use `ollama pull` to download without starting chat, `ollama stop` to unload a running model, and `ollama rm` to delete its stored files. Stopping does not reclaim disk space, while removing requires another download if the model is needed again.
Read the exact name from `ollama ls` before any destructive command. After `stop`, confirm that `ollama ps` no longer lists the model; after `rm`, confirm that it is absent from `ollama ls`.
Before removal, preserve the source URL, tag, license notes, and any Modelfile needed to reproduce the setup. If disk space does not change as expected, inspect the documented model location and application state rather than deleting unknown cache directories manually.
| Goal | Command | Deletes model file |
|---|---|---|
| Download only | ollama pull gemma3:4b | No |
| List downloads | ollama ls | No |
| List loaded models | ollama ps | No |
| Unload | ollama stop gemma3:4b | No |
| Remove | ollama rm gemma3:4b | Yes |
Troubleshoot in a fixed order
Separate download, model-loading, first-output, and sustained-generation delays. Keep the model and prompt fixed while you close memory-heavy applications, start a new conversation, and then compare a smaller model.
For download failures, check free storage, network controls, certificates, and HTTPS proxy configuration. For connection failures, verify the existing app or service before opening a firewall; for GPU issues, check official support, drivers, server logs, and VRAM use before reinstalling.
Keep a three-step incident note: exact command and time, observed error or log line, then one changed condition. Retest the same command after each change and verify that faster output has not introduced factual or formatting failures.
- Compare first and repeated runs
- Inspect PROCESSOR and CONTEXT in `ollama ps`
- Start a new conversation
- Try a smaller model
- Read the official troubleshooting logs
Verify that the workload is local
Ollama documents both local and cloud features. A model tag ending in `:cloud`, web search, or a connected application can use a different data path from local inference, so read the exact model name before entering sensitive material.
For a local-only setup, the official FAQ documents `disable_ollama_cloud` in `~/.ollama/server.json` and the `OLLAMA_NO_CLOUD=1` environment setting. Apply the setting to the existing app or service and restart it instead of launching a duplicate server.
Retest with harmless data while the network is disconnected, then inspect logs, backup folders, plugins, and synchronization separately. A successful offline prompt verifies local inference, but it does not prove that every derived file or connected tool is isolated.
# ~/.ollama/server.json
{
"disable_ollama_cloud": true
}Record a baseline before updates
Write down the installation date, OS and driver versions, Ollama version, CPU/GPU and memory, exact model tag, and context shown by `ollama ps`. A model name without its tag is not enough to reproduce the file or behavior.
Save one public source and five known-answer prompts with the observed quality, first-response wait, repeated-run time, and placement. Preserve failures as well as successes, including the condition that caused an out-of-memory or connection error.
After an Ollama, driver, or model update, rerun the same baseline before returning to important work. If quality, speed, or data-path checks regress, the record gives you a reason to pause adoption instead of treating a changed result as normal.
Frequently asked questions
Why is only the first response slow?
The model may be loading from storage into memory. Repeated requests can be faster while it stays loaded. For a practical check, follow the “Prepare the computer” section, change one condition at a time, and record the result.
Does `ollama stop` delete the model?
No. It unloads the model from memory. Use `ollama rm` to delete the stored model.
Can I expose port 11434 to the internet?
Do not expose it without authentication, network controls, TLS, rate limits, and a security review. Local and `:cloud` models have different data paths; verify the model name. For a practical check, follow the “Run the first model” section, change one condition at a time, and record the result.
Primary sources
Check the original documentation for version-specific details.
Ollama Quickstart Ollama CLI Reference Ollama FAQ Ollama Linux Ollama Troubleshooting Ollama GPU Support Ollama Model Library