VERIFICATIONCommands and behavior checked against the current Ollama Quickstart, CLI, FAQ, and troubleshooting documentation.

30-SECOND SUMMARY

What to take away

  • Use the official installer for your operating system.
  • Check CPU/GPU loading with `ollama ps` before diagnosing performance.
  • Local and `:cloud` models have different data paths; verify the model name.
DIAGNOSTIC 01

When generation feels slow

Change one condition at a time.

  1. 01
    REPEAT

    First run or every run?

  2. 02
    PROCESSOR

    Check ollama ps

  3. 03
    MEMORY

    Close heavy apps

  4. 04
    MODEL

    Try one size smaller

Use it this way Record loading delay and sustained generation separately.
SECTION 01

Prepare the computer

Confirm the operating system version, CPU architecture, free storage, and installation permissions. Model downloads can consume many gigabytes even when the Ollama application itself is small, and supported acceleration depends on the operating system, hardware, and drivers.

Use the official download page, connect a laptop to power, and close memory-heavy applications before the first large download. On managed equipment, obtain installation and model-use approval before creating local caches.

Reserve space for more than one download: the operating system, model cache, logs, and a second test model all need headroom. Record the OS version, free storage, RAM, and CPU/GPU name before installation; this makes later compatibility and performance troubleshooting much faster.

SECTION 02

Install and verify the service

On macOS and Windows, run the official installer and confirm that the Ollama application is active. On Linux, follow the official Linux guide; the install script, manual installation, and service setup are documented there, while service details can vary by distribution.

Open Terminal or PowerShell and run `ollama` or `ollama -v`. If a server is not already managed by the desktop app or system service, `ollama serve` can expose the startup error directly. Do not start a duplicate server when one is already listening.

Verify two independent signals: the version or help output proves the CLI is available, and a model request followed by `ollama ps` proves the local server responds. If either fails, preserve the message and consult the official OS-specific log location before reinstalling.

# Verify the CLI
ollama -v

# Start manually only when no app or service is already running
ollama serve
SECTION 03

Run the first model

The first `run` downloads the model, loads it, and enters chat mode, so several kinds of delay are mixed together. Start with a short public question that is easy to verify and exit with `/bye`.

The example below uses `gemma3:4b`, but model availability and requirements can change. Confirm the exact local tag in the official model library rather than guessing a name or accidentally choosing a `:cloud` model.

Run the same prompt twice. The first response may include download or load work; the second is a better indication of normal interactive use. Success means the answer completes, `ollama ls` includes the exact tag, and another run starts without downloading the model again.

ollama run gemma3:4b

>>> Explain local AI in three sentences.

/bye
SECTION 04

Confirm where the model is loaded

`ollama ls` lists downloaded models. `ollama ps` shows models currently in memory, their context allocation, and the processor split. A model may run partly on the GPU and partly on the CPU when VRAM is limited.

Read PROCESSOR as a placement report, not a speed score: `100% GPU` is fully GPU-resident, `100% CPU` uses system-memory inference, and a split means both are used. None is automatically an error if the task still meets its quality and latency target.

Run a short prompt and inspect `ollama ps` immediately, because an unloaded model may not appear. Record the exact model tag, PROCESSOR, CONTEXT, first-response wait, and repeated-run time; another model can produce a different placement on the same computer.

ollama ls
ollama ps
SECTION 05

Manage downloads, memory, and deletion

Use `ollama pull` to download without starting chat, `ollama stop` to unload a running model, and `ollama rm` to delete its stored files. Stopping does not reclaim disk space, while removing requires another download if the model is needed again.

Read the exact name from `ollama ls` before any destructive command. After `stop`, confirm that `ollama ps` no longer lists the model; after `rm`, confirm that it is absent from `ollama ls`.

Before removal, preserve the source URL, tag, license notes, and any Modelfile needed to reproduce the setup. If disk space does not change as expected, inspect the documented model location and application state rather than deleting unknown cache directories manually.

GoalCommandDeletes model file
Download onlyollama pull gemma3:4bNo
List downloadsollama lsNo
List loaded modelsollama psNo
Unloadollama stop gemma3:4bNo
Removeollama rm gemma3:4bYes
SECTION 06

Troubleshoot in a fixed order

Separate download, model-loading, first-output, and sustained-generation delays. Keep the model and prompt fixed while you close memory-heavy applications, start a new conversation, and then compare a smaller model.

For download failures, check free storage, network controls, certificates, and HTTPS proxy configuration. For connection failures, verify the existing app or service before opening a firewall; for GPU issues, check official support, drivers, server logs, and VRAM use before reinstalling.

Keep a three-step incident note: exact command and time, observed error or log line, then one changed condition. Retest the same command after each change and verify that faster output has not introduced factual or formatting failures.

  • Compare first and repeated runs
  • Inspect PROCESSOR and CONTEXT in `ollama ps`
  • Start a new conversation
  • Try a smaller model
  • Read the official troubleshooting logs
SECTION 07

Verify that the workload is local

Ollama documents both local and cloud features. A model tag ending in `:cloud`, web search, or a connected application can use a different data path from local inference, so read the exact model name before entering sensitive material.

For a local-only setup, the official FAQ documents `disable_ollama_cloud` in `~/.ollama/server.json` and the `OLLAMA_NO_CLOUD=1` environment setting. Apply the setting to the existing app or service and restart it instead of launching a duplicate server.

Retest with harmless data while the network is disconnected, then inspect logs, backup folders, plugins, and synchronization separately. A successful offline prompt verifies local inference, but it does not prove that every derived file or connected tool is isolated.

# ~/.ollama/server.json
{
  "disable_ollama_cloud": true
}
SECTION 08

Record a baseline before updates

Write down the installation date, OS and driver versions, Ollama version, CPU/GPU and memory, exact model tag, and context shown by `ollama ps`. A model name without its tag is not enough to reproduce the file or behavior.

Save one public source and five known-answer prompts with the observed quality, first-response wait, repeated-run time, and placement. Preserve failures as well as successes, including the condition that caused an out-of-memory or connection error.

After an Ollama, driver, or model update, rerun the same baseline before returning to important work. If quality, speed, or data-path checks regress, the record gives you a reason to pause adoption instead of treating a changed result as normal.

FAQ

Frequently asked questions

Why is only the first response slow?

The model may be loading from storage into memory. Repeated requests can be faster while it stays loaded. For a practical check, follow the “Prepare the computer” section, change one condition at a time, and record the result.

Does `ollama stop` delete the model?

No. It unloads the model from memory. Use `ollama rm` to delete the stored model.

Can I expose port 11434 to the internet?

Do not expose it without authentication, network controls, TLS, rate limits, and a security review. Local and `:cloud` models have different data paths; verify the model name. For a practical check, follow the “Run the first model” section, change one condition at a time, and record the result.

Primary sources

Check the original documentation for version-specific details.

Ollama Quickstart Ollama CLI Reference Ollama FAQ Ollama Linux Ollama Troubleshooting Ollama GPU Support Ollama Model Library

READ NEXT

LM Studio Setup: Run Your First Local LLMConnect Open WebUI to Ollama