13k.euES
Menu

How-to · Local AI & Hardware

How to run an LLM locally on Windows, Mac or Linux

By 13k.eu editorsUpdated and checked 4 min read

Leer en español

Short answer

Install Ollama (irm https://ollama.com/install.ps1 | iex on Windows, curl -fsSL https://ollama.com/install.sh | sh on macOS and Linux), download a model that fits your memory with ollama pull, chat with ollama run, and check with ollama ps that it shows 100% GPU. Other apps can then use it through the local API at http://localhost:11434.

A dark mechanical keyboard with three edge-lit glass tiles floating above it

Prices, limits and features change often. We date every figure and link to its source: check the vendor's page before you buy or build. How we make money.

This guide runs a language model on your own computer with Ollama, a free, open-source tool for Windows, macOS and Linux. With a local model, your prompts are processed on your machine and there is no subscription; Ollama's cloud models are optional and can be switched off. This guide uses the commands from Ollama's own documentation and README, checked on October 1, 2026.

Before you start: will it fit?

Two things limit what you can run:

  • Memory. The model has to fit in your GPU's memory (VRAM) or, on Apple Silicon Macs, in the shared system memory. An 8-billion-parameter model needs roughly 6 to 8 GB; our VRAM calculator gives the figure for any model, quantization and context length.
  • Disk space. On Windows, Ollama needs at least 4 GB for the program itself, and each model adds its download size on top: from 523 MB for the smallest Qwen3 model to 142 GB for the largest.

You can run models without a GPU, on the processor alone, but they will be much slower. On Intel Macs, Ollama runs on the CPU only.

Step 1: install Ollama

Windows (Windows 10 22H2 or newer). Open PowerShell and run:

irm https://ollama.com/install.ps1 | iex

Or download OllamaSetup.exe from ollama.com. It installs in your user folder without administrator rights, runs in the background and adds the ollama command to your terminal. If you have an NVIDIA card, Ollama needs driver 551.61 or newer.

macOS (Sonoma 14 or newer). In Terminal:

curl -fsSL https://ollama.com/install.sh | sh

Or download Ollama.dmg and drag the app to Applications. On first launch it offers to link the ollama command into /usr/local/bin.

Linux. The same script installs it:

curl -fsSL https://ollama.com/install.sh | sh

For NVIDIA GPUs you also need the CUDA drivers (nvidia-smi should list your card); for AMD, ROCm 7. Ollama's Linux guide also explains how to run it as a systemd service so it starts with the computer.

Check that it works:

ollama -v

Step 2: choose a model

Ollama's library lists each model with its download size. A few starting points, from smallest to largest:

Model Command Download size Good for
Llama 3.2 3B ollama pull llama3.2 2.0 GB Older or low-memory machines
Qwen3 4B ollama pull qwen3:4b 2.5 GB 6 to 8 GB of VRAM
Qwen3 8B ollama pull qwen3:8b 5.2 GB 8 GB of VRAM with a short context
Gemma 4 E4B ollama pull gemma4 6.6 GB 8 to 12 GB of VRAM
Qwen3 14B ollama pull qwen3:14b 9.3 GB 12 to 16 GB of VRAM
gpt-oss 20B ollama pull gpt-oss:20b 14 GB 16 GB of VRAM
Qwen3 32B ollama pull qwen3:32b 20 GB 24 GB of VRAM

The download size is close to the memory the weights need; the context adds more on top. If you are between two sizes, start with the smaller model.

Step 3: chat with it

ollama run qwen3:8b

You get a prompt where you can type questions. To paste several lines at once, wrap them in triple quotes ("""). Other commands you will use:

Command What it does
ollama ls Lists downloaded models
ollama ps Shows loaded models and whether they run on the GPU
ollama stop qwen3:8b Unloads a model from memory
ollama rm qwen3:8b Deletes a downloaded model

Step 4: check that it runs on the GPU

While a model is loaded, run ollama ps. The PROCESSOR column says "100% GPU" when the whole model is in video memory. If it shows a split between CPU and GPU, part of the model is running on the processor and answers will be slower: try a smaller model or a smaller quantization.

Step 5: give it more context (optional)

By default, Ollama uses a 4K context when your GPU has less than 24 GiB of VRAM. That is fine for short questions, but agents, coding tools and long documents need more; Ollama recommends at least 64,000 tokens for those tasks. You can raise it with the slider in the Ollama app's settings, or when starting the server:

OLLAMA_CONTEXT_LENGTH=64000 ollama serve

A longer context needs more memory. Check ollama ps again afterwards.

Step 6: use it from other apps

Ollama runs an API on your computer at http://localhost:11434. A minimal request:

curl http://localhost:11434/api/chat -d '{
  "model": "qwen3:8b",
  "messages": [{ "role": "user", "content": "Why is the sky blue?" }],
  "stream": false
}'

Apps built for OpenAI's API can use Ollama by setting the base URL to http://localhost:11434/v1, and Anthropic-style clients can use http://localhost:11434. Ollama supports a subset of each API, so check its compatibility pages if a feature is missing. The command line can also connect coding agents to Ollama, for example ollama launch claude for Claude Code.

Where the models are stored, and how to move them

  • Windows: %HOMEPATH%\.ollama. To store models on another drive, create a user environment variable called OLLAMA_MODELS with the new folder, then quit and reopen Ollama.
  • macOS: ~/.ollama.

Troubleshooting

  • Squares instead of progress bars on Windows 10: change your terminal font; Ollama uses Unicode characters that some older fonts lack.
  • AMD Radeon RX 6000 cards on Windows: some do not expose ROCm 7 with current drivers. Ollama enables Vulkan by default as the fallback.
  • Logs: on Windows, %LOCALAPPDATA%\Ollama (server.log); on macOS, ~/.ollama/logs; on Linux with the service, journalctl -e -u ollama.

Prefer a graphical app?

LM Studio does the same job with a desktop interface for browsing and chatting with models. The license is different, which matters at work: see Ollama vs LM Studio.

What we checked

Every command and requirement on this page comes from Ollama's documentation, README and CLI reference, and the download sizes from Ollama's library pages, all checked on October 1, 2026. We have not timed downloads or measured speeds on specific hardware.

What we checked

  • Install commands: irm https://ollama.com/install.ps1 | iex (Windows) and curl -fsSL https://ollama.com/install.sh | sh (macOS and Linux); REST API example at http://localhost:11434/api/chat. (Ollama (GitHub), )
  • CLI commands run, pull, ls, ps, stop, rm, serve and launch; multiline input with triple quotes. (Ollama (GitHub docs), )
  • Windows 10 22H2 or newer, NVIDIA driver 551.61 or newer, at least 4 GB for the binaries, no administrator rights, models in the .ollama folder of your user profile (%HOMEPATH%), OLLAMA_MODELS to move them, logs in the Ollama folder of %LOCALAPPDATA%, Vulkan fallback for some RX 6000 cards. (Ollama, )
  • macOS Sonoma (14) or newer; Apple M-series use CPU and GPU, Intel Macs CPU only; models in ~/.ollama and logs in ~/.ollama/logs. (Ollama, )
  • Linux install script, ollama -v to verify, CUDA drivers checked with nvidia-smi, ROCm 7 for AMD, systemd service and journalctl logs. (Ollama, )
  • Default context of 4K below 24 GiB of VRAM, at least 64,000 tokens recommended for agents and coding tools, OLLAMA_CONTEXT_LENGTH=64000 ollama serve, and ollama ps showing 100% GPU. (Ollama, )
  • OpenAI-compatible endpoints at http://localhost:11434/v1 (a subset of the API). (Ollama, )
  • Anthropic-style clients at http://localhost:11434. (Ollama, )
  • Cloud models are optional and cloud features can be disabled. (Ollama, )
  • Download sizes: qwen3 0.6b 523 MB, 4b 2.5 GB, 8b 5.2 GB, 14b 9.3 GB, 32b 20 GB, 235b 142 GB. (Ollama, )
  • Download size of gpt-oss:20b: 14 GB. (Ollama, )
  • Download size of gemma4 (latest, e4b): 6.6 GB. (Ollama, )
  • Download size of llama3.2 (latest, 3b): 2.0 GB. (Ollama, )

What may change

  • Install commands, minimum system versions and driver requirements change with new Ollama releases.
  • Model tags and download sizes in Ollama's library change when models are updated.
  • The default context length depends on Ollama's version and your VRAM.

Frequently asked questions

Do I need a GPU to run an LLM locally?

No, but it helps a lot. Ollama can run models on the processor alone, and on Intel Macs it only uses the CPU, but generation is much slower than on a GPU or an Apple Silicon Mac.

Which model should I start with?

One that fits your memory with room to spare. With 8 GB of VRAM, Qwen3 8B (5.2 GB download) works with a short context; with 16 GB, gpt-oss 20B (14 GB). Our VRAM calculator gives the figure for other models and settings.

How do I know the model runs on my GPU?

Run ollama ps while it is loaded. The PROCESSOR column shows 100% GPU when the whole model is in video memory; a CPU/GPU split means part of it runs on the processor.

Can other apps use my local model?

Yes. Ollama serves an API at http://localhost:11434, with OpenAI-compatible endpoints at /v1 and support for Anthropic-style clients, each covering a subset of the original API.

Sources

  1. ollama/ollama: README (install commands) and MIT license, Ollama (GitHub). Accessed October 1, 2026.
  2. Ollama CLI reference: run, pull, ls, ps, stop, rm, launch, Ollama (GitHub docs). Accessed October 1, 2026.
  3. Ollama on Windows: system requirements, install location and models, Ollama. Accessed October 1, 2026.
  4. Ollama on macOS: system requirements, Ollama. Accessed October 1, 2026.
  5. Ollama on Linux: install, service and GPU drivers, Ollama. Accessed October 1, 2026.
  6. Ollama: context length defaults and settings, Ollama. Accessed October 1, 2026.
  7. Ollama: OpenAI compatibility, Ollama. Accessed October 1, 2026.
  8. Ollama: Anthropic compatibility, Ollama. Accessed October 1, 2026.
  9. Ollama cloud: models, usage and data handling, Ollama. Accessed October 1, 2026.
  10. Ollama library: qwen3 (tags and download sizes), Ollama. Accessed October 1, 2026.
  11. Ollama library: gpt-oss (tags and download sizes), Ollama. Accessed October 1, 2026.
  12. Ollama library: gemma4 (tags and download sizes), Ollama. Accessed October 1, 2026.
  13. Ollama library: llama3.2 (tags and download sizes), Ollama. Accessed October 1, 2026.

Spotted an error or an outdated price? Tell us and we will fix it.

Change history

  • : First published, with the install commands for Windows, macOS and Linux from Ollama's README and docs.

Next review: .

Part of our Local AI & Hardware guide.