---
title: "Run your own LLM on a VPS with Ollama, no GPU needed"
description: "How to run an LLM on a VPS without a GPU using Ollama: RAM needed for 7B-33B models, install, the API on port 11434, an SSH tunnel, Open WebUI and measuring speed."
url: https://tihost.io/en/blog/ollama-on-vps
language: en
section: "Guides"
published: 2026-10-05
updated: 2026-10-05
publisher: Tihost (https://tihost.io)
---

# Run your own LLM on a VPS with Ollama, no GPU needed

> **In short:** Ollama runs open language models on a VPS without a graphics card, on the CPU alone. Per the Ollama README, 7B models need 8 GB of RAM, 13B models 16 GB and 33B models 32 GB. Generation on a CPU is noticeably slower than on a GPU, so such a server suits chatbots, summarization and classification of modest volumes of text.

**Key takeaways:**

- Ollama runs open large language models (LLMs) on a VPS without a graphics card: it installs with one command, `curl -fsSL https://ollama.com/install.sh | sh`, and runs as a systemd service.
- Per the Ollama README, 7B models need at least 8 GB of RAM, 13B models 16 GB and 33B models 32 GB.
- On a CPU, Ollama generates text noticeably slower than on a GPU; the real speed in tokens per second is the `eval rate` line printed by `ollama run --verbose`.
- The Ollama API listens only on `127.0.0.1:11434` and has no authentication: never open port 11434 to the internet - reach it through an SSH tunnel or a password-protected reverse proxy.
- Tihost plans have no GPUs; Power configurations with 16-32 GB of RAM suit 13B-33B models.

> This guide is written for Ubuntu 22.04/24.04 and Debian 12 on a KVM VPS. Do the [first server setup](https://tihost.io/en/blog/vps-first-setup) before you start: a sudo user, key-based login and UFW enabled.

## How much RAM does an LLM need in Ollama?

Memory is the main limit for an LLM on a server without a GPU: the whole model is loaded into RAM, and if there is not enough, the model will not start. The Ollama README gives guidelines: 8 GB of RAM for 7B models, 16 GB for 13B and 32 GB for 33B. By default Ollama downloads quantized models (4 bits per weight), so a 7B model takes about 4-5 GB on disk. A long context adds memory use on top of that.

| Model size | RAM per the Ollama README | Tihost configuration |
| --- | --- | --- |
| 1-3B | Not stated; a 3B model file is about 2 GB | Starter with 4 GB of RAM |
| 7B | 8 GB | Advanced: 4 vCPU / 8 GB, 12 GB for headroom |
| 13B | 16 GB | Power: 8 vCPU / 16 GB |
| 33B | 32 GB | Power: 16 vCPU / 32 GB (Poland) |

*The 7B-33B rows are guidelines from the Ollama README; the 1-3B row is an estimate based on model file size.*

## How slow is an LLM without a GPU?

Generation on a CPU is noticeably slower than on a GPU, and slower the bigger the model: a 3B model answers faster than a 13B one on the same server. Speed depends on the CPU, clock, memory and the model itself, so measure it on your own server (step 3) rather than trusting someone else's numbers. An LLM on a CPU makes sense when:

- you need a chatbot or assistant for a small team, where an answer in a few seconds is acceptable;
- you need to summarize, classify or tag modest volumes of text in the background - support requests or reviews, for example;
- data must not go to an external API and the model has to run on your own server;
- you want a sandbox for experimenting with prompts and models without paying per token.

## Step 1. Install Ollama

First check how much RAM, how many cores and how much free disk the server has - models take gigabytes. Then install Ollama with the official script: it downloads the program, creates an `ollama` user and a systemd service that starts at boot.

```bash
free -h
nproc
df -h /
```

```bash
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
sudo systemctl status ollama
```

The script runs with root privileges, so you may want to download and read it first. With the script install, models are stored in `/usr/share/ollama/.ollama/models`.

## Step 2. Download and run a model

`ollama run <model>` downloads a model from the library at [ollama.com](https://ollama.com) and opens a chat in the terminal; type `/bye` to exit. Start with a small model such as `llama3.2:3b`: it downloads quickly and shows what the server can do.

```bash
ollama run llama3.2:3b
ollama list
ollama ps
```

`ollama list` shows downloaded models, `ollama ps` the ones loaded in memory. After 5 idle minutes Ollama unloads a model from memory, so the first answer after a pause takes longer while the model loads again. `ollama rm <model>` deletes models you no longer need.

## Step 3. Measure generation speed

The `--verbose` flag prints statistics after every answer. The key line is `eval rate`: how many tokens per second the model generates on your server. `prompt eval rate` is how fast the prompt is read, `load duration` is the time to load the model into memory.

```bash
ollama run llama3.2:3b --verbose
```

Ask a few questions typical of your task and compare `eval rate` across model sizes - that way you pick the largest model that still answers fast enough. How to evaluate the server's CPU and disk in general is covered in [VPS benchmarking](https://tihost.io/en/blog/vps-benchmark).

## Step 4. Call the model through the API

The Ollama service exposes an HTTP API on `127.0.0.1:11434` for scripts and bots on the same server. With `"stream": false` the answer arrives as one complete JSON object instead of token by token:

```bash
curl http://localhost:11434/api/generate -d '{
  "model": "llama3.2:3b",
  "prompt": "Explain what a VPS is in two sentences",
  "stream": false
}'
```

The answer text is in the `response` field. The same way you can connect a [Telegram bot on a VPS](https://tihost.io/en/blog/telegram-bot-on-vps) or your own app to Ollama: requests go to a local address and never leave the server.

## Step 5. Outside access: an SSH tunnel instead of an open port

The Ollama API has no authentication: anyone who reaches port 11434 can use the model and load your server. So do not change the bind address to `0.0.0.0` and do not open 11434 in UFW. To reach it from your own computer, forward the port over [SSH](https://tihost.io/en/blog/ssh-connect-to-vps):

```bash
# on your own computer, not on the server
ssh -N -L 11434:localhost:11434 user@SERVER_IP
# while the tunnel is open, the API is available locally
curl http://localhost:11434/api/tags
```

> If other people or services need the model, put an nginx reverse proxy with HTTPS and authentication in front of Ollama - setting up nginx and a certificate is covered in [nginx and Let's Encrypt](https://tihost.io/en/blog/nginx-letsencrypt-https).

## Step 6. The Open WebUI web interface

Open WebUI is a ChatGPT-style web interface for Ollama models: chat history, model selection, multiple users. It runs in Docker (see [installing Docker](https://tihost.io/en/blog/install-docker-ubuntu-debian)). The host-network variant sees Ollama on `127.0.0.1` and serves the interface on port 8080:

```bash
docker run -d --network=host \
  -v open-webui:/app/backend/data \
  -e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
  --name open-webui --restart always \
  ghcr.io/open-webui/open-webui:main
```

With UFW enabled, port 8080 is closed from outside - open the interface through an SSH tunnel (`ssh -N -L 8080:localhost:8080 user@SERVER_IP`, then `http://localhost:8080` in the browser). Create an account right away: the first one registered becomes the administrator.

## Which Tihost VPS to pick for Ollama?

Tihost plans have no GPUs - models run on AMD Ryzen 9 CPUs. For 7B models, [Advanced](https://tihost.io/en/blog/advanced-vps) with 8-12 GB of RAM works; for 13B and 33B, [Power](https://tihost.io/en/blog/power-vps) with 16-32 GB. The Ryzen 9 9950X in Poland supports AVX-512 and DDR5 memory, which speed up neural network math on a CPU - more in [Ryzen 9 5950X vs 9950X](https://tihost.io/en/blog/amd-ryzen-9-5950x-vs-9950x).

| Configuration | Germany | Finland | Poland |
| --- | --- | --- | --- |
| 8 vCPU / 16 GB RAM / 300 GB NVMe | $32.00 | $32.00 | $40.00 |
| 12 vCPU / 24 GB RAM / 350 GB NVMe | $49.00 | $49.00 | $61.00 |
| 16 vCPU / 32 GB RAM / 450 GB NVMe | - | - | $77.50 |

*Price per 30 days, USD. A dash means the configuration is not sold in that location.*

**Launch a server in 2 minutes.** AMD Ryzen 9, NVMe and DDoS protection in Germany, Finland and Poland. Pay with crypto or card. [Order a Server](https://tihost.io/login)

## FAQ

### Can I run an LLM on a VPS without a GPU?

Yes. Ollama runs open language models on the CPU alone; generation is slower than on a GPU, but it is enough for chatbots, summarization and classification of modest volumes of text.

### How much RAM does a 7B model need in Ollama?

Per the Ollama README, a 7B model needs at least 8 GB of RAM, a 13B model 16 GB and a 33B model 32 GB. A long context adds memory use on top of these figures.

### How do I measure a model's speed in tokens per second?

Run the model with `ollama run <model> --verbose` and ask a question: the `eval rate` line in the statistics after the answer shows generation speed in tokens per second on your server.

### Can I expose the Ollama API to the internet?

Better not: the Ollama API on port 11434 has no authentication. For outside access use an SSH tunnel or an nginx reverse proxy with HTTPS and a password, and keep port 11434 closed.

### Does Tihost offer a VPS with a GPU?

No, Tihost plans are CPU-only: AMD Ryzen 9 without a GPU. The largest configuration is 16 vCPU and 32 GB of RAM, enough for models up to 33B by the Ollama README guidelines.

### How do I add a ChatGPT-like web interface to Ollama?

Run Open WebUI in Docker with `--network=host` and `OLLAMA_BASE_URL=http://127.0.0.1:11434`. The interface opens on port 8080; reach it from outside through an SSH tunnel.

---

Updated 2026-10-05 · https://tihost.io/en/blog/ollama-on-vps
