- Ollama runs open large language models (LLMs) on a VPS without a graphics card: it installs with one command,
curl -fsSL https://ollama.com/install.sh | sh, and runs as a systemd service. - Per the Ollama README, 7B models need at least 8 GB of RAM, 13B models 16 GB and 33B models 32 GB.
- On a CPU, Ollama generates text noticeably slower than on a GPU; the real speed in tokens per second is the
eval rateline printed byollama run --verbose. - The Ollama API listens only on
127.0.0.1:11434and has no authentication: never open port 11434 to the internet - reach it through an SSH tunnel or a password-protected reverse proxy. - Tihost plans have no GPUs; Power configurations with 16-32 GB of RAM suit 13B-33B models.
How much RAM does an LLM need in Ollama?
Memory is the main limit for an LLM on a server without a GPU: the whole model is loaded into RAM, and if there is not enough, the model will not start. The Ollama README gives guidelines: 8 GB of RAM for 7B models, 16 GB for 13B and 32 GB for 33B. By default Ollama downloads quantized models (4 bits per weight), so a 7B model takes about 4-5 GB on disk. A long context adds memory use on top of that.
| Model size | RAM per the Ollama README | Tihost configuration |
|---|---|---|
| 1-3B | Not stated; a 3B model file is about 2 GB | Starter with 4 GB of RAM |
| 7B | 8 GB | Advanced: 4 vCPU / 8 GB, 12 GB for headroom |
| 13B | 16 GB | Power: 8 vCPU / 16 GB |
| 33B | 32 GB | Power: 16 vCPU / 32 GB (Poland) |
How slow is an LLM without a GPU?
Generation on a CPU is noticeably slower than on a GPU, and slower the bigger the model: a 3B model answers faster than a 13B one on the same server. Speed depends on the CPU, clock, memory and the model itself, so measure it on your own server (step 3) rather than trusting someone else's numbers. An LLM on a CPU makes sense when:
- you need a chatbot or assistant for a small team, where an answer in a few seconds is acceptable;
- you need to summarize, classify or tag modest volumes of text in the background - support requests or reviews, for example;
- data must not go to an external API and the model has to run on your own server;
- you want a sandbox for experimenting with prompts and models without paying per token.
Step 1. Install Ollama
First check how much RAM, how many cores and how much free disk the server has - models take gigabytes. Then install Ollama with the official script: it downloads the program, creates an ollama user and a systemd service that starts at boot.
free -h
nproc
df -h /curl -fsSL https://ollama.com/install.sh | sh
ollama --version
sudo systemctl status ollamaThe script runs with root privileges, so you may want to download and read it first. With the script install, models are stored in /usr/share/ollama/.ollama/models.
Step 2. Download and run a model
ollama run <model> downloads a model from the library at ollama.com and opens a chat in the terminal; type /bye to exit. Start with a small model such as llama3.2:3b: it downloads quickly and shows what the server can do.
ollama run llama3.2:3b
ollama list
ollama psollama list shows downloaded models, ollama ps the ones loaded in memory. After 5 idle minutes Ollama unloads a model from memory, so the first answer after a pause takes longer while the model loads again. ollama rm <model> deletes models you no longer need.
Step 3. Measure generation speed
The --verbose flag prints statistics after every answer. The key line is eval rate: how many tokens per second the model generates on your server. prompt eval rate is how fast the prompt is read, load duration is the time to load the model into memory.
ollama run llama3.2:3b --verboseAsk a few questions typical of your task and compare eval rate across model sizes - that way you pick the largest model that still answers fast enough. How to evaluate the server's CPU and disk in general is covered in VPS benchmarking.
Step 4. Call the model through the API
The Ollama service exposes an HTTP API on 127.0.0.1:11434 for scripts and bots on the same server. With "stream": false the answer arrives as one complete JSON object instead of token by token:
curl http://localhost:11434/api/generate -d '{
"model": "llama3.2:3b",
"prompt": "Explain what a VPS is in two sentences",
"stream": false
}'The answer text is in the response field. The same way you can connect a Telegram bot on a VPS or your own app to Ollama: requests go to a local address and never leave the server.
Step 5. Outside access: an SSH tunnel instead of an open port
The Ollama API has no authentication: anyone who reaches port 11434 can use the model and load your server. So do not change the bind address to 0.0.0.0 and do not open 11434 in UFW. To reach it from your own computer, forward the port over SSH:
# on your own computer, not on the server
ssh -N -L 11434:localhost:11434 user@SERVER_IP
# while the tunnel is open, the API is available locally
curl http://localhost:11434/api/tagsStep 6. The Open WebUI web interface
Open WebUI is a ChatGPT-style web interface for Ollama models: chat history, model selection, multiple users. It runs in Docker (see installing Docker). The host-network variant sees Ollama on 127.0.0.1 and serves the interface on port 8080:
docker run -d --network=host \
-v open-webui:/app/backend/data \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
--name open-webui --restart always \
ghcr.io/open-webui/open-webui:mainWith UFW enabled, port 8080 is closed from outside - open the interface through an SSH tunnel (ssh -N -L 8080:localhost:8080 user@SERVER_IP, then http://localhost:8080 in the browser). Create an account right away: the first one registered becomes the administrator.
Which Tihost VPS to pick for Ollama?
Tihost plans have no GPUs - models run on AMD Ryzen 9 CPUs. For 7B models, Advanced with 8-12 GB of RAM works; for 13B and 33B, Power with 16-32 GB. The Ryzen 9 9950X in Poland supports AVX-512 and DDR5 memory, which speed up neural network math on a CPU - more in Ryzen 9 5950X vs 9950X.
| Configuration | Germany | Finland | Poland |
|---|---|---|---|
| 8 vCPU·16 GB RAM·300 GB NVMe | $32.00 | $32.00 | $40.00 |
| 12 vCPU·24 GB RAM·350 GB NVMe | $49.00 | $49.00 | $61.00 |
| 16 vCPU·32 GB RAM·450 GB NVMe | - | - | $77.50 |
AMD Ryzen 9, NVMe and DDoS protection in Germany, Finland and Poland. Pay with crypto or card.
FAQ
Can I run an LLM on a VPS without a GPU?
Yes. Ollama runs open language models on the CPU alone; generation is slower than on a GPU, but it is enough for chatbots, summarization and classification of modest volumes of text.
How much RAM does a 7B model need in Ollama?
Per the Ollama README, a 7B model needs at least 8 GB of RAM, a 13B model 16 GB and a 33B model 32 GB. A long context adds memory use on top of these figures.
How do I measure a model's speed in tokens per second?
Run the model with ollama run <model> --verbose and ask a question: the eval rate line in the statistics after the answer shows generation speed in tokens per second on your server.
Can I expose the Ollama API to the internet?
Better not: the Ollama API on port 11434 has no authentication. For outside access use an SSH tunnel or an nginx reverse proxy with HTTPS and a password, and keep port 11434 closed.
Does Tihost offer a VPS with a GPU?
No, Tihost plans are CPU-only: AMD Ryzen 9 without a GPU. The largest configuration is 16 vCPU and 32 GB of RAM, enough for models up to 33B by the Ollama README guidelines.
How do I add a ChatGPT-like web interface to Ollama?
Run Open WebUI in Docker with --network=host and OLLAMA_BASE_URL=http://127.0.0.1:11434. The interface opens on port 8080; reach it from outside through an SSH tunnel.