Ollama + Open WebUI VPSYour open AI models, on your own server in France
- Ollama and Open WebUI deployed together
- Open models running on your VPS
- On the processor, or an NVIDIA L4 GPU as an option
- Hosted in France, 24/7 support
Pick your Ollama VPS
Ollama and Open WebUI are deployed automatically on every plan on this page. Without a GPU, the model runs on the processor and is loaded into memory: only plans with at least 4 vCPU, 8 GB of RAM and 60 GB of disk are offered. On the High performance range, the NVIDIA L4 GPU option can be added from the console.
Ubuntu26.04 · 24.04 · 22.04
Debian13 · 12 · 11
VPS delivered in 10 seconds
- HP-8Processor4 vCPUAMD EPYC 9554Memory8 GBDDR5Storage80 GB NVMeor 160 GB SSDNetwork1 GbpsChoose€12.79excl. VAT/month€15.99Save 38€/year
- HP-16Most popularProcessor8 vCPUAMD EPYC 9554Memory16 GBDDR5Storage160 GB NVMeor 320 GB SSDNetwork1 GbpsChoose€25.59excl. VAT/month€31.99Save 77€/year
- HP-32Processor16 vCPUAMD EPYC 9554Memory32 GBDDR5Storage320 GB NVMeor 640 GB SSDNetwork1 GbpsChoose€51.19excl. VAT/month€63.99Save 154€/year
- HP-64Processor32 vCPUAMD EPYC 9554Memory64 GBDDR5Storage640 GB NVMeor 1280 GB SSDNetwork1 GbpsChoose€102.39excl. VAT/month€127.99Save 307€/year
Every plan has everything you need, and more
- Ollama + Open WebUI deployed
- Automatic HTTPS
- Unlimited traffic
- DDoS protection
- 24/7 French-speaking support
Why Ollama on an OnetSolutions VPS
Recent processors, memory and NVMe storage under every Ollama server.
Intel Xeon or AMD EPYC, NVMe storage
Up to 1 Gbps
OnetSolutions API
Backups and snapshots
Your local AI in three steps
Nothing to install: you order, Ollama and Open WebUI are deployed on your server with an HTTPS address.
- 01
Pick your VPS
Choose a plan with at least 8 GB of RAM: Ollama + Open WebUI is already selected as a One Click App in the order.
- 02
Ollama and Open WebUI deploy
Automatically once the server is created. Only Open WebUI gets an HTTPS address; Ollama stays internal to the server.
- 03
Download a model
Create the administrator account, then type a model name in the Open WebUI model selector: it is downloaded from the Ollama library.
Ollama and Open WebUI, an AI that runs on your server
Ollama downloads and runs open models, Open WebUI gives you the interface. Both run on your VPS, on the processor or on an optional NVIDIA L4 GPU.
Your team
Browser, desktop or mobile
Open WebUI on your VPS
Accounts, history and documents in a datacenter in France
Ollama on the same VPS
The model runs on the server's processor and memory
Models downloaded from the interface
Type a model name in the Open WebUI model selector, or go to Admin > Connections: Ollama downloads it for you.
Ollama is not exposed
In this deployment only Open WebUI gets an HTTPS address. Ollama can only be reached by Open WebUI, inside the server.
No API provider
With a local model, your questions and the answers are computed on your VPS: they are not sent to any external AI service.
Chat with your documents
Upload files or build knowledge bases: Open WebUI finds the relevant passages and hands them to the model.
Multiple users
Roles, groups and permissions in Open WebUI: everyone has their own account and history, the administrator decides who uses which models.
Library of open models
Llama, Gemma, Qwen, Mistral and many more are published in the Ollama library, in versions of different sizes.
Local and API side by side
For tasks too heavy for the processor, add an OpenAI-compatible API in Open WebUI and keep your local models for the rest.
Several models at once
Open WebUI queries several models in parallel within one conversation, so you can compare their answers before choosing.
How much memory for which model?
Without the GPU option, Ollama loads the model into the VPS RAM. The model file, its context, Open WebUI and the 1 GB reserved for the system and the platform must all fit together: RAM decides how large a model you can run.
| Model (Ollama library) | Download | Recommended plan |
|---|---|---|
| gemma3:1b | 815 MB | From 8 GB of RAM |
| llama3.2:3b | 2.0 GB | From 8 GB of RAM |
| qwen3:4b | 2.5 GB | From 8 GB of RAM |
| mistral:7b | 4.4 GB | 16 GB of RAM recommended |
| qwen3:8b | 5.2 GB | 16 GB of RAM recommended |
| qwen3:14b | 9.3 GB | 32 GB of RAM recommended |
Sizes taken from ollama.com/library in September 2026. The recommended plans are a guide: a longer context or several simultaneous users need more memory.
NVIDIA L4 GPU option on the High performance range
High performance VPS accept an NVIDIA L4 Tensor Core GPU with 24 GB of VRAM, attached and detached from the console. The L4 is on Ollama's list of supported cards: inference is accelerated and larger models become usable.
- €0.70 excl. VAT per hour, billed by the minute, no commitment
- One VPS reboot when the GPU is attached
- GPU passthrough: standard NVIDIA drivers and CUDA
- The One Click App deployment starts Ollama without GPU access: install the NVIDIA driver and the NVIDIA Container Toolkit with root access, then give the GPU to the Ollama service
- General Purpose and Performance ranges: no GPU option
Without the GPU option, Ollama runs in processor-only mode: answers come noticeably slower than on a graphics card, and slower still as the model grows. The VPS then suits small quantised models and light use; for larger or faster models, add the GPU option on a High performance plan or connect Open WebUI to an API.
Ollama is open source under the MIT licence. Open WebUI is published under the Open WebUI License, a BSD-style licence with a clause that forbids removing the "Open WebUI" branding beyond 50 users over a rolling 30 days, unless the publisher agrees. Each model has its own licence, shown on its ollama.com page. No model is preinstalled.
What you can do with Ollama on a VPS
Open models on your own server, for the tasks where a small model is enough.
Private assistant
Rephrase, summarise, translate or draft with a small open model, without sending your texts to an external service.
- Llama
- Mistral
- Qwen
Questions about your documents
Procedures, internal notes, documentation: the model answers from the files you uploaded.
Sensitive data
For texts that must not leave your infrastructure, the model runs on your VPS in a datacenter in France.
Try open models
Download several models, ask them the same question and compare the answers before adopting one.
Local and API
Keep a local model for everyday work and add an OpenAI-compatible API for heavier requests.
- Mistral AI
- Anthropic
- OpenRouter
One team, one server
Every team member has an Open WebUI account; the administrator approves sign-ups and splits access by group.
Ollama and Open WebUI run on your VPS, our console runs the server and the application
Your VPS and the application are managed from the OnetSolutions console, without SSH, and you keep root access to the server.
- Status, logs and environment variables of each service
- Your domain, with an automatic HTTPS certificate
- Updates detected and applied from the console
- Daily backup of the application volumes, 7-day retention
- System
- Ubuntu 26.04 LTS
- IP address
- 51.210.24.8
- Uptime
- 34 j
- Size
- 4 vCPU · 8 Go
A team behind your server
Technicians who know your VPS, reachable at any hour.
24/7 French-speaking support
Open a ticket from the console, day or night: a person answers you about your server.
Help center
One Click Apps guides: deployment, domains, backups and updates.
30-day money-back guarantee
Try Ollama on your VPS: if it doesn't suit you, you get refunded.
Manage your VPS with AI
Our MCP server connects Claude, Cursor or any Model Context Protocol client to your OnetSolutions account: ask about your servers and their status in natural language, without leaving your assistant.
- Read-only: the assistant looks, it changes nothing
- Signs in with your OnetSolutions account
- Claude
- Cursor
- Any MCP client
Ollama FAQ
What to know before you order, and right after.
What is Ollama + Open WebUI?
Ollama is open source software that downloads and runs open AI models. Open WebUI is a web interface that connects to it to chat with those models, manage users and work on your documents. This One Click App deploys both together on your VPS, already connected.
Do I need a GPU? How fast does the model answer?
Without a GPU, Ollama computes on the server's processor, which is noticeably slower than a graphics card and slower still as the model grows: plan for small quantised models, from 1 to 14 billion parameters depending on the plan's RAM. On a High performance plan you can attach an NVIDIA L4 GPU with 24 GB of VRAM (€0.70 excl. VAT/hour, billed by the minute, no commitment), which Ollama supports. The One Click App deployment starts Ollama without GPU access: you need to install the NVIDIA driver and the NVIDIA Container Toolkit, then give the GPU to the Ollama service. You can also add an OpenAI-compatible API in Open WebUI.
Why at least 8 GB of RAM and 4 vCPU?
Without a GPU, the whole model is loaded into the VPS memory, next to Open WebUI and the 1 GB reserved for the system and the One Click Apps platform. The starter model in Ollama's documentation, Gemma 4 E2B, is a 7.2 GB download and Ollama recommends 8 GB of memory for it. Small models fit in 8 GB, 7 to 8 billion parameter models are more comfortable with 16 GB, and the computation runs on the plan's vCPU: that is why this page only offers plans with at least 4 vCPU, 8 GB of RAM and 60 GB of disk.
Is a model included?
No model is preinstalled. Once the administrator account is created, type a model name in the Open WebUI model selector (for example llama3.2:3b) or open Admin > Connections: it is downloaded from the Ollama library. Each model has its own licence, shown on its ollama.com page.
Do my conversations leave the server?
With a local Ollama model, no: your questions and the answers are computed on your VPS, in a datacenter in France, and Open WebUI stores the history on the same server. If you add an external API in Open WebUI, the messages sent to that model are processed by its provider.
How do I access the application after ordering?
Once the VPS is created, Ollama and Open WebUI are deployed automatically. Open the application in the One Click Apps section of the console and follow its HTTPS address. Create your account right away: the first account created becomes the administrator. Then add your domain in the Domains tab: the HTTPS certificate is issued automatically.
Are my data and models backed up?
The application volumes, including the one where Ollama stores downloaded models, are backed up every day at 3 a.m. with 7 days of retention, on the VPS disk: plan your disk space accordingly. These backups protect you from a mistake in the application, not from losing the server: for that, add the VPS backup option (from €2.99 excl. VAT/month) or take a snapshot.
How do updates work?
Every night, the console checks whether new Ollama and Open WebUI versions have been published and shows it on the application page. Nothing is updated without your action: one click on "Update" is enough, and the data volumes are kept.
Who maintains the application?
Ollama and Open WebUI are developed by their respective publishers. OnetSolutions provides the server, the deployment and the console tools; configuring the application, choosing models and using them is up to you. Our support covers the VPS: server, network, storage and availability.
Are Ollama and Open WebUI paid?
No. Ollama is open source under the MIT licence. Open WebUI is free, under the Open WebUI License: a BSD-style licence with a clause that forbids removing the "Open WebUI" branding beyond 50 users over a rolling 30 days, unless the publisher agrees. The prices on this page are those of the VPS.
Launch your local AI
Pick your plan: the server arrives with Ollama and Open WebUI deployed, ready to download your first models.
- 30-day money-back guarantee
- Hosted in France
- 24/7 French-speaking support
€15.99with the annual commitment
- 4 vCPU AMD EPYC 9554
- 8 GB DDR5
- 80 GB NVMe
- 1 Gbps