You write a message to a commercial assistant, and that text travels to the provider’s servers, where it may be logged, reviewed, or used according to their terms. There is a direct alternative: Open WebUI over Ollama. On your own server, you can set up a private AI chat whose only boundary is your machine. You install Docker, launch a container, download a model, and converse with an AI from your browser that never leaves your network.
What actually changes with a private AI chat
The first change is where inference happens: your conversations stay on your equipment instead of crossing the internet to third parties.
The first change is where inference happens: your conversations stay on your equipment instead of crossing the internet to third parties.
Your conversations stop traveling to third parties
In a hosted service, every prompt and every document you attach crosses the internet and passes through third-party infrastructure. With Ollama, the model runs on your CPU or GPU, and inference happens entirely on your equipment; neither text nor files leave the machine. If you cut the network connection, the chat continues to work exactly the same.
This is especially important when dealing with sensitive material: contracts, proprietary code, internal notes, or customer data. With self-hosting, you don’t have to read the fine print because you write the fine print.
Cost shifts from subscriptions to hardware
A private AI chat shifts the expense: instead of a monthly fee per user or token, you pay in electricity and via the equipment you already own. For intensive use, this arithmetic becomes favorable very quickly. For sporadic use, the savings are smaller, and the real motivation is data control.
No third-party retention policies weigh on the equation, and no change in commercial terms can make your installation more expensive overnight. Before buying new hardware, try it with what you have on hand: a standard laptop can already run small models.
Open WebUI: what it is and what it adds
Open WebUI adds the face to Ollama's local engine: a web chat interface that communicates with its API.
A web chat interface for Ollama
Ollama is the engine: it downloads and executes models locally and exposes an API on your machine. Open WebUI is the face: a web chat interface deployed as a Docker container that communicates with that API. The result looks like any commercial assistant, except that the entire circuit lives on your server.
This separation of roles has a practical advantage: you can update the interface without touching the models, and change models without touching the interface.
History, multi-user support, and documents
The interface saves conversation history locally, so you can resume a thread from weeks ago without depending on a cloud. It supports multiple accounts: the first one created becomes the administrator, and from there, others and their permissions are managed. It also allows you to upload documents to query them in the chat, turning the installation into an assistant for your own files.
The set is completed by model management, system prompts, and reusable templates, all from within the interface. The next logical step is to see it in action: the Docker installation is covered in the following section.
Install with Docker in minutes
Before installing with Docker, ensure your machine has Docker running, free space, and, for speed, a GPU with drivers.
Prerequisites
You need a machine with Docker running, a few gigabytes of free space for the model, and, if you want speed, a GPU with its drivers. Installing Ollama on Linux is reduced to one line:
curl -fsSL https://ollama.com/install.sh | sh
For macOS and Windows, there are official installers on the project website. If you prefer not to install anything separately, there is an Open WebUI image tag that includes Ollama inside.
Launch the container
The documented command for the most common case, with Ollama installed on the host, is this:
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data --name open-webui --restart always ghcr.io/open-webui/open-webui:main
Each piece serves a purpose: the port mapping exposes the interface on 3000, the volume preserves users and conversations between restarts, and the –add-host flag allows the container to locate Ollama on the host machine. If you work with Compose, these same parameters translate into a docker-compose.yaml without complication.
First access and Ollama connection
Open http://localhost:3000 (or the server IP) and create your account: the first one becomes the administrator. In the connection settings, the interface usually detects Ollama at http://host.docker.internal:11434; if it doesn’t appear, type it in manually. From that panel, you can also manage models without touching the terminal.
Packaged routes if you want a shortcut
Setting it up manually teaches you how each piece fits together, but it’s not mandatory. In the d0a1 packs, there are packaged installations of self-hosted stacks, designed for those who prefer a more direct route than docker run. Use them as a starting point and customize later.
Check that the container is still up after restarting the server: the –restart always flag takes care of launching it without you doing anything.
Models: download, choose, and tune
Start with a small variant: pull your first model, like llama3, from the terminal or the models panel.
Your first model
From the terminal, ollama pull llama3 downloads the model and makes it ready; the interface itself also allows you to do this from its models panel. Names like mistral or qwen2.5 are common alternatives in the Ollama library. Start with a small variant: the download, test, and adjustment cycle will be a matter of minutes, not hours.
What your hardware can handle
Small models run on the RAM of a normal laptop and respond reasonably well on a CPU. Large models require a GPU with ample memory, and the speed difference between these two worlds is noticeable in every response. With NVIDIA cards, Ollama uses the GPU with standard drivers; on Mac with Apple silicon, unified memory is a major advantage.
System prompts and model profiles
Open WebUI data and Ollama models live in separate locations, and copying both is your backup.
Each chat allows you to set a system prompt that defines tone and behavior, and save that setting as a template for reuse. With several models downloaded, you can alternate between a fast one for short questions and a more capable one for serious work. This combination of profiles is one of the least advertised advantages of self-hosting.
Download a small one today and use it for a couple of days; only then decide if it’s worth downloading a large model.
Privacy, data, and daily operation
Open WebUI data and Ollama models live in separate locations, and copying both is your backup.
Where everything lives
Open WebUI data (users, conversations, uploaded documents) resides in the volume you mounted when creating the container. The models remain under the Ollama directory on the host. Copying both is your backup: a tar of the volume and the models directory, with no cloud services involved.
The beauty of a private AI chat is exactly that: the complete set of data fits in a directory that you control.
Expose the service wisely
The interface requires a username and password, but that doesn’t make it a service ready for the open internet. My recommendation is to keep it on the local network or access it via a VPN like WireGuard or Tailscale; if you publish it, place it behind a reverse proxy with TLS. The pattern is the same as with any self-hosted web application: minimal exposed surface.
Update without losing anything
Both Ollama and Open WebUI expose APIs, so any script on your end can talk to the models.
Updating is reduced to two steps: pulling the new image and recreating the container.
docker pull ghcr.io/open-webui/open-webui:main docker rm -f open-webui
After relaunching the initial docker run with the same parameters, the volume data is still there: history, accounts, and documents intact. Create a tar of the volume beforehand, so you have a copy if something goes wrong.
Add that tar to your monthly backup routine and your history will be safe regardless of any update.
From manual chat to automation
Both Ollama and Open WebUI expose APIs, so any script on your end can talk to the models.
Bring the models into your own systems
Ollama exposes a REST API on port 11434, and Open WebUI opens its own API with a format compatible with OpenAI, so any script on your end can talk to the models. That is the leap from occasional chatting to integrated use in your workflows. If your destination is a CMS, there is a guide for setting up local AI in WordPress with Ollama and Qdrant that covers that case from start to finish.
From chat to agent
The next step is for the AI to execute tasks instead of just responding. For this, there is a stack to automate tasks with Hermes Agent and docker compose, which fits alongside the Open WebUI container without interference. Both can share the same Ollama, so the downloaded models are utilized twice.
Before automating anything, launch a curl against the API from another machine on your network; if it responds, you have the base set up.
Where to run it: your machine or a server
For testing and personal use, your own machine is enough, though phone access, outside access, and nightly shutdowns can be inconvenient.
On your own machine
For testing and personal use, a laptop or home desktop is more than enough. The inconvenience arises when you want to access it from your phone, from outside, or when the computer shuts down at night. A private chat loses some of its appeal if it’s only available occasionally.
On an always-on server
To keep it always available, a dedicated server is advisable. A VPS for local AI keeps the container awake and accessible from anywhere, without depending on your laptop being open. Before choosing a size, hardware and prices are compared at d0a1.es/vps/ to avoid overpaying; large models require RAM and GPU that not all plans include, so check that point carefully. For the full decision across usage patterns, the AI self-hosting guide compares all three routes with real numbers.
Decide based on your use: if you only consult it during work hours from home, your equipment is enough; if you want it always at hand, move the stack as-is to the server and repeat the installation.
Where to start
The short path
If something breaks
If the container cannot find Ollama, check the connection by pointing to http://host.docker.internal:11434 and verify the –add-host flag at startup. If the page doesn’t load, check that port 3000 is free with docker ps and change the mapping if necessary. When responses are slow, try a smaller model variant or a machine with a GPU. And if you want to start from scratch, delete the volume and run the command again, knowing you will lose your history.
Set aside ten minutes today for the first four steps and leave the fifth for when you have a real document at hand. When that first message responds without leaving your network, the stack is set up; everything else (profiles, agents, servers) are optional extensions of the same container.
