You open Telegram, type a question into your bot’s private chat, and within a few seconds, a model running on your own VPS responds. No external API has read that conversation, and no SaaS has stored it in a queue — the same trade-off behind the choice in our local AI vs SaaS comparison. The piece that makes this isolation possible is long polling: the bot only opens outgoing connections, meaning you don’t need a domain, certificates, or firewall rules. If you already have a VPS with Docker, the entire stack consists of two services in a compose file and a short Python script. To size the machine with real numbers, our self-hosting decision guide compares VPS, rented GPUs and your own hardware by usage pattern.
Bot Architecture: Long Polling over Webhooks
A Telegram bot has two ways to receive messages, and this choice dictates the entire deployment. A webhook reverses the flow: Telegram sends an HTTPS POST to a public URL you register, requiring a domain, a reverse proxy with a certificate, and port 443 open. Long polling maintains the natural flow: your script asks the API, and Telegram holds the response until something arrives. For a first deployment on a personal VPS, the latter is the sensible choice.
How getUpdates and the Offset Work
The getUpdates method accepts a timeout parameter in seconds. With a high timeout, Telegram keeps the HTTP request open until an update arrives or the time expires, and the response contains all accumulated messages. Your code processes that batch, saves the update_id of the last one, adds one to it, and calls the API again with that offset. The result is a near real-time loop using a single outgoing connection to api.telegram.org.
Wrapped inside a while loop with a generous try/except block, the bot survives network outages: it logs the error, waits a few seconds, and retries. The server listens for nothing from the Internet, and the exposed surface area is reduced solely to the bot token.
When a Webhook Makes Sense
Webhooks shine when the bot serves many people in parallel, because each message arrives instantly and you don’t maintain permanent open connections. For a personal bot or a small group, that extra infrastructure is technical debt: more pieces to monitor and a public entry point to protect. There is one practical detail to know from the start: the API does not allow both methods simultaneously, so if you test a webhook, you will need to call deleteWebhook before getUpdates will respond again.
The Docker Compose Stack: Ollama and the Bot
The stack consists of two services: ollama, which serves the model on internal port 11434, and bot, which runs your Python script. Compose creates an internal network between both containers, and its DNS resolves the name ollama, so your code calls http://ollama:11434 without publishing that port to the outside world. The only connection leaving the server is the bot’s HTTPS traffic to Telegram, and the model remains isolated within the Docker network.
The Minimal docker-compose.yml
This compose file is enough to get started:
services:
ollama:
image: ollama/ollama
volumes:
- modelos:/root/.ollama
restart: unless-stopped
bot:
build: ./bot
env_file: .env
depends_on:
- ollama
restart: unless-stopped
volumes:
modelos:
Note that no service declares ports: nothing listens to the Internet, and isolation is handled by the internal Compose network. The volume persists downloaded models between restarts, and the restart: unless-stopped directive brings both containers up if the VPS reboots, which is exactly what an always-available service requires. The bot’s Dockerfile needs nothing exotic: it starts from a python:3.12-slim image, installs requests via pip, and launches the script.
Environment Variables and Bot Token
The token is generated by BotFather in Telegram using the /newbot command; it is best to store it in a .env file outside the repository. Also define the name of the model the bot will use in that file, such as MODEL=llama3.2:3b, so you can change models without touching the code. Avoid writing the token in the logs: anyone who obtains it can completely control your bot.
The Message Loop with the Ollama API
The /api/chat call with requests
The Ollama API exposes /api/chat with a JSON object containing the model and a list of messages with system, user, and assistant roles. Each call is stateless: your bot forwards the full history in each turn because Ollama does not retain conversations between requests. This is the minimum call using requests:
import json
import requests
OLLAMA = "http://ollama:11434/api/chat"
def responder(historial, modelo):
r = requests.post(
OLLAMA,
json={"model": modelo, "messages": historial, "stream": False},
timeout=120,
)
r.raise_for_status()
return r.json()["message"]["content"]
The high timeout is intentional: the first request after startup loads the model into RAM, and on a VPS without a GPU, this can take significantly longer than subsequent requests. Ollama keeps the model loaded for about five minutes after each use—a behavior configurable via the keep_alive parameter—so "warm" responses arrive quickly.
Streaming with stream: true
If you set stream to true, Ollama responds in NDJSON: one JSON line per text fragment and a final line with done set to true. With the stream parameter of requests enabled, your bot iterates through these lines and accumulates the content of the message field:
def responder_stream(historial, modelo):
payload = {"model": modelo, "messages": historial, "stream": True}
with requests.post(OLLAMA, json=payload, stream=True, timeout=300) as r:
r.raise_for_status()
for linea in r.iter_lines():
if linea:
yield json.loads(linea)["message"]["content"]
To replicate a live-typing effect, edit the Telegram message with editMessageText every two or three accumulated fragments; editing it with every single token wastes your request quota without providing a visible benefit. The simplest option is to send a typing action with sendChatAction while the model generates and publish the full text upon completion. Start with the second option and add edited streaming once the bot is fully functional.
Conversation Memory per Chat
RAM History and Context Window
The simplest memory is a dictionary with the chat_id as the key and the list of messages as the value. Before each call, you add the user’s turn, and after the response, the assistant’s turn. The limit is set by the model’s context window: when the history grows too large, trim the oldest turns and keep the system prompt along with the most recent exchanges. Setting a maximum number of turns is simpler and more robust than counting tokens for a personal bot.
Lightweight Persistence with SQLite
A RAM dictionary dies with every redeploy of the container, so dump the history to SQLite with a messages table that stores chat_id, role, content, and date. Mount the database in a Docker volume so it survives bot container rebuilds. Add a /reset command to delete rows for that chat: users appreciate being able to start from scratch, and you control the growth of the table.
Isolate by chat from day one and avoid sending logs containing message content to external services: the context must remain on your server. If you are debating between this approach and delegating to a provider, the comparison between local AI and SaaS makes it clear who controls the data in each case.
Telegram Limits and Rate Limiting
Limits Documented by the Bot API
Telegram does not publish a rigid table of limits, but its bot FAQ highlights two concrete rules: do not send more than one message per second to the same chat, and keep bulk sends below approximately 30 messages per second. Additionally, there is a hard limit of 4096 characters per text message, which requires splitting long responses before sending. These figures are from official documentation, though Telegram warns that dynamic limits vary based on server load.
Sending Queue and 429 Handling
When you exceed the allowed rate, the API responds with code 429 and a parameters field including retry_after in seconds. Your code must read this value, wait, and retry rather than hammering the API. For a personal bot, it is sufficient to centralize sends in a function that waits as needed between calls and handles 429 errors in one place. If you prefer not to build your own queue, the python-telegram-bot library encapsulates polling, queues, and retries, although requests and a short function are more than enough.
VPS Deployment: From Compose to Server
Hardware requirements for this stack
Setup and Maintenance
This design keeps the entire cycle within your server: messages rest in your SQLite, the model runs in your RAM, and the only external footprint is traffic with Telegram. The same compose file allows you to grow toward a vector store to search through your own documents—as seen in how to connect Ollama to WordPress—or toward an agent that executes tasks, such as a local agent powered by Hermes. None of these extensions require rewriting the bot: they are simply more services on the same network. If you still need a machine for this, you can deploy your bot on a VPS today; a basic plan carries the two-container stack around the clock.
For the first deployment, SSH into your server, create the compose directory, and run docker compose up -d: in one minute, you’ll have a bot listening for messages. If you don’t have a machine yet, you can deploy your bot on a VPS and leave it running 24/7; a basic plan handles the two-container stack without breaking a sweat. To size it with numbers, the Self-hosting AI guide: VPS, costs, and hardware compares options by usage pattern.
