Blog

Connecting Local AI: Building a Tool-Enabled Stack with MCP, Ollama, Qdrant, and WordPress

In November 2024, Anthropic released the Model Context Protocol (MCP), an open specification for connecting language models to external tools and data. OpenAI incorporated protocol compatibility into its Agents SDK during 2025, and the catalog of MCP servers has grown steadily since then. If you already have Ollama running models on your machine, Qdrant for embeddings, and WordPress for content, MCP replaces the "hand-crafted glue" between the three with a common contract. This article sets up that connection using verifiable components: the official Qdrant MCP server, the mcpo proxy, and a tool-supported agent.

What Exactly Does MCP Solve?

The core issue MCP addresses is the M×N integration problem.

The core issue MCP addresses is the M×N integration problem.

The M×N Integration Problem

The underlying issue is the M×N integration problem: with N applications and M data sources, every pair requires its own connector. MCP reduces this equation to M+N. A server exposes its capabilities once, and any compatible client discovers them dynamically at startup. There are three primitives: tools (functions the model invokes), resources (data the client reads), and prompts (reusable templates).

Two Transports: stdio and HTTP

The specification defines two transports: stdio for local processes and HTTP for networked services. There is a practical detail worth knowing before following tutorials: the March 2025 revision of the spec replaced the original HTTP+SSE transport with Streamable HTTP, so examples predating that date may fail with current clients.

Before writing code, check the modelcontextprotocol/servers repository on GitHub. It includes reference servers for the file system, Git, and PostgreSQL, helping you calibrate the granularity of a well-designed tool. Keeping the catalog on your own server is the bet that the local AI vs. SaaS comparison explains with hard numbers. Keeping that catalog on your own server is the bet our local AI vs SaaS comparison explains with numbers.

The Role of Each Piece in the Stack

First up is Ollama, the inference engine that downloads and runs models locally and exposes an API on port 11434.

Ollama: The Inference Engine

Ollama is the inference engine: it downloads and runs models locally and exposes an API on port 11434, with an OpenAI-compatible endpoint at /v1. It has supported tool calling since mid-2024, although support depends on the model: Llama 3.1, Qwen 2.5, or Mistral Nemo include variants trained for function calling. If the chosen model doesn’t know how to invoke tools, the rest of the setup is just decoration.

Qdrant: Long-Term Memory

Qdrant provides the long-term memory. It is a vector database written in Rust, with a REST API on port 6333 and gRPC on 6334; its team maintains an official MCP server in the qdrant/mcp-server-qdrant repository. That server exposes two tools: qdrant-store, which saves text alongside its embeddings, and qdrant-find, which searches by similarity.

WordPress: The REST API as Common Ground

WordPress does not understand MCP out of the box, but it does expose a mature REST API that has accepted application passwords since version 5.6. In a previous article, we manually assembled the Ollama and Qdrant duo within WordPress; what changes now is the connection layer, not the components. The first practical check: send a request to /api/chat and verify that your model lists the available tools, as this is the requirement that determines if everything else works.

mcpo: From stdio to REST

Since most MCP servers operate as local stdio processes, web services need a translator like mcpo to convert them into REST APIs.

Why a Translator is Necessary

Most MCP servers operate as local processes communicating via stdio—a format convenient for desktop clients but awkward for web services. mcpo, a proxy from the Open WebUI team, resolves this friction: it converts an MCP server into a REST API documented with OpenAPI, where each tool appears as an endpoint and documentation is automatically generated at /docs.

uvx mcpo --port 8000 -- uvx mcp-server-qdrant

One Container, One Startup Line

With that line, mcpo starts the Qdrant MCP server as a child process and publishes its tools on port 8000. The pattern is identical for any other server: you change the command following the — and the proxy regenerates the endpoints. Start the proxy, open http://localhost:8000/docs, and run a test call with curl before bringing the agent into the scene; if the endpoint responds, the transport layer is ready.

Stack Architecture with Docker Compose

The stack fits into four containers—Ollama, Qdrant, the Qdrant MCP server, and mcpo—plus a WordPress MCP server, all within an internal network.

Four Containers and an Internal Network

The set fits into four containers: Ollama on 11434, Qdrant on 6333, the Qdrant MCP server, and mcpo on 8000. Added to this group is a second MCP server dedicated to WordPress, which wraps the REST API and exposes operations such as searching posts, creating drafts, or listing categories. Persistent volumes for Ollama models and Qdrant storage go in the same compose file. It is the same service pattern we previously set up in the guide for local AI in WordPress with Ollama and Qdrant, with the MCP server as the new piece. This is the same service pattern we already built in our guide to connecting Ollama to WordPress, with the MCP server as the new piece.

The Agent: MCP Client and Local Model

Sitting on top is the agent: a program with an MCP client and access to the Ollama model that decides which tool to invoke at each step. If you prefer starting from a pre-built orchestrator, the community publishes Docker Compose implementations ready for this architecture. Map out the ports and environment variables before writing a single line of the compose file: knowing who speaks to whom avoids debugging connection strings blindly. An agent like the one from Hermes Agent with Docker Compose fits here without architectural changes. An agent like Hermes Agent fits right in without any architecture changes.

Step-by-Step Setup

The power-on order matters: base services first, MCP servers next, and the agent last.

Base Services, MCP Servers, and Agent

The power-on order matters: first the base services, then the MCP servers, and finally the agent.

  1. Spin up Ollama with the ollama/ollama image, mount the model volume, and download a tool-supported model: docker exec -it ollama ollama pull qwen2.5:7b.
  2. Start Qdrant with the qdrant/qdrant image and map port 6333.
  3. Launch the Qdrant MCP server with uvx mcp-server-qdrant, specifying QDRANT_URL=http://qdrant:6333 and the collection name in the COLLECTION_NAME variable.
  4. Place mcpo in front to expose the tools via REST.
  5. In the agent, register the MCP server and verify that the tools list arrives complete.

Application Passwords in WordPress

For the WordPress side, generate an application password from the user profile and choose between: an MCP plugin from the official directory that exposes editing operations, or a generic MCP server configured against the /wp-json/wp/v2 endpoints. Apply the principle of least privilege: an author role is enough to create drafts, and you do not need administrator credentials to write content. The final validation is a round-trip cycle: save a text with qdrant-store, retrieve it with qdrant-find, and verify that the agent cites the exact content.

An End-to-End Example Flow

To illustrate, here is how a request moves from Qdrant retrieval to a WordPress draft.

From Qdrant to a WordPress Draft

Imagine this request to the agent: "Retrieve the notes from Qdrant most similar to this draft and create a draft in WordPress with the summary and suggested internal links." The model breaks this down into two chained calls: qdrant-find with the draft text, and, with the results now in context, the post-creation tool from the WordPress server. Your work happens beforehand, in the configuration of permissions, collections, and roles.

Every New Server Multiplies the Catalog

The real benefit appears when expanding the catalog. If you add a Git or file system MCP server, the agent discovers the new tools at startup, and you can ask it to do things like "create a branch with the revised draft." This is the contract that separates MCP from a custom integration: tools are described in a format that any compatible client can read without additional work.

Security and Real-World Limits

Because an MCP server acts with your permissions, apply least privilege by creating minimum-scope credentials and documenting which tool touches which resource.

Least Privilege Permissions

An MCP server acts with your permissions. If you give it an administrator application password, every tool can write across the entire site, so create minimum-scope credentials and document which tool touches which resource. If you expose mcpo outside your machine, put a reverse proxy with authentication in front of it: an automatically generated API inherits the combined scope of all its tools, and it is advisable to review the proxy documentation before opening any ports.

A Young Protocol: What Might Change

The protocol is still evolving. The March 2025 revision changed an entire transport, and not all servers publish which version of the spec they are using, so mixing clients and servers from different eras requires verification. Factor in the context cost: every registered tool adds its definition to the prompt, an overhead that is noticeable in 7B models with tight windows. Log every tool call during the first few weeks of use; the emerging pattern will tell you which servers are redundant and which confuse the model with ambiguous definitions.

Start With a Single Server

Avoid spinning up five MCP servers on day one, since too many tools paralyze an agent; start with one server this week.

One Server This Week

The temptation with MCP is to spin up five servers on the first day, and the usual result is an agent paralyzed by too many tools. The qdrant-store and qdrant-find cycle is the best entry point: two tools, one container, and one proxy, with visible results in the first session. When that flow works without surprises, add the WordPress server and then whatever real usage demands.

Where Value Grows With Each New Server

The protocol’s value grows with every server added: new tools become available to any MCP client in the stack without repeating integrations. The next concrete step is to run the Qdrant MCP server with uvx, validate the store-find cycle against your collection, and let the stack run for a week before adding the WordPress piece. When the catalog grows and your laptop can’t keep up, a VPS for your MCP stack keeps Ollama, Qdrant, and the agent always on. And when the catalog outgrows your laptop, a VPS for your MCP stack keeps Ollama, Qdrant and the agent running around the clock.

Try it in 3 minutes

Install Ollama + Open WebUI on your computer, free and private.