An AI that knows your own company's data can sit on your own side.
The previous chapter put the information in order. The preparation is the work and the AI is the last step — and this chapter is that last step. The Independence part closes by laying AI on top of everything stood up so far. Foundation, gate, documents, code, mail, meetings — an AI grounded in the data piled there, on your own side. The pgvector enabled back in 2-03 finally pays off here.
Three reasons to hold your own AI
- Keep data in — confidential internal documents are never handed to another company's API
- Always-on is cheap — classification, summarization, and extraction run continuously at zero marginal cost
- Grounded in your data — answers take internal documents, code, and history into account
Stand up the model — North Mini Code on Ollama
To start easily, Ollama. It stands up an open-weight model in one line and serves it as an API.
The first one to load is North Mini Code (Cohere) — an open-weight (Apache 2.0) agentic coding model. It is exactly the tool this series centers on: the builder has the AI write the code. A 30B MoE with only 3B active, it is light enough to run with low latency even on local hardware.
It lives on a separate server, not on the one machine from 2-02. Install it the official way, as a systemd service, without Docker (2-02). Listen only on the internal network, reachable from the 2-02 machine alone.
curl -fsSL https://ollama.com/install.sh | sh # official installer; sets up a systemd service
ollama pull north-mini-code-1.0 # open-weight, runs on your own side
You can hit it on OpenRouter's free tier to try it out, but in production you run it yourself, and neither code nor data leaves your side. Cohere is one corner of sovereign AI, alongside Europe's Aleph Alpha (→ blog 027).
For RAG and chat, load a general model (Qwen and the like) and an embedding model separately on the same Ollama. As volume grows, move to the higher-throughput vLLM (on PyPI). Stand one up first, swap as needed.
Put contents into the 2-03 pgvector with RAG
This is the payoff from 2-03. Turn internal documents, code, and mail into embeddings (vectors), put them in pgvector, pull the fragments closest to a question, and have the model answer. This is RAG, retrieval-augmented generation.
# 1) embed a document and put it in the 2-03 pgvector
emb = embed(text) # a local embedding model
pg.execute("INSERT INTO docs(body, embedding) VALUES (%s, %s)", [text, emb])
# 2) pull the closest fragments and have the model answer
hits = pg.execute(
"SELECT body FROM docs ORDER BY embedding <=> %s LIMIT 5", [embed(q)])
answer = llm(f"Answer based on the following sources:\n{hits}\n\nQuestion: {q}")
The table for which 2-03 only had the vessel ready now gets its contents. An AI that answers from your own real data, with citations, stands up on your side.
The AI does not bypass the gate
One thing has to be decided as design, up front. *RAG searches only within the documents the person asking can open* — write this into the spec from day one. Feed the whole company's documents to an AI, and an employee without clearance who asks a question gets, as the answer, the contents of documents they could never open. The gate (2-05) and the document store's permissions (2-07) must not be bypassed here.
Two tiers are enough. By default, what goes into RAG is only the shared knowledge everyone can read, prepared in 2-15 — deciding the scope at preparation time is the most reliable control. If permissioned documents have to be searchable, carry each document's permission on its pgvector rows and filter the search by that person's token. The sources then come only from documents that person can open.
Search, too, happens inside the key. The AI does not bypass the gate.
Stand up Open WebUI as the window people use
The window people use is Open WebUI — a screen resembling ChatGPT or Copilot,
connected to the model and the RAG you stood up. It is on PyPI, so
uv tool install open-webui installs it (no Docker), and it sits on the same AI
server as Ollama. The entrance from outside is Caddy on the 2-02 machine, which
hands traffic, behind the gate, to the AI server's internal port.
ai.example.com { reverse_proxy <the AI server's internal address>:8080 }
Draw the line between in-house and borrowed, honestly
Open models have reached practical sufficiency. But *for the hardest judgment and large-scale code generation, a frontier model (the AI subscribed to in 2-02) is still stronger*. This is the same shape as the outbound mail relay (2-08) and Cloudflare (2-11).
- Keep in-house — processing of confidential data, always-on classification, summarization, RAG (the real body of control)
- Borrow — hard judgment and heavy generation go out to a frontier model's API
Control stays on your side; capability is borrowed to the extent needed. Keeping everything in-house is not the goal. Hold the data and the daily processing in your own hands, and send only the hardest parts out.
And the range you can hold yourself widens as hardware advances. AMD's Ryzen AI Max PRO 400 series (Q3 2026, from ASUS, HP, and Lenovo) carries up to 192GB of unified memory, of which up to 160GB can be allocated as VRAM, and runs 300B-class models locally (AMD's announcement, 2026). Once a machine with this is your AI server, the heavy document RAG that today can only be borrowed comes home — the line of what you borrow recedes year by year.
The value of your own AI is not maximizing cleverness. It is embedding AI into the everyday without letting go of your data.
How to check you are done
This chapter is done when these five hold.
- Opening
ai.example.comin a browser shows the screen only after passing the 2-05 gate's login - Asking about an internal document returns an answer with citations, and opening a citation shows that the document really exists
- Asking the same thing from an account without clearance does not produce that document's contents in the answer
- The list of models stood up on Ollama holds North Mini Code, the general model, and the embedding model
- While the AI writes code, no traffic leaves the internal machine
curl -s localhost:11434/api/tags # the list of models answers
psql -c "SELECT count(*) FROM docs;" # how many documents are in pgvector
curl -sI https://ai.example.com | head -1 # the window's page answers
What the human holds
Values the human supplies
- The window's domain name (
ai.example.com) and the DNS entry that points it at the machine - The scope of documents that go into RAG — by default, the shared knowledge everyone can read, prepared in 2-15
- The database holding pgvector, and its password (2-03)
- The gate (2-05) token settings, so the search is filtered by the asker's permissions
- The API key for the frontier model you borrow, and the line on what may be sent there
Actions the AI states before performing
- Sending an internal document out to an external API
- Widening the scope that goes into RAG (adding permissioned documents)
- Deleting or rebuilding the data in the pgvector tables
- Exposing the window somewhere visible from outside the company
- Starting to run a billed API as always-on processing
Versions checked, and when
- Ollama (the official install script), North Mini Code 1.0 (Cohere, Apache 2.0, a 30B MoE with 3B active; checked in the Ollama library on 2026-10-05), a general model (Qwen and the like) and an embedding model
- Open WebUI 0.11 (PyPI), vLLM 0.31 (PyPI), pgvector (2-03), Caddy (2-11)
- The AMD Ryzen AI Max PRO 400 series is Q3 2026, from ASUS, HP, and Lenovo (AMD's announcement)
- This procedure was written on 2026-07-16 and reviewed on 2026-10-06
- If a version has moved, have the AI confirm the official procedure before proceeding
Summary — what has been stood up
Your own AI, on top of everything.
- Ollama / vLLM — start with North Mini Code (Cohere, an open-weight coding model), add a general model alongside for RAG. On a separate server, installed the official way
- RAG (pgvector) — fill the 2-03 vessel with internal data and answer with citations
- Open WebUI — a ChatGPT-like window, behind the gate; installed from PyPI on the AI server
- In-house versus borrowed — data and always-on processing in-house, the hard parts borrowed from a frontier model
From 2-02 through 2-16, we replaced Microsoft 365 and the vendor packages under the core systems, one at a time, with OSS. The one machine handed to the AI, the foundation, the gate, documents, code, mail, meetings, booking, the web, the API, the information preparation, and the AI — none of it was written; it was stood up.
As written in 1-05 — the effect of OSS is greater than the effect of AI. The generic is already shared with the world.
Once it is all stood up, what comes next is how it runs. The next chapter draws the line on how far the AI you have placed is allowed to go — not autonomously, and with whatever is settled frozen into code and commands. That closes the Independence part.