What “private” should mean
Private RAG means your documents, the embeddings made from them, the questions people ask and the answers generated all stay inside your network. That is a stricter test than “we host the app ourselves”: a single call to an external embedding or language-model API breaks it.
The components to self-host
- Ingestion: parsers for PDF, Office files and email, plus OCR for scans.
- Chunking: splitting documents into passages that keep their meaning and their source.
- Embedding model: turns passages and questions into vectors.
- Vector database: for example Qdrant or Milvus, both of which you can run yourself.
- Language model: an open-weight model, often quantised, running on your own GPU.
- Optional reranker: reorders retrieved passages before generation.
- API and interface: where permissions are enforced and answers are shown with their sources.
Put access control in retrieval
Store each passage's permissions as metadata and filter on them when searching, so the model only ever sees passages the person is allowed to read. Filtering after generation is too late, because the answer may already contain the restricted text.
Close the door on outbound traffic
- Block outbound network access from the RAG services by default and allow only what is required.
- Turn off telemetry in libraries and tools.
- Download model files once, verify them and store them internally, so nothing fetches weights from the internet at runtime.
- Keep logs inside the perimeter, and decide how much query text you really need to store.
Evaluate before you tune
Collect a set of real questions from the people who will use the system, together with the documents that answer them. First measure whether the right passages are retrieved. Many poor answers come from retrieval, not from the language model, and changing the model will not fix them. Keep the set versioned so you can compare changes fairly.
Check the model licence
Open-weight models come with licences, and some include conditions on use. Confirm that your use case is permitted before you build on a model, and record which version you deployed.
Plan for operations
- Re-index when documents change, and re-embed everything if you change the embedding model.
- Back up the vector database and the configuration.
- Watch answer quality over time with the same evaluation set.
- Decide who owns updates to the model and the pipeline.
