Back to Article List

Self-hosted RAG with Ollama, Qdrant and EmbeddingGemma 2

Self-hosted RAG with Ollama, Qdrant and EmbeddingGemma 2

By the end of this guide one Linux server will read a folder of your documents, turn every passage into a vector with EmbeddingGemma 2, store those vectors in Qdrant and answer questions with a local model that cites the passages it used.

Nothing goes to an outside API, which is the reason to build RAG (retrieval-augmented generation, an LLM answering from documents you hand it) yourself instead of renting it.

Google released EmbeddingGemma 2 on 6 October 2026, and Ollama had it in its library the same day. It's a good fit for a server: the text-only build is small enough to sit next to a chat model without fighting it for memory, and it reads 8,192 tokens of context per input, four times what the first EmbeddingGemma accepted.

Everything below assumes Ubuntu 24.04, Ollama already installed and a sudo user.

Server requirements for a private RAG stack

Three things run on the box. Ollama serves two models, the embedder and the chat model. Qdrant stores vectors and runs the similarity search. A Python script of about 70 lines glues them together.

If Ollama isn't on the server yet, follow the Ollama install guide for Ubuntu first. Its default leaves the API on 127.0.0.1:11434. Keep it there: the script talks to Ollama locally, and nothing in this setup needs the port open to the internet.

The embedder is cheap: the bf16 text build is 574 MB on disk. The chat model sets the hardware bar. This guide uses gemma4:e4b, which Ollama lists at 6.6 to 9.5 GB depending on the build, comfortable on a 16 GB GPU and slow on CPU. The Ollama hardware requirements guide has the RAM and VRAM tiers for bigger models. On a CPU-only VPS, drop to gemma4:e2b and expect answers measured in seconds per sentence.

Qdrant's memory is easy to estimate. A 768-dimension float32 vector is 768 × 4 bytes, so about 3 KB. A hundred thousand chunks is roughly 300 MB of raw vectors before the HNSW index and payloads are added on top. Most document collections a team would point at this never get close to that.

The whole stack fits on one GPU VPS with an RTX PRO 2000 Blackwell card. Its 16 GB of VRAM holds the embedder and a mid-sized chat model at the same time, so neither gets unloaded between an ingest and a question.

Pull EmbeddingGemma 2 in Ollama

Check the Ollama version first. The model is a day old, so an older install may refuse to load it; if the pull or the first embed call fails, upgrading Ollama is the first fix to try.

ollama -v
ollama pull embeddinggemma-2:270m-bf16-text

The tag matters more than usual. EmbeddingGemma 2 is one 740M-parameter model with a 270M text component, a 170M vision encoder and a 300M audio encoder. Ollama packages slices of it:

  • 270m: text only, the one this guide uses
  • 440m: text plus the vision encoder, listed with image input
  • 570m: text plus the audio encoder, though the Ollama tags page lists only text input for it
  • 740m (also latest): the whole model, listed with text and image input

Each size comes in bf16, mxfp8 and nvfp4 builds, and the short tags (270m, latest) carry the same digests as the nvfp4 ones. This guide pins the bf16 tag instead, the only build that matches the precision Google's model card asks for: bfloat16 or float32, because in float16 the model "returns NaN or silently degraded embeddings rather than raising an error." A quantized default may be fine, but I haven't tested the nvfp4 builds on a CPU-only Linux server, and silent degradation is the kind of fault nobody spots until retrieval quality is already bad.

Confirm the model answers and returns the expected size:

curl -s http://127.0.0.1:11434/api/embed \
  -d '{"model": "embeddinggemma-2:270m-bf16-text", "input": "test"}' \
  | python3 -c "import sys, json; print(len(json.load(sys.stdin)['embeddings'][0]))"

It prints 768. That number becomes the size of the Qdrant collection, so write it down.

Task prefixes EmbeddingGemma 2 expects

Every input should start with a short instruction, and queries and documents get different ones. The EmbeddingGemma 2 model card on Hugging Face lists seven query tasks and one document format. On documents it says:

Documents with a real title should be formatted as title: {title} | text: {content}. Use title: none when no title is available.

Queries take a task name. task: search result | query: is the general search form and task: question answering | query: is tuned for questions. People type questions into a RAG box, so the script uses the second one. Swap it for the search form if your users paste keywords instead.

Ollama passes the text through to the model, and the library page describes the prefixes as a feature without saying who adds them. Check the template on your own install:

ollama show embeddinggemma-2:270m-bf16-text --template

If the output is empty or a bare {{ .Prompt }}, nothing is added for you and the prefixes are the script's job, which is how the script below is written. Skipping them raises no error. Retrieval just gets worse, by an amount you'd only measure with a test set.

Run Qdrant on localhost with an API key

Qdrant's shipped config.yaml sets host: 0.0.0.0 and leaves api_key empty. Run the container as-is with a plain -p 6333:6333 and anyone who finds port 6333 can read, overwrite or delete your index. Two changes fix that: publish the port on 127.0.0.1 only and set a key.

sudo apt update && sudo apt install -y docker.io
export QDRANT_API_KEY=$(openssl rand -hex 32)
echo "QDRANT_API_KEY=$QDRANT_API_KEY" >> ~/.rag-env

sudo docker run -d --name qdrant --restart unless-stopped \
  -p 127.0.0.1:6333:6333 \
  -v qdrant_storage:/qdrant/storage \
  -e QDRANT__SERVICE__API_KEY=$QDRANT_API_KEY \
  qdrant/qdrant:v1.19.2

The version is pinned on purpose. Qdrant's advisory GHSA-f632-vm87-2m2f describes an arbitrary file write through the /logger endpoint that needed only read-only access, in versions 1.9.3 to 1.15.5. The 1.19.2 notes add a fix that stops read-only keys from working on the internal gRPC API. Pinning means you upgrade on purpose, after reading the notes, instead of whenever the container happens to restart.

curl -s -H "api-key: $QDRANT_API_KEY" http://127.0.0.1:6333/collections

An empty collection list means Qdrant is up and accepting the key. The same request without the header should be refused.

Write the ingestion and query script

Python 3.12 ships with Ubuntu 24.04. Give the script its own virtual environment:

sudo apt install -y python3-venv
mkdir -p ~/rag && cd ~/rag
python3 -m venv .venv
.venv/bin/pip install qdrant-client requests

Save this as ~/rag/rag.py:

import os, sys, uuid, pathlib, requests
from qdrant_client import QdrantClient, models

OLLAMA = os.environ.get("OLLAMA_URL", "http://127.0.0.1:11434")
EMBED_MODEL = os.environ.get("EMBED_MODEL", "embeddinggemma-2:270m-bf16-text")
CHAT_MODEL = os.environ.get("CHAT_MODEL", "gemma4:e4b")
COLLECTION = "docs"
DIM = 768

qdrant = QdrantClient(url=os.environ.get("QDRANT_URL", "http://127.0.0.1:6333"),
                      api_key=os.environ["QDRANT_API_KEY"])

def embed(texts):
    r = requests.post(f"{OLLAMA}/api/embed",
                      json={"model": EMBED_MODEL, "input": texts,
                            "options": {"num_ctx": 8192}}, timeout=600)
    r.raise_for_status()
    return r.json()["embeddings"]

def chunks(text, size=1500):
    buf = ""
    for para in text.split("\n\n"):
        if buf and len(buf) + len(para) > size:
            yield buf.strip()
            buf = ""
        buf += para + "\n\n"
    if buf.strip():
        yield buf.strip()

def ingest(folder):
    if not qdrant.collection_exists(COLLECTION):
        qdrant.create_collection(COLLECTION, vectors_config=models.VectorParams(
            size=DIM, distance=models.Distance.COSINE))
    for path in sorted(pathlib.Path(folder).rglob("*")):
        if path.suffix not in (".md", ".txt"):
            continue
        parts = list(chunks(path.read_text(errors="ignore")))
        for start in range(0, len(parts), 32):
            batch = parts[start:start + 32]
            vectors = embed([f"title: {path.stem} | text: {p}" for p in batch])
            qdrant.upsert(COLLECTION, points=[
                models.PointStruct(
                    id=str(uuid.uuid5(uuid.NAMESPACE_URL, f"{path}#{start + i}")),
                    vector=v,
                    payload={"source": str(path), "text": p})
                for i, (p, v) in enumerate(zip(batch, vectors))])
        print(f"{path}: {len(parts)} chunks")

def ask(question, k=5):
    vector = embed([f"task: question answering | query: {question}"])[0]
    hits = qdrant.query_points(COLLECTION, query=vector, limit=k).points
    context = "\n\n".join(f"[{n}] ({h.payload['source']})\n{h.payload['text']}"
                          for n, h in enumerate(hits, 1))
    r = requests.post(f"{OLLAMA}/api/chat", json={
        "model": CHAT_MODEL, "stream": False, "options": {"num_ctx": 8192},
        "messages": [
            {"role": "system", "content": "Answer only from the numbered context. "
             "Cite sources as [n]. If the context does not contain the answer, say so."},
            {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}]},
        timeout=600)
    r.raise_for_status()
    print(r.json()["message"]["content"])
    for n, h in enumerate(hits, 1):
        print(f"[{n}] {h.score:.3f} {h.payload['source']}")

if __name__ == "__main__":
    if sys.argv[1] == "ingest":
        ingest(sys.argv[2])
    else:
        ask(" ".join(sys.argv[2:]))

A few choices in there are deliberate.

Chunks are packed from whole paragraphs up to about 1,500 characters, with no overlap. Overlap is the default in most RAG tutorials, and it has a cost. One developer's post in an r/Rag thread on chunk overlap put a number on it: 512-token chunks with 25% overlap left 70% duplicate content in the top five results. Paragraph boundaries keep a thought together. A long paragraph can overshoot the limit, which is why embed() passes num_ctx 8192: Ollama loads the embedder with its server default otherwise and cuts longer input without an error, while EmbeddingGemma 2 itself reads up to 8,192 tokens.

Point IDs are UUIDs derived from the file path and chunk number. Run ingest twice and the second run overwrites the same points instead of doubling them. That only holds while a file keeps its chunk count. If a document shrinks, its old trailing chunks stay behind until you delete them, covered at the end.

The collection uses cosine distance, and Qdrant normalizes vectors on upload for cosine collections. Ollama's /api/embed already returns normalized vectors, so this costs nothing now and saves you later if you truncate dimensions.

num_ctx is set to 8192 in the chat call as well. Five chunks of 1,500 characters plus the instructions come to around 2,000 tokens, which fits Ollama's current 4,096-token default, but a few oversized paragraphs push past it and older installs default to 2,048. The Ollama context window guide covers how those defaults are picked. A window shorter than the prompt gets the prompt cut, which looks like a model ignoring its sources.

Ingest a folder of documents

ollama pull gemma4:e4b
cd ~/rag
set -a; . ~/.rag-env; set +a
.venv/bin/python rag.py ingest /srv/docs

The script reads .md and .txt files and prints one line per file with its chunk count. PDFs and Word files need converting to text first; pdftotext from the poppler-utils package handles most PDFs. The Python client warns "Api key is used with an insecure connection." On 127.0.0.1 nothing leaves the machine, so it's safe to ignore here and a real problem the day QDRANT_URL points at another host.

Ask a question

.venv/bin/python rag.py ask "how long do we keep nightly backups"

You get the model's answer with [n] markers, then the five retrieved chunks with their cosine scores and source files. Read the scores before blaming the chat model. If the right file isn't in the list, no prompt will fix the answer, and the problem sits in chunking, prefixes or the documents themselves.

Store fewer dimensions with Matryoshka truncation

The first 512, 256 or 128 numbers of an EmbeddingGemma 2 vector still work as a smaller embedding, because the model was trained that way (Matryoshka representation learning, nested sizes inside one vector). Google's model card gives the cost on MTEB:

DimensionsMultilingualEnglishCode
76861.3668.4678.68
51261.1768.4177.24
25660.4167.7876.18
12857.8965.6871.41


Those are Google's own numbers, not a test on your documents. At 256 dimensions the raw vectors take a third of the memory (the HNSW graph and the stored chunk text don't shrink), for under one point on the multilingual and English scores and about 2.5 on code. Ollama's embed endpoint takes a dimensions field in the request body, so the change is "dimensions": 256 in embed() and DIM = 256, applied to a new collection. The model card says a truncated vector has to be re-normalized before cosine similarity; a Qdrant cosine collection does that on upload, but if you ever compare vectors yourself in NumPy, normalize them first.

For a few thousand documents, 768 is the sensible setting. At 3 KB a vector, memory only becomes the limit somewhere in the millions of chunks.

Fix "Vector dimension error" after changing the model

A Qdrant collection's size is fixed when you create it. Point the script at a different embedding model, or change dimensions, and the upsert fails with an error shaped like this:

Wrong input: Vector dimension error: expected dim: 768, got 1536

That one is easy because it fails loudly. The quiet version is worse: two models that both output 768 dimensions fit the same collection without complaint, but their vector spaces have nothing to do with each other, so a query embedded by the new model lands next to random chunks embedded by the old one. Open WebUI's RAG documentation says the same about its own index: change the embedding model and you re-embed everything.

The safe migration is a second collection. Set COLLECTION = "docs_v2", ingest everything again with the new model, compare a handful of known questions against both, then switch the name in the script. The old collection stays as the rollback until you delete it.

Remove chunks for deleted or shortened files

Every point carries its source path in the payload, so cleaning up one file is a filtered delete. From ~/rag, after set -a; . ~/.rag-env; set +a, start .venv/bin/python and paste:

from rag import qdrant, COLLECTION
from qdrant_client import models

qdrant.delete(COLLECTION, points_selector=models.FilterSelector(filter=models.Filter(
    must=[models.FieldCondition(key="source", match=models.MatchValue(value="/srv/docs/old-runbook.md"))])))

Run that before re-ingesting a file that got shorter, and the stale tail chunks go with it.

Both services keep their ports on localhost. For one person, an SSH tunnel to 11434 and 6333 is enough to run the script from a laptop. For a team, the reverse proxy and VPN setups in the Ollama VPS hosting guide work for Qdrant's port too, with the API key still required on every request.

Questions?

Can I use pgvector instead of Qdrant?

Yes. If you already run PostgreSQL, the vector extension stores the same 768-dimension vectors in a table column, and backups become part of your existing pg_dump. Swap the Qdrant calls for an INSERT and an ORDER BY embedding undefinedlt;=undefinedgt; $1 LIMIT 5 query. Qdrant is the easier start on a fresh server because it needs no schema and does payload filtering out of the box.

About the author

Søren K

I am Danish and work as an independent technical writer. Before that I studied software development in Aarhus and worked as a developer for a little under two years. I was competent at it. I did not particularly want to do it forever (or at all if I'm honest).

Since 2023 I have not had a permanent address. I usually stay somewhere for at least six weeks - three months is better. Work is mostly Linux and infrastructure, and sometimes automation. I take on documentation projects too, but I prefer work where I can install the thing myself and understand it before writing. I climb indoors, shoot 35 mm film and read a lot on trains. I do not keep a travel blog... There are already enough travel blogs.

Race towards the future

Unrivaled speed meets competitive pricing

Ready in seconds 7-day money-back guaranteeA risk-free way to try LumaDock. Covers the GPU VPS plan on your first order. Cancel anytime
Számlázási ciklus

GPU.T4

$149.00 Save  13 %
$129.00 havonta
  • Dedikált GPU
  • Tesla T4

  • 16 GB GDDR6vRAM
  • 2560CUDA CORES
  • Virtuális szerver
  • 8 vCPUAMD EPYC
  • 32 GBECC MEMÓRIA
  • 250 GB NVMeTÁRHELY
  • Korlátlan sávszélesség
  • IPv4 & IPv6 mellékelve Az IPv6 támogatás jelenleg nem érhető el Franciaországban, Finnországban vagy Hollandiában.

GPU.ADA4000SFF

$249.00 Save  20 %
$199.00 havonta
  • Dedikált GPU
  • RTX 4000 SFF Ada

  • 20 GB GDDR6 ECCvRAM
  • 6144CUDA CORES
  • Virtuális szerver
  • 16 vCPUAMD EPYC
  • 64 GBECC MEMÓRIA
  • 350 GB NVMeTÁRHELY
  • Korlátlan sávszélesség
  • IPv4 & IPv6 mellékelve Az IPv6 támogatás jelenleg nem érhető el Franciaországban, Finnországban vagy Hollandiában.

GPU.PRO4000SFF

$299.00 Save  20 %
$239.00 havonta
  • Dedikált GPU
  • RTX PRO 4000 Blackwell

  • 24 GB GDDR7 ECCvRAM
  • 8960CUDA CORES
  • Virtuális szerver
  • 16 vCPUAMD EPYC
  • 64 GBECC MEMÓRIA
  • 400 GB NVMeTÁRHELY
  • Korlátlan sávszélesség
  • IPv4 & IPv6 mellékelve Az IPv6 támogatás jelenleg nem érhető el Franciaországban, Finnországban vagy Hollandiában.

GPU.PRO4500

$499.00 Save  20 %
$399.00 havonta
  • Dedicated GPU
  • RTX PRO 4500 Blackwell

  • 32 GB GDDR7 ECCvRAM
  • 10496CUDA CORES
  • Virtual Server
  • 16 vCPUAMD EPYC
  • 64 GBECC MEMORY
  • 450 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

GPU.PRO5000

$699.00 Save  20 %
$559.00 havonta
  • Dedicated GPU
  • RTX PRO 5000 Blackwell

  • 48 GB GDDR7 ECCvRAM
  • 14080CUDA CORES
  • Virtual Server
  • 32 vCPUAMD EPYC
  • 96 GBECC MEMORY
  • 500 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

GPU.PRO6000

$1,199.00 Save  19 %
$969.00 havonta
  • Dedicated GPU
  • RTX PRO 6000 Blackwell

  • 96 GB GDDR7 ECCvRAM
  • 24064CUDA CORES
  • Virtual Server
  • 32 vCPUAMD EPYC
  • 128 GBECC MEMORY
  • 650 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

*VAT excluded.

INCLUDED WITH EVERY PLAN

✓ No setup fees ✓1 Gbps network
✓ Free server monitoring ✓ Firewall management ✓24/7 support ✓ KVM virtualization