Back to Article List

How to self-host Cloudflare Clef-flash on a GPU VPS

How to self-host Cloudflare Clef-flash on a GPU VPS

Clef-flash is a 9B decision model that Cloudflare released on 1 October 2026 under Apache 2.0, and it answers the same request a Jev client already sends: a state, a few typed questions, a probability for every allowed answer. Since Ollama 0.35.1 it is also one ollama pull away, and the 11 GB build fits a 16 GB card. This article installs it that way on Ubuntu 24.04, keeps the port off the internet, then covers the full-precision route with Cloudflare's own Python for long threads and screenshots, and puts the result next to Jev and Laya so you can pick.

If you already run Laya, the short version is this. Laya is 421 million parameters, reads 512 tokens of English and answers in single-digit milliseconds. Clef-flash is 9 billion parameters, reads up to 64k tokens, accepts images and wants a GPU. They solve the same problem at two different sizes, and the choice is mostly about how long your state is.

What Clef and Clef-flash are

Cloudflare trained two models. Clef is post-trained from Qwen3.8-27B, Clef-flash from Qwen3.5-9B. In both cases the Qwen model is frozen and the work is done by rank-256 LoRA adapters plus a small transformer that Cloudflare calls the joint schema head. The Qwen model runs one prefill pass over the state and the questions; the head reads the final hidden states and scores every option of every question in the same pass. A per-question softmax turns the scores into probabilities. Nothing is generated token by token, which is why the median latency Cloudflare reports for Clef-flash is 38.8 ms and not the seconds a chat model would take to write "billing".

Three question types, the same three Jev and Laya use, and the Cloudflare/clef-flash card documents all of them. noul is a yes/no question and returns the probability of true. choice takes a dictionary of named options with a one-line description each and returns the winner, a confidence and the full distribution. score takes an ordered list of levels and returns the expected level plus a legend. The card states that the API is fully compatible with Jev and SystemOne, and the request format in the shipped code matches: model, state, questions, with optional images and videos.

The bigger Clef carries Cloudflare's headline benchmarks and runs hosted on Workers AI as @cf/cloudflare/clef. At full precision it is a 96 GB card job, and the FAQ has the Ollama numbers for it. Clef-flash is the model most teams will put on a VPS, so it is the model this article installs.

Which GPU VPS Clef-flash needs

Two routes, two budgets. Ollama's library build of Clef-flash is 11 GB on disk and a little more in the GPU once loaded, and it is what a 16 GB card runs. The full-precision weights are 18.2 GB in BF16, which is what Cloudflare evaluated and what their own code loads; that is a 24 GB card. Cloudflare's Michelle Chen told The Register that a single request at the full 64k window takes 41 GB, so the whole context window is a 48 GB job on either route. Most support states are a few thousand tokens, which is why the two smaller cards get most of the work.

Between the floor and the ceiling sits the length of the state you allow. On Ollama that is OLLAMA_CONTEXT_LENGTH; in Cloudflare's code it is max_length, 16,384 by default, with max_state_tokens to cut the state first. LumaDock's GPU VPS lineup runs from 16 GB to 96 GB on Blackwell cards with full PCIe passthrough; the 16 GB plan runs route 1, the 24 GB and 32 GB ones run route 2 for tickets and threads and the 48 GB one runs the full window. I like to have something measured before writing a recommendation, and the thing to measure here is memory, so both routes below print nvidia-smi after the first request.

Route 1: Run Clef-flash with Ollama on Ubuntu 24.04

Ollama added decision models in version 0.35.0 and Clef in 0.35.1, both at the end of September 2026, with the same /v1/systemone endpoint the hosted Jev API uses. The library build of Clef-flash is an 11 GB 8-bit file, images included. On a fresh Ubuntu 24.04 VPS with a 16 GB or larger card the whole route is four steps.

Step 1: Install Ollama and pull the model

nvidia-smi
curl -fsSL https://ollama.com/install.sh | sh
ollama --version
ollama pull clef-flash

nvidia-smi has to show your card before anything else; if it fails, run sudo ubuntu-drivers install, reboot and come back. The install script detects the NVIDIA driver, puts the binary in place and registers an ollama systemd service that starts at boot, so there is no unit file to write on this route. The version line has to read 0.35.1 or later; an older Ollama pulls the model and then answers /v1/systemone with a 400, because the runner that scores options is missing. The pull is about 11 GB. clef-flash:9b-q8_0 is the explicit tag for the same build if you prefer to pin it.

Step 2: Make the first decision

curl -s http://localhost:11434/v1/systemone \
  -H "content-type: application/json" \
  -d '{
    "model": "clef-flash",
    "state": "Hi, I was billed twice this month for the same server. Refund one of them or I cancel.",
    "questions": {
      "department": {
        "type": "choice",
        "instructions": "Which team should handle the message?",
        "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"}
      },
      "urgency": {
        "type": "score",
        "instructions": "How soon does this need an answer?",
        "criteria": ["Can wait", "This week", "Today"]
      },
      "churn_risk": {
        "type": "noul",
        "instructions": "Is the customer threatening to cancel?"
      }
    }
  }'

The first call loads 11 GB into the GPU, so it takes longer than the ones after it. The response has model, answers and usage. Under answers.department you get choice, probabilities for both options and a confidence; under answers.urgency a score between 0 and 2 (a probability-weighted average, so 1.7 is a legal answer), a legend mapping index to label, the probabilities and a confidence; under answers.churn_risk a single number, the probability of true. For this ticket, billing should win and churn_risk should sit close to 1.

The two shapes of criteria are the thing to get right. A choice question takes a dictionary of option key to description, 2 to 255 options; a score question takes a list of 2 to 26 levels, lowest first, and the index in that list is the score. A noul question needs no criteria. Ollama's docs allow 1 to 64 questions per request and a 64 KiB body without images.

Screenshots go in an images array of base64 strings (PNG, JPEG or WebP; URLs are not accepted) and are shared by every question in the request, with the body limit raised to 32 MiB. A screenshot of the broken checkout page next to the ticket text is the case where Clef-flash does something Jev cannot.

Step 3: Set the context length and keep the model loaded

Two defaults will bite on a server. Ollama loads models with a 4,096-token context window, and the decision endpoint never truncates: a state that does not fit comes back as a 400. And a model is unloaded after five minutes without a request, so the first ticket after a quiet half hour pays the load time again.

sudo systemctl edit ollama.service

In the editor, add:

[Service]
Environment="OLLAMA_CONTEXT_LENGTH=16384"
Environment="OLLAMA_KEEP_ALIVE=-1"
sudo systemctl daemon-reload
sudo systemctl restart ollama
nvidia-smi --query-gpu=memory.used,memory.total --format=csv

16,384 tokens covers a long support thread and matches the default in Cloudflare's own code. The cost is memory: the cache for a longer window sits on top of the 11 GB of weights, and the nvidia-smi line after the next request is where you see how much. On a 16 GB card, watch that number before going higher; the model's full 64k window is a route 2 question. OLLAMA_KEEP_ALIVE=-1 keeps Clef-flash resident, which is what you want on a box that exists to answer decisions. OLLAMA_NUM_PARALLEL stays at its default of 1: one request at a time, which keeps the memory budget honest.

Step 4: Keep port 11434 off the public internet

Ollama has no authentication. It binds to 127.0.0.1 by default, which is why the curl above worked from the same machine and nothing else could reach it. To let your application server call it, bind to all interfaces and then let the firewall decide who gets in.

sudo systemctl edit ollama.service

Add one more line under the two from step 3:

Environment="OLLAMA_HOST=0.0.0.0:11434"
sudo systemctl daemon-reload
sudo systemctl restart ollama
sudo ufw allow from 203.0.113.10 to any port 11434 proto tcp
sudo ufw status

Replace 203.0.113.10 with the address of the server that will send decisions, and keep the default deny for everything else. An Ollama port open to the internet is a free GPU for whoever scans it first. For a client outside your network, put nginx or Caddy in front with TLS and basic auth, or route the traffic over WireGuard; Ollama's own FAQ shows the nginx block.

Route 2: Full precision with Cloudflare's own code

Take this route when the state is long, when images matter to the decision or when you want the BF16 weights Cloudflare evaluated rather than an 8-bit build. It needs a 24 GB card, a Python environment, an 18 GB download and a small server that Cloudflare did not ship. The work is in the server; the model code itself is three function calls.

Everything below runs as a dedicated clef user on the same Ubuntu 24.04 VPS, with the NVIDIA driver already checked in route 1. If you skipped route 1, run nvidia-smi first and install the driver with sudo ubuntu-drivers install if it is missing.

Step 1: Create the user and install the Python packages

sudo useradd -r -m -d /opt/clef -s /usr/sbin/nologin clef
sudo -u clef python3 -m venv /opt/clef/.venv
sudo -u clef /opt/clef/.venv/bin/pip install "torch>=2.11,<2.12" "transformers==5.10.2" huggingface_hub safetensors pillow fastapi "uvicorn[standard]"
sudo -u clef /opt/clef/.venv/bin/python -c "import torch; print(torch.cuda.is_available(), torch.cuda.get_device_name(0))"

The torch wheel on PyPI for Linux includes CUDA, so there is no separate toolkit to install; it is a large download. The last line must print True and your card's name. False with a working nvidia-smi means the driver is older than the CUDA build in the wheel; the PyTorch install page lists wheels built against older CUDA versions.

transformers==5.10.2 is pinned on purpose. joint_schema_model.py imports Qwen3_5ForConditionalGeneration from transformers, and that class exists only in recent releases. fastapi and uvicorn are for the server in step 4 and are not part of Cloudflare's release.

Step 2: Download the model

sudo -u clef /opt/clef/.venv/bin/hf download Cloudflare/clef-flash --local-dir /opt/clef/models/clef-flash
ls -lh /opt/clef/models/clef-flash

Expect about 18 GB of sharded safetensors for the Qwen model plus joint_head.safetensors, joint_head_config.json, joint_schema_model.py and the tokenizer and processor configs. The model card uses snapshot_download from Python instead; the hf command does the same thing and leaves the files where systemd can find them.

Step 3: Make the first decision from Python

Save this as /opt/clef/first.py. It is the model card's example with a hosting ticket as the state.

import json, sys
sys.path.insert(0, "/opt/clef/models/clef-flash")
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model("/opt/clef/models/clef-flash", device="cuda")

request = {
    "model": "clef-flash",
    "state": "Hi, I was billed twice this month for the same server. Refund one of them or I cancel.",
    "questions": {
        "department": {
            "type": "choice",
            "instructions": "Which team should handle the message?",
            "criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
        },
        "urgency": {
            "type": "score",
            "criteria": ["Can wait", "This week", "Today"],
        },
        "churn_risk": {
            "type": "noul",
            "instructions": "Is the customer threatening to cancel?",
        },
    },
}

print(json.dumps(systemone(model, processor, request), indent=2))
sudo -u clef /opt/clef/.venv/bin/python /opt/clef/first.py
nvidia-smi --query-gpu=memory.used,memory.total --format=csv

The first run loads 18 GB into the GPU, which takes a while from NVMe and longer from a cold page cache. The output is a JSON object with model, answers and usage. Under answers.department you get choice, confidence and probabilities for the two options; under answers.urgency a score (the expected level index), a confidence, a legend mapping index to label and the probabilities; under answers.churn_risk the probability of true. usage.output_tokens is always 0, because nothing was generated.

The second command prints what the process left in the GPU. With a short state the number sits a little above the 18.2 GB of weights, and it grows with the state you send; that is the figure to compare against your card before you raise the limit in step 4.

Note the two shapes of criteria. A choice question takes a dictionary of option name to description. A score question takes a list, and the index in that list is the score. Mixing them up raises a ValueError before anything reaches the GPU.

Step 4: Put a /v1/systemone server in front of it

This is the part Cloudflare did not ship. Save it as /opt/clef/serve.py.

import os, sys
sys.path.insert(0, "/opt/clef/models/clef-flash")
from fastapi import FastAPI, HTTPException, Request
from joint_schema_model import load_release_model, systemone

MODEL_DIR = "/opt/clef/models/clef-flash"
MAX_LENGTH = int(os.environ.get("CLEF_MAX_LENGTH", "16384"))
API_KEY = os.environ.get("CLEF_API_KEY")

model, processor = load_release_model(MODEL_DIR, device="cuda")
app = FastAPI()

@app.post("/v1/systemone")
async def decide(req: Request):
    if API_KEY and req.headers.get("authorization") != f"Bearer {API_KEY}":
        raise HTTPException(status_code=401, detail="bad or missing bearer token")
    body = await req.json()
    body.setdefault("model", "clef-flash")
    try:
        return systemone(model, processor, body, max_length=MAX_LENGTH)
    except ValueError as e:
        raise HTTPException(status_code=422, detail=str(e))

@app.get("/health")
def health():
    return {"ok": True}

Two decisions in that file are deliberate. The handler is async and calls the synchronous systemone directly, so requests are processed one at a time on the GPU; that is what the memory budget in the VRAM section assumes, and the simplest way to never see an out-of-memory error from two long states arriving together. And CLEF_MAX_LENGTH is the knob from the VRAM section, exposed as an environment variable so you can lower it on a 24 GB card without editing code. Requests here follow the same JSON as the Ollama route, with one addition: Cloudflare's code truncates a state that runs past the limit, where Ollama refuses it.

sudo -u clef env CLEF_API_KEY=change-me /opt/clef/.venv/bin/uvicorn --app-dir /opt/clef serve:app --host 0.0.0.0 --port 8000

Then from a second terminal:

curl -s localhost:8000/v1/systemone \
  -H "Authorization: Bearer change-me" \
  -H "content-type: application/json" \
  -d '{"state":"The checkout page returns a 502 for every customer since 09:10.","questions":{"department":{"type":"choice","criteria":{"billing":"Payments or invoices","technical":"Bugs or outages"}},"outage":{"type":"noul","instructions":"Is a service down?"}}}'

Did technical win and did outage come back near 1? Then the server is doing its job. If instead you get CUDA out of memory in the uvicorn log, lower CLEF_MAX_LENGTH to 8192 or 4096 and restart; your tickets are shorter than that anyway.

Step 5: Keep it running with systemd

sudo mkdir -p /etc/clef
echo "CLEF_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/clef/serve.env
sudo chmod 600 /etc/clef/serve.env

Write /etc/systemd/system/clef-serve.service:

[Unit]
Description=Clef-flash System One server
After=network-online.target
Wants=network-online.target

[Service]
User=clef
WorkingDirectory=/opt/clef
EnvironmentFile=/etc/clef/serve.env
Environment=CLEF_MAX_LENGTH=16384
ExecStart=/opt/clef/.venv/bin/uvicorn --app-dir /opt/clef serve:app --host 0.0.0.0 --port 8000
Restart=on-failure
RestartSec=5
TimeoutStartSec=600

[Install]
WantedBy=multi-user.target

TimeoutStartSec=600 matters. systemd's default start timeout is 90 seconds and loading 18 GB of weights can take longer than that on the first boot, after which the unit would be killed mid-load and restarted in a loop.

sudo systemctl daemon-reload
sudo systemctl enable --now clef-serve
journalctl -u clef-serve -f
sudo ufw allow from 203.0.113.10 to any port 8000 proto tcp

Replace 203.0.113.10 with the address of the application server that will call it. The server speaks plain HTTP; for clients outside your network put Caddy or nginx in front for TLS, exactly as you would for Laya or for Ollama.

Point a Jev or Laya client at your server

Change the base URL to http://YOUR_SERVER_IP:11434 for Ollama, with any bearer token (Ollama ignores it, and TypeSafe's SDK insists on one), or to http://YOUR_SERVER_IP:8000 and the key in /etc/clef/serve.env for route 2. The request body is the same JSON the hosted Jev API and laya-serve accept, with these differences to check before you switch traffic:

  • model field: both routes treat model as required, and on Ollama it has to be clef-flash (or the tag you pulled). The route 2 server fills it in when a client omits it; Ollama does not, so a Laya client that leaves it out gets a 400.
  • Length: Ollama refuses a state that does not fit the context window with a 400, like Laya does with a 422. Cloudflare's code truncates it silently instead. On route 2, log usage.input_tokens if you want to catch it.
  • Confidence: Ollama documents its confidence as one minus normalised entropy, the same definition Laya uses. Jev's is (n*p_max - 1)/(n - 1), and the formula in Cloudflare's own code is not documented on the card. Three sources, one field name, at least two definitions; retune any threshold you gate on before you carry it across.
  • Images: Ollama takes base64 strings in images. Cloudflare's Python API takes PIL images, and the route 2 server passes the JSON body through unchanged, so it does not accept images over HTTP; add a base64 field and decode it to PIL if you need that there.

Everything else, question ids, the three types, the shape of answers, is the same, which is the point of the SystemOne format.

Where llama.cpp and vLLM stand

Ollama runs on llama.cpp underneath, so it is reasonable to expect plain llama.cpp to do the same thing. Not yet. Hugging Face lists several GGUF conversions of Clef-flash, including one from ggml-org, and that card says the model needs llama.cpp pull request #29831, "model: add support for clef decision model". As of 8 October 2026 the PR is an open draft: no vision, one sequence per batch and "make GGUF" still on its to-do list. Mainline llama.cpp loads the Qwen weights as a chat model, with no decision head. Ollama shipped its own scoring runner rather than wait, which is why route 1 works today and llama-server does not; our Ollama vs llama.cpp guide covers how the two relate. Expect the gap to close once the PR lands, since Ollama tracks upstream closely.

vllm serve Cloudflare/clef-flash, which the model card lists under deployment, has the same gap from the other side. vLLM loads the Qwen weights and exposes an OpenAI-compatible completions endpoint on port 8000; nothing in the card or in joint_schema_model.py wires joint_head.safetensors into it, and a Jev client posting to /v1/systemone there gets a 404.

Clef-flash vs Jev vs Laya: Context, latency and cost

Clef-flashJevLaya
Where it runsYour GPU, 16 GB and upTypeSafe's hosted APIYour CPU or GPU
Parameters9B (Qwen3.5-9B plus the decision head)Not disclosed421M (ModernBERT-large); 322M multilingual variant
Context64k tokens; the shipped code defaults to 16,38432k tokens512 tokens English; 1,024 multilingual, raisable to 8,192
InputsText, JSON, images, video framesText, JSONText, JSON
Median latency, Cloudflare's run38.8 ms (p95 122.4)524.1 ms (p95 536.0), over the network5.8 ms (p95 222.5)
What you pay forThe GPU, whatever the volumeInput tokens ($0.042 per million at launch); output is freeThe server, whatever the volume
Weights11 GB 8-bit on Ollama, 18.2 GB in BF16, Apache 2.0Closed808 MB English, 647 MB multilingual, Apache 2.0
ServerOllama, or the small one in route 2Hosted POST /v1/systemonelaya-serve

The latency row needs a footnote. Cloudflare measured all three, the two open models on its own hardware and Jev as a hosted API, so Jev's figure includes the trip over the network and the other two do not; TypeSafe quotes 70 to 500 ms end to end. Laya's own card quotes 32.8 ms per question on the GPU the Laya authors used. None of these numbers is yours. What holds across both sources is the ratio: Laya answers in a fraction of the time, and Clef-flash answers in a time that is still far below any generative model.

Accuracy is where the cards disagree in a way I cannot resolve from reading. On Cloudflare's run of the Jev Decision Index, Clef-flash scores 90.93 macro-F1 on BANKING77 and Laya scores 14.29. Laya's own card reports 0.425 on the same dataset, with Jev at 0.870. Different prompts, different checkpoints or a zero-shot setup would explain the gap; the two pages do not say, and I have not reproduced either figure. Read both as vendor numbers.

The same Cloudflare table puts Clef-flash next to Jev, and here the result is split. Clef-flash leads on tool and API selection (98.76 vs 95.75 on BFCL, 93.11 vs 88.19 on API-Bank), on BANKING77 (90.93 vs 79.74) and on phishing detection (75.05 vs 62.55). Jev leads on When2Call (80.97 vs 65.58), on CLINC150 with out-of-scope queries (89.27 vs 66.77) and by a wide margin on GPQA Diamond (78.3 vs 51.0), the benchmark closest to reasoning. Cloudflare ran Jev itself, through the API, so these are a competitor's numbers about Jev.

On Laya, what the two cards agree on is that a 421M encoder trained for typed decisions does well on the four workflows it was fine-tuned for and badly on 77-way intent classification it has never seen, and a 9B model with 64k of context does well on both because it has read the whole ticket.

So which one? Jev when the state is text, fits in 32k and you would rather pay per token than run a GPU. Laya when the state fits in 512 tokens, the labels are few and latency or cost is the constraint: routing short tickets, scoring a form field, gating a webhook. Clef-flash when the state is a long thread, a JSON record with history, a screenshot or a decision that depends on more than a sentence. I think most teams that start with Laya end up wanting both on the same box, Laya for the first cheap pass and Clef-flash for the cases Laya is unsure about. Laya listens on 8000 and Ollama on 11434, so the two run side by side without a change.

Clef-flash vs Jev: When self-hosting pays off

On cost alone, it rarely does. At Jev's launch price a million decisions on a 500-token ticket come to about $21 in input tokens, and the smallest GPU plan that holds Clef-flash costs more than that each month; the GPU only wins on price somewhere above five million such decisions a month. If your states are short text and the bill is the only question, stay on the API.

The reasons to run Clef-flash yourself are the ones a price per token cannot fix. A state longer than Jev's 32k window, such as a full thread with its logs. An image in the decision, which Jev does not take. Tickets and customer records that you would rather not send to a third party, whatever its retention policy says. And the weights themselves: Apache 2.0 means you can fine-tune Clef-flash on your own labels, which with Jev depends on what TypeSafe offers. I would add one more, smaller one. A hosted model can change under the same name, and our look at Jev at launch already found the docs and the launch post quoting different builds at different prices; a file on your disk does not move.

Where Clef-flash is weaker than Laya

It needs a GPU. Laya's smallest useful deployment is a CPU VPS; Clef-flash's is a 16 GB card, and the build that fits there is 8-bit, which Cloudflare has not published accuracy figures for. Its latency is roughly seven times Laya's in Cloudflare's own table. Full precision means a 24 GB card and a server you write yourself. And the card says nothing about languages, so for a Spanish or Romanian ticket queue the multilingual Laya checkpoint is a known quantity and Clef-flash is a test you have to run yourself.

Where Laya is weaker than Clef-flash

512 tokens. That is roughly two paragraphs of a support email, and the English checkpoint reads nothing past it. The multilingual checkpoint goes to 8,192 with max_len raised, which covers a long ticket but not a thread. Laya's card also notes that accuracy drops above about 20 options and that score is its weakest type, and its base checkpoints land near chance on typed decisions zero-shot, so you fine-tune or you use the typed-decisions checkpoint for its four workflows. Clef-flash takes a screenshot of the broken checkout page in the same request as the ticket, reads a 30-message thread whole and, on Cloudflare's run, holds its accuracy on a 77-label intent set. If the decision needs the whole context, Laya cannot see it.

A hosting company's own queue makes the split concrete. "Is this order fraud" from a two-line signup record is a Laya question. "Which of these 40 KB of logs, screenshots and replies explains why the customer's server is unreachable, and does it need a human now" is a Clef-flash question, and until this month it was a question for a model that writes a paragraph before it answers.

Questions?

Can I run Laya and Clef-flash on the same GPU VPS?

Yes. laya-serve uses under 5 GB of VRAM with all three checkpoints loaded and listens on port 8000 by default, so move it to another port with LAYA_PORT or run the Clef-flash server on 8001. On a 32 GB card both fit with the default CLEF_MAX_LENGTH; on a 24 GB card lower it to 8192 first. Route short states to Laya and send the ones it answers with low confidence to Clef-flash.

Race towards the future

Unrivaled speed meets competitive pricing

Ready in seconds 7-day money-back guaranteeA risk-free way to try LumaDock. Covers the GPU VPS plan on your first order. Cancel anytime
Ciclo di fatturazione

GPU.T4

$149.00 Save  13 %
$129.00 Mensile
  • GPU dedicata
  • Tesla T4

  • 16 GB GDDR6vRAM
  • 2560CUDA CORES
  • Server virtuale
  • 8 vCPUAMD EPYC
  • 32 GBMEMORIA ECC
  • 250 GB NVMeDISCO
  • Banda illimitata
  • IPv4 & IPv6 inclusi Il supporto IPv6 al momento non è disponibile in Francia, Finlandia o nei Paesi Bassi.

GPU.ADA4000SFF

$249.00 Save  20 %
$199.00 Mensile
  • GPU dedicata
  • RTX 4000 SFF Ada

  • 20 GB GDDR6 ECCvRAM
  • 6144CUDA CORES
  • Server virtuale
  • 16 vCPUAMD EPYC
  • 64 GBMEMORIA ECC
  • 350 GB NVMeDISCO
  • Banda illimitata
  • IPv4 & IPv6 inclusi Il supporto IPv6 al momento non è disponibile in Francia, Finlandia o nei Paesi Bassi.

GPU.PRO4000SFF

$299.00 Save  20 %
$239.00 Mensile
  • GPU dedicata
  • RTX PRO 4000 Blackwell

  • 24 GB GDDR7 ECCvRAM
  • 8960CUDA CORES
  • Server virtuale
  • 16 vCPUAMD EPYC
  • 64 GBMEMORIA ECC
  • 400 GB NVMeDISCO
  • Banda illimitata
  • IPv4 & IPv6 inclusi Il supporto IPv6 al momento non è disponibile in Francia, Finlandia o nei Paesi Bassi.

GPU.PRO4500

$499.00 Save  20 %
$399.00 Mensile
  • Dedicated GPU
  • RTX PRO 4500 Blackwell

  • 32 GB GDDR7 ECCvRAM
  • 10496CUDA CORES
  • Virtual Server
  • 16 vCPUAMD EPYC
  • 64 GBECC MEMORY
  • 450 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

GPU.PRO5000

$699.00 Save  20 %
$559.00 Mensile
  • Dedicated GPU
  • RTX PRO 5000 Blackwell

  • 48 GB GDDR7 ECCvRAM
  • 14080CUDA CORES
  • Virtual Server
  • 32 vCPUAMD EPYC
  • 96 GBECC MEMORY
  • 500 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

GPU.PRO6000

$1,199.00 Save  19 %
$969.00 Mensile
  • Dedicated GPU
  • RTX PRO 6000 Blackwell

  • 96 GB GDDR7 ECCvRAM
  • 24064CUDA CORES
  • Virtual Server
  • 32 vCPUAMD EPYC
  • 128 GBECC MEMORY
  • 650 GB NVMeSTORAGE
  • Unmetered bandwidth
  • IPv4 & IPv6IPv6 is currently unavailable in France, Finland or the Netherlands. included

*VAT excluded.

INCLUDED WITH EVERY PLAN

✓ No setup fees ✓1 Gbps network
✓ Free server monitoring ✓ Firewall management ✓24/7 support ✓ KVM virtualization