blog.back_article_list

High availability for n8n: Workers, webhooks and Postgres

High availability for n8n: Workers, webhooks and Postgres

n8n Community edition runs a single main process. A second main needs an Enterprise license, and no load balancer, restart policy or container orchestrator changes that, so most of what follows is about making everything around that one process survive a failure and getting the main itself back within a minute when it goes. That is a less exciting definition of high availability than two editors behind a balancer, but it's the honest one for a self-hosted instance, and it covers the failures that happen in practice: a worker dying mid-run, a server reboot, a Postgres disk filling up, an upgrade that takes longer than planned.

Multi-main setup and the Enterprise license

Multi-main is the mode where several main processes run at once, elect a leader through Redis and share the editor, API and webhook traffic behind a load balancer with sticky sessions. It's switched on with N8N_MULTI_MAIN_SETUP_ENABLED=true, the leader key lives in Redis for N8N_MULTI_MAIN_SETUP_KEY_TTL seconds (default 10) and followers check it every N8N_MULTI_MAIN_SETUP_CHECK_INTERVAL seconds (default 3). It requires queue mode, Postgres, the same n8n version on every instance and a license: the queue mode documentation lists it as Enterprise, self-hosted only, not available on n8n Cloud. Without the key the variable does nothing useful. I'm stating it this flatly because a lot of older guides describe running two editor containers as if it were a configuration choice.

What survives a main process restart in queue mode

Quite a lot, as long as the instance is in queue mode. Workers hold their own connections to Redis and Postgres, so an execution that's halfway through when the main restarts finishes normally and writes its result. Jobs already sitting in Redis get picked up as workers free up. Schedule Triggers are the loss: they fire from the main's memory, a run whose moment passes during the restart is skipped, and the durable scheduler added in 2.36 exists to close that gap by recording runs in the database. New production webhooks are the other loss, since the main is what receives them, and that one has a Community edition fix.

In regular mode none of this applies. One process, one restart, everything in flight is gone. If a single instance matters enough that you're reading about availability, the n8n queue mode compose setup is the first move and the rest of this page assumes you've made it. It takes an afternoon, including the Postgres migration if you're still on SQLite.

Webhook processors behind a load balancer

An n8n process started with the webhook command accepts production webhooks and enqueues them without involving the main, and you can run as many as you want on Community. Two of them on two servers behind a proxy that routes /webhook/ and /webhook-waiting/ to the pair means a Stripe event or a form submission during a main restart still lands in Redis and runs when a worker gets to it. The main comes back, the editor comes back and nobody outside notices. The n8n webhook processor and worker scaling guide has the routing rules and the compose service for a processor. The availability point is narrower: webhook ingress is the one front-door service you can make redundant without a license, and for most automations it's the front door that matters.

Health checks and restart policies in Docker Compose

The main answers on /healthz (process up) and /healthz/readiness (database connected and migrated), and per the n8n monitoring docs those are always on for the main and off on workers unless QUEUE_HEALTH_CHECK_ACTIVE=true. A compose healthcheck against readiness, with restart: unless-stopped, gets a wedged main restarted by Docker without anyone paging you. The fragment below shows all four services; merge it into your existing file.

services:
  postgres:
    image: postgres:17
    restart: unless-stopped
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U n8n -d n8n"]
      interval: 10s
      timeout: 5s
      retries: 5
      start_period: 20s

  redis:
    image: redis:7
    restart: unless-stopped
    command: redis-server --requirepass ${REDIS_PASSWORD} --appendonly yes
    healthcheck:
      test: ["CMD", "redis-cli", "-a", "${REDIS_PASSWORD}", "--no-auth-warning", "ping"]
      interval: 10s
      timeout: 5s
      retries: 5

  n8n-main:
    image: n8nio/n8n:2.38.5
    restart: unless-stopped
    stop_grace_period: 2m
    healthcheck:
      test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:5678/healthz/readiness || exit 1"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 90s
    depends_on:
      postgres:
        condition: service_healthy
      redis:
        condition: service_healthy

  n8n-worker:
    image: n8nio/n8n:2.38.5
    restart: unless-stopped
    command: worker --concurrency=10
    stop_grace_period: 5m
    environment:
      QUEUE_HEALTH_CHECK_ACTIVE: "true"
      N8N_GRACEFUL_SHUTDOWN_TIMEOUT: "300"
    healthcheck:
      test: ["CMD-SHELL", "wget -qO- http://127.0.0.1:5678/healthz || exit 1"]
      interval: 30s
      timeout: 10s
      retries: 3
      start_period: 60s
    depends_on:
      n8n-main:
        condition: service_healthy

The start_period on the main matters more than it looks. Migrations on a big executions table can take a couple of minutes on an upgrade, readiness returns non-200 the whole time and without the grace period Docker counts those as failures, hits retries and restarts the container in the middle of a migration. Ninety seconds is where I landed after one upgrade looped three times before I read the log properly, and on a large database I'd go higher.

The worker healthcheck is the one I forgot the variable for once, the endpoint returned nothing the compose healthcheck marked the worker unhealthy and Docker restarted it every 90 seconds for most of a day. It kept processing jobs between restarts so nothing failed loudly. If a worker's uptime in docker compose ps never passes two minutes, that's what it is.

Postgres availability: streaming replication or managed Postgres

Postgres is the state. Everything else can be rebuilt from a compose file and an encryption key, so this is where the availability budget belongs. The built-in option is streaming replication: a primary with wal_level = replica, a standby created with pg_basebackup, a replication user allowed in pg_hba.conf and the standby following the WAL stream a few milliseconds behind. What the PostgreSQL 17 standby documentation walks through and replication itself doesn't give you is failover: when the primary dies, something has to promote the standby and something has to point n8n at the new address, and that something is Patroni, repmgr or a person with a runbook.

I use a managed Postgres for the instances I care about and keep replication for the ones I don't, because failover automation is where I've seen more outages than the database itself ever caused. A managed service gives you a stable connection string that survives the failover it performs for you. n8n copes with the reconnect: it pings the database every DB_PING_INTERVAL_SECONDS (2), counts failures up to DB_PING_MAX_FAILURES_BEFORE_RECOVERY (3) and then enters a recovery loop with backoff between DB_RECOVERY_BACKOFF_MIN_MS (1000) and DB_RECOVERY_BACKOFF_MAX_MS (30000), reconnecting when the endpoint answers again. Set DB_POSTGRESDB_SSL_ENABLED=true and the CA in DB_POSTGRESDB_SSL_CA for a managed endpoint, since it'll be reached over a network you don't own.

Whichever route, the Postgres server lives on a second machine over a private network or on the managed service, never on the same VPS as the main, otherwise the main's host going down takes the state with it and the whole exercise is moot. The n8n private networking layout has the two-server split with the firewall rules. A managed database gives you the same split without the second server to patch.

Redis persistence and Redis Cluster

Redis holds the queue, and the queue is work that has been accepted but not done. Append-only persistence with fsync every second is the baseline and it bounds a crash to one second of lost enqueues. For redundancy of Redis itself, the queue-mode env reference has QUEUE_BULL_REDIS_CLUSTER_NODES, a comma-separated host:port list for connecting to Redis Cluster, and QUEUE_BULL_REDIS_TLS for encrypted connections. I found no Sentinel-specific variable on that page, and I haven't run n8n against a Sentinel setup, so I can't say if Bull's Sentinel support is reachable through n8n's config or not. On the deployments I run, Redis is a single AOF-backed container next to Postgres on the datastore server, and a Redis loss means at most one second of webhooks, which the senders retry anyway.

Upgrades with one main process

An upgrade is a short full outage on Community, and planning for it beats pretending otherwise. Main and workers must run the same image tag, so you don't upgrade the workers first and let them soak; you take everything down together. The sequence I use:

docker compose pull
docker compose stop n8n-worker        # drains, up to stop_grace_period
docker compose up -d n8n-main         # runs migrations on the new version
docker compose up -d n8n-worker

Webhooks fail for the duration and schedules that fall inside it are skipped, so the window goes at the quietest hour. From the pull finishing to readiness returning 200 takes about 45 seconds on our WHMCS automation instance, with a 400 MB Postgres; migrations are what stretch it. A blue/green variant, where the new version is started as a second compose project against a restored copy of the database and traffic is switched at the proxy, works for testing the upgrade but doesn't help production: the executions that ran on the old stack during the test aren't in the new database. Pin the tag, upgrade monthly, read the release notes for the versions you're skipping over.

Backups as the disaster recovery plan, RTO and RPO

For a single-main deployment the recovery plan is a backup and a rehearsed restore, and the two numbers to write down are how much data you can lose (RPO) and how long the rebuild takes (RTO). A nightly pg_dump gives an RPO of up to 24 hours; WAL archiving or a managed service's point-in-time recovery brings it to minutes. RTO is the time to provision a VPS, restore the dump, put the encryption key in place and start compose, and with the n8n backup and restore guide open next to you that's 15 to 20 minutes if you've done it before and an afternoon if you haven't.

A short aside on the numbers people quote. Five nines is 5 minutes and 15 seconds of downtime a year; a single monthly upgrade at 45 seconds spends 9 of those minutes. Three nines is 8 hours and 46 minutes a year, and a single-main queue-mode instance with health checks, a replicated database and a tested restore can sit inside that comfortably. If a contract says five nines, the answer is the Enterprise license and multi-main, or a different architecture where n8n isn't in the synchronous path.

Automate faster, for less

Bring your winning ideas to life with AMD power, NVMe speed and unmetered bandwidth.