Back to Article List

Scale n8n workers: Concurrency, webhook processors and Redis

Scale n8n workers: Concurrency, webhook processors and Redis - Scale n8n workers: Concurrency, webhook processors and Redis

An n8n worker runs 10 executions at once by default, and a single worker on a 2 vCPU VPS drains most queues before anyone notices they existed. So the question behind "how many workers do I need" is usually a different one: what kind of executions are they, and what is the box doing while they run. This guide is about that question, about taking webhook traffic off the main process, about the Redis and shutdown settings that decide if a restart loses work, and about when the answer stops being a bigger VPS. It assumes queue mode is already running; if it isn't, the n8n queue mode setup with Docker Compose gets you there and this page picks up after it.

How many n8n workers do you need

Concurrency is per worker, so total parallel capacity is workers multiplied by the --concurrency value. What that capacity costs depends on the workflow. An execution that calls three HTTP APIs and writes a row spends nearly all its time waiting on the network, and a worker at concurrency 10 handling nothing but those barely registers on a CPU core. An execution with a Code node chewing through 20,000 items is CPU-bound, and ten of those on a 2 vCPU box just take turns.

Memory is the harder limit on a VPS. A worker container idles at roughly 300 MB on my instances and climbs with the size of the item arrays passing through it, so ten concurrent executions each holding a few thousand JSON rows can push one worker past 1 GB. Code nodes make this worse in a specific way in 2.x: a task runner executes them, and with internal mode (the default) that runner lives as a child process inside the worker container, with its own concurrency cap of 5 from N8N_RUNNERS_MAX_CONCURRENCY and its own heap. Two Code-heavy executions on the same worker can be waiting on the runner while eight slots sit idle.

Starting points I use, and they are starting points:

VPSWorkers x concurrencyFits
2 vCPU, 4 GB1 x 10API glue, webhooks, light Code nodes
4 vCPU, 8 GB2 x 5 or 2 x 10Mixed, some batch processing
8 vCPU, 16 GB3 x 10ETL passes, Code-heavy transforms, AI agents with tools

Then watch n8n_scaling_mode_queue_jobs_waiting on the main's metrics endpoint. Waiting staying above zero for a few minutes while CPU has headroom means add a worker. Waiting above zero while CPU is pegged means the workers are fine and the server is small, and another worker on the same box only adds contention. The other signal is the executions list showing runs that took much longer than the workflow's own steps explain, which is queue time.

Postgres connections are part of this arithmetic. Every n8n process opens its own pool, sized by DB_POSTGRESDB_POOL_SIZE with a default of 2, so a main and six workers hold around 14 connections, well under Postgres's default of 100. Raise the pool size on the workers only if you see acquisition timeouts in the logs (DB_CONNECTION_ACQUISITION_TIMEOUT_MS, 30 seconds by default).

N8N_CONCURRENCY_PRODUCTION_LIMIT on the main process

N8N_CONCURRENCY_PRODUCTION_LIMIT caps how many production executions run at once and defaults to -1, meaning no cap. The docs say it applies in both regular and scaling modes. On workers I don't touch it and let --concurrency do the job; on the main I set it to a small number like 5, because in queue mode the main still runs manual executions (unless OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS is on), and a colleague running three test executions of a big import at the same time as I do is enough to stall the editor on a 2 vCPU main.

Dedicated webhook processors with n8n webhook

The main process receives every webhook, and a burst of them competes with the editor, the schedule triggers and the API. Past a certain rate the fix is a webhook processor: n8n launched with the webhook command, which does nothing except accept production webhooks and push jobs into Redis. It needs the same environment as a worker (EXECUTIONS_MODE=queue, the encryption key, Postgres and Redis), and unlike workers it scales horizontally on the Community edition, so two of them on two servers is allowed without a license. In compose it's a copy of the worker service with a different command and a published port:

  n8n-webhook:
    image: n8nio/n8n:2.38.5
    restart: unless-stopped
    command: webhook
    environment:
      <<: *n8n-env
      N8N_WEBHOOK_URL: https://n8n.example.com/
    ports:
      - "127.0.0.1:5680:5678"

The load balancer or reverse proxy then has to split traffic by path. The queue mode documentation gives the rule: /webhook/* and /webhook-waiting/* go to the webhook processor pool, /webhook-test/* and everything else go to the main, and the main should not be in the webhook pool. In nginx that is two upstreams and three locations:

upstream n8n_main {
    server 127.0.0.1:5678;
}
upstream n8n_webhook {
    server 127.0.0.1:5680;
    # server 10.20.0.12:5680;   second processor on another server
}

server {
    listen 443 ssl;
    http2 on;
    server_name n8n.example.com;

    location /webhook/ {
        proxy_pass http://n8n_webhook;
        include /etc/nginx/snippets/n8n-proxy.conf;
    }
    location /webhook-waiting/ {
        proxy_pass http://n8n_webhook;
        include /etc/nginx/snippets/n8n-proxy.conf;
    }
    location / {
        proxy_pass http://n8n_main;
        include /etc/nginx/snippets/n8n-proxy.conf;
    }
}

The included snippet holds the usual forwarded headers and the WebSocket upgrade lines, the same ones from the n8n behind nginx guide, and they belong on all three locations. Once processors exist you can set N8N_DISABLE_PRODUCTION_MAIN_PROCESS=true on the main so it stops registering production webhooks itself, which makes a misrouted webhook fail loudly at the proxy and never reach the main.

Large webhook responses have their own limit since 2.34: N8N_WEBHOOK_RESPONSE_RELAY_SIZE_MAX caps the response a worker sends back through the main at 64 MiB, and N8N_WEBHOOK_RESPONSE_RELAY_OFFLOAD_ENABLED=true makes the worker park anything larger in binary storage and skip the relay. Most webhooks respond with a few hundred bytes and never touch either.

Redis memory and persistence for the n8n queue

Redis holds job metadata, not execution data, so it stays small. On our instance INFO memory reports about 12 MB used with a few thousand completed jobs retained, and I have never seen it above 100 MB on any n8n deployment. A maxmemory of 256 MB is generous. What matters more is the eviction policy: leave it at noeviction so Redis returns an error when full and keeps every queue key, because a dropped key is an execution that never runs and nothing logs it.

Persistence is the append-only file with fsync every second, which bounds the loss on a hard crash to one second of enqueued jobs. Redis on the same VPS as the workers is fine at every size I have run; the network hop to a second server costs more latency per job than Redis itself ever will, and the reason to move it is usually that Postgres is moving too and you want the datastores together.

Lock duration, stalled jobs and graceful shutdown

Three variables decide what happens to an execution when its worker dies. QUEUE_WORKER_LOCK_DURATION (60000 ms) is how long a worker holds the lease on a job, QUEUE_WORKER_LOCK_RENEW_TIME (10000 ms) is how often it renews, and QUEUE_WORKER_STALLED_INTERVAL (30000 ms) is how often workers scan for jobs whose lease expired. If a worker is killed mid-execution, its jobs sit locked until the lease runs out, a surviving worker finds them at the next stalled check and the executions run again from the first node. n8n has no record of which node they died at, so a workflow that charges a card or sends an email before its last node needs to be idempotent or needs the side effect guarded. QUEUE_WORKER_MAX_STALLED_COUNT, which older tutorials mention, was removed in 2.0 and does nothing now.

Graceful shutdown is what makes a restart not count as a death. N8N_GRACEFUL_SHUTDOWN_TIMEOUT defaults to 30 seconds: on SIGTERM the worker stops taking jobs, finishes what it has and exits, or gets killed at the deadline. Docker's own default is 10 seconds, and that one bit me. docker compose down sends SIGTERM, waits 10 seconds and sends SIGKILL, so the n8n setting never had a chance. On a Sunday upgrade last winter that killed a 40-second import halfway through and it ran twice after the stalled check. Both numbers have to agree, and the compose side is stop_grace_period:

  n8n-worker:
    image: n8nio/n8n:2.38.5
    command: worker --concurrency=10
    stop_grace_period: 5m
    environment:
      <<: *n8n-env
      N8N_GRACEFUL_SHUTDOWN_TIMEOUT: "300"
      QUEUE_HEALTH_CHECK_ACTIVE: "true"

Set the n8n timeout to the longest execution you expect plus margin, and the grace period to at least that. A rolling worker restart with those in place is docker compose stop n8n-worker, which drains, then docker compose up -d n8n-worker; the main and the other workers keep going throughout.

Durable scheduler for Schedule Trigger workflows

Schedule Triggers fire from the main process's memory, so a main restart at 02:00 skips the 02:00 run. n8n 2.36 made the durable scheduler generally available (preview from 2.32) as the fix: runs are recorded in the database and claimed from there, they survive restarts and a missed run inside N8N_SCHEDULER_MISFIRE_GRACE (60 seconds by default) can be caught up according to a per-workflow misfire policy. It's off by default and needs two switches, N8N_SCHEDULER_ENABLED=true and N8N_USE_WORKFLOW_PUBLICATION_SERVICE=true, on the main; without the second one Schedule Triggers stay on the in-memory scheduler. Poll triggers have their own flag, N8N_SCHEDULER_POLL_TRIGGERS_ENABLED, which the durable scheduler docs mark as not fully stable. I have it running on a staging main and it does what it says after a restart; how it behaves over months with a few hundred scheduled workflows, I can't tell you yet, so production here is still on the in-memory scheduler.

One large VPS or several small servers

A single 8 vCPU VPS running main, three workers, Postgres and Redis is simpler than four small servers and, for most n8n deployments, faster: every hop between containers is loopback, Postgres round-trips are sub-millisecond and there is one firewall and one compose file. Vertical scaling stops paying when one of two things happens. The first is CPU contention between Postgres and the workers during batch jobs, which shows as executions slowing down while the queue is empty. The second is wanting a worker restart or a worker host failure to not touch the database. Either one is the moment to put Postgres and Redis on a second server, or to put the workers there, over a private network; the n8n private networking guide has that layout with the firewall rules. Workers on another host need only the shared environment block and a route to Postgres and Redis; they don't need the main to be reachable.

Past roughly 16 vCPU of workers, or when Code nodes run for minutes at a time on every execution, a VPS with shared cores starts to cost more in noisy-neighbour variance than in money, and a dedicated server for the worker pool is the usual next step. Few n8n deployments get there. The ones that do are almost always running ETL or AI agent tool loops, and they tend to be better served by moving that one class of workflow to its own workers than by growing everything.

Automate faster, for less

Bring your winning ideas to life with AMD power, NVMe speed and unmetered bandwidth.