Back to Article List

n8n troubleshooting: Common errors on a self-hosted VPS

n8n troubleshooting: Common errors on a self-hosted VPS

Before you search for an error string, get the logs into a state where you can read them. Most n8n problems on a VPS announce themselves in the first 30 lines of container output, and half the tickets I've answered on the forum were solved by someone pasting those lines instead of describing the symptom. The list below is organised by the message you'll see, verbatim where the string is stable across versions, by symptom where it isn't. Everything assumes n8n 2.x in Docker Compose; the fixes name the variable or config line, not "check your settings".

Read the n8n logs first

The container writes to stdout. Follow it with:

docker compose logs -f --tail=200 n8n

For more detail, set N8N_LOG_LEVEL=debug and restart; the default is info. N8N_LOG_FORMAT=json gives one JSON object per line, which is what you want if the logs go to Loki or anything else that parses them, and N8N_LOG_OUTPUT=file with N8N_LOG_FILE_LOCATION writes to disk inside the container if stdout isn't convenient. Two more for specific cases: DB_LOGGING_ENABLED=true with DB_LOGGING_OPTIONS=query prints every SQL statement (loud, turn it off after), and CODE_ENABLE_STDOUT=true sends console.log from Code nodes to the container log so you can see what a runner is doing.

Two endpoints tell you the state of the process without reading anything: /healthz answers 200 when n8n is up, and /healthz/readiness answers 200 only once the database is connected and migrated. A main that answers the first and not the second is alive and disconnected from Postgres, which narrows things down before you open a single log line.

Startup and database errors

Mismatching encryption keys

Full line: Mismatching encryption keys. The encryption key in the settings file /home/node/.n8n/config does not match the N8N_ENCRYPTION_KEY env var. n8n wrote a random key to that file the first time it started without one, and now your compose file sets a different value. The container refuses to start, which is the right call, since the credentials in the database are encrypted with the old key.

If the instance has credentials you care about, the env var is the wrong one: read the key from the config file (docker compose exec -u node n8n cat /home/node/.n8n/config) and put that value in N8N_ENCRYPTION_KEY. If it's a fresh install with nothing in it, delete the volume and start again with the env var set. Moving to a new key with existing credentials is a separate procedure, covered in the n8n encryption key rotation guide. Either way, write the key down somewhere outside the server before you touch anything else.

Credentials could not be decrypted

Same root cause, seen later: n8n started fine (no env var conflict, because the config file was new or gone) but the database came from an instance with a different key. Typical after restoring a Postgres dump onto a new server without restoring the n8n_data volume, or after the volume was recreated. The fix is the original key; there is no recovery without it. Workflows still open and run, only the credentials are lost, so worst case you re-enter every API token by hand.

Permissions 0644 for n8n settings file /home/node/.n8n/config are too wide

In 1.x this was a warning. In 2.x N8N_ENFORCE_SETTINGS_FILE_PERMISSIONS defaults to true and the file is expected to be 0600. Fix the mode from inside the container:

docker compose exec -u node n8n chmod 600 /home/node/.n8n/config

You'll see it on bind mounts from Windows, on some NFS shares and on volumes copied over with a tool that didn't preserve modes. Setting N8N_ENFORCE_SETTINGS_FILE_PERMISSIONS=false makes the message go away and is the documented escape hatch for filesystems that can't hold the mode; on a Linux VPS with a normal volume, chmod and keep the enforcement.

There was an error initializing DB: connect ECONNREFUSED

The log shows There was an error initializing DB followed by connect ECONNREFUSED 127.0.0.1:5432 or another address. n8n can't reach Postgres. If the address is 127.0.0.1, DB_POSTGRESDB_HOST is set to localhost, which inside the n8n container means the n8n container, not the Postgres one. Use the compose service name (postgres, or whatever you called it). If the address is the right container IP and the connection is still refused, Postgres hasn't finished starting; add a healthcheck to the Postgres service and depends_on with condition: service_healthy on n8n, so it waits instead of racing. If Postgres is on another server, the same message with that server's IP is a firewall or listen_addresses problem on the Postgres side.

SQLITE_BUSY: database is locked

Two processes are writing to database.sqlite at once. The common cause on a VPS is a backup or a manual sqlite3 session holding a write lock while n8n runs, and the second is running two n8n containers against the same volume by mistake after a compose rename. Queue mode on SQLite produces it constantly, because workers and main all write, and n8n's own docs tell you not to run queue mode on SQLite. If you're seeing it under normal single-instance load, the database has outgrown the engine and the answer is switching n8n from SQLite to Postgres. It has been the fix every time I've seen this on a single instance.

FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory

Node ran out of heap. On a 2.x instance this is nearly always one execution loading too much into memory at once: an HTTP Request node pulling a 300 MB response, a Code node building an array of a million objects, a Loop Over Items with the batch size left at 1 across a huge input. Find the workflow in the executions list (it's the one that was running when the container restarted) and reduce what it holds: paginate the request, split the work into a sub-workflow so memory is released per call, move binary data to filesystem mode. Raising the ceiling is a stopgap, but it's a useful one while you fix the workflow:

NODE_OPTIONS=--max-old-space-size=4096

That's in megabytes and it applies to the main process; the memory issues page in the n8n docs lists the same causes with a bit more on workflow design. Code node work runs on task runners in 2.x, so a runner that blows up shows a task error in the execution instead of killing main, and its ceiling is N8N_RUNNERS_MAX_OLD_SPACE_SIZE.

Editor and login errors

Connection lost

The full banner reads: Connection lost: You have a connection issue or the server is down. n8n should reconnect automatically once the issue is resolved. The server is almost never down. The editor keeps a push connection open to /rest/push, over WebSocket by default (N8N_PUSH_BACKEND=websocket), and the reverse proxy in front of n8n isn't passing the upgrade. For nginx the missing lines are:

proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_buffering off;

The complete server block, with the buffering and timeout settings that go with it, is in the n8n Nginx reverse proxy tutorial. Caddy passes WebSockets without being told, which is one reason it's the proxy in the n8n VPS template (it deploys with Docker and Caddy, and I've never seen this banner on one). Cloudflare in front of either works for WebSockets on all plans; a corporate proxy or an older load balancer between you and the server sometimes doesn't, and switching to N8N_PUSH_BACKEND=sse gets the editor working over plain HTTP at the cost of a few features that need the two-way channel. The solved thread on the n8n community forum is the canonical reference for this one. If none of it helps, the browser's dev tools show the /rest/push request and its status code, which tells you where the upgrade died.

Your n8n server is configured to use a secure cookie

The rest of the message says you're visiting over an insecure URL or with Safari. You opened http://<ip>:5678 and n8n set a cookie with the Secure flag, which the browser drops on plain HTTP, so the login never sticks. Two fixes. The right one: put TLS in front and open the HTTPS URL. The quick one for a test box on a private network: N8N_SECURE_COOKIE=false. Don't use the second on anything reachable from the internet, since it downgrades the session cookie for every user.

Code node and task runner errors

In 2.x every Code node runs on a task runner, either a child process of n8n (internal mode, the default) or a separate n8nio/runners container (external mode). The errors in this group didn't exist in 1.x setups without runners, so a lot of older forum answers point at settings that no longer apply.

Task request timed out after 60 seconds

Followed by Your Code node task was not matched to a runner within the timeout period. The Code node asked the broker for a runner and none picked the task up in 60 seconds (N8N_RUNNERS_TASK_REQUEST_TIMEOUT). In internal mode that means the runner process is dead or stuck, and restarting the container fixes it; if it keeps happening, the runner is saturated by earlier tasks, and N8N_RUNNERS_MAX_CONCURRENCY (default 5) is the setting to raise. In external mode check that the runners container is up, that N8N_RUNNERS_TASK_BROKER_URI points at the n8n container on port 5679, that N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0 is set on the n8n side (it defaults to 127.0.0.1, which a sidecar can't reach) and that both containers carry the same N8N_RUNNERS_AUTH_TOKEN. In queue mode every worker needs its own runners sidecar; a worker without one produces this error for every Code node it picks up.

Code node stops after five minutes

The task itself ran and was killed at 300 seconds. That's N8N_RUNNERS_TASK_TIMEOUT, in seconds, and the docs describe it as the maximum time a task can run before the runner stops it and restarts. Raise it if the work is legitimately long, or break the input into batches with Loop Over Items so each task is short. I raised ours to 900 for one workflow that reshapes a large export, then went back to the default after moving that job into a sub-workflow, since a five minute ceiling has caught two runaway loops since.

access to env vars denied

The expression or Code node used $env.SOMETHING and got [ERROR: access to env vars denied]. N8N_BLOCK_ENV_ACCESS_IN_NODE defaults to true in 2.x, so $env is blocked everywhere in workflows unless you set it to false. My preference is to leave the block on and put the value in a credential (for secrets) or in a Set node right after the trigger (for plain config). Opening $env exposes every variable in the container, including the database password and the encryption key, to anyone who can edit a workflow.

Cannot find module 'crypto'

A Code node called require('crypto') and the runner refused it. Built-in Node modules are blocked in the runner unless listed in NODE_FUNCTION_ALLOW_BUILTIN; npm packages go in NODE_FUNCTION_ALLOW_EXTERNAL. In internal mode the variable goes on the n8n container, in external mode on the runners container:

NODE_FUNCTION_ALLOW_BUILTIN=crypto

A comma-separated list allows several, and * allows all of them, which I wouldn't do on a shared instance. The full runner variable list, both the n8n-side and the runner-side halves, is on the task runner environment variables page. If the module you want is crypto for an HMAC check on a webhook, the Crypto node does that without any of this.

Workflow and webhook errors

The workflow has issues and cannot be executed for that reason. Please fix them first

A node has a required parameter empty or a credential selected that no longer exists, and the editor is refusing to run until you fix it. The node is marked with a red triangle in the canvas, but on a large workflow it hides inside a collapsed area or behind a sticky note. Two ways to find it fast: open each node's settings and look for the red field, or export the workflow and search the JSON for "issues". The credential case shows up most after importing a workflow from another instance, where the credential id in the JSON doesn't exist on the target; open the node, pick the credential from the dropdown, save.

The requested webhook is not registered

The call reached n8n and n8n has nothing listening at that path and method. Check which URL you called. Test webhooks (/webhook-test/...) only exist while you've pressed "Listen for test event" in the editor and go away after one call; production webhooks (/webhook/...) only exist while the workflow is published. The other frequent cause is a method mismatch, since a Webhook node set to POST doesn't answer a GET at the same path. If the workflow is published and the URL is right, look at what the Webhook node shows as its production URL: a wrong domain there means N8N_WEBHOOK_URL is unset or stale, and the webhook URL fix for n8n behind a reverse proxy walks through the variables that build it. Fix the variable, restart, and the URL shown in the node updates on its own.

Webhook path already in use when publishing a workflow

Publishing fails and the message names another workflow that already registered the same path and method. Two Webhook nodes on one instance can't share POST /webhook/orders. Change the path on one of them; if you duplicated a workflow to test a change, that's the copy. On older versions the wording was "Workflow could not be activated", and you'll still find it under that name on the forum.

413 Request Entity Too Large

Either the proxy or n8n rejected the body. nginx caps request bodies at 1 MB by default (client_max_body_size), and n8n caps JSON payloads at 16 MiB (N8N_PAYLOAD_SIZE_MAX, in MiB) and form-data file uploads at 200 MiB (N8N_FORMDATA_FILE_SIZE_MAX). If the 413 page looks like nginx's, raise client_max_body_size in the server block; if it comes back as JSON from n8n, raise the n8n variable. Set both to the same number so you only debug this once.

Queue mode errors

Executions stuck in Starting soon or Running

In queue mode an execution shows "Starting soon" (the newer label for queued) until a worker takes it. Stuck there for minutes means no worker is consuming: workers are down, or they're pointed at a different Redis than main (compare QUEUE_BULL_REDIS_HOST, QUEUE_BULL_REDIS_DB and QUEUE_BULL_PREFIX on both), or they connect to a different database and can't find the workflow. docker compose logs n8n-worker-1 tells you within seconds which one.

Stuck in "Running" is the other case: a worker picked up the job and died mid-way (out of memory, killed by a restart, Redis connection dropped) and the job's lock expired without anyone finishing it. The QUEUE_WORKER_LOCK_DURATION (60,000 ms) and QUEUE_WORKER_STALLED_INTERVAL (30,000 ms) settings decide how fast a stalled job is noticed; the execution stays in the list as running until you stop it from the UI. What I couldn't tell you is how a stalled job behaves once it's been retried and stalls again, because I've only ever seen it happen once per job on our setup and the docs don't describe a retry limit since QUEUE_WORKER_MAX_STALLED_COUNT was removed in 2.0. The whole main-workers-Redis arrangement is in the n8n queue mode setup guide, including how to check n8n_scaling_mode_queue_jobs_waiting on the metrics endpoint to see the backlog as a number. A backlog with idle workers is a configuration mismatch; a backlog with busy workers is a sizing problem.

n8n monitoring and alerting

Every error above is easier to catch when a graph shows the moment it started. N8N_METRICS=true exposes a Prometheus endpoint with queue depth, memory and event loop lag, and an alert on any of those fires before the first support message does. The full setup, from scrape config to alert rules, is in the n8n monitoring guide with Prometheus and Grafana. Set the queue depth alert first; it catches the most.

Automate faster, for less

Bring your winning ideas to life with AMD power, NVMe speed and unmetered bandwidth.