Nodal-Agents

Troubleshooting & FAQ

The most common Nodal-Agents failures — symptom, cause, and fix — for connectors, models, jobs, and the local stack.

A central list of the failures people actually hit, grouped by area. Each entry is symptom → cause → fix.

Connectors & OAuth

redirect_uri_mismatch when connecting an OAuth connector

Symptom. The OAuth popup (Google / Notion / Airtable) errors out with redirect_uri_mismatch or "redirect URI does not match."

Cause. The redirect URI registered in your OAuth app does not exactly match the one Nodal-Agents uses. The redirect URI is built from your dashboard's own origin plus a per-provider callback path (e.g. http://localhost:3000/api/oauth/google-oauth/callback). A trailing slash, http vs https, or a different port all count as a mismatch. This is the single most common connector failure.

Fix. Open the Credentials wizard (/credentials), copy the Authorized redirect URI it shows (there's a Copy button), and paste it verbatim into your OAuth app's redirect-URI field at the provider. If your dashboard's port rotated (see below), the origin changed — re-copy and update the registered URI. See Connecting a tool.

Notion "object not found"

Symptom. Notion tools return "object not found" even though the page exists and the token is valid.

Cause. Notion only exposes pages explicitly shared with the integration. A fresh integration sees nothing until you grant it each page.

Fix. For every Notion page (or database) the agent should reach, open it → the ··· menu → Connections → select your integration. Then retry.

Airtable invalid_scope

Symptom. The Airtable OAuth flow fails with invalid_scope.

Cause. The four required scopes weren't saved on the integration before you started the flow. Checking the boxes isn't enough — the integration has to be saved/updated so the scopes persist.

Fix. At airtable.com/create/oauth check data.records:read, data.records:write, schema.bases:read, schema.bases:write, then save/update the integration so the scopes stick, and retry.

Airtable also rejects raw LAN-IP redirect URIs (e.g. 192.168.x.x). If your dashboard runs on a LAN IP, run this one flow from http://localhost:3000. The resulting credential then works from any host.

Models

Model "Test connection" fails

Symptom. Adding or testing an LLM key reports connection_failed with a status line like <provider> responded 401: … or … responded 404: ….

Cause and fix — the message tells you which:

  • 401 / 403 — wrong or missing API key. Re-paste the key.
  • 404 / "model not found" — wrong base URL or model id. For openai-compatible and ollama, the base URL is required — the test fails with baseUrl is required for <provider> if it's blank. Check the model id matches what the provider actually serves.
  • Connection refused / timeout — the base URL points nowhere reachable (see local-model section below).

The test hits the provider's models endpoint (e.g. /models for OpenAI-compatible, /api/tags for Ollama) with your credentials; on success it reports the number of models available.

Local model (LM Studio / Ollama) not reachable

Symptom. A local model times out or the connection is refused, in the test or mid-job.

Cause. The base URL doesn't point at a running local server, or points at the wrong port.

Fix. Use the right base URL and make sure the server is up:

  • LM Studio — http://localhost:1234/v1 (start the local server in LM Studio's Developer tab).
  • Ollama — http://localhost:11434 (ollama serve running, model pulled).

On Windows, prefer 127.0.0.1 over localhost if a connection that "should work" times out — localhost can resolve to IPv6 (::1) first while the server only listens on IPv4.

Jobs

token_budget_exceeded

Symptom. A job fails with token_budget_exceeded.

Cause. Cumulative tokens for the job crossed the per-job ceiling — a runaway loop or a very large context. The default is 1,500,000 tokens.

Fix. Investigate the loop first (a job burning 1.5M tokens is usually stuck, not legitimately large). To genuinely raise the ceiling, set MAX_TOTAL_TOKENS_PER_JOB in the runner's environment. Leveraging prompt caching (supported models) also stretches the budget, since the budget counts effective input tokens.

agent_budget_exceeded (the agent's budget stopped the job)

Symptom. A job fails with agent_budget_exceeded, or code_task refuses to start with the same code. The result ends with [stopped: agent budget — …].

Cause. The agent spent its daily or monthly ceiling: API calls of every provider and coding CLI runs, counted together.

Fix. Raise or clear the ceiling on the agent's Settings tab, in Budget (0 = none), or wait for the day or the month to end.

cost_budget_exceeded or run_time_exceeded (the run budget stopped the job)

Symptom. A job fails with cost_budget_exceeded or run_time_exceeded. Its result keeps what it wrote, then a [stopped: run budget — …] line.

Cause. The run crossed the workspace's run budget: its cost (default $2.00 per run) or its working time (no limit by default).

Fix. Raise or clear the ceiling in Settings → Safety → Run budget (0 = none). The MAX_COST_PER_JOB_USD environment variable is no longer read.

Local stack

"Worker secret missing" → runner 403, jobs stuck pending

Symptom. Dashboard tasks never start — they sit in pending. Settings flags a missing worker secret, or the web logs show [sendTaskAction] WORKER_SECRET missing — cannot ping runner.

Cause. The web app wakes the runner with a POST /api/worker call authenticated by a shared WORKER_SECRET bearer token. If the secret is missing or mismatched between web and runner, the runner answers 403 invalid_worker_secret (or 500 server_misconfiguration if it has no secret configured) and the job never gets picked up.

Fix. Ensure the same WORKER_SECRET is set for both the web and runner processes. In the managed CLI stack this is wired automatically; if you run the processes yourself, set it in both environments.

Port already in use / ports rotated unexpectedly

Symptom. Startup complains about a port, or you notice the dashboard came up on a different port than expected (default ports: web 3000, runner 3001, Postgres 25432).

Cause. A previous run was closed abruptly, leaving processes of ours holding the ports — or the OS has reserved a port (Windows Hyper-V/WinNAT excluded ranges), or something else of yours is simply listening there.

On startup the CLI cleans up only what it can show is its own: a PID it recorded itself, or one descending from it. If a port is still unbindable it rotates to a free neighbour and persists the new port in ~/.nodalai/config.json.

Fix. Usually nothing — let it rotate; the ready message prints the actual URL. If the dashboard origin changed, update any registered OAuth redirect URIs to match. To reset, run nodal-agents down and start again.

"Port conflict with a process that is not ours"

Symptom. Startup stops with:

Port conflict with a process that is not ours:
  - :3000 is held by pid 56980, which Nodal-Agents did not start

Cause. Something that is not Nodal-Agents is listening on a port it wants — very often your own dev server on :3000.

This refusal is deliberate. Earlier versions killed whatever held the port and started anyway, which meant silently destroying someone else's work to free a port we merely wanted. Refusing is recoverable; killing is not.

Fix. Either stop that process yourself, or give Nodal-Agents different ports in ~/.nodalai/config.json.

"Cleaned up N process(es) left by the previous session"

Symptom. A start prints this line before doing anything else.

Cause. The previous session ended without a clean shutdown — most often Ctrl+C. On Windows, Ctrl+C on a .cmd launcher makes the console destroy the whole process group at once, the CLI's own cleanup routine included, in the middle of its work. Some background processes then outlive it.

This message is the fix working, not a fault. The process tree is recorded while the services are healthy, so the next start can sweep whatever survived — including workers that hold no port and are therefore invisible to any port scan. Each process is matched by creation time as well as PID, so a number Windows has since recycled is never mistaken for one of ours.

Fix. Nothing to do. To avoid it entirely, stop with nodal-agents down rather than Ctrl+C.

A Codex task wrote files in read mode

Symptom. code_task ran with provider: "codex" and mode: "read" — which the approval card describes as analysis only — and files were modified anyway.

Cause. Your own ~/.codex/config.toml disabled the sandbox, and Nodal was loading it. One key was enough: [windows] sandbox = "elevated" turned --sandbox read-only into no confinement at all. Measured 2026-08-21 — same directory, same command, one flag apart: without --ignore-user-config the CLI spawned PowerShell and wrote the file; with it, "the filesystem is read-only" and nothing was written.

Fixed since. Nodal now always passes --ignore-user-config, so no setting in that file can weaken the confinement stated on the approval card. Your subscription still authenticates — auth lives outside the file. The side effect is that model and reasoning-effort defaults set there no longer apply; set them per agent in Nodal instead.

To verify on your machine, or after a Codex upgrade:

node scripts/probe-codex-sandbox.mjs

It attempts a real write with the arguments Nodal passes. Exit 0 means confined, 1 means not, 2 means it could not tell — and it never reports "confined" without positive proof that a command was attempted and refused.

"Could not read the process table"

Symptom. A yellow warning during shutdown, naming a timeout or an error from Get-CimInstance / Get-WmiObject.

Cause. Windows would not enumerate its processes in time. Heavily loaded, containerised, or locked-down machines do this — GitHub's own Windows CI runners cannot do it at all.

Consequence, stated plainly. The tree-kill falls back to taskkill /T, and a background worker may survive it. The next start cleans up what it recorded, so this costs you a message rather than a stuck port.

Fix. Nothing required. If it happens on every stop and you would rather not see it, stopping with nodal-agents down on a quieter machine avoids it.

Orphaned Postgres holds the shared-memory block

Symptom. Postgres won't start; the log mentions a pre-existing shared memory block is still in use.

Cause. A previous Postgres crashed without releasing its Windows shared-memory section. That section is keyed to the data dir, not the port — so rotating to another port doesn't help; the new postmaster reattaches the same key and dies with that FATAL error.

Fix. The CLI tries a graceful pg_ctl stop -m fast (which releases the SHM cleanly) and then a direct kill of the orphan postmaster. If it still survives, the CLI aborts with an actionable message:

  1. Run nodal-agents down.
  2. Kill the orphan postmaster (Stop-Process -Id <pid> -Force on Windows, or kill -9 <pid>).
  3. If it persists, reboot to clear the orphan kernel object.

No data is lost — the pg-data directory is preserved.

Node.js too old

Symptom. nodal-agents errors out at install or boot with syntax/engine errors.

Cause. An old Node.js runtime.

Fix. Use Node 22 or newer. Check with node -v and upgrade if it's below 22. See Self-hosting.

npm skipped the database binaries

Symptom. The install ended with npm warn allow-scripts, and nodal-agents up stops straight away on a line naming @embedded-postgres/<platform>.

Cause. Recent npm versions do not run a package's install scripts until you approve them. @embedded-postgres/<platform> carries the Postgres binaries Nodal boots, and its install step finishes preparing them. Skipped, the embedded database has nothing to start.

Fix. Approve that one package and install again. The name is the one npm printed:

npm approve-scripts @embedded-postgres/windows-x64
npm install -g nodal-agents@latest

The other packages npm lists (@whiskeysockets/baileys, protobufjs, tesseract.js) run fine without their scripts. See Getting Started.

On Windows the gate is harmless. @embedded-postgres/windows-x64 ships the Postgres executables in its tarball, and the list of links its install step would recreate is empty (2 bytes, [], against 1100 bytes in the Linux package). So a Windows install that fails after that warning fails for another reason: send the output of nodal-agents up rather than assuming the gate.

On this page