Troubleshooting & FAQ
The most common Nodal-Agents failures — symptom, cause, and fix — for connectors, models, jobs, and the local stack.
A central list of the failures people actually hit, grouped by area. Each entry is symptom → cause → fix.
Connectors & OAuth
redirect_uri_mismatch when connecting an OAuth connector
Symptom. The OAuth popup (Google / Notion / Airtable) errors out with
redirect_uri_mismatch or "redirect URI does not match."
Cause. The redirect URI registered in your OAuth app does not exactly
match the one Nodal-Agents uses. The redirect URI is built from your dashboard's
own origin plus a per-provider callback path (e.g.
http://localhost:3000/api/oauth/google-oauth/callback). A trailing slash,
http vs https, or a different port all count as a mismatch. This is the
single most common connector failure.
Fix. Open the Credentials wizard (/credentials), copy the
Authorized redirect URI it shows (there's a Copy button), and paste it
verbatim into your OAuth app's redirect-URI field at the provider. If your
dashboard's port rotated (see below), the origin changed — re-copy and update the
registered URI. See Connecting a tool.
Notion "object not found"
Symptom. Notion tools return "object not found" even though the page exists and the token is valid.
Cause. Notion only exposes pages explicitly shared with the integration. A fresh integration sees nothing until you grant it each page.
Fix. For every Notion page (or database) the agent should reach, open it → the ··· menu → Connections → select your integration. Then retry.
Airtable invalid_scope
Symptom. The Airtable OAuth flow fails with invalid_scope.
Cause. The four required scopes weren't saved on the integration before you started the flow. Checking the boxes isn't enough — the integration has to be saved/updated so the scopes persist.
Fix. At airtable.com/create/oauth check
data.records:read, data.records:write, schema.bases:read,
schema.bases:write, then save/update the integration so the scopes stick,
and retry.
Airtable also rejects raw LAN-IP redirect URIs (e.g. 192.168.x.x). If your dashboard runs on a
LAN IP, run this one flow from http://localhost:3000. The resulting credential then works from
any host.
Models
Model "Test connection" fails
Symptom. Adding or testing an LLM key reports connection_failed with a
status line like <provider> responded 401: … or … responded 404: ….
Cause and fix — the message tells you which:
- 401 / 403 — wrong or missing API key. Re-paste the key.
- 404 / "model not found" — wrong base URL or model id. For
openai-compatibleandollama, the base URL is required — the test fails withbaseUrl is required for <provider>if it's blank. Check the model id matches what the provider actually serves. - Connection refused / timeout — the base URL points nowhere reachable (see local-model section below).
The test hits the provider's models endpoint (e.g. /models for
OpenAI-compatible, /api/tags for Ollama) with your credentials; on success it
reports the number of models available.
Local model (LM Studio / Ollama) not reachable
Symptom. A local model times out or the connection is refused, in the test or mid-job.
Cause. The base URL doesn't point at a running local server, or points at the wrong port.
Fix. Use the right base URL and make sure the server is up:
- LM Studio —
http://localhost:1234/v1(start the local server in LM Studio's Developer tab). - Ollama —
http://localhost:11434(ollama serverunning, model pulled).
On Windows, prefer 127.0.0.1 over localhost if a connection that "should
work" times out — localhost can resolve to IPv6 (::1) first while the server
only listens on IPv4.
Jobs
token_budget_exceeded
Symptom. A job fails with token_budget_exceeded.
Cause. Cumulative tokens for the job crossed the per-job ceiling — a runaway loop or a very large context. The default is 1,500,000 tokens.
Fix. Investigate the loop first (a job burning 1.5M tokens is usually stuck,
not legitimately large). To genuinely raise the ceiling, set
MAX_TOTAL_TOKENS_PER_JOB in the runner's environment. Leveraging prompt caching
(supported models) also stretches the budget, since the budget counts effective
input tokens.
agent_budget_exceeded (the agent's budget stopped the job)
Symptom. A job fails with agent_budget_exceeded, or code_task refuses to
start with the same code. The result ends with [stopped: agent budget — …].
Cause. The agent spent its daily or monthly ceiling: API calls of every provider and coding CLI runs, counted together.
Fix. Raise or clear the ceiling on the agent's Settings tab, in Budget
(0 = none), or wait for the day or the month to end.
cost_budget_exceeded or run_time_exceeded (the run budget stopped the job)
Symptom. A job fails with cost_budget_exceeded or run_time_exceeded. Its
result keeps what it wrote, then a [stopped: run budget — …] line.
Cause. The run crossed the workspace's run budget: its cost (default $2.00 per run) or its working time (no limit by default).
Fix. Raise or clear the ceiling in Settings → Safety → Run budget (0 =
none). The MAX_COST_PER_JOB_USD environment variable is no longer read.
Local stack
"Worker secret missing" → runner 403, jobs stuck pending
Symptom. Dashboard tasks never start — they sit in pending. Settings flags
a missing worker secret, or the web logs show
[sendTaskAction] WORKER_SECRET missing — cannot ping runner.
Cause. The web app wakes the runner with a POST /api/worker call
authenticated by a shared WORKER_SECRET bearer token. If the secret is missing
or mismatched between web and runner, the runner answers 403 invalid_worker_secret
(or 500 server_misconfiguration if it has no secret configured) and the job
never gets picked up.
Fix. Ensure the same WORKER_SECRET is set for both the web and runner
processes. In the managed CLI stack this is wired automatically; if you run the
processes yourself, set it in both environments.
Port already in use / ports rotated unexpectedly
Symptom. Startup complains about a port, or you notice the dashboard came up on a different port than expected (default ports: web 3000, runner 3001, Postgres 25432).
Cause. A previous run was closed abruptly, leaving processes of ours holding the ports — or the OS has reserved a port (Windows Hyper-V/WinNAT excluded ranges), or something else of yours is simply listening there.
On startup the CLI cleans up only what it can show is its own: a PID it
recorded itself, or one descending from it. If a port is still unbindable it
rotates to a free neighbour and persists the new port in
~/.nodalai/config.json.
Fix. Usually nothing — let it rotate; the ready message prints the actual
URL. If the dashboard origin changed, update any registered OAuth redirect URIs
to match. To reset, run nodal-agents down and start again.
"Port conflict with a process that is not ours"
Symptom. Startup stops with:
Port conflict with a process that is not ours:
- :3000 is held by pid 56980, which Nodal-Agents did not startCause. Something that is not Nodal-Agents is listening on a port it wants —
very often your own dev server on :3000.
This refusal is deliberate. Earlier versions killed whatever held the port and started anyway, which meant silently destroying someone else's work to free a port we merely wanted. Refusing is recoverable; killing is not.
Fix. Either stop that process yourself, or give Nodal-Agents different ports
in ~/.nodalai/config.json.
"Cleaned up N process(es) left by the previous session"
Symptom. A start prints this line before doing anything else.
Cause. The previous session ended without a clean shutdown — most often
Ctrl+C. On Windows, Ctrl+C on a .cmd launcher makes the console destroy
the whole process group at once, the CLI's own cleanup routine included, in the
middle of its work. Some background processes then outlive it.
This message is the fix working, not a fault. The process tree is recorded while the services are healthy, so the next start can sweep whatever survived — including workers that hold no port and are therefore invisible to any port scan. Each process is matched by creation time as well as PID, so a number Windows has since recycled is never mistaken for one of ours.
Fix. Nothing to do. To avoid it entirely, stop with nodal-agents down
rather than Ctrl+C.
A Codex task wrote files in read mode
Symptom. code_task ran with provider: "codex" and mode: "read" — which
the approval card describes as analysis only — and files were modified anyway.
Cause. Your own ~/.codex/config.toml disabled the sandbox, and Nodal was
loading it. One key was enough: [windows] sandbox = "elevated" turned
--sandbox read-only into no confinement at all. Measured 2026-08-21 — same
directory, same command, one flag apart: without --ignore-user-config the CLI
spawned PowerShell and wrote the file; with it, "the filesystem is read-only"
and nothing was written.
Fixed since. Nodal now always passes --ignore-user-config, so no setting in
that file can weaken the confinement stated on the approval card. Your
subscription still authenticates — auth lives outside the file. The side effect
is that model and reasoning-effort defaults set there no longer apply; set them
per agent in Nodal instead.
To verify on your machine, or after a Codex upgrade:
node scripts/probe-codex-sandbox.mjsIt attempts a real write with the arguments Nodal passes. Exit 0 means confined, 1 means not, 2 means it could not tell — and it never reports "confined" without positive proof that a command was attempted and refused.
"Could not read the process table"
Symptom. A yellow warning during shutdown, naming a timeout or an error from
Get-CimInstance / Get-WmiObject.
Cause. Windows would not enumerate its processes in time. Heavily loaded, containerised, or locked-down machines do this — GitHub's own Windows CI runners cannot do it at all.
Consequence, stated plainly. The tree-kill falls back to taskkill /T, and a
background worker may survive it. The next start cleans up what it recorded, so
this costs you a message rather than a stuck port.
Fix. Nothing required. If it happens on every stop and you would rather not
see it, stopping with nodal-agents down on a quieter machine avoids it.
Orphaned Postgres holds the shared-memory block
Symptom. Postgres won't start; the log mentions a pre-existing shared memory block is still in use.
Cause. A previous Postgres crashed without releasing its Windows shared-memory section. That section is keyed to the data dir, not the port — so rotating to another port doesn't help; the new postmaster reattaches the same key and dies with that FATAL error.
Fix. The CLI tries a graceful pg_ctl stop -m fast (which releases the SHM
cleanly) and then a direct kill of the orphan postmaster. If it still survives,
the CLI aborts with an actionable message:
- Run
nodal-agents down. - Kill the orphan postmaster (
Stop-Process -Id <pid> -Forceon Windows, orkill -9 <pid>). - If it persists, reboot to clear the orphan kernel object.
No data is lost — the pg-data directory is preserved.
Node.js too old
Symptom. nodal-agents errors out at install or boot with syntax/engine
errors.
Cause. An old Node.js runtime.
Fix. Use Node 22 or newer. Check with node -v and upgrade if it's
below 22. See Self-hosting.
npm skipped the database binaries
Symptom. The install ended with npm warn allow-scripts, and nodal-agents up stops straight away on a line naming @embedded-postgres/<platform>.
Cause. Recent npm versions do not run a package's install scripts until you
approve them. @embedded-postgres/<platform> carries the Postgres binaries
Nodal boots, and its install step finishes preparing them. Skipped, the embedded
database has nothing to start.
Fix. Approve that one package and install again. The name is the one npm printed:
npm approve-scripts @embedded-postgres/windows-x64
npm install -g nodal-agents@latestThe other packages npm lists (@whiskeysockets/baileys, protobufjs,
tesseract.js) run fine without their scripts. See
Getting Started.
On Windows the gate is harmless. @embedded-postgres/windows-x64 ships the Postgres executables
in its tarball, and the list of links its install step would recreate is empty (2 bytes, [],
against 1100 bytes in the Linux package). So a Windows install that fails after that warning fails
for another reason: send the output of nodal-agents up rather than assuming the gate.
Related
- Connecting a tool — the full connector setup flow
- Connectors & MCP
- Self-hosting