Nodal-Agents
Reference

Operating & observability

Watch runs, read cost and budget caps, control job retention, and tail logs.

Once agents are running, the dashboard and the CLI give you everything you need to watch what they do, cap what they cost, and reclaim disk.

Watching runs

Logs → Activity (/logs) is the list of runs: one folded row per run, unfoldable onto its own tool calls, filterable by agent, by tool name or by job id, 50 to a page. It is the fleet-wide audit trail and the run list at once — three screens used to tell the same story, each in its own way. /jobs now lands here.

Its second tab, Service logs, shows the runner's and the web app's own log output in the app, without going to a terminal.

Opening a row gives you the run page, which is the same page for every origin — a schedule, a project, the dashboard chat, an MCP client. Its header card carries seven figures:

FigureWhat it counts
CostThe dollar cost the provider reported, cumulative across resumes
DurationWall-clock, when it was measured
Input tokensTotal input, cache included
Output tokensTotal output
Cache readsTokens served from a cached prefix
Files changedWhat the run wrote, constated — see Proof
ActivityHow long ago something last happened on this run

A figure that is genuinely zero prints 0; a dash means the figure is absent. The two were indistinguishable before 0.8.11, which made "this run cost nothing" and "nobody measured this run" look the same.

Below the card: what was delivered, the review, the proof, the files, and the whole activity trail. A running job can be cancelled from the header.

What a delegation costs in lost cache

A sequential delegation always outlasts a provider's prompt cache — five minutes for an Anthropic ephemeral prefix — so the parent re-pays its whole prefix at fresh-input rate on every resume. That is by construction, not a bug, and it is roughly a fifth of a long run's bill.

Nothing billed changes, but the run now shows it: the figure is read from the llm_calls rows themselves, on six conditions, and never estimated from a percentage. A provider that does not report cache writes triggers nothing at all rather than being given a guessed cache lifetime.

Cost and budget caps

The runner stops a run that crosses one of these ceilings, between two turns. It keeps what the run already wrote: the result holds that text, then a line such as [stopped: run budget — $2.40 spent, ceiling $2.00, turn 3].

CeilingWhere it is setDefaultError code
Cost of one runSettings → Safety → Run budget$2.00 (0 = none)cost_budget_exceeded
Working time of one runSettings → Safety → Run budgetnonerun_time_exceeded
Tokens of one runMAX_TOTAL_TOKENS_PER_JOB on the runner1500000 (1.5M)token_budget_exceeded
What one agent spends in a day or a monththe agent's Settings → Budgetnoneagent_budget_exceeded

The agent's budget counts every API call of that agent, whatever the provider, and the cost its coding CLI runs report. Days and months start at midnight and on the 1st in the workspace time zone. Once a ceiling is reached, each of the agent's runs stops before its next call until the window ends, and a coding CLI run does not start.

The working time counts only the time the run spends working: waiting for an approval or for a delegated agent does not count. While a time ceiling is set, one model call never waits for its first token longer than half of what remains, with a 60 s floor.

The cost is what the provider reports per call, or, when it reports nothing, the tokens times the model's list price from the catalogue. A model with neither stays at $0, and the token ceiling is then the backstop.

Before 0.9.3 the cost ceiling was the MAX_COST_PER_JOB_USD environment variable. It is no longer read: a runner that still has it writes max_cost_env_ignored in its log.

These are last-resort backstops. The anti-loop guards baked into the runtime — max delegation depth, max tool calls per turn, max chain length, and a no-progress detector — stop most runaways long before a budget cap is reached.

Job retention

By default nothing is ever deleted — job history grows until you prune it. To cap database growth, set RETENTION_DAYS on the runner:

ValueEffect
0 (default)Disabled. No automatic deletion.
N > 0Each cron tick, terminal jobs (completed, failed, cancelled) whose completed_at is older than N days are pruned. Their tool_calls and approval_requests are removed via ON DELETE CASCADE.

Deleting job history is irreversible. Start with a large value (e.g. 90) and reduce over time.

Reading logs

Service logs are written to ~/.nodalai/logs/. Tail them from the CLI:

nodal-agents logs runner      # the runner (LLM loop, cron, delivery)
nodal-agents logs web         # the dashboard server
nodal-agents logs postgres    # the embedded database
nodal-agents logs             # print all three (no follow)

A single service follows live (tail -f style) until Ctrl+C; with no argument all three are printed once. If Postgres refuses to start, set NODALAI_PG_LOG=1 before nodal-agents up for verbose Postgres startup output.

The cluster keeps its own log, separate from the three above, in ~/.nodalai/logs/postgres/<data-dir>/postgresql-YYYY-MM-DD.log — one subdirectory per data directory, one file per day. That is where a backend crash is written, and it is what to read when a stack goes down without an explanation on the runner's side.

For the in-app view, use the dashboard's Logs page: Activity for the runs and their tool calls, Service logs for the same runner and web output without a terminal.

Undoing a run

Before any tool that can modify a workspace runs, Nodal snapshots that folder into a shadow git store. Listing and restoring live in the CLI, on purpose — deciding that a run went wrong is a judgement, and an agent able to roll itself back could roll back the evidence:

nodal-agents checkpoints                    # snapshots for the current directory
nodal-agents checkpoints restore 4f3a9c21   # put it back (the replaced state is saved first)

Read Proof for what a checkpoint does not cover, and what a refused write tells you.

On this page