Nodal-Agents
Guides

Shell commands

Let an agent run shell commands in its workspace — with a human-in-the-loop approval gate by default.

The command-execution skill unlocks the run_command built-in tool, which runs a shell command in the agent's workspace and returns stdout, stderr, and the exit code. Use it for installing dependencies, running scripts, build steps, or any CLI tool.

By default, a command asks for your approval before it runs. The agent proposes the command; you approve or reject it in the dashboard. What can run without asking is set per agent, in two places on its Autonomy tab: whether ordinary commands need you at all, and what the agent may not do with a shell, kind of action by kind of action.

Enable shell commands for an agent

  1. Go to Skills in the dashboard and find Command execution in the built-in library (or your assigned skills list).
  2. Assign it to the agent.
  3. That's it. The agent can now call run_command. On its next job it will pause for your approval before running any command.

How approval works — you approve a command before it runs

When the agent calls run_command, the runner creates an approval request and suspends the job. You see the exact command in the dashboard and approve or reject it. If you approve, the runner executes the command and the job resumes with the output. If you reject, the agent receives an error and can try a different approach.

Where you approve a command: the Approvals page of the dashboard, which lists every command waiting for you. By default, nothing an agent proposes runs until you approve that call. Two things let commands through without that question: a Run without asking choice on the agent's Run commands row, optionally confined to one of its folders — the card's Approve for this project writes exactly that choice for one folder; and, when the agent has no choice set, the workspace's autonomy level set to Autonomous, gate destructive, under which an ordinary command runs unasked. In both cases, the agent's shell checklist still applies: a kind of action set to Ask me asks, one set to Never is refused.

The Run commands section of the Autonomy tab says in one sentence what really happens for that agent, given its rule and the workspace level.

On a Telegram-bound agent, the agent gets a bounded number of turns to tell you via chat that it is waiting — so you are not left wondering why it went quiet.

The Run commands row: three choices

On the agent's Autonomy tab, Run commands offers the same three choices as every other tool, in every auth mode:

  • Run without asking: the agent runs commands immediately, after you confirm the choice once. Commands are still logged, and the kinds of action on the agent's shell checklist still follow their setting.
  • Ask for approval: every command waits for you.
  • Block: the agent cannot run commands.

With no choice set, none is lit and the workspace autonomy decides; the sentence above the row says what that means for this agent. Let the workspace autonomy decide removes a choice. Before 0.9.3 this row was a "Yolo" switch, whose Off position could not block and did not always mean "asks".

Only the workspace owner can change it, both in the dashboard and on the approval card. That was already true on both sides, which is why the separate LAN-only master switch was removed in 0.8.7: it was a second lock for the same key, and it only existed outside local-trust — a red button that works in one auth mode is not a red button.

The brake: Settings → Safety → Auto-run brake

What survived, and became its only job, is the emergency brake. It is workspace-wide, inactive by default, and it applies in every auth mode.

While it is engaged, on every job:

  • any auto_approve rule on a code-execution tool is dropped, and
  • if no tool-specific rule then remains, an explicit require_approval rule is injected, so a blanket wildcard (*) auto-approve cannot sweep the tool back in.

Nothing is deleted from the database: releasing the brake re-arms your rules exactly as they were. The brake also outranks the agent's autonomy level — a fully_autonomous ROOT cannot auto-approve past it.

The tools it covers are the ones that can end up spawning something on your machine: run_command, run_skill_script, skill_file_write, code_task, create_mcp and attach_mcp (a stdio server is a local subprocess), and declare_verification — because declaring a proof command is causing it to run later, at finalisation, outside any approval flow. The list lives in one place in the code, so a tool added to the autonomy guard cannot be forgotten here.

What an agent may not do with a shell

The agent's Autonomy tab lists, under What it may do with a shell, six kinds of action Nodal recognises in what the agent asks to run. Each one is Allowed, Ask me or Never:

Kind of actionRecognised by
Run code written into a commandpython -c, node -e, bash -c, a script piped into an interpreter…
Delete files or discard changesrm, del, Remove-Item, git reset --hard, git clean…
Install software or packagesnpm, pip, winget, choco, brew install…
Download files from the internetwget, curl -o, git clone, Invoke-WebRequest…
Stop other programs or serviceskill, taskkill, Stop-Process, systemctl…
Change system settings, permissions or disksicacls, chmod, format, diskpart…
  • Ask me waits for you, and the approval card names the kinds of action it saw.
  • Never refuses it. The agent is told what it may not do, so it does not look for a way around.
  • The list applies at every autonomy level, and under Run without asking too: it only ever adds a question or a refusal, it never removes one. A rule that blocks the shell tool, and the hard floor below, still win.
  • A new agent is autonomous: it asks only for what leaves its workspace or cannot be undone. Without asking, it downloads into its workspaces and runs code, inline (node -e, python -c, curl … | bash) or in a script. Inline code has the same power as a script the agent writes with the file tools and then runs, which never asked. It asks before it deletes, installs, stops a program, changes system settings, or downloads outside its workspaces. When a download replaces a file in a workspace, the checkpoint taken just before keeps the old version.
  • What you set on an agent is kept as you set it. Set either row (inline code, downloads) to Ask me to be asked again, or to Never to refuse it.
  • A download runs without asking only into the job's workspaces, and into the model and image stores of the programs that keep one. Nodal reads where the line writes (curl -o/-O/--output-dir, wget -O/-P, Invoke-WebRequest -OutFile, Start-BitsTransfer -Destination, aria2c -d/-o, git clone URL DIR, pip download -d, hf download --local-dir, > file), glued values included (curl -sLo/x), from the folder the line is in at that point: the folder it starts in, then each cd, pushd or Push-Location that comes before it. A link in a workspace is followed to where it points, even when nothing is there yet. A place outside the workspaces asks, and the approval card names it. So does a place the text does not say (a variable, ~, a sub-shell, a folder left with popd).
  • This reading is a net, not a wall. Nodal reads the common forms (curl -o, wget -O…), and it is not watertight: what programs write on their own is not bounded. That covers a copy (cp), a redirection it cannot place, a config file (curl -K, aria2c -i), a glued git -C/path, or a link made in the same line. Bounding that takes an operating-system sandbox (#628).
  • A download into a program's own store (comfy model download, ollama pull, docker pull, podman pull, hf download without --local-dir) names no path and runs without asking unless downloads are set otherwise. That is deliberate: an agent that makes images fetches the models it is missing. The store lies outside the workspaces and no checkpoint covers it. Set downloads to Ask me to be asked for these too.
  • An agent already set to Run without asking everywhere when this list appeared kept what it could do.
  • The checklist also judges a declared proof (declare_verification), because its steps run later without asking again.
  • Nodal recognises each kind by the program started (rm, npm i, curl … > file), including what bash -c, cmd /c or $(…) wrap. A mention in an argument (git commit -m "rm old refs") is not the action.

What this list does not do. It reads the text the agent asks to run, not what a program does once it runs. Code the agent runs, inline or from a script, can read, delete or download anything your user account can reach without the list seeing it. Where a download lands is read; the rest of the shell is not kept inside the agent's folders, only the file tools are confined to them. Keeping what a program does inside a folder is the job of an operating-system sandbox, which Nodal does not have yet (#628).

Folded under Advanced, the same tab keeps the older list of programs an agent may start, for people who know exactly what to write there.

The safety floor — commands that never auto-run

Shell commands split into two tiers of danger, and each tier is handled differently.

The hard floor: refused even if you click Approve

A small, fixed set of commands is machine-wide destructive: wiping the disk root, formatting a drive, repartitioning, or powering off/rebooting the machine. These patterns can never run, with Run without asking or not, approved or not:

  • rm -rf (or Remove-Item -Recurse -Force, del /s /q, rd /s /q) against /, ~, a Windows drive root, or a wildcard
  • a fork bomb
  • mkfs (any filesystem)
  • dd ... of=/dev/... or a shell redirection to a raw device (> /dev/sda)
  • a Windows format <drive>:
  • diskpart, or a destructive PowerShell disk cmdlet (Format-Volume, Clear-Disk, Initialize-Disk)
  • shutdown / reboot / halt / poweroff (or the PowerShell equivalents, Stop-Computer / Restart-Computer)
  • init 0 / init 6 (Unix runlevel shutdown/reboot)

Hiding one behind a wrapper (sudo, cmd /c, powershell -Command "…") does not help: the classifier unwraps those before checking.

They always force a human decision, and even if you click Approve, the command still does not execute. A single tap should never be able to greenlight machine-wide destruction. When one is refused, the agent (and you) get a plain-language explanation of why, so nothing fails silently.

Destructive or heavy: gated, but approvable

Everything else dangerous is not on the hard floor. It follows the agent's shell checklist: Ask me by default, and once you approve it, it actually runs. This class covers ordinary destructive actions (deleting files, killing a process, installing a package, recursive permission changes, force-pushing git) and a download outside the agent's workspaces.

Code in a command runs without asking by default. Code passed on the command line to an interpreter (python -c "…", node -e "…", perl -e, ruby -e, php -r, sh -c, bash -c, powershell -Command), code piped into a bare interpreter (curl … | bash, echo '…' | python) and awk programs that shell out (awk 'BEGIN{system(…)}') all count as Run code written into a command. Their code is not read: it has the same power as a script file (python my_script.py, node build.js, bash deploy.sh), which runs like any ordinary command, whoever wrote it. Set that row to Ask me or Never to gate it. What the code does inside the program is not read, and a deletion it wraps (bash -c "rm -rf x") still asks as a deletion.

This floor is deliberately narrow: only the 8 machine-wide destroyers above are un-bypassable. The floor is the last-resort circuit breaker, not a general safety net.

Command environment

  • Shell: cmd.exe on Windows, /bin/sh on Unix. Shell features (&&, |, quoting) work.
  • Working directory: the agent's workspace root. Pass a workspace-relative cwd to run somewhere else inside the workspace — the path cannot point outside it. That sets where the command starts, not what it can reach: see what this list does not do.
  • Default timeout: 300 seconds. Pass timeout_seconds to override. On timeout the command and its child processes are killed.
  • Output cap: approximately 100 000 characters per stream. For verbose commands, redirect output to a file in the workspace and read it back with the file tools.
  • Non-interactive only: commands that wait for input will hang until the timeout. Use non-interactive flags (e.g. npm install --yes, --no-input).

Tips

Batch multi-step work into one call. A compound command like npm install && node build.js && node run.js is one run_command call — and in approval mode, one approval instead of three.

A non-zero exit code is data. The exit code and stderr are returned to the agent; it can read them, correct the command, and retry. A failing command does not crash the job.

Stop when done. Once the command gives the agent what it needs, it should deliver the result with return_result and end the turn. Re-running or double-checking wastes approvals and adds latency.

On this page