Shell commands
Let an agent run shell commands in its workspace — with a human-in-the-loop approval gate by default.
The command-execution skill unlocks the run_command built-in tool, which
runs a shell command in the agent's workspace and returns stdout, stderr, and the
exit code. Use it for installing dependencies, running scripts, build steps, or
any CLI tool.
By default, a command asks for your approval before it runs. The agent proposes the command; you approve or reject it in the dashboard. What can run without asking is set per agent, in two places on its Autonomy tab: whether ordinary commands need you at all, and what the agent may not do with a shell, kind of action by kind of action.
Enable shell commands for an agent
- Go to Skills in the dashboard and find Command execution in the built-in library (or your assigned skills list).
- Assign it to the agent.
- That's it. The agent can now call
run_command. On its next job it will pause for your approval before running any command.
How approval works — you approve a command before it runs
When the agent calls run_command, the runner creates an approval request and
suspends the job. You see the exact command in the dashboard and approve or
reject it. If you approve, the runner executes the command and the job resumes
with the output. If you reject, the agent receives an error and can try a
different approach.
Where you approve a command: the Approvals page of the dashboard, which lists every command waiting for you. By default, nothing an agent proposes runs until you approve that call. Two things let commands through without that question: a Run without asking choice on the agent's Run commands row, optionally confined to one of its folders — the card's Approve for this project writes exactly that choice for one folder; and, when the agent has no choice set, the workspace's autonomy level set to Autonomous, gate destructive, under which an ordinary command runs unasked. In both cases, the agent's shell checklist still applies: a kind of action set to Ask me asks, one set to Never is refused.
The Run commands section of the Autonomy tab says in one sentence what really happens for that agent, given its rule and the workspace level.
On a Telegram-bound agent, the agent gets a bounded number of turns to tell you via chat that it is waiting — so you are not left wondering why it went quiet.
The Run commands row: three choices
On the agent's Autonomy tab, Run commands offers the same three choices as every other tool, in every auth mode:
- Run without asking: the agent runs commands immediately, after you confirm the choice once. Commands are still logged, and the kinds of action on the agent's shell checklist still follow their setting.
- Ask for approval: every command waits for you.
- Block: the agent cannot run commands.
With no choice set, none is lit and the workspace autonomy decides; the sentence above the row says what that means for this agent. Let the workspace autonomy decide removes a choice. Before 0.9.3 this row was a "Yolo" switch, whose Off position could not block and did not always mean "asks".
Only the workspace owner can change it, both in the dashboard and on the
approval card. That was already true on both sides, which is why the separate
LAN-only master switch was removed in 0.8.7: it was a second lock for the same
key, and it only existed outside local-trust — a red button that works in one
auth mode is not a red button.
The brake: Settings → Safety → Auto-run brake
What survived, and became its only job, is the emergency brake. It is workspace-wide, inactive by default, and it applies in every auth mode.
While it is engaged, on every job:
- any
auto_approverule on a code-execution tool is dropped, and - if no tool-specific rule then remains, an explicit
require_approvalrule is injected, so a blanket wildcard (*) auto-approve cannot sweep the tool back in.
Nothing is deleted from the database: releasing the brake re-arms your rules
exactly as they were. The brake also outranks the agent's autonomy level — a
fully_autonomous ROOT cannot auto-approve past it.
The tools it covers are the ones that can end up spawning something on your
machine: run_command, run_skill_script, skill_file_write, code_task,
create_mcp and attach_mcp (a stdio server is a local subprocess), and
declare_verification — because declaring a proof command is causing it to run
later, at finalisation, outside any approval flow. The list lives in one place in
the code, so a tool added to the autonomy guard cannot be forgotten here.
What an agent may not do with a shell
The agent's Autonomy tab lists, under What it may do with a shell, six kinds of action Nodal recognises in what the agent asks to run. Each one is Allowed, Ask me or Never:
| Kind of action | Recognised by |
|---|---|
| Run code written into a command | python -c, node -e, bash -c, a script piped into an interpreter… |
| Delete files or discard changes | rm, del, Remove-Item, git reset --hard, git clean… |
| Install software or packages | npm, pip, winget, choco, brew install… |
| Download files from the internet | wget, curl -o, git clone, Invoke-WebRequest… |
| Stop other programs or services | kill, taskkill, Stop-Process, systemctl… |
| Change system settings, permissions or disks | icacls, chmod, format, diskpart… |
- Ask me waits for you, and the approval card names the kinds of action it saw.
- Never refuses it. The agent is told what it may not do, so it does not look for a way around.
- The list applies at every autonomy level, and under Run without asking too: it only ever adds a question or a refusal, it never removes one. A rule that blocks the shell tool, and the hard floor below, still win.
- A new agent is autonomous: it asks only for what leaves its workspace or
cannot be undone. Without asking, it downloads into its workspaces and
runs code, inline (
node -e,python -c,curl … | bash) or in a script. Inline code has the same power as a script the agent writes with the file tools and then runs, which never asked. It asks before it deletes, installs, stops a program, changes system settings, or downloads outside its workspaces. When a download replaces a file in a workspace, the checkpoint taken just before keeps the old version. - What you set on an agent is kept as you set it. Set either row (inline code, downloads) to Ask me to be asked again, or to Never to refuse it.
- A download runs without asking only into the job's workspaces, and into
the model and image stores of the programs that keep one. Nodal reads
where the line writes (
curl -o/-O/--output-dir,wget -O/-P,Invoke-WebRequest -OutFile,Start-BitsTransfer -Destination,aria2c -d/-o,git clone URL DIR,pip download -d,hf download --local-dir,> file), glued values included (curl -sLo/x), from the folder the line is in at that point: the folder it starts in, then eachcd,pushdorPush-Locationthat comes before it. A link in a workspace is followed to where it points, even when nothing is there yet. A place outside the workspaces asks, and the approval card names it. So does a place the text does not say (a variable,~, a sub-shell, a folder left withpopd). - This reading is a net, not a wall. Nodal reads the common forms (
curl -o,wget -O…), and it is not watertight: what programs write on their own is not bounded. That covers a copy (cp), a redirection it cannot place, a config file (curl -K,aria2c -i), a gluedgit -C/path, or a link made in the same line. Bounding that takes an operating-system sandbox (#628). - A download into a program's own store (
comfy model download,ollama pull,docker pull,podman pull,hf downloadwithout--local-dir) names no path and runs without asking unless downloads are set otherwise. That is deliberate: an agent that makes images fetches the models it is missing. The store lies outside the workspaces and no checkpoint covers it. Set downloads to Ask me to be asked for these too. - An agent already set to Run without asking everywhere when this list appeared kept what it could do.
- The checklist also judges a declared proof (
declare_verification), because its steps run later without asking again. - Nodal recognises each kind by the program started (
rm,npm i,curl … > file), including whatbash -c,cmd /cor$(…)wrap. A mention in an argument (git commit -m "rm old refs") is not the action.
What this list does not do. It reads the text the agent asks to run, not what a program does once it runs. Code the agent runs, inline or from a script, can read, delete or download anything your user account can reach without the list seeing it. Where a download lands is read; the rest of the shell is not kept inside the agent's folders, only the file tools are confined to them. Keeping what a program does inside a folder is the job of an operating-system sandbox, which Nodal does not have yet (#628).
Folded under Advanced, the same tab keeps the older list of programs an agent may start, for people who know exactly what to write there.
The safety floor — commands that never auto-run
Shell commands split into two tiers of danger, and each tier is handled differently.
The hard floor: refused even if you click Approve
A small, fixed set of commands is machine-wide destructive: wiping the disk root, formatting a drive, repartitioning, or powering off/rebooting the machine. These patterns can never run, with Run without asking or not, approved or not:
rm -rf(orRemove-Item -Recurse -Force,del /s /q,rd /s /q) against/,~, a Windows drive root, or a wildcard- a fork bomb
mkfs(any filesystem)dd ... of=/dev/...or a shell redirection to a raw device (> /dev/sda)- a Windows
format <drive>: diskpart, or a destructive PowerShell disk cmdlet (Format-Volume,Clear-Disk,Initialize-Disk)shutdown/reboot/halt/poweroff(or the PowerShell equivalents,Stop-Computer/Restart-Computer)init 0/init 6(Unix runlevel shutdown/reboot)
Hiding one behind a wrapper (sudo, cmd /c, powershell -Command "…") does
not help: the classifier unwraps those before checking.
They always force a human decision, and even if you click Approve, the command still does not execute. A single tap should never be able to greenlight machine-wide destruction. When one is refused, the agent (and you) get a plain-language explanation of why, so nothing fails silently.
Destructive or heavy: gated, but approvable
Everything else dangerous is not on the hard floor. It follows the agent's shell checklist: Ask me by default, and once you approve it, it actually runs. This class covers ordinary destructive actions (deleting files, killing a process, installing a package, recursive permission changes, force-pushing git) and a download outside the agent's workspaces.
Code in a command runs without asking by default. Code passed on the
command line to an interpreter (python -c "…", node -e "…", perl -e,
ruby -e, php -r, sh -c, bash -c, powershell -Command), code piped
into a bare interpreter (curl … | bash, echo '…' | python) and awk
programs that shell out (awk 'BEGIN{system(…)}') all count as Run code
written into a command. Their code is not read: it has the same power as a
script file (python my_script.py, node build.js, bash deploy.sh), which
runs like any ordinary command, whoever wrote it. Set that row to Ask me or
Never to gate it. What the code does inside the program is not read, and a
deletion it wraps (bash -c "rm -rf x") still asks as a deletion.
This floor is deliberately narrow: only the 8 machine-wide destroyers above are un-bypassable. The floor is the last-resort circuit breaker, not a general safety net.
Command environment
- Shell:
cmd.exeon Windows,/bin/shon Unix. Shell features (&&,|, quoting) work. - Working directory: the agent's workspace root. Pass a workspace-relative
cwdto run somewhere else inside the workspace — the path cannot point outside it. That sets where the command starts, not what it can reach: see what this list does not do. - Default timeout: 300 seconds. Pass
timeout_secondsto override. On timeout the command and its child processes are killed. - Output cap: approximately 100 000 characters per stream. For verbose commands, redirect output to a file in the workspace and read it back with the file tools.
- Non-interactive only: commands that wait for input will hang until the
timeout. Use non-interactive flags (e.g.
npm install --yes,--no-input).
Tips
Batch multi-step work into one call. A compound command like
npm install && node build.js && node run.js is one run_command call — and in
approval mode, one approval instead of three.
A non-zero exit code is data. The exit code and stderr are returned to the agent; it can read them, correct the command, and retry. A failing command does not crash the job.
Stop when done. Once the command gives the agent what it needs, it should
deliver the result with return_result and end the turn. Re-running or
double-checking wastes approvals and adds latency.
Related
- Skill reference: Command execution
- Concepts: Agents
- Concepts: Proof — the snapshot taken before a command writes
- Concepts: Connectors & MCP
Automations
Schedule an agent to run a task on a recurring cron schedule, with optional notifications on the channel of your choice.
Coding with an agent
Hand a real development task to Claude Code or Codex running under your own subscription — read-only by default, approved by you, and traceable afterwards.