Coding with an agent
Hand a real development task to Claude Code or Codex running under your own subscription — read-only by default, approved by you, and traceable afterwards.
The code-task skill unlocks the code_task built-in tool. It hands a complete
development task — analyse, review, debug, implement — to a coding CLI that is
already installed and logged in on this machine: Claude Code or Codex.
The important part is where the work happens. The CLI is itself a full coding agent: it explores the working directory, reasons about it, and comes back with an answer. It runs under your own subscription, not through an API key, so a long investigation costs you nothing beyond what you already pay for.
Read-only by default, and every run needs your approval unless you turn Yolo on for that agent.
Before you start
The CLI has to exist on the machine and be signed in. Nodal-Agents does not install it and cannot log in for you.
claude --version # Claude Code
codex --version # CodexIf either is missing, install it and sign in first. Each provider row on the agent's edit page carries a Test button that runs a real probe and reports exactly what failed — a missing binary and an expired login look different there. The onboarding flow runs the same probe for both providers.
Enable it for an agent
- Open Skills in the dashboard and find Coding CLI (Claude Code / Codex) in the built-in library.
- Assign it to the agent.
- On the agent's edit page, choose which providers it may use and what each one defaults to.
Both providers are allowed by default. Restricting an agent to one is a real choice, not cosmetic: it decides whose subscription gets spent.
Read mode and write mode
code_task runs in one of two modes, and the difference is enforced by the CLI
itself rather than by us:
| Mode | What the CLI can do | Enforced by |
|---|---|---|
read (default) | Explore and answer. No file changes, no shell. | Claude Code hides its write tools; Codex runs its sandbox read-only. |
write | Change files inside the workspace. | The same mechanisms, opened up. |
An agent that only needs to answer "where is X handled" never needs write. The
mode travels with the approval request, so you see which one you are approving.
The two mechanisms are not equally strong. Claude Code removes the write tools from the model, so there is nothing to escape. Codex sandboxes them at the OS level, which is a real boundary but a more fragile one — as the section below explains, it was silently disabled on this machine until 2026-08-21.
Codex runs ignore your ~/.codex/config.toml
Nodal starts Codex with --ignore-user-config. Your subscription still
authenticates — auth lives outside that file — but nothing else in it applies.
This is a security boundary, not a preference. Two things were coming through that file into every Nodal-spawned run:
- Your personal MCP servers, with the credentials in their env blocks. The
narrower fix, an empty
mcp_serversoverride, does not work: an empty TOML table merges with your config instead of replacing it, so the servers came through anyway. - Your sandbox settings. A single key —
[windows] sandbox = "elevated"— was enough to turn--sandbox read-onlyinto no confinement at all. Same directory, same command, one flag apart: without--ignore-user-configthe CLI spawned a shell and wrote the file; with it, "the filesystem is read-only" and nothing was written.
That second one is the important one, and it is not Windows-specific in principle: any setting in a file Nodal does not control could weaken the confinement its approval card promises, on any OS, without a trace.
The trade-off, stated plainly: defaults you set in config.toml — model,
reasoning effort, features — no longer apply to Nodal runs. Set the model and
effort per agent in Nodal if you do not want the CLI's built-in defaults.
To check confinement yourself, on any machine, after any Codex upgrade:
node scripts/probe-codex-sandbox.mjsIt attempts a real write with the exact arguments Nodal passes and reports whether the bytes landed.
Approval
The agent proposes the task; the job suspends; you see the full task text, the mode, the working directory and the agent's stated purpose before anything runs.
Approve and the run starts. Reject and the agent gets an error it can act on.
Yolo mode skips this gate per agent, on the Autonomy tab — the same control and the same caveats as shell commands.
The budget
A coding run counts in the agent's budget, set on the agent's Settings
tab: a daily and a monthly ceiling in USD (0 = none). The same ceilings count
the agent's API calls to any provider, so one number says what the agent may
spend, whatever does the work.
Two things are worth knowing about what a coding run adds:
- It is notional. Claude Code reports a cost even when the work is covered by your subscription, and that reported figure is what counts. It is a brake on runaway usage, not a bill.
- Codex reports no cost at all, so its runs add nothing. What bounds them is the per-call timeout. An agent that already spent its day on API calls is still stopped, Codex or not.
When a ceiling is reached, code_task refuses to start and says so, and the
agent's runs stop between two turns, keeping what they wrote, until the day or
the month ends. Before 0.9.3 this was a daily cap on coding runs only, on the
Autonomy tab; an agent that had it keeps it as its daily ceiling.
Watching the work
The Code tab is gone in 0.9.0. /code redirects to Workspaces, and what it carried now
lives in two places that were already there.
The project's page (/spaces/<id>) lists every session and conversation that
touched the project, newest first, with its folder and its proof in a panel
docked to the right.
The run page is what a session row opens, and it is the same page for every origin. It answers what you actually check after the fact:
- Which files were touched, with the diff of the one you select. When the
project folder is a git repository, that list is read from
git statusbefore and after the run, so a command that wrote ten files without naming one is still credited. See Proof. - The activity trail, turn by turn — what the CLI did, in order.
- Token accounting, including cache reads. On a Claude run the cache line is usually dominant, which is worth seeing rather than guessing at.
- The proof — the commands the agent declared to check its own work, with their exit codes and output.
Code is not done until it is reviewed
Write-mode work comes back marked review-required, and the skill instructs the agent never to conclude "done" straight after a successful write.
This is deliberate, and it comes from watching harnesses that let a coder certify its own work: the agent that wrote the change is the worst judge of whether the change is right. Either the agent requests a review, or you read the diff on the run page. Both are fine. Skipping both is not.
When a review does happen, its verdict is a typed record, not the reviewer's prose: the thread and the run page read that record, so a reviewer whose first sentence says one thing and whose verdict says another cannot mislead the screen. Asking for the same review twice, with nothing completed in between, is refused rather than run.
Related
- Coding CLI skill reference — the exact guidance injected into the agent's prompt
- Proof — what the run wrote, what proved it, and the snapshot taken first
- Projects — the folders this work lands in
- Shell commands — the same approval model, for arbitrary commands
- Workspaces — where an agent's files live