Learning loop
How agents automatically write and refine reusable skills after a substantial job.
The learning loop is an opt-in feature that lets agents improve themselves over time. After a qualifying job completes, a lightweight reflection pass reviews the transcript and either patches an existing skill or, as a last resort, creates a new one. A separate curator periodically consolidates the growing skill library.
The feature is off by default and must be enabled per workspace, on the
dashboard's Learned Skills page (Agents → Learned Skills). The same page
carries the assignment mode: auto-assign what an agent writes, or hold it for
your approval. There is also a global kill-switch (REFLECTION_ENABLED=false)
that disables it for every workspace regardless of their flag.
Tier 1 — Reflection
The reflection pass fires fire-and-forget after a job reaches completed
status. It never blocks the job response.
Gates (all must pass)
- Global kill-switch not set to
false. - Job status is
completed(failed/blocked jobs are skipped). - Job channel is not
chatorreflection(no interactive turns, no recursion). - The job is substantial enough: it made at least
REFLECTION_MIN_TOOL_ITERStool calls (default: 10), counted across the whole transcript rather than LLM turns. A short heartbeat or cron run with only a couple of tool calls is not worth reflecting on, even if it spanned several turns. - The entity has
reflection_enabled = true. - A per-entity rolling-hour throttle slot is available (
REFLECTION_MAX_PER_HOUR, default: 6).
What the pass does
The reflection model receives a compacted version of the transcript (tool-result
bodies are truncated to ~2 000 chars each) and sees the entity-wide skill
library — not just the skills assigned to this one agent. Every skill is marked
with its provenance: [agent], [user], or [system].
The model has three tools:
skill_view— read the full content of an existing skill before patching itupdate_skill— patch an existing agent-authored skill (the default action)create_skill— author a new skill (last resort only)
Patch-first by design. The system prompt instructs the model to extend an
existing umbrella skill before creating anything new. The reflection pass itself is
bounded by REFLECTION_MAX_TURNS (default: 3), and an optional
REFLECTION_MAX_NEW_SKILLS_PER_PASS cap limits new skills per pass — once hit, the
model is steered to patch instead. The model can only patch skills with
created_by = 'agent' — user- and system-authored skills are provenance-sandboxed
and cannot be modified.
Anti-lesson filter. The prompt includes an explicit filter that blocks the model from persisting environment failures, negative tool claims, transient errors that resolved on retry, and one-off task narratives. Only durable, reusable techniques are written. A no-op pass (no tool call) is the correct and common outcome.
When the entity's skill_assignment_mode is auto, newly created skills are
automatically assigned to the authoring agent. With the default approval mode they
are queued for the entity owner to review.
Tier 2 — Curator
The curator is a separate periodic pass that keeps the skill library from growing unbounded. It operates on the entity's agent-authored, active skills and looks for clusters of narrow, overlapping skills that would be better served by a single broader umbrella.
When it finds a genuine cluster, it:
- Creates an umbrella skill (via
create_skill). - Archives each narrow skill the umbrella replaces (via
archive_skill).
Archiving is the maximum destructive action — nothing is deleted, and the entity owner can recover archived skills. The curator refuses to archive user- or system-authored skills. A no-op pass is correct and common when the library is already well-structured.
Viewing learned skills
The dashboard's Learned Skills page (/learned-skills) lists all skills with
created_by = 'agent' for your entity, showing patch counts and last-used dates.
You can review, edit, approve, or archive them from there.
Tunables
These runner environment variables (defined in apps/runner/src/env.ts) gate and
bound both passes. The defaults are sensible for most self-hosters; the whole
feature still ships off until you enable reflection_enabled per entity.
| Variable | Default | Controls |
|---|---|---|
REFLECTION_ENABLED | (unset) | Global kill-switch — set to false to disable reflection + curator for every entity. |
REFLECTION_MIN_TOOL_ITERS | 10 | Minimum tool-call iterations in the job's transcript before it's "substantial" enough to reflect on. No env default; falls back to this constant when unset. |
REFLECTION_MAX_TURNS | 3 | Max LLM turns inside a single reflection pass. |
REFLECTION_MAX_PER_HOUR | 6 | Per-entity rolling-hour cap on reflection passes. |
REFLECTION_MAX_NEW_SKILLS_PER_PASS | 2 | Hard cap on new skills the reflection pass may create in one run. No env default; falls back to this constant when unset. |
REFLECTION_MODEL | (unset) | Model id to run the reflection/curator passes on; falls back to the agent's model. |
CURATOR_STALE_DAYS | 30 | Days before an agent skill transitions active → stale. |
CURATOR_ARCHIVE_DAYS | 90 | Days before a stale agent skill transitions stale → archived. |
CURATOR_MIN_SKILLS | 5 | Min agent-created active skills per entity to trigger LLM consolidation. |
CURATOR_INTERVAL_DAYS | 7 | Min days between LLM consolidation passes per entity. |
Related pages
- Skills — the broader skills system the learning loop writes into
- Skills reference — catalog of built-in skills
- Memory — the complementary system for persisting facts (not skills) across jobs