Nodal-Agents
Concepts

Memory

How an agent remembers something for later, and how the fact comes back next time it works.

Agents accumulate facts in a persistent agent_memory table. These facts survive across sessions, restarts, and model swaps — whenever a job runs, relevant memories are already waiting in the system prompt before the first LLM call.

When you tell an agent to remember this, to keep something for later, or to do it this way next time, a memory is what you are asking for: the fact is written once, and read back on its own every time it is needed.

What a memory contains

Each record holds a fact (plain text), a category, an importance score (1–5), optional skill_tags, and timestamps for creation and last access.

Importance can be locked by the user (importance_locked in the DB). When locked, the curator's usage-driven re-scoring must leave the value untouched: a user-pinned importance always wins over the agent's guess or the curator's correction.

The four categories are:

CategoryIntended use
preferenceHow the user or agent prefers to work
contextBackground facts about the entity or domain
outcomeResults of past tasks worth remembering
learned_ruleDiscovered rules, constraints, or patterns

Facts can also carry an optional valid_to expiry and an archived flag so stale knowledge is never injected silently.

Auto-injection into every job — the fact is there next time

Before a job starts, the runner selects memories from the entity-wide pool and injects them as a frozen block inside the agent's system prompt. No tool call is needed — the agent already has context.

Selection works in two stages:

  1. Candidate fetch — up to 200 non-archived, non-expired memories for the entity (MAX_CANDIDATES in packages/memory/src/inject.ts). The ordering depends on the current job:
    • Task-relevant ranking. When the job's task is tokenizable, candidates are ordered by ts_rank against the task text (full-text relevance), with importance and recency only as a tie-break. A fact relevant to this task can surface even if it sits outside the importance/recency top tier.
    • Fallback ranking. When the task has no usable tokens, candidates are ordered by importance DESC, then last_accessed_at DESC.
  2. Budget pack — a greedy selector fits as many as possible under a character budget derived from agents.memoryTokenBudget. In the relevance path, the FTS order is authoritative and is not re-sorted by importance. A single oversized fact is skipped rather than truncated.

Memory is entity-scoped, not agent-scoped — every agent in the same entity sees the same pool. This matches the query_memory tool's behaviour: knowledge follows the user across agents, not the individual agent.

Searching memories during a job

Agents can call query_memory explicitly when they need to dig deeper.

By default, search is keyword-only. The runner ships with EMBEDDING_PROVIDER set to keyword, so out of the box there is no embedding step — semantic search does not run unless you explicitly enable it (see below).

  • Keyword (ILIKE) — the default path. Splits the query into words and matches them against the fact text. Needs no embedding client and works with zero extra configuration.
  • Vector (semantic) — available only when an embedding provider is configured. The query is embedded and matched by cosine similarity against stored embeddings (threshold: 0.5 by default). If the embedding call fails or returns no results, the search falls back to keyword.

Both paths update last_accessed_at on every returned row, which feeds back into the injection ranking for future jobs.

Enable semantic memory

Semantic (vector) search is opt-in. To turn it on, set these runner environment variables (defined in apps/runner/src/env.ts):

VariableDefaultPurpose
EMBEDDING_PROVIDERkeywordSet to ollama or openai to enable embeddings
EMBEDDING_MODEL(unset)The embedding model id to use
EMBEDDING_BASE_URL(unset)Override the embedding endpoint (e.g. a local Ollama)

With EMBEDDING_PROVIDER left at keyword, the model/base-URL values are unused and query_memory runs the keyword path only.

Writing memories — remember this, for later

Agents write memories with save_memory — the tool behind "remember that I prefer short answers" and "keep that for later". Duplicate detection runs on a normalized hash of the fact text — an identical (non-archived) fact already in the store is rejected before it reaches the DB, so the pool stays clean.

Content is sanitized before storage to guard against injection payloads being embedded into the entity-wide pool.

  • Agents — how agents are configured, including the memory budget field
  • Learning loop — the related feature that writes reusable skills (not facts) after a job completes

On this page