This is the abridged developer documentation for Manykind Developer Docs # What is Manykind > Build agents on data your users own. Manykind is an agent platform built on one inversion: a customer’s data stays with the customer. Every customer holds a [**sovereign data vault**](/vaults/overview/), owned by their wallet the way a crypto wallet owns coins, and your agents and workflows work **inside** it. You hold none of that data. There is nothing to import, nothing to migrate, and no store of user data of your own to secure. ## Two sides and a connection | | Who owns it | What it is | | -------------- | ------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Agent** | You, the developer | What you build and ship: a [workflow, an LLM agent, or a code agent](/agents/overview/). Your agents live in your [Agent Service](/get-started/agent-service/), the account that also holds their cubbies, widgets, datasets, and team. | | **Vault** | Your customer: a person or an organization | Their data, their [Memory Bank](/vaults/memory-bank/), their [connectors](/vaults/connectors/) to outside systems, their [members](/vaults/members/). | | **Connection** | Signed by the vault owner | A scoped, revocable agreement that lets one agent run on that vault’s data. | An agent does nothing with a customer’s data until the vault owner [connects](/vaults/connections-and-consent/) it. The connection names which part of the vault it may use, and the owner can withdraw it at any time. ## What you ship * **Workflows.** A typed graph of steps: triggers, model calls, agents, people, cubby reads and writes, connector actions. Build one in the ROC Workflow Builder or in code. A workflow is an agent, so it publishes, deploys, and connects the same way. * **LLM agents.** Instructions, a model, tools, and connections, created in ROC; or an A2A agent you already run elsewhere. * **Code agents.** TypeScript that reacts to events in the vault. All of them call models the platform serves, keep shared working state in [cubbies](/agents/building-blocks/cubbies/), file what they learn into the vault’s [Memory Bank](/vaults/memory-bank/), and can ship [widgets](/agents/building-blocks/widgets/) as their UI. Your own apps sit around them. A web app, a mobile app, or a server job can talk to a vault directly through the [Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) to onboard a customer’s data or read results back. ## What makes Manykind different * **Your agent is the guest.** In a conventional app a user’s data becomes yours the moment they hand it over. Here it never leaves their vault: you get access, not custody. * **Consent is already built.** Every vault that connects your agent has a wallet-signed, scoped agreement on record. Revocation is a signature, not a support ticket. * **No lock-in to defend.** The data lives in the customer’s vault, not in your app, so they can run your agent beside others and keep their history if they switch. *Switch any agent, keep every memory.* * **An open standard.** Manykind is the first implementation of the **Self-Sovereign Context Protocol (SSCP)**: you build against a specification, not one vendor’s product. ## Start building 1. [**How it fits together**](/get-started/how-it-fits-together/): the two sides, the connection, and where each part of these docs fits 2. [**Quickstart**](/get-started/quickstart/): build and run a workflow 3. [**Agents**](/agents/overview/): workflows, LLM agents, and code agents 4. [**Vaults**](/vaults/overview/): what your customer owns, and how your apps work with it Using a coding agent? Point it at our `llms.txt`: see [Build with AI](/reference/build-with-ai/). ## Related * [Next: How it fits together](/get-started/how-it-fits-together/) * [Install the CLI](/get-started/install/) * [Glossary](/reference/glossary/) # Build a widget > Build a widget, a static web page that reads your Agent Service's cubbies and the vault's Memory Bank through window.WidgetRuntime, run it with cef dev, and ship it with an agent or with cef widget push. A **widget** is a static web page (an entry HTML file plus sibling JS, CSS, and assets) that shows what your agents did and lets people send data back. It runs `@cef-ai/widget-runtime`, which signs the reader in and exposes `window.WidgetRuntime`. Your page never handles keys or endpoints. A widget reads two places: | Read | Declared as | Reads | | --------------- | ------------ | ------------------------------------------------------------------------- | | A cubby | `sql` query | One of the Agent Service’s cubbies, the same database `ctx.cubby` writes. | | The Memory Bank | `tool` query | The vault’s Memory Bank: `search`, `get`, `neighbours`, or `countByType`. | ## Two ways to ship one | Ship it | When | Command | | ----------------- | ----------------------------------------------------------------------------------- | ------------------------------------------------------------------- | | **With an agent** | The widget shows that agent’s work. | Declare it in the agent’s `widgets[]`; `cef push` uploads it. | | **On its own** | The widget shows the whole service (a runs board, a report), or a workflow pins it. | A widget subproject with its own `package.json`; `cef widget push`. | ## Option A: with an agent ### 1. Declare it cef.config.ts ```ts widgets: [ { id: "hello", name: "Hello", description: "Recent messages.", cubbyAlias: "history", kind: "custom", queries: [ { id: "recent", label: "Recent messages", sql: "SELECT text, ts FROM messages ORDER BY ts DESC LIMIT 20" }, ], dir: "./widgets/hello", entry: "index.html", }, ], ``` | Field | Meaning | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `id` | Unique within the agent. | | `name`, `description` | Labels. | | `cubbyAlias` | The default cubby for `sql` queries. | | `kind` | `custom` for your own page, or a built-in kind: `list`, `record`, `dashboard`, `submit`, `conversation`, `composite`. See [Visualize vault data](/agents/building-blocks/visualize-vault-data/). | | `config` | The built-in kind’s configuration. | | `queries` | Named reads: `{ id, label?, sql?, cubby?, tool?, limit?, timeoutMs? }`. Set exactly one of `sql` and `tool`. `cubby` picks another service cubby for `sql`. | | `events` | Event types the widget uses. | | `dir` | The built widget directory. `cef build` copies it as is; build it yourself first if it needs a build step. | | `entry` | The entry file in `dir`. Use `index.html` so the widget is served at its directory URL. | ### 2. Write the page widgets/hello/index.html ```html Hello
Loading…
``` Do not add the runtime script or the manifest yourself: `cef dev` and `cef build` inject both into ``. widgets/hello/app.js ```js (function () { var app = document.getElementById("app"); function render(result) { var i = result.columns.indexOf("text"); app.innerHTML = result.rows.length ? "" : "

No messages yet.

"; } function showConnect() { app.innerHTML = ''; document.getElementById("cta").onclick = function () { window.WidgetRuntime.connectAgent().then(load); }; } function load() { window.WidgetRuntime.query("recent").then(render).catch(function (err) { if (err && err.name === "AgentNotConnectedError") showConnect(); else app.textContent = "Error: " + (err && err.message ? err.message : String(err)); }); } load(); })(); ``` `query(id, params?)` resolves to `{ columns, rows, meta: { rowCount } }`. Rows are arrays in column order. Until the reader’s vault has connected the agent, `query` rejects with `AgentNotConnectedError`; `connectAgent()` runs the connection. ### 3. Run it locally ```bash cef dev hello --as-pubkey ``` `cef dev` serves the widget on `127.0.0.1` with the runtime injected and reloads the browser when files change. The page opens top-level, so the reader signs in with the Manykind passkey wallet and the widget reads their vault. Without `--as-pubkey`, the widget has no full agent id and `connectAgent()` and `query()` fail. | Flag | Default | Meaning | | --------------- | --------------------- | ----------------------------------------- | | `[widgetId]` | first declared widget | Which widget to serve. | | `--port ` | a free port | Listen port. | | `--host ` | `127.0.0.1` | Listen host. | | `--env ` | `dev` | Which environment’s endpoints to bake in. | | `--no-watch` | — | Do not reload on change. | ### 4. Ship it `cef build` copies the runtime into each widget, writes the widget’s manifest, and injects both into the entry HTML. `cef build` logs the runtime version it vendors; it is the version the installed CLI resolves, not your project’s pin. `cef push` then uploads each widget directory next to the bundle. See [Push and deploy](/agents/ship/push-and-deploy/). ## Option B: on its own A widget subproject is an ordinary package: its own `package.json`, its own build, output in `dist/`. The `cef` block in `package.json` declares the widget: ```json { "name": "@acme/runs-board", "version": "0.1.0", "scripts": { "build": "vite build" }, "cef": { "id": "runs-board", "name": "Runs board", "entry": "index.html", "scope": "default", "cubbyAlias": "runs", "kind": "custom", "queries": [ { "id": "latest", "label": "Latest runs", "sql": "SELECT * FROM runs ORDER BY started_at DESC LIMIT 50" } ] } } ``` ```bash npm run build cef widget push --bucket --as-pubkey ``` `cef widget push` writes the manifest into the built entry file and publishes the directory under the service bucket’s `widgets` root. The id defaults to the last segment of the package name, and the version to the package version. Every kind except `custom` must render into an element with `id="app"`; the push refuses an entry without one. All flags are in the [CLI reference](/reference/cli/#cef-widget-push). The service’s widgets appear on the **Widgets** page in ROC, and a workflow can pin one; see [Pin it](/agents/building-blocks/widgets/#pin-it). ## How the reader is signed in | Opened | Identity | | --------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | Framed by ROC or another host | The host supplies the identity over a `postMessage` handshake and signs for the widget. A host can also name the vault and the run to show. | | Top-level (a direct link, or `cef dev`) | The reader signs in with the Manykind passkey wallet. A reader with an active session is signed in silently; otherwise the widget shows a sign-in control. | A direct link can name the vault and the run in its fragment: `…/index.html#vault=&run=`. Naming a vault grants nothing: every read is still checked against the reader’s own access. To frame a widget in your own page, mount `createWidgetHost` from `@cef-ai/widget-runtime`. See the [widget-runtime reference](/reference/widget-runtime/#host-contract). ## Related * [Next: Visualize vault data](/agents/building-blocks/visualize-vault-data/) * [Widgets](/agents/building-blocks/widgets/) * [Onboard data with a widget](/agents/building-blocks/onboard-data-with-a-widget/) * [Widget runtime reference](/reference/widget-runtime/) # Cubbies > The SQLite databases an Agent Service's workflows and code agents share — what a cubby is, how a workflow's cubby steps and a code agent's ctx.cubby use it, and when to use the Memory Bank instead. A **cubby** is a SQLite database your Agent Service declares. Every workflow and code agent the service publishes shares it by alias, and it lives in each vault that connects them. It is where shared working state goes: pipeline status, intermediate results, rows a widget renders. ## Shape | Property | What it means | | ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Owned by the service** | Every agent the service publishes reads and writes the same cubby by alias. An agent of a different service cannot reach it. | | **One per vault** | A service’s cubby is global within a vault: one database per vault, service, and alias, whichever scope triggered the run. A run in vault A never sees vault B’s rows. | | **SQLite** | Real tables, indexes, and `CHECK` constraints, declared by migrations. The `sqlite-vec` extension is available for vector search. | | **Durable** | Survives every run and every redeploy. Disconnecting an agent keeps the data. | ## Declare one A cubby is an alias and a directory of numbered `.sql` migrations. Keep one directory per cubby under `cubbies/`; the directory name is the alias. ```plaintext ticket-triage/ ├── cef.config.ts └── cubbies/ └── history/ ├── 001-init.sql └── 002-add-customer.sql ``` ```sql -- cubbies/history/001-init.sql CREATE TABLE IF NOT EXISTS history_messages ( id TEXT PRIMARY KEY, text TEXT NOT NULL, ts INTEGER NOT NULL ); ``` Publish the declarations to your Agent Service’s bucket: ```bash cef cubby push --bucket ``` | Option | What it does | | -------------------------- | ---------------------------------------------------------------------- | | `--bucket ` | Required. The Agent Service’s bucket, the same one you push agents to. | | `--dir ` | Directory holding one subdirectory per cubby. Default `./cubbies`. | | `--alias ` | Push only this cubby. | | `--cubby-version ` | Version to publish each declaration under. Default `0.1.0`. | Credentials are the same as for `cef push`: `$CEF_DDC_ACCESS_TOKEN` or `--access-token`. See [CLI reference](/reference/cli/). You can also declare cubbies in ROC: the Agent Service’s **Cubbies** page has **New cubby** and, per cubby, **Add migration**. Each cubby shows its schema version and published version. An agent or workflow can also declare a cubby itself, with `cubbies: [{ alias, migrations }]` in its config. `cef push` then declares any alias the bucket does not have yet, but never changes an existing declaration: schema changes always go through `cef cubby push`. `cef cubby push` creates no database: a vault’s copy is created, and its migrations applied, the first time an agent or step touches the cubby (see [When migrations run](/agents/building-blocks/cubby-schema/#when-migrations-run)). A cubby belongs to the Agent Service, not to one workflow, so prefix table names with the workflow or feature that owns them (`triage_tickets`, `triage_escalations`), or give each workflow its own alias. Migrations are forward-only; the rules, reserved aliases, and when migrations run are on [Cubby schema and migrations](/agents/building-blocks/cubby-schema/). ## Use it from a workflow Two step kinds reach a cubby: `cubbyQuery` runs a `SELECT` and puts the rows on the carried item; `cubbyExec` runs a write. Values go in `args` with `?` placeholders, never in the SQL: ```ts { id: "save", kind: "cubbyExec", params: { alias: "history", sql: "INSERT OR IGNORE INTO history_messages (id, text, ts) VALUES (?, ?, ?)", args: ["={{ $runId }}", "={{ $json.text }}", "={{ $json.ts }}"], }, } ``` Parameters and the `{{ $runId }}` convention are on [Steps: Cubby query and Cubby exec](/agents/workflows/steps/#cubby-query-and-cubby-exec). ## Use it from a code agent `ctx.cubby(alias).query(sql, params)` returns rows; `.exec(sql, params)` returns `{ changes, lastInsertRowid }`: ```ts await ctx.cubby("history").exec( "INSERT INTO history_messages (id, text, ts) VALUES (?, ?, ?) ON CONFLICT(id) DO NOTHING", [event.eventId, text, Date.now()], ); const recent = await ctx.cubby("history").query<{ text: string }>( "SELECT text FROM history_messages ORDER BY ts DESC LIMIT 20", ); ``` See [Write an agent](/agents/code-agents/overview/#3-use-ctx). ## Other readers | From | How | | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | A widget | Named `sql` queries in the widget declaration. See [Widgets](/agents/building-blocks/widgets/). | | Your app | `vault.service(agentServicePubkey).cubby(alias).query(sql, params)` with the Vault SDK. | | ROC | **Cubbies** → **Inspect data**: the **Tables** list with row counts, a **SQL runner** for queries, and the **Migrations** applied. In the Workflow Builder, the canvas side rail’s **Cubbies** button opens the same view beside the workflow. | ## Write idempotently A Task can run more than once and an event can be delivered twice, in a workflow and in a code agent alike. Key rows with `PRIMARY KEY` or `UNIQUE` (the run id, the event id, a natural key) and write with `ON CONFLICT` or `INSERT OR IGNORE`. ## Cubby or Memory Bank | | Cubby | [Memory Bank](/vaults/memory-bank/) | | ------------------- | ------------------------------------------------------------------ | --------------------------------------------------- | | What goes in | Working state: status, intermediate results, rows a widget shows | Durable conclusions the organization should keep | | Who owns the schema | Your Agent Service | The platform: records and relations | | Who can read it | Your service’s agents and widgets; the vault’s members through ROC | Anyone the vault’s grants and privacy classes allow | | Classification | None | Every record and relation carries a privacy class | Work happens in a cubby; conclusions go to the Memory Bank. ## Limits | Limit | Value | When you hit it | | ----------------- | ----- | ------------------------------------------------------------------------------------------------------------------------------ | | Size of one cubby | 10 GB | Writes are refused: `413 QUOTA_EXCEEDED`, “storage quota exceeded”. Delete or archive old rows; keep large content as objects. | ## Related * [Next: Cubby schema and migrations](/agents/building-blocks/cubby-schema/) * [Steps: Cubby query and Cubby exec](/agents/workflows/steps/#cubby-query-and-cubby-exec) * [Memory Bank](/vaults/memory-bank/) * [Your Agent Service](/get-started/agent-service/) # Cubby schema and migrations > The schema and migration guide for cubbies: where migrations live, when they run, the rules that keep every vault's copy in step, rebuilding a table, table conventions, and vector search. A cubby’s schema is a directory of numbered SQL migrations. The same rules apply whether a workflow’s cubby steps or a code agent’s `ctx.cubby` uses it. What a cubby is and how each kind of agent reads and writes it is on [Cubbies](/agents/building-blocks/cubbies/). ## Where migrations live | Declared by | Location | Published with | | ------------------------- | ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------ | | One agent or workflow | `cubbies: [{ alias, migrations: "./migrations/" }]` in `cef.config.ts` | `cef push` (inside the agent’s manifest) | | The Agent Service | `cubbies//*.sql` in the service project | `cef cubby push --bucket `. See [Cubbies](/agents/building-blocks/cubbies/#declare-one) for its options. | | The Agent Service, in ROC | **Cubbies** → **New cubby** (an alias and its first migration), then **Add migration** per cubby | ROC | Aliases are 1 to 64 letters, digits, `-`, or `_`. Avoid `memory`: the Memory Bank is `ctx.memory`, and a cubby with that name only causes confusion (`cef build` warns about it). Do not use `runs`: it is the workflow runner’s own state. ## When migrations run | Declared | Migrations run | | ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | In an agent’s or workflow’s `cubbies` | When a vault connects it. A failing migration fails the connect with `CUBBY_PROVISION_FAILED`, and nothing is left behind. | | On the Agent Service | Publishing creates no database. The platform creates a vault’s copy the first time an agent or step touches the cubby, and applies every pending migration then. | | Nowhere | An alias nobody declared is created empty on first use. | ## Migration rules Files are named `NNN-description.sql`. The leading number is the migration’s version; a file that does not start with digits fails the build. ```plaintext cubbies/crm/ ├── 001-init.sql ├── 002-crm-notes.sql └── 003-crm-contacts-email-index.sql ``` The platform records the highest applied version in a `schema_migrations` table in each cubby. On each touch it runs, in order, only the files whose version is greater than that. **Never edit an applied migration.** The platform tracks migrations by version number, not by content, so an edited file is silently skipped in every vault that already applied it. Its new SQL runs only in vaults that have never seen that version, and your vaults drift apart. **Make every change a new, higher-numbered file.** A file numbered at or below the highest applied version is skipped too, so do not fill gaps or renumber. **Create new tables only in new files.** Add a table with a new migration, never by extending `001-init.sql`. **Prefix table names.** Every agent in the service shares the cubby, and two agents that both create `messages` collide. Prefix tables with the agent or feature they belong to: `crm_contacts`, `sensor_readings`. Each migration runs in a transaction with its version record. A migration that fails rolls back and is retried on the next touch. ## Change a table SQLite cannot alter a `CHECK` constraint or drop most constraints in place. Rebuild the table in a new migration: ```sql -- 004-crm-contacts-status.sql CREATE TABLE crm_contacts_new ( id TEXT PRIMARY KEY, email TEXT NOT NULL, status TEXT NOT NULL CHECK (status IN ('new', 'active', 'won', 'lost')), updated_at TEXT NOT NULL ); INSERT INTO crm_contacts_new SELECT id, email, status, updated_at FROM crm_contacts; DROP TABLE crm_contacts; ALTER TABLE crm_contacts_new RENAME TO crm_contacts; ``` Adding a nullable column needs only `ALTER TABLE … ADD COLUMN`. ## Table conventions * **Text keys.** Use `TEXT PRIMARY KEY` with an id you control, such as the event id or a natural key, so redelivered events upsert instead of duplicating. * **References without `FOREIGN KEY`.** Store the referenced id in a `TEXT` column and join in your code. This avoids cascade surprises when a migration rebuilds a table. * **JSON for evolving shapes.** Store nested documents as JSON in a `TEXT` column suffixed `_json`, and read them with SQLite’s JSON functions. * **Closed states.** Enforce a state machine with `CHECK (status IN (…))`; widen it with a rebuild. * **Idempotent writes.** `INSERT … ON CONFLICT DO NOTHING`, `ON CONFLICT DO UPDATE`, or `INSERT OR IGNORE`. ## Vector search The cubby loads the `sqlite-vec` extension. Store vectors as `BLOB`, convert JSON-array text with `vec_f32(?)`, and rank in SQL: ```sql -- 005-crm-note-embeddings.sql CREATE TABLE crm_note_embeddings ( note_id TEXT PRIMARY KEY, embedding BLOB NOT NULL, model TEXT NOT NULL, dims INTEGER NOT NULL ); ``` ```ts await ctx.cubby("crm").exec( "INSERT INTO crm_note_embeddings(note_id, embedding, model, dims) VALUES (?, vec_f32(?), ?, ?) ON CONFLICT(note_id) DO UPDATE SET embedding = excluded.embedding", [noteId, JSON.stringify(vector), "my-embedder", vector.length], ); const nearest = await ctx.cubby("crm").query( "SELECT note_id FROM crm_note_embeddings ORDER BY vec_distance_cosine(embedding, vec_f32(?)) LIMIT 5", [JSON.stringify(queryVector)], ); ``` Pass vectors as JSON text, not as typed arrays. ## Related * [Next: Widgets](/agents/building-blocks/widgets/) * [Cubbies](/agents/building-blocks/cubbies/) * [Steps: Cubby query and Cubby exec](/agents/workflows/steps/#cubby-query-and-cubby-exec) * [Event streams](/agents/code-agents/event-streams/) * [CLI reference](/reference/cli/#cef-cubby-push) # Models > How workflows and code agents use models — the catalogue, binding an alias to a model, writing its input, calling it from a model step or from ctx.models, metering, and failures. Building blocks are what your agents use besides their own logic: **models**, **cubbies**, and **widgets**. Each belongs to your Agent Service, and each page here shows how a workflow and a code agent use it. Agents and workflows call models without holding an API key, picking a provider, or knowing where a GPU lives. You bind an **alias** to a model from the platform’s catalogue, call the alias, and the platform routes the request to a server that is serving that model, meters the call, and returns the output. ## Where models come from The platform serves a curated catalogue. Each model is described by a `model.json` in content-addressed storage, at: ```text //models///model.json ``` The `model.json` names the model, its alias, its tasks, its input and output schema, and whether it streams. The catalogue spans text generation, speech-to-text, embeddings, vision, and other tasks. What a model accepts and returns is defined by its own schema, not by a fixed request shape. Open **Models** in ROC’s side navigation to browse it. Each card shows the model’s name, version, task, alias (as `ctx.models.`), price, and whether it is available; search by name, alias, task, or version, and filter by availability (**Active** or **Pending**). Open a model to see its input and output schema and try it. The models in your own Agent Service’s bucket are also listed by the API: ```plaintext GET /api/v1/agent-services/:asPubKey/models ``` ## Declare an alias A workflow and a code agent declare models the same way: a `models` map in `cef.config.ts`, from alias to the model’s `model.json` URL. ```ts import { defineAgent } from "@cef-ai/agent-sdk/config"; export default defineAgent({ id: "summarizer", version: "1.0.0", entry: "./src/agent.ts", models: { llm: "https://cdn.ddc-dragon.com//models///model.json", }, }); ``` `defineWorkflow` takes the same `models` field. | Rule | Detail | | ---- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | URL | The path must be `//models///model.json`, with an integer bucket. `cef build` reads the bucket, name, and version from it and refuses any other shape. | | Key | Must equal the `alias` written inside that `model.json`, which is often not the name in its URL. A key that does not match fails at the first call with `model alias not declared: ''`. | | Use | A model step or a `ctx.models.` reference to an alias that is not in `models` fails `cef build`. | The alias is the stable handle; the binding lives in config. Swapping a model is a config change and a redeploy, not a code change. ## In a workflow A **model step** names an alias, resolves its request from the carried item, calls the model once, and puts the answer on the item under `into`: ```ts { id: "classify", kind: "model", params: { alias: "llm", input: { prompt: "=Classify this ticket:\n{{ $json.text }}", max_tokens: 64 }, into: "classification", }, } ``` Its parameters, structured output, per-entry calls, and when to use an Agent step instead are on [Steps: Model](/agents/workflows/steps/#model). ## In a code agent A code agent calls a declared model with `ctx.models[alias].infer(input)`. **Type it.** Run `cef typegen` after you declare or change `models`. It reads each declared `model.json` and writes `.cef/generated.d.ts`, which types `ctx.models.` from the model’s input and output schemas, and it writes `cef.lock.json`. **Call it.** ```ts @OnEvent("document.added") async onDocument(event: Event<{ text: string }>, ctx: Context) { const out = (await ctx.models.llm.infer({ messages: [ { role: "system", content: "Summarize in two sentences." }, { role: "user", content: event.payload.text }, ], temperature: 0, max_tokens: 400, })) as { text: string }; await ctx.vault.publish("document.summarized", { summary: out.text }); } ``` | Method | Behavior | | --------------- | ------------------------------------------------------------------------------------------------------------------------------------- | | `infer(input)` | Sends `input` to the model and resolves with the model’s output. The input shape is the model’s own, as declared in its `model.json`. | | `stream(input)` | Returns an async iterable. In the agent sandbox it yields the complete output once. It does not stream tokens. | To pass a vault object to a model that fetches it, mint a short-lived URL with `ctx.vault.objects.presignedUrl(path)` and send the URL. **Choose the model per deployment.** Declare a `modelAlias` param so a deployment can switch models without a rebuild: ```ts models: { small: "https://cdn.ddc-dragon.com//models///model.json", large: "https://cdn.ddc-dragon.com//models///model.json", }, params: { model: { type: "modelAlias", default: "small", enum: ["small", "large"] }, }, ``` ```ts const alias = String(ctx.params.model); const out = await ctx.models[alias].infer({ prompt: event.payload.text, max_tokens: 400 }); ``` `enum` must list declared aliases. A deployment that sets a value outside `enum` falls back to the default. See [Push and deploy](/agents/ship/push-and-deploy/#deployment-records). **Test without a cluster.** Stub the alias in the test harness: ```ts import { testAgent, createModelMock } from "@cef-ai/testing"; const h = testAgent(Summarizer, { models: { llm: { infer: async () => "A short summary.", stream: async function* () {} } }, }); ``` `createModelMock()` returns a handle with `.expect(input).respond(output)` and a `.calls` array for assertions. See the [testing reference](/reference/testing/). ## Write the input The input is the model’s own schema; read it on the model’s page in ROC, or rely on the types `cef typegen` generates. An ASR model, for example, takes an audio URL. The platform’s language models take this input and return `{ text }`, plus `tool_calls` when the model calls a tool and `usage` when the model reports it: | Field | Default | Meaning | | ------------------------------------------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `messages` | — | Chat messages `{ role, content }`. Either `messages` or `prompt` is required. | | `prompt` | — | A single user prompt. | | `max_tokens` | `256` | Output budget. Set it explicitly: long output, JSON in particular, is cut off at the budget. | | `temperature` | `0.7` | Sampling temperature. | | `top_p`, `top_k`, `stop` | — | Standard sampling controls. | | `frequency_penalty`, `presence_penalty`, `repetition_penalty` | — | Repetition controls. | | `response_format` | — | Constrain output: `"json"`, `{ type: "json_object" }`, or `{ type: "json_schema", schema }`. See [Structured output](/agents/code-agents/structured-output/). | | `tools` | — | Tool declarations, rendered into the prompt by the model’s chat template. | | `image` | — | Image URL, for vision models. | ## Metering Every inference call is metered against the run that made it. Each call records its duration and, when the model reports them, input tokens, output tokens, and GPU units on the Task. GPU units count toward the connection’s **Compute limit**, which the vault owner can set; once it is reached, new work stops. See [Spend limits](/agents/ship/connect-to-a-vault/#spend-limits). ## Failures Inference crosses the network. A model can be pending, busy, or slow. * **In a workflow**, a model call that fails in transport (a 5xx, a timeout, a dropped connection) is retried up to 3 attempts in total, 0.4 s and then 1.6 s apart. A 4xx fails the step at once. Treat a failed model step as a normal failure path. * **In a code agent**, a failed call rejects `infer`. Catch it and degrade gracefully, and write a model’s output to a [cubby](/agents/building-blocks/cubbies/) as soon as you have it, so a retried Task does not pay for the same call twice. In tests, `createModelMock()` from `@cef-ai/testing` scripts a model’s answers for both. See [Test a workflow](/agents/workflows/test/) and [Test and debug](/agents/code-agents/test-and-debug/). ## Limits | Limit | Value | | --------------------------------- | ----------------------------------------------------- | | Model call attempts in a workflow | 3, transport failures only, 400 ms then 1600 ms apart | ## Related * [Next: Cubbies](/agents/building-blocks/cubbies/) * [Steps: Model](/agents/workflows/steps/#model) * [Code agents](/agents/code-agents/overview/) * [Structured output](/agents/code-agents/structured-output/) # Onboard data with a widget > Capture form input and audio in a widget, upload it into the reader's vault, publish an event for your agent, and show the result, with the built-in submit kind or a custom page. An onboarding widget gets data into a vault. It captures input, stores large files as vault objects, publishes an event that your agent handles, and shows the agent’s result. Everything is written as the signed-in reader, into their vault; your agent sees it only because the vault has connected it. This page assumes the setup from [Build a widget](/agents/building-blocks/create-a-widget/). ## The `submit` kind `submit` renders a form and an audio control, then runs capture → upload → publish → poll → show result. You write no upload or signing code. ```ts widgets: [ { id: "record-call", name: "Record a call", cubbyAlias: "calls", kind: "submit", queries: [ { id: "result", label: "Result", sql: "SELECT status, summary FROM calls_sessions WHERE session_id = ?" }, ], config: { kind: "submit", title: "Record a call", submitLabel: "Submit", form: [ { name: "account", label: "Account", type: "text", required: true }, { name: "stage", label: "Stage", type: "select", options: ["Discovery", "Demo", "Close"] }, ], audio: { mode: "both", required: true, segmentSeconds: 25, softCapSeconds: 1800 }, event: { type: "call.recorded", audioField: "audio_urls", formEnvelope: "meta" }, result: { query: "result", poll: { intervalMs: 3000, timeoutMs: 300000, doneWhen: { column: "status", equals: "done" } }, render: "summary", summary: [{ label: "Summary", path: "summary" }], }, }, dir: "./widgets/record-call", entry: "index.html", }, ], ``` | Config | Meaning | | ------------------------------------------------------------------- | ---------------------------------------------------------------------------- | | `title`, `intro`, `submitLabel`, `processingLabel`, `connectPrompt` | Copy. Defaults: `Submit`, `Processing…`, `Connect your account to continue.` | | `form[]` | `{ name, label, type: "text" \| "number" \| "select", options?, required? }` | | `audio.mode` | `upload` (pick a file), `record` (microphone), or `both`. | | `audio.segmentSeconds` | Length of each uploaded chunk. | | `audio.softCapSeconds` | Warn on recordings longer than this. | | `event.type` | The event to publish. | | `event.audioField` | Payload field that receives the list of audio URLs. | | `event.formEnvelope` | Nest form values under this field; omit to merge them into the payload. | | `result.query` | A declared query, called with the session id as its only parameter. | | `result.poll` | `intervalMs`, `timeoutMs`, and `doneWhen: { column, equals }`. | | `result.render` | `message` or `summary` (labelled `path`s from the row). | ### What happens on submit 1. **Upload.** The audio is split into WAV segments and uploaded as objects into the reader’s vault, in the widget’s scope. Each segment gets a short-lived URL that a model can fetch. 2. **Publish.** The widget publishes `event.type` with `context` set to a new session id. The payload is: ```json { "schema_version": 1, "session_id": "", "audio_urls": ["…"], "meta": { "account": "…", "stage": "…" } } ``` 3. **Poll.** It runs `result.query` with the session id until the `doneWhen` column matches or `timeoutMs` passes. Your agent’s side: handle `call.recorded`, pass the URLs to a model (see [Models](/agents/building-blocks/models/#in-a-code-agent)), and write a row with `session_id`, `status = 'done'`, and `summary` into the cubby. Declare the payload in `eventSchemas` so the contract is in the manifest. `conversation` is the turn-by-turn variant: it publishes a start event, one event per recorded turn, and an optional confirm event, and polls a turns query. See the [widget-runtime reference](/reference/widget-runtime/#kinds). ## Custom page, no audio `publish(type, payload, context?, options?)` writes an event as the reader: ```js document.getElementById("save").onclick = function () { window.WidgetRuntime.publish("contact.added", { schema_version: 1, name: document.getElementById("name").value, }).then(function (r) { console.log("published", r.eventId); }); }; ``` To start a workflow, target it. A workflow handles `workflow.start`, not your trigger’s own event type, and an untargeted `workflow.start` would start every workflow connected in the scope: ```js await window.WidgetRuntime.publish("workflow.start", input, requestId, { target: asPubkey + ":my-workflow", }); ``` Then follow the run with `subscribe(requestId, …)`; see [Visualize vault data](/agents/building-blocks/visualize-vault-data/#live-updates). ## Test it ```bash cef dev record-call --as-pubkey ``` Sign in, submit, and watch the result appear once the deployed agent writes its row. ## Related * [Next: Push and deploy](/agents/ship/push-and-deploy/) * [Data onboarding](/vaults/build-on-a-vault/data-onboarding/) * [Build a widget](/agents/building-blocks/create-a-widget/) * [Event streams](/agents/code-agents/event-streams/) * [Widget runtime reference](/reference/widget-runtime/) # Visualize vault data > Render cubby rows and Memory Bank records in a widget, either with the built-in list, record, and dashboard kinds or with a custom page, and keep it live with subscribe. A read-only widget runs named queries and renders the rows. Start with a built-in kind; write a custom page when you need your own layout or charts. This page assumes the setup from [Build a widget](/agents/building-blocks/create-a-widget/). ## Built-in kinds Set `kind` and a matching `config`; the runtime renders into `
` and handles the empty and not-connected states. Your entry HTML is only the mount point. | Kind | Renders | `config` | | ----------- | -------------------------- | -------------------------------------------------------------------------------------------- | | `list` | One query as a list | `{ kind: "list", query, item: { title, subtitle?, meta? }, empty?, limit? }` | | `record` | One row as labelled fields | `{ kind: "record", query, fields: [{ label, column, format? }], empty? }` | | `dashboard` | Several panels | `{ kind: "dashboard", panels: [{ title, query, render, value?, label?, columns?, item? }] }` | `render` is `metric`, `bar`, `table`, or `list`. Item and field values are column names, optionally with a format: `"created_at:reltime"`. Formats: `text`, `multiline`, `date`, `reltime`, `number`. ```ts widgets: [ { id: "insights", name: "Insights", cubbyAlias: "analysis", kind: "dashboard", queries: [ { id: "total", label: "Total", sql: "SELECT COUNT(*) AS n FROM analysis_findings" }, { id: "by_tag", label: "By tag", sql: "SELECT tag, COUNT(*) AS n FROM analysis_findings GROUP BY tag ORDER BY n DESC LIMIT 5" }, { id: "recent", label: "Recent", sql: "SELECT title, created_at FROM analysis_findings ORDER BY created_at DESC LIMIT 10" }, ], config: { kind: "dashboard", panels: [ { title: "Total findings", query: "total", render: "metric", value: "n" }, { title: "Top tags", query: "by_tag", render: "bar", value: "n", label: "tag" }, { title: "Recent", query: "recent", render: "list", item: { title: "title", subtitle: "created_at:reltime" } }, ], }, dir: "./widgets/insights", entry: "index.html", }, ], ``` `composite` combines panels of any kind behind a menu, with a first-run gate; see the [widget-runtime reference](/reference/widget-runtime/#kinds). ## Custom page ```js function render(result) { var tag = result.columns.indexOf("tag"), n = result.columns.indexOf("n"); document.getElementById("app").innerHTML = result.rows.map(function (r) { return "
" + String(r[tag]) + ": " + String(r[n]) + "
"; }).join(""); } window.WidgetRuntime.query("by_tag").then(render).catch(function (err) { if (err && err.name === "AgentNotConnectedError") { /* show Connect → connectAgent() */ } }); ``` ### Parameters Positional parameters bind to `?` in the declared SQL: ```js window.WidgetRuntime.query("since", [Date.now() - 7 * 24 * 3600 * 1000]).then(render); ``` The page can only run declared queries; it never sends SQL. ## Read the Memory Bank Declare a `tool` query instead of `sql`: ```ts queries: [ { id: "find", label: "Search", tool: "search", limit: 20 }, { id: "record", label: "Record", tool: "get" }, { id: "links", label: "Neighbours", tool: "neighbours", limit: 50 }, { id: "counts", label: "By type", tool: "countByType" }, ], ``` | Tool | First parameter | | ------------------- | --------------- | | `search` | The match text. | | `get`, `neighbours` | A record id. | | `countByType` | None. | ```js window.WidgetRuntime.query("find", ["renewal"]).then(render); ``` The reader sees only the records their access allows. ## Live updates Follow a stream with `subscribe`; each poll’s new events arrive as one ordered batch: ```js var stop = window.WidgetRuntime.subscribe( "run:123", { types: ["analysis.updated"], intervalMs: 2500 }, function (events) { window.WidgetRuntime.query("recent").then(render); }, ); // later: stop(); ``` | Option | Default | Meaning | | -------------- | -------------- | ---------------------------------------------- | | `types` | all | Event types to deliver. | | `intervalMs` | `2500` | Poll interval. | | `maxBackoffMs` | `30000` | Ceiling for the wait while polls keep failing. | | `onError` | `console.warn` | Called on a failed poll. | Polling stops when you call the returned function or when the reader signs out. `onIdentityChange(cb)` reports sign-in changes (`status`: `anon`, `connecting`, `ready`). ## Related * [Next: Onboard data with a widget](/agents/building-blocks/onboard-data-with-a-widget/) * [Build a widget](/agents/building-blocks/create-a-widget/) * [Cubby schema and migrations](/agents/building-blocks/cubby-schema/) * [Memory Bank](/vaults/memory-bank/) * [Widget runtime reference](/reference/widget-runtime/) # Widgets > The UI an Agent Service ships — browser screens that act as the signed-in person, read the service's cubbies and the vault's Memory Bank, and how a workflow and a code agent each use one. A **widget** is a screen your Agent Service publishes. It runs in the browser, acts as **the signed-in person**, and reads and writes the vault that person is working in. There is no backend of yours behind it and no copy of the data to hold. ## What a widget does * **Shows results.** Render rows from the service’s [cubbies](/agents/building-blocks/cubbies/) and records from the vault’s [Memory Bank](/vaults/memory-bank/). A widget reads from exactly those two places: a named query with `sql`, or one with `tool`. * **Takes input.** Record audio, upload a file, submit a form, and write it into the vault. * **Starts and follows work.** Publish an event, or `workflow.start` targeted at one workflow, then follow that stream’s events as they arrive. Because a widget acts as the person, everything it does lands under that person’s own access to the vault, never under a credential of yours. ## Where a widget comes from | Source | Declared in | Published by | | --------------------- | ----------------------------------------------- | -------------------------- | | **An agent’s widget** | `widgets: [...]` in the agent’s `cef.config.ts` | `cef push`, with the agent | | **A service widget** | Its own widget project | `cef widget push` | ROC’s **Widgets** page lists every screen the service publishes. ## With a workflow A workflow’s cubby steps write rows; a widget is the screen onto them. ### Pin it Pin a service widget to a workflow and it opens from the end of the workflow, handed the run you are looking at, so it shows exactly that run’s rows. Build and publish the widget first: see [Build a widget](/agents/building-blocks/create-a-widget/). Every workflow canvas ends with an **Outcome** card. 1. Click **Pin a widget** on the Outcome card. The picker lists the widgets your Agent Service has published. 2. Pick one. The card shows the widget’s name; its menu has **Open**, **Change…**, and **Unpin**. 3. Deploy. The pin is saved with the workflow’s canvas as `pinnedWidgetId`. The same pinned widget appears in the **Summary** panel of a run. Opening it from a run passes that run to the widget. ![The Outcome card with a pinned widget and its menu open](/shots/outcome-card.png) A pin belongs to the Builder canvas. A workflow pushed from a repo has no canvas, so it carries no pin. ### Read the run it was opened for When a person opens the pinned widget from a run, the host hands the widget a context: | Field | Value | | --------- | ---------------------------------------------------------------------------------------------- | | `runId` | The run’s id, `run-`. The same value as `{{ $runId }}` in the workflow. | | `vaultId` | The vault the run happened in. The widget reads that vault’s data, not the viewer’s own vault. | ```js const { context } = await window.WidgetRuntime.identity(); const runId = context?.runId; const result = await window.WidgetRuntime.query("by-run", [runId]); ``` with a query declared on the widget such as: ```sql SELECT ticket_id, category, priority FROM triage_tickets WHERE run_id = ? ORDER BY ticket_id ``` This works because the workflow’s cubby steps write `{{ $runId }}` into a `run_id` column; see [Steps: Cubby query and Cubby exec](/agents/workflows/steps/#cubby-query-and-cubby-exec). A share link to a widget carries the same values in its URL fragment: `#vault=&run=`. Without a `runId` (opened directly, not from a run), show the newest rows instead. ### Start runs from it Publish `workflow.start` with `{ target: ":" }`, then subscribe to the stream for `workflow.completed`. ```js await WidgetRuntime.publish("workflow.start", input, requestId, { target: `${agentServicePubkey}:${workflowId}`, }); const stop = WidgetRuntime.subscribe( requestId, { types: ["workflow.completed"] }, (events) => events.forEach(render), ); ``` ## With a code agent Ship the widget with the agent: declare it in the agent’s `widgets`, and `cef push` uploads it next to the bundle. It reads the cubbies the agent writes with `ctx.cubby`, and publishes the events the agent’s `@OnEvent` handlers handle. Until the reader’s vault has connected the agent, its queries fail with `AgentNotConnectedError`; `connectAgent()` runs the connection. See [Build a widget](/agents/building-blocks/create-a-widget/#option-a-with-an-agent). ## The runtime `@cef-ai/widget-runtime` is the browser half of every widget. It boots from the manifest the CLI bakes, signs the reader in, and exposes `window.WidgetRuntime`: | Member | Does | | ---------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | `query` | Run a named query: `sql` against one of the service’s cubbies, or `tool` (`search`, `get`, `neighbours`, `countByType`) against the vault’s Memory Bank. | | `publish` | Publish an event into the widget’s scope; pass `{ target }` to address one agent or workflow. | | `subscribe` | Follow one stream of the scope; new events arrive as ordered, de-duplicated batches. | | `connect`, `connectAgent`, `agentStatus` | Connect an agent and check its connection. | | `identity` | The resolved identity, and the vault and run a host named. | The runtime also renders config-driven kinds (`record`, `list`, `dashboard`, `submit`, `conversation`, `composite`) and carries the audio capture and upload path. Fully custom pages use the same surface. A widget framed by ROC or your own page gets the reader’s identity from the host; opened by direct link, it signs the reader in itself. See [How the reader is signed in](/agents/building-blocks/create-a-widget/#how-the-reader-is-signed-in). ## Build one | Task | Page | | ----------------------------------------------- | --------------------------------------------------------------------------------- | | Declare, write, run locally, and ship a widget | [Build a widget](/agents/building-blocks/create-a-widget/) | | Render cubby rows and Memory Bank records, live | [Visualize vault data](/agents/building-blocks/visualize-vault-data/) | | Capture input and files into the vault | [Onboard data with a widget](/agents/building-blocks/onboard-data-with-a-widget/) | | Open it from a workflow’s runs | [Pin it](#pin-it) | ## Related * [Next: Build a widget](/agents/building-blocks/create-a-widget/) * [Widget runtime reference](/reference/widget-runtime/) * [Cubbies](/agents/building-blocks/cubbies/) # Event streams > Feed a stream of events into a vault scope with the vault SDK, handle each one in a code agent, keep handlers idempotent, and follow the stream from a client. Events are how data enters a vault and how agents react to it. A **stream** is every event in a scope that shares one `context` value. You never create a stream: you publish events with a `context`, and they form one. Each connected agent handles the events of one stream in one Job. A new `context` starts a new Job. ## 1. Publish events From a client or server, use [`@cef-ai/vault-sdk`](/reference/vault-sdk/): ```ts import { VaultSDK, KeypairWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: vaultApiUrl, signer: await KeypairWallet.fromSeed(seed) }); const vault = await sdk.vault.current(); const { eventId } = await vault.scope("default").publish({ type: "reading.received", context: "sensor:42", payload: { schema_version: 1, celsius: 21.5 }, }); ``` | Field | Required | Meaning | | --------------------------------------------------- | -------- | --------------------------------------------------------------------------------------------------------------------------------- | | `type` | yes | The event type. A code agent handles it with `@OnEvent("")`. | | `context` | yes | The stream key. Events that share it form one stream. | | `payload` | yes | Your data. | | `target` | no | `:`. Delivers the event only to that agent. Without it, every connected agent that handles the type receives it. | | `role` | no | `"source"`, `"user"`, or `"agent"`. | | `correlationId`, `metadata`, `parents`, `timestamp` | no | Tracing, labels, causal lineage, and the source time. | `publish` throws if the vault rejects the event. Publishing returns before any agent runs: dispatch is asynchronous. An agent publishes with `ctx.vault.publish(type, payload, opts?)`. The event joins the stream of the event being handled. ## 2. Handle them ```ts import { Engagement, OnEvent, type Context, type Event } from "@cef-ai/agent-sdk"; type Reading = { schema_version: number; celsius: number }; @Engagement({ id: "default", goal: "Track sensor readings" }) export default class SensorAgent { @OnEvent("reading.received") async onReading(event: Event, ctx: Context) { if (event.payload?.schema_version !== 1) { console.warn("unknown schema_version", event.payload); return; } await ctx.cubby("readings").exec( "INSERT INTO sensor_readings(event_id, context, celsius, at) VALUES (?, ?, ?, ?) ON CONFLICT(event_id) DO NOTHING", [event.eventId, event.context, event.payload.celsius, event.timestamp], ); } } ``` `event` carries `type`, `payload`, `timestamp` (ISO 8601), `context`, `role`, and, when present, `from`, `eventId`, and `parents`. Give each payload a `schema_version` and ignore versions you do not know. Old and new shapes then coexist safely during a rollout. ## Make handlers idempotent Delivery is at-least-once, so the same event can arrive twice. Make every write safe to repeat: * Key rows on `event.eventId` or a natural key, and write with `ON CONFLICT … DO NOTHING`, `ON CONFLICT … DO UPDATE`, or `INSERT OR IGNORE`. * Keep no state in memory between events. Each event can run in a fresh isolate. Read what you need from the cubby when the handler starts. ## Who receives an event The vault delivers an event to every agent connected in that scope that handles its type, unless the event names a `target`. LLM agents receive only events targeted at them. An event no connected agent handles reaches nobody, and nothing reports an error. ## Follow a stream ```ts const sub = vault.scope("default").subscribe( { context: "sensor:42", types: ["reading.alert"] }, (event) => console.log(event.type, event.payload), { intervalMs: 1000 }, ); // later sub.unsubscribe(); ``` `subscribe` polls the stream, every second by default, and delivers each event once. `subscribeAll({ types }, handler)` follows every stream in the scope. To read history, use `vault.scope(s).stream(context).events.list()`. A widget follows a stream with `WidgetRuntime.subscribe`; see [Visualize vault data](/agents/building-blocks/visualize-vault-data/#live-updates). ## Test with the Sandbox In ROC, open **Sandbox**, pick the agent and a vault, and use the simulator: **Single event** sends one event, **Data stream** sends a sequence, and **Replay** replays events from the vault into a new context. ## Failure modes | Symptom | Cause | Fix | | ---------------------- | --------------------------------------------------------------------------------------- | ------------------------------------------------------ | | Event accepted, no Job | No connected agent handles that type in that scope, or the event targets another agent. | Connect the agent; check the type string and `target`. | | Duplicate rows | A redelivered event. | Make the write idempotent. | | `PAYLOAD_TOO_LARGE` | The event is too big. | Upload the data to vault objects and publish its path. | ## Limits | Limit | Value | | -------------------------- | ------ | | Publish request body | 1 MiB | | Events per publish call | 100 | | Event retention per stream | 7 days | See [Runs and events](/agents/how-agents-run/#limits). ## Related * [Next: Structured output](/agents/code-agents/structured-output/) * [Runs and events](/agents/how-agents-run/) * [Write an agent](/agents/code-agents/overview/) * [Cubby schema and migrations](/agents/building-blocks/cubby-schema/) * [Vault SDK reference](/reference/vault-sdk/) # Write an agent > Write a code agent in TypeScript: engagement classes, the ctx surface, defineAgent, and schedules, then build, push, and deploy it with the cef CLI. This group covers code agents: writing one, handling event streams, getting structured output from a model, and testing and debugging it. How a code agent calls a model is on [Models](/agents/building-blocks/models/#in-a-code-agent). A **code agent** is TypeScript you write with `@cef-ai/agent-sdk`. You declare what it uses in `cef.config.ts`, write its handlers as **engagement** classes, then ship it with `cef build` → `cef push` → `cef deploy`. It runs in a sandboxed V8 isolate, inside the vault of whoever connected it. ## Code agent, LLM agent, or workflow | Choose | When | | ---------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | A **workflow** | The work is a sequence of steps (trigger, model, agent, person, cubby, connector) that you want to see, edit in the ROC Workflow Builder, and evaluate. | | An **LLM agent** | A prompt describes the work: instructions, a model, and tools are enough. See [LLM agents](/agents/llm-agents/overview/). | | A **code agent** | You need logic a graph or a prompt does not express well, such as custom parsing, loops over external APIs, or a conversation with its own state machine. | A workflow can call a code agent as one of its steps, as it calls an LLM agent, so you can combine them. See [Agents and workflows](/agents/overview/). ## 1. Scaffold ```bash npx @cef-ai/cli init my-agent cd my-agent npm install @cef-ai/agent-sdk@5.8.0 npm install -D @cef-ai/cli@2.8.0 @cef-ai/testing@3.3.5 ``` `cef init` writes a hello-world project: `cef.config.ts`, `src/agent.ts`, `migrations/`, `deployments/`, a widget, and a test. The flags are listed in the [CLI reference](/reference/cli/#cef-init). Your `tsconfig.json` needs `"experimentalDecorators": true`. The scaffold sets it. ## 2. Write an engagement An **engagement** is a class whose methods handle events. When an event reaches a connected agent, the platform picks one engagement and creates a **Job** for it. That engagement handles every event in the Job until the Job ends. src/agent.ts ```ts import { Engagement, OnEvent, OnStart, OnClose, type Context, type Event } from "@cef-ai/agent-sdk"; @Engagement({ id: "default", goal: "Reply to a user message and remember it" }) export default class HelloAgent { @OnStart async onStart() { console.info("started"); } @OnEvent("user_message") async onMessage(event: Event<{ text?: string }>, ctx: Context) { const text = typeof event.payload?.text === "string" ? event.payload.text.trim() : ""; if (!text) { await ctx.vault.publish("reply", { text: 'send { "text": "..." }' }); return; } const now = Date.now(); await ctx.cubby("history").exec( "INSERT OR IGNORE INTO messages(id, text, ts) VALUES (?, ?, ?)", [event.eventId ?? String(now), text, now], ); await ctx.vault.publish("reply", { text: `you said: ${text}` }); } @OnClose async onClose(_ctx: Context, reason: string) { console.info("closing", { reason }); } } ``` | Decorator | Applies to | Effect | | --------------------------- | ---------- | -------------------------------------------------------------------------------------------------------------------------------------- | | `@Engagement({ id, goal })` | class | Names the engagement. `id` must match `[a-zA-Z][a-zA-Z0-9_-]*`. | | `@OnEvent("type")` | method | Handles events of that type. The argument must be a string literal. The method receives `(event, ctx)`. | | `@OnStart` | method | Runs once when the Job starts, before the first event. | | `@OnClose` | method | Runs when the Job ends. Receives `(ctx, reason)`, where `reason` is `"revoked"`, `"idle_timeout"`, `"closed_by_agent"`, or `"failed"`. | Handlers must not throw on bad input, so check the payload first, as in the example. The same event can be delivered more than once, so write cubby rows idempotently. See [Event streams](/agents/code-agents/event-streams/#make-handlers-idempotent). ### Several engagements Declare each engagement as its own class and list them in `cef.config.ts` (`engagements: [{ id, entry }, …]`). Put these decorators on a class to control when the platform picks it: | Decorator | Effect | | ------------------------ | -------------------------------------------------------------------------------------------------- | | `@Condition("cel expr")` | A CEL selection expression. You can apply it more than once; every condition must hold. | | `@Priority(n)` | A lower value wins. | | `@Weight(n)` | Splits traffic by probability among engagements with the same priority. | | `@Limit(n, per)` | Caps how often the engagement runs. `per` is `"day"`, `"connection/day"`, or `"connection/month"`. | | `@Params({ … })` | Param values for this engagement. | Only one engagement may declare `@OnStart`, and only one may declare `@OnClose`. `cef build` fails if two do. ## 3. Use `ctx` `ctx` gives you the platform. Everything else is a global in the sandbox: log with `console.*`, call HTTP with `fetch`, and use `globalThis.crypto` for WebCrypto. | Member | What it does | | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ctx.vault.publish(type, payload, opts?)` | Publishes an event into the scope, in the same stream as the event that triggered it. `opts` takes `target` (`:`, the one agent the event is for), `title`, `description`, and `correlation`. | | `ctx.vault.objects` | `upload`, `get`, `head`, `presignedUrl`, and `list` on the vault’s object storage. There is no `delete`: only the vault owner can delete objects. | | `ctx.cubby(alias)` | A SQLite store with `query(sql, params)` and `exec(sql, params)`. The cubby belongs to the Agent Service, so all of the service’s agents share it. See [Cubby schema and migrations](/agents/building-blocks/cubby-schema/). | | `ctx.memory` | The vault’s Memory Bank: `upsert`, `update`, `setPrivacy`, `delete`, `relation`, `search`, `get`, `neighbours`, and `countByType`. Every write needs a `privacy` value: `"public"`, `"internal"`, `"private"`, or `"restricted"`. | | `ctx.models[alias]` | The models you declared. Call them with `infer(input)`. See [Models](/agents/building-blocks/models/#in-a-code-agent). | | `ctx.params` | The effective params for this Job (read-only). | | `ctx.settings` | The values the vault owner entered when they connected the agent (read-only). | | `ctx.self` | `agentId`, `vaultId`, `scope`, `context`, `jobId`, and `taskId`. To address a sibling agent, take your own `agentId` and swap in the sibling’s alias. | | `ctx.close(reason?)` | Ends the Job. `@OnClose` then receives `"closed_by_agent"`. | ## 4. Declare the agent in `cef.config.ts` ```ts import { defineAgent } from "@cef-ai/agent-sdk/config"; export default defineAgent({ id: "my-agent", version: "0.1.0", entry: "./src/agent.ts", idleTimeout: "30m", card: { name: "My agent", description: "Replies to messages and remembers them." }, cubbies: [{ alias: "history", migrations: "./migrations/history" }], params: { temperature: { type: "number", default: 0.3, min: 0, max: 1 }, }, settings: [ { key: "apiKey", type: "secret", required: true, label: "API key" }, ], eventSchemas: { user_message: { type: "object", properties: { text: { type: "string" } }, required: ["text"] }, }, }); ``` | Field | Meaning | | ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `id` | The alias. The full agent id is `:`. | | `version` | Semver for this build. | | `agentServicePubkey` | Optional. The Agent Service pubkey as shown in ROC. When you set it, `cef build` writes the full identity into the manifest. `--as-pubkey` on `cef push` overrides it. | | `entry` / `engagements` | Either one entry file, or a list of `{ id, entry, goal?, condition?, priority?, weight?, limit?, params?, enabled? }`. Set exactly one of the two. | | `card` | `name` and `description`, plus optional `iconUrl` and `capabilities`. ROC shows the card. | | `cubbies` | `{ alias, migrations? }`. `migrations` is a directory of numbered `*.sql` files. | | `models` | A map from alias to a DDC `model.json` URL. Run `cef typegen` after changing it. See [Models](/agents/building-blocks/models/#in-a-code-agent). | | `params` | A map from name to `{ type, default, min?, max?, enum? }`. `type` is `"number"`, `"string"`, `"boolean"`, or `"modelAlias"`. | | `settings` | `{ key, type, required?, label?, description?, default? }`. `type` is `"string"`, `"number"`, `"boolean"`, `"url"`, or `"secret"`. | | `schedules` | Recurring triggers. See [Schedules](#schedules). | | `widgets` | UIs shipped with the agent. See [Build a widget](/agents/building-blocks/create-a-widget/). | | `requiredScopes` | The vault scopes the agent may be connected into. Defaults to `["default"]`. A connection that names any other scope is refused with `MANIFEST_INVALID`. `"*"` allows any scope. | | `idleTimeout` | A duration such as `"500ms"`, `"30s"`, `"15m"`, `"1h"`, or `"2d"`. The Job ends after this long with no activity. Defaults to `"30m"`; `"0s"` turns it off. | | `eventSchemas` | JSON Schemas for the events the agent owns. | The cubby’s schema lives in its migrations directory: ```sql -- migrations/history/001-init.sql CREATE TABLE messages ( id TEXT PRIMARY KEY, text TEXT NOT NULL, ts INTEGER NOT NULL ); ``` **Params and settings.** Params are your own settings for tuning the agent. For each Job, a param starts at its manifest default; the deployment can override it, and the engagement can override that. If the result is outside the declared `min`, `max`, or `enum`, the param falls back to its default. Settings are values the vault owner enters when they connect the agent. `type: "secret"` only labels the form field. Nothing encrypts or hides the value. ### Schedules A schedule makes the vault publish an event to the agent on a cron schedule. Each connected vault runs its own copy, and the vault owner can pause it or change its timing. ```ts schedules: [ { id: "daily-digest", cron: "0 8 * * 1-5", timezone: "Europe/Berlin", eventType: "digest.run", payload: { window: "24h" } }, ], ``` | Field | Rule | | ----------- | ------------------------------------------------------------------------------------------------------------- | | `id` | Keep it stable across versions. Renaming it replaces the schedule. | | `cron` | Five fields: minute, hour, day of month, month, day of week. | | `timezone` | An IANA timezone name. Defaults to UTC. | | `eventType` | Must be an event the agent handles. Otherwise `cef build` refuses the config. | | `payload` | Merged into the payload of every event the schedule publishes, alongside the `schedule` block the vault adds. | Schedules fire at 1-minute granularity at most, and a fire more than 4 minutes late is skipped, never backfilled. See [Schedule limits](/agents/workflows/triggers/#schedule-limits). ## 5. Test ```ts import path from "node:path"; import { testAgent } from "@cef-ai/testing"; import HelloAgent from "../src/agent.js"; const h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: path.resolve("migrations/history") }], }); const events = await h.dispatch({ type: "user_message", payload: { text: "hi" } }); // events contains { type: "reply", … } const rows = await h.cubby("history").query("SELECT text FROM messages"); h.dispose(); ``` See the [testing reference](/reference/testing/). ## 6. Build, push, deploy ```bash cef build # dist//bundle.js + manifest.json cef push --bucket --as-pubkey cef deploy --as-pubkey # applies deployments/ ``` A deployed agent still does nothing until a vault owner connects it. See [Push and deploy](/agents/ship/push-and-deploy/) and [Connect to a vault](/agents/ship/connect-to-a-vault/). ## What `cef build` rejects * An `@OnEvent` argument that is not a string literal. * `ctx.models.` for an alias that is not in `models`. * Imports of Node built-ins (`fs`, `path`, `net`, `http`, `child_process`, `crypto`, `buffer`, `stream`, and others) and of `pg`, `mysql`, `mysql2`, `mongodb`, `redis`, `ioredis`, `sqlite3`, `better-sqlite3`, `ws`, `node-fetch`, and `axios`. * `@OnEvent("__start__")` or `@OnEvent("__close__")`. Use `@OnStart` and `@OnClose` instead. * A schedule whose `eventType` the agent does not handle. A `ctx.cubby` alias that is not in `cubbies` only produces a warning, because the Agent Service may declare that cubby. The [eslint plugin](/reference/eslint-plugin/) shows the same checks in your editor. ## Related * [Next: Event streams](/agents/code-agents/event-streams/) * [Agents and workflows](/agents/overview/) * [Agent SDK reference](/reference/agent-sdk/) * [Push and deploy](/agents/ship/push-and-deploy/) * [LLM agents](/agents/llm-agents/overview/) # Structured output > Get a typed object from an LLM: constrain generation with response_format, validate with Zod, and repair once, so agent code receives a checked value or a typed error. When an agent needs an object rather than prose, constrain the model with `response_format`, then validate what comes back. Constraining makes a valid result likely; validating makes it certain. ## 1. Define the shape ```ts import { z } from "zod"; const Verdict = z.object({ verdict: z.enum(["positive", "neutral", "negative"]), score: z.number().int().min(0).max(10), summary: z.string(), }); type Verdict = z.infer; ``` ## 2. Constrain generation Platform LLMs accept `response_format` and use guided decoding to make the output match it: | `response_format` | Constrains output to | | -------------------------------------------------- | -------------------------------------- | | `"json"` or `{ type: "json_object" }` | Any valid JSON. | | `{ type: "json_schema", schema }` | JSON that matches `schema`. | | `{ type: "json_schema", json_schema: { schema } }` | The same, in the OpenAI request shape. | Convert the Zod schema with [`zod-to-json-schema`](https://www.npmjs.com/package/zod-to-json-schema). `$refStrategy: "none"` inlines definitions so the schema stands alone: ```ts import { zodToJsonSchema } from "zod-to-json-schema"; const responseFormat = { type: "json_schema" as const, schema: zodToJsonSchema(Verdict, { $refStrategy: "none" }), }; ``` ## 3. Call, validate, repair once ````ts import type { Context } from "@cef-ai/agent-sdk"; import { z } from "zod"; import { zodToJsonSchema } from "zod-to-json-schema"; type Message = { role: "system" | "user" | "assistant"; content: string }; export class StructuredOutputError extends Error { constructor(message: string, readonly attempts: number, readonly lastText: string) { super(message); } } function parseJson(text: string): unknown { return JSON.parse(text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```$/, "")); } export async function callStructured( ctx: Context, alias: string, schema: T, messages: Message[], maxTokens: number, ): Promise> { const response_format = { type: "json_schema" as const, schema: zodToJsonSchema(schema, { $refStrategy: "none" }), }; let turn = messages; let lastText = ""; for (let attempt = 1; attempt <= 2; attempt++) { const out = (await ctx.models[alias].infer({ messages: turn, response_format, max_tokens: maxTokens, temperature: 0, })) as { text: string }; lastText = out.text; const parsed = schema.safeParse((() => { try { return parseJson(out.text); } catch { return undefined; } })()); if (parsed.success) return parsed.data; turn = [ ...messages, { role: "assistant", content: out.text }, { role: "user", content: "That did not match the required schema. Reply with only the corrected JSON object." }, ]; } throw new StructuredOutputError("structured output failed after one repair", 2, lastText); } ```` Call it from a handler: ```ts const result = await callStructured(ctx, "llm", Verdict, [ { role: "system", content: "Classify the message. Reply in JSON." }, { role: "user", content: event.payload.text }, ], 1024); ``` ## Set `max_tokens` The default output budget is 256 tokens. A JSON object cut off at the budget fails to parse with `Unexpected end of JSON input`, which looks like a schema problem but is a budget problem. Size `max_tokens` to the largest output your schema allows. ## Repair in code before you repair with the model A repair call costs another inference. Fix mechanical violations in code first, and keep the model repair as the last step: | Layer | Fix | | ----------------------- | -------------------------------------------------------------------------------------------------------------------- | | 1. Parse and validate | Zod `safeParse`. | | 2. Deterministic repair | Snap to the nearest enum value, clamp numbers into range, trim arrays to their maximum length. Log what you changed. | | 3. Cross-field checks | Warn on contradictions; do not block. | | 4. Model repair | One round trip, as above. Never more than one per object. | When a run extracts several objects, return the ones that succeeded and mark the gaps instead of failing the whole run. ## When to use it Use it when downstream code reads fields from the result. For free text such as summaries or replies, call the model directly and treat the output as a string. ## Related * [Next: Test and debug](/agents/code-agents/test-and-debug/) * [Models](/agents/building-blocks/models/) * [Models](/agents/building-blocks/models/) * [Agent SDK reference](/reference/agent-sdk/) # Test and debug > Test a code agent in-process with testAgent and testPlatform from @cef-ai/testing, run the local loop, read the agent's logs, and match common failure signals to causes. Test a code agent in-process first: `@cef-ai/testing` runs your engagement classes against an in-memory vault and SQLite cubbies built from your real migrations, with no network. Go to the platform only when the bug needs a real model, real vault data, or a real trigger, and read the agent’s logs there. ```sh pnpm add -D @cef-ai/testing vitest ``` ## The local loop | Loop | Command | Catches | | ----------------------------------------- | -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Handler logic | `vitest` + `testAgent` | Payload checks, cubby SQL, what the handler publishes, model-output handling. `console.*` prints to your terminal. | | Several agents, or your app with an agent | `vitest` + `testPlatform` | Events routed between agents, Jobs per stream, Memory Bank and object writes. | | Build checks | `cef build` | Banned imports, non-literal `@OnEvent` types, undeclared model aliases, unhandled schedule events. See [What `cef build` rejects](/agents/code-agents/overview/#what-cef-build-rejects). | | Bundle routing | `cef inspect dist/` | A handler the manifest routes to that the bundle does not export. Exits with code `2`, so it fits CI. | `cef test` is not implemented; run `vitest` directly. ## Test one agent with testAgent `testAgent(Agent, opts)` builds a harness around one agent class. `dispatch` runs one event through the matching `@OnEvent` handler and resolves to the events the handler published. The example tests the agent from [Write an agent](/agents/code-agents/overview/#2-write-an-engagement). test/agent.test.ts ```ts import path from "node:path"; import { afterEach, describe, expect, it } from "vitest"; import { testAgent, type TestHarness } from "@cef-ai/testing"; import HelloAgent from "../src/agent.js"; const HISTORY = path.resolve("migrations/history"); describe("hello agent", () => { let h: TestHarness | undefined; afterEach(() => h?.dispose()); it("replies and stores the message", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); const events = await h.dispatch({ type: "user_message", payload: { text: "hi" } }); expect(events.map((e) => e.payload)).toEqual([{ text: "you said: hi" }]); const rows = await h.cubby("history").query<{ text: string }>("SELECT text FROM messages"); expect(rows.map((r) => r.text)).toEqual(["hi"]); }); it("answers bad input without writing", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); await h.dispatch({ type: "user_message", payload: {} }); expect(await h.cubby("history").query("SELECT * FROM messages")).toEqual([]); }); it("runs the lifecycle hooks", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); await h.start(); await h.close("idle_timeout"); expect(h.lifecycle).toEqual({ state: "closed", terminalReason: "idle_timeout" }); }); }); ``` An event type no handler declares is ignored, as on the platform. A handler that throws rejects `dispatch`, so assert on bad input explicitly. ### Stub what leaves the agent | Dependency | Stub it with | | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ctx.models[alias]` | `models: { llm: createModelMock() }`, then `llm.expect(input).respond(output)`. An unregistered input throws `[modelMock] no expectation for input …`. `llm.calls` lists every call. See [Models](/agents/building-blocks/models/#in-a-code-agent). | | Global `fetch` | `h.fetchMock().when({ url, method? }).reply(status, body?)`. A request with no stub throws `[fetchMock] no stub for `. `assertCalled({ url })` checks a call was made. | | `ctx.settings`, `ctx.params` | The `settings` and `params` options. Both are frozen, as on the platform. | | `ctx.self` | `self: { agentId, vaultId, scope, context }`. The default is a valid synthetic identity, so code that derives a sibling’s id works. | | Time | `clock: { now }`, then `h.advanceTime(ms)`. | `testAgent` does not back `ctx.memory` or `ctx.vault.objects`: calling them throws. Test an agent that uses either with `testPlatform`. ### Replay a recorded session `h.replay("test/fixtures/session.jsonl")` feeds a file of events, one JSON object per line with at least `type`, through the harness. It resolves to the events published during the replay, a row count per cubby table, and the duration. Use it to pin an agent’s behavior over a whole stream. Keep fixtures synthetic, never copied from customer data. `h.snapshot()` and `h.restore(snap)` capture and restore cubbies, the clock, and the publish log, so several tests can start from one prepared state. ## Test agents together with testPlatform `testPlatform({ agents })` runs agents in a simulated vault: a publish routes to the connected agents that handle its type and resolves once they have run, `ctx.memory` writes to an in-memory Memory Bank, and `ctx.vault.objects` to an in-memory object store (`p.objects`). ```ts import path from "node:path"; import { expect, it } from "vitest"; import { testPlatform } from "@cef-ai/testing"; import HelloAgent from "../src/agent.js"; it("replies in the same stream", async () => { const p = testPlatform({ agents: { hello: { source: HelloAgent, cubbies: [{ alias: "history", migrations: path.resolve("migrations/history") }] }, }, }); await p.vault.agents.connect({ agentId: "hello" }); await p.vault.scope("default").publish({ type: "user_message", context: "c-1", payload: { text: "hi" } }); const page = await p.vault.scope("default").stream("c-1").events.list({ types: ["reply"] }); expect(page.items.map((e) => e.payload)).toEqual([{ text: "you said: hi" }]); p.dispose(); }); ``` Replace an agent you do not want in the test with `mockAgent`. `p.orchestrator.jobs.list()` shows the Jobs the publishes opened, and `p.runInCubby(agentId, alias, fn)` reads an agent’s cubby. Every option is in the [testing reference](/reference/testing/#testplatform). The simulator does not verify signatures or consent, run real models, or enforce platform limits you do not inject. Those need the platform. ## Read logs Inside an agent, `console.*` is the logging API. On the platform, output is captured and shipped to the log store, not to a terminal. * `log` and `info` record at info level; `warn`, `error`, and `debug` keep their own levels. * Records are labelled with the agent, version, Job, and Task, so you can narrow to one run. **In ROC**, open the agent from **Agents** and select **Logs**. Filter by **Time range** and **Level**, and narrow to one Job or Task. **With the Vault SDK**, read one Task’s logs from a vault you own (useful in an integration test): ```ts const page = await vault.jobs.get(jobId).tasks.logs(taskId); ``` Pages carry `{ logs, total, offset, limit, hasMore }`; `limit` defaults to 100 and is capped at 1000. Logs are a short-term aid, not a record: * Logs are batched before they ship, so a line can appear a moment after the code that wrote it, and a batch that cannot be delivered is dropped. * Only the last 24 hours are queryable. Anything you must explain later belongs in a cubby row. ## Common failure signals | You see | Likely cause | Fix | | ----------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | No Job opens after you publish | The agent is not connected in that scope, or no `@OnEvent` handles the event’s type. An event no connection handles is dropped without an error. | Connect the agent; check the event type against the handlers `cef inspect` prints. | | Connect fails with `MANIFEST_INVALID` | The scope is not in the agent’s `requiredScopes`. | Connect into a declared scope, or push a version that declares it. | | Connect fails with `CUBBY_PROVISION_FAILED` | A cubby migration failed. | Fix it in a new migration file and push a new version. See [Cubby schema](/agents/building-blocks/cubby-schema/#when-migrations-run). | | Task fails with `execute_failed` | The handler threw, or the call could not complete. | Read the Task’s logs; validate the payload before using it. | | Task fails with `execute_timeout` | The handler ran past its time budget. | Do less per event. | | Task fails with `agent_threw_non_error` or `agent_invalid_result` | The handler threw a non-`Error` value, or returned a value that does not survive JSON serialization. | Throw `Error`s; return plain data. | | `model alias not declared: ''` | The `models` key does not match the alias inside the model’s `model.json`. | Declare it under the model’s own alias. See [Models](/agents/building-blocks/models/#declare-an-alias). | | Two rows for one event | The event was delivered more than once and the write is not idempotent. | Key rows on the event id; use `ON CONFLICT` or `INSERT OR IGNORE`. See [Event streams](/agents/code-agents/event-streams/#make-handlers-idempotent). | | `@OnClose` receives `idle_timeout` before the work is done | The Job had no activity for `idleTimeout`. | Raise `idleTimeout` in `cef.config.ts`. | | Task fails with `gpu_units_ceiling` | The connection reached its compute limit. | See [Spend limits](/agents/ship/connect-to-a-vault/#spend-limits). | | Logs are empty | The run is older than 24 hours, or the batch was dropped. | Record what you need later in a cubby row. | Every error code is in [Errors](/reference/errors/). Debugging a workflow run is on [Monitor runs](/agents/workflows/monitor-runs/#why-a-run-failed-parked-or-stalled). ## Related * [Next: Models](/agents/building-blocks/models/) * [Write an agent](/agents/code-agents/overview/) * [`@cef-ai/testing` reference](/reference/testing/) * [LLM agents](/agents/llm-agents/overview/) * [Errors](/reference/errors/) # Runs and events > How work starts and runs on Manykind — events in a vault scope, the Jobs and Tasks they create, how engagements and workflow runs advance, and the sandbox your code runs in. Everything on Manykind starts with an **event** landing in a vault **scope**. The platform turns events into **runs**: a **Job** per agent and stream, with one **Task** per event. A workflow run is a Job too. ## Events An event is one typed message published into a scope. It is the only way work enters the platform. | Field | Meaning | | -------------------------------------- | -------------------------------------------------------------------------------------- | | `type` | What the event is, e.g. `user_message`, `workflow.start`. | | `context` | The stream it belongs to. Events with the same `context` form one stream, and one Job. | | `payload` | The body. | | `role` | `source`, `user`, or `agent`. | | `from` | Who published it, when the vault knows: a person’s key or an agent id. | | `target` | The agent the event is for, `:`. | | `correlationId`, `metadata`, `parents` | Correlation, free-form metadata, and the ids of the events that caused this one. | | `timestamp` | Set by the vault when absent. | ### Where events come from | Source | How | | ----------- | -------------------------------------------------------------------------------------------------------------------------------- | | A person | Runs a workflow in ROC, or acts in a widget. | | Your app | `vault.scope(name).publish({ type, context, payload })` with the [Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/). | | A schedule | The vault fires the agent’s declared schedules (cron). See [Triggers](/agents/workflows/triggers/). | | A webhook | An outside system calls the workflow’s webhook URL with a key. See [Triggers](/agents/workflows/triggers/). | | A connector | A Slack or Telegram message arrives through a vault [connector](/vaults/connectors/). | | An agent | `ctx.vault.publish(type, payload, { target })` from a running agent. See [Event streams](/agents/code-agents/event-streams/). | An event reaches an agent only through an **active connection** on that scope. Without `target`, an agent’s publish goes to every agent subscribed in the scope; name the recipient to hand work to one agent. An LLM agent receives another agent’s output only when it is targeted. ## Jobs and Tasks | Noun | What it is | | -------- | ----------------------------------------------------------------------------------------------------- | | **Job** | A long-lived run, keyed by vault, agent, scope, and `context`. The record of a workflow run is a Job. | | **Task** | One event’s worth of work inside a Job. | When an event arrives, the platform finds or opens the Job for that agent and stream, and adds a Task. By default a Job runs one Task at a time, in order. * **Job states:** `queued`, `throttled`, `processing`, `active`, `completed`, `failed`, `cancelled`, `dead`. * **Task states:** `pending`, `dispatched`, `running`, `completed`, `failed`. A Task carries an `attempt` count and, on failure, an `error` with `code`, `message`, and `retryable`. * **Idle timeout.** A Job with no new activity for the agent’s `idleTimeout` (default `30m`; `"0s"` disables it) ends with reason `idle_timeout`. An agent can end its own Job with `ctx.close(reason)`. * **Initiator.** A Job records the person who started it. A Job started by a schedule or a connector has none. Read runs from outside with the Vault SDK: `vault.jobs.list()`, `vault.jobs.get(jobId).tasks.list()`, and `.tasks.logs(taskId)` for the log lines your code wrote. In ROC, open a workflow to see its runs. See [Monitor runs](/agents/workflows/monitor-runs/). ## How a code agent runs A [code agent](/agents/overview/#code-agents) is a set of engagements. For a new Job the platform **selects one engagement** (conditions, then priority, then weight) and pins it for the Job’s life. Each Task calls the pinned engagement’s `@OnEvent` handler for the event’s type. `ctx.params` is resolved per Job: manifest defaults, then deployment values, then the engagement’s own. `ctx.settings` holds the vault owner’s values from connect time. ## How a workflow runs A [workflow](/agents/overview/#workflows) is dispatched exactly like a code agent; its bundle is the platform’s workflow runner. A trigger event opens the run, and the runner walks the graph: each step’s output feeds the next, edges with conditions choose the path, and steps that wait (people, agents, connector actions) resume the run when their answer arrives as an event. ## The sandbox Code agents, and the runner itself, run in a JavaScript sandbox (a V8 isolate) with **no Node.js runtime**. | Available | Not available | | ------------------------------------------------ | --------------------------- | | `fetch` | `process`, `Buffer` | | `globalThis.crypto` (`randomUUID()`, `subtle.*`) | `node:*` built-in modules | | `TextEncoder`, `TextDecoder`, typed arrays | Database and socket clients | | `console.*` (captured as task logs) | A file system | `cef build` refuses an agent entry file that imports a banned module: `fs`, `path`, `net`, `http`, `https`, `crypto`, `buffer`, `stream`, `os`, `process`, `child_process`, and other Node built-ins (with or without the `node:` prefix), plus packages such as `pg`, `redis`, `ws`, `axios`, and `node-fetch`. The build checks the entry file only; a Node import deeper in your dependency graph fails at runtime instead. Use Web globals: `globalThis.crypto.randomUUID()` for ids, `fetch` for HTTP, `Uint8Array` and `TextEncoder` for bytes. Three rules follow: * **Nothing in memory survives.** The same isolate may not handle the next event. Keep state in a [cubby](/agents/building-blocks/cubbies/) or the [Memory Bank](/vaults/memory-bank/), and read it at the start of each handler. * **A handler can run again.** A failed Task is retried. Make every write idempotent: key rows with `PRIMARY KEY` or `UNIQUE` and use upserts. * **No secrets in the bundle.** Your bundle is published to content-addressed storage and is readable by whoever can resolve it. To reach an outside system with credentials, use a vault [connector](/vaults/connectors/), whose secrets are sealed in the vault. ## Limits | Limit | Value | When you hit it | | ----------------------------- | ----------- | --------------------------------------------------------------------------------- | | Publish request body | 1 MiB | `413 PAYLOAD_TOO_LARGE`. Store large data as a vault object and publish its path. | | Events per publish call | 100 | `413 PAYLOAD_TOO_LARGE`. Publish in batches of 100 or fewer. | | Event retention per stream | 7 days | Older events are dropped. Keep what you need later in a cubby or the Memory Bank. | | Finished Jobs and their Tasks | Kept 7 days | Older runs are deleted. | ## Related * [Next: Workflows](/agents/workflows/overview/) * [Agents and workflows](/agents/overview/) * [Event streams](/agents/code-agents/event-streams/) * [Triggers](/agents/workflows/triggers/) * [Monitor runs](/agents/workflows/monitor-runs/) # Bring your own (A2A) > Register an A2A agent you already run elsewhere, from its Agent Card with cef push --kind external or Import agent in ROC, and learn how the platform calls it, maps its task states, and turns its answers into vault events. Bring an LLM agent you already run: any [A2A](https://a2a-protocol.org) agent on your own infrastructure, in any language or framework. The platform stores only the URL of its Agent Card and reads the card each time it calls the agent. Apart from that, it behaves like an [LLM agent created in ROC](/agents/llm-agents/overview/): it appears in ROC, a vault owner connects it, a workflow’s Agent step can call it, and its answers land in the vault as events. Use one when the agent already exists as a service, needs runtimes or dependencies the platform does not offer, or belongs to another team. Create an [LLM agent in ROC](/agents/llm-agents/overview/) when instructions and a model are enough, and write a [code agent](/agents/code-agents/overview/) when the logic is yours to run on the platform. The platform’s own name for an agent hosted elsewhere is **external**: the CLI registers it with `--kind external`, and ROC’s Fleet dashboard and Sandbox switcher mark it with an **External** badge. ## Register it ```bash cef push --kind external \ --card https://agent.example.com \ --bucket \ --as-pubkey ``` `cef push` fetches the card and checks that it declares an endpoint. It then writes a manifest for the agent into your Agent Service’s bucket. No code is built or uploaded. | Flag | Default | Meaning | | --------------------------- | ------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `--card ` | required | The Agent Card URL. If you give a bare origin, `/.well-known/agent-card.json` is appended. | | `--bucket ` | required | The Agent Service’s DDC bucket. | | `--as-pubkey ` | none | The Agent Service pubkey, which forms the agent id `:`. Without it the manifest has no agent id, and you have to push again with it. | | `--alias ` | slug of the card’s `name` | The registry alias. It must not contain `:`. | | `--agent-version ` | `1.0.0` | The version to register. | | `--scope ` | `default` | The vault scopes the agent asks for. The vault owner still grants them when they connect. | | `--idle-timeout ` | `30m` | How long a conversation survives with no events. `0` is refused. | Credentials and environment flags (`--secret-phrase`, `--access-token`, `--subject-phrase`, `--env`, `--preset`, `--endpoint`, `--cdn`) work as they do for a code agent. See the [CLI reference](/reference/cli/#cef-push). Flags that only apply to a code agent (`--agent`, `--out`, `--vault`, `--vault-scope`, `--vault-api`, `--vault-token`) are refused with `--kind external`, and `--card` is refused without it. You can also register one in ROC: open **Agents**, click **Import agent**, and fill in **Card URL or Bridge name**, **Registry name**, and **Description**. In the **Agents** list its row reads **runs elsewhere**. ## What the card must declare | Card field | Used for | | ------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------ | | An interface with a JSON-RPC binding (`supportedInterfaces[].protocolBinding: "JSONRPC"` or 0.3’s `interfaces[].transport`), or a top-level `url` | Where the platform sends calls. A card that declares no endpoint is refused at push. | | `name`, `description` | The default alias and the listing shown in ROC. | | `securitySchemes` / `security` | Decides which credential the platform sends. See [Authentication](#authentication). | The card is read again every time the platform calls the agent, so you can change endpoints and skills without pushing again. ## How the platform calls it * **Transport:** JSON-RPC 2.0 over HTTP POST to the card’s endpoint. The methods are `SendMessage` and `GetTask`. If your server answers `-32601`, the platform retries once with `message/send` and `tasks/get`. * **No streaming:** the platform sends the message with `returnImmediately: true` and then polls `GetTask` until the task reaches a final state. A turn times out after 15 minutes. * **Message:** the first non-empty value among the event payload’s `text`, `message`, `prompt`, `question`, and `body` becomes a text part. The rest of the payload becomes a data part. * **Metadata:** each message carries `vaultId`, `jobId`, `taskId`, `scope`, `context`, `engagement`, `eventType`, and `onBehalfOf`. `onBehalfOf` is the public key of the person whose event started the turn; it is empty for schedules and for events from other agents. ### Context and Jobs One **Job** is one A2A `contextId`. The first turn sends no `contextId`. The platform stores the one your agent returns and sends it on every later turn in that Job. Each turn is a new A2A task. The Job ends after `--idle-timeout` with no events. The next event then starts a new Job with a new context, so your agent sees it as a new conversation. Set the timeout to cover the human gaps in your flow: a person reading, or a reply to a question your agent asked. ### State mapping | Your task state | What happens | | ----------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `TASK_STATE_COMPLETED` | The turn succeeds and the platform publishes `agent.answered`. | | `TASK_STATE_INPUT_REQUIRED`, `TASK_STATE_AUTH_REQUIRED` | The turn ends and `agent.answered` is published with `payload.state` set to that state. The person’s reply arrives as the next event, which starts a new task in the same context. | | `TASK_STATE_FAILED`, `TASK_STATE_CANCELED`, `TASK_STATE_REJECTED` | The turn fails and is retried. When retries run out, the platform publishes `agent.answered` with `payload.error` and `payload.code`. | | `TASK_STATE_SUBMITTED`, `TASK_STATE_WORKING` | The platform keeps polling. On timeout, the retry rejoins the same remote task instead of sending the message again. | A reply that is a bare Message rather than a Task counts as completed. ## How answers reach the vault The platform publishes your answer into the vault on your agent’s behalf: | Event | When | Payload | | ---------------- | ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `agent.answered` | After every turn, including failures | `agentId`, `agent` (alias), `text`, `a2aTaskId`, `state`, `jobId`, `taskId`; plus `data` (structured parts), `usage`, `conversation`, and `actions` when present; `error` and `code` on failure | | `agent.progress` | While a turn runs | Progress steps | The answer’s `parents` is the event that asked. If the asking event carried `metadata.correlation`, the answer echoes it back unchanged. This is how a [workflow Agent step](/agents/workflows/steps/#agent) matches your answer to the step that asked. ### Answering a workflow step A workflow agent step publishes an event targeted at your agent (`target: :`). Events reach an agent registered this way only when they are targeted at it; untargeted events in the scope are not delivered. Your reply becomes `agent.answered`, and the step reads it. Return the fields the step needs in a data part, which arrives as `payload.data`. ## Authentication If your card lists `oauth2`, `openIdConnect`, or an `http` `bearer` scheme with `bearerFormat: "JWT"` as a required security scheme, the platform sends `Authorization: Bearer `. The execution token is a JWT signed by the Agent Service’s key. It names the vault, scope, agent, Job, and task, and the person the turn is on behalf of. Verify it before you act. If the platform cannot mint the token, the call fails; it never falls back to an unauthenticated call. ## Spend ceilings A vault owner can cap an agent registered this way with an `a2aTokens` ceiling on the connection. The platform counts the token usage your agent reports and refuses new turns once the count passes the ceiling. See [Connect to a vault](/agents/ship/connect-to-a-vault/#spend-limits). ## Related * [Next: Write an agent](/agents/code-agents/overview/) * [LLM agents](/agents/llm-agents/overview/) * [Steps: Agent](/agents/workflows/steps/#agent) * [Push and deploy](/agents/ship/push-and-deploy/) * [CLI reference](/reference/cli/#cef-push) # LLM agents > Create an LLM agent in ROC from instructions, a model, tools, connections, and scopes; use it as a workflow step; and shape its answer so later steps can read its fields. This group covers LLM agents: agents defined by instructions and a model rather than by code. Create one in ROC, or bring an A2A agent you already run elsewhere. An **LLM agent** is a language model with a job description. You give it instructions, a model, the tools and connections it may use, and the vault scopes it declares; it works the request through and answers. It is a peer of a [code agent](/agents/code-agents/overview/): both are agents in your [Agent Service](/get-started/agent-service/), and a [workflow](/agents/workflows/overview/) uses either as an Agent step. ## What defines one | Field | What it sets | | ---------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Name** | What you call it. Spaces and capitals are fine. Its alias, which workflow steps refer to, is fixed once the agent is created. | | **What it does** | One sentence. The agent picker shows it. | | **Instructions** | What it should do, and, if it should answer JSON, the shape. | | **Model** | How capable a model it needs. Leave it on the deployment’s default, or pick one; a cheaper model suits mechanical work. | | **Thinking** | **On** (the deployment’s own behaviour) or **Off** (answer without thinking first). Leave it on when the agent must decide what to do next; turn it off when the step hands it what it needs and asks for a specific answer, which is several times faster. **Off** also applies to any agent it calls mid-turn. | | **Scopes** | The vault scopes it declares. A workflow can use it as a step only if it declares that workflow’s scope. Set when you create the agent. | | **Tools** | Granted to the agent directly: `Bash`, `Read`, `Grep`, `Glob`, `Edit`, `Write`, `WebFetch`, `WebSearch`. None by default, which is right for most workflow steps: the agent works from what the step hands it and from its connections. | | **Connections** | Accounts already authorised on the Agent Bridge. Choosing one lets the agent use it; the Bridge keeps the credential and the agent never sees it. | ## Create one In the Workflow Builder: 1. Select an **Agent** step and choose **Choose an agent…**. 2. Choose **New agent**. 3. Fill the form, then choose **Create agent**. The step now uses the new agent. ![The New agent form with Parameters and Answer columns](/shots/new-llm-agent.png) You can also open **Agents** in the side navigation and choose **New agent**; the form is the same. **Create agent** does two things, and the form names whichever is running: 1. **Creates it on the Agent Bridge**, the platform’s host for LLM agents, and publishes it there. 2. **Registers it with your Agent Service**: a manifest in the service’s bucket that points at the A2A card the Bridge serves for it. The platform calls an LLM agent over A2A, the same way it calls one you [bring yourself](/agents/llm-agents/bring-your-own/). In the **Agents** list its row reads **runs elsewhere**, and the Builder’s agent picker marks it **EXT**. ## Edit one On an Agent step that uses it, choose **Open**. The same form opens with the agent’s current definition; change it and choose **Save agent**. Every workflow that uses the agent picks the change up on its next turn. **Change** on the step picks a different agent instead. Scopes are not editable after creation. If the Agent Bridge mounts its agents read-only, the form shows the definition but cannot save it. ## Use it in a workflow An LLM agent is used through an [Agent step](/agents/workflows/steps/#agent). The step sends it a request (**What to ask it**), parks the run, and continues when the agent answers. Before a workflow runs on a vault, every agent it calls is connected to that vault too; deploying and running from ROC connects them with the workflow. ## Its answer An LLM agent answers in text. The runner turns the answer into fields on the carried item: * The whole answer is kept as `text` and as `output`. * If the answer contains a JSON object, bare, fenced in ` ```json `, or wrapped in prose, its fields are added too. So a later step or branch can read a field such as `$json.score` only if the agent answered JSON with a `score` key, and it answers JSON only if something asked it to. Ask in one of two places: * **In the instructions.** The form’s **Answer** column reads the instructions as you type and lists the fields they ask for. It checks nothing: it is the instructions read back. * **On the step**, with **Answer it should give** (a JSON example such as `{"score": 8, "recommendation": "advance"}`). With **Ask the agent for this shape** on (the default), the shape is appended to the request, and an answer missing any of its keys fails the step, naming the missing keys. Declare the shape on every step whose output a branch or a later step depends on. An agent asked for prose and wired to a branch stalls the run one step later, after doing its job well. ```text You screen candidates against our hiring policy. Answer as JSON and nothing else: {"score": 8, "recommendation": "advance", "why": "…"} ``` ## Related * [Next: Bring your own (A2A)](/agents/llm-agents/bring-your-own/) * [Steps: Agent](/agents/workflows/steps/#agent) * [Code agents](/agents/code-agents/overview/) * [Quickstart](/get-started/quickstart/) # Agents and workflows > What runs on Manykind — workflows, LLM agents, and code agents — how they relate, how a workflow uses the other two as steps, and how to choose. This group covers what your Agent Service runs: how runs happen, the three kinds of agent (workflows, LLM agents, code agents), the building blocks they use, and how you ship them to a customer’s vault. Your [Agent Service](/get-started/agent-service/) is your account on Manykind, like a cloud account: it holds your agents and workflows, the cubby schemas, widgets, and datasets they use, and the members who work on them. Everything in it that runs on a vault’s data is an **agent**, and there are three kinds, side by side: | Kind | What it is | Where you build it | | -------------- | ------------------------------------------------------------------------------------------- | ------------------------------------------------------- | | **Workflow** | A typed graph of steps: triggers, models, agents, people, cubbies, connector actions. | ROC Workflow Builder, or code with `defineWorkflow` | | **LLM agent** | Instructions, a model, tools, connections, and scopes; the model works the request through. | ROC (**New agent**), or any A2A agent you run elsewhere | | **Code agent** | TypeScript engagement classes that react to events. | Code, with `@cef-ai/agent-sdk` | All three share one path: you publish them to your Agent Service, deploy a version, and a vault owner [connects](/vaults/connections-and-consent/) them. **A workflow uses the other two as steps.** Its Agent step asks an LLM agent or a code agent and waits for the answer. LLM agents and code agents are peers: either can answer a step, and either can also run on its own, on events in the vault. In `cef.config.ts` the kind is `kind: "workflow" | "internal" | "external"`: `internal` is a code agent, and `external` an A2A agent you run elsewhere. It is a classification for surfaces such as ROC; dispatch treats them the same. ## Workflows A **workflow is an agent**: same publish, deploy, and connect path, same agent id, same consent. What it adds: * **A typed graph.** Steps and the edges between them, with conditions on edges and typed inputs and outputs. The platform’s workflow runner walks the graph; you write no handlers. * **A UI.** Build and edit it in the ROC Workflow Builder, run it, and watch each run step by step. * **Plug-ins as steps.** Triggers, connectors, models, cubbies, widgets, people, LLM agents, and code agents are steps you drop into the graph. | Step family | Step kinds | Covered in | | ----------- | ----------------------------------------------- | ---------------------------------------------------------------------------------- | | Start | `trigger` (event, schedule, webhook, connector) | [Triggers](/agents/workflows/triggers/) | | Think | `model`, `agent` | [Model](/agents/workflows/steps/#model), [Agent](/agents/workflows/steps/#agent) | | People | `human` | [People](/agents/workflows/people/) | | Flow | `branch`, `join`, `split`, `aggregate` | [Steps](/agents/workflows/steps/), [Items](/agents/workflows/split-and-aggregate/) | | State | `cubbyQuery`, `cubbyExec` | [Cubby steps](/agents/workflows/steps/#cubby-query-and-cubby-exec) | | Memory | `remember`, `relate`, `recall` | [Memory Bank](/vaults/memory-bank/) | | Act | `action` (connector), `publish` | [Connector action](/agents/workflows/steps/#connector-action) | | Shape | `transform`, `code`, `output` | [Steps](/agents/workflows/steps/) | The graph is **data**, not compiled code. Authored in code, it is baked into the manifest as a default; the Workflow Builder writes the deployed version. A workflow can start in git and be edited in ROC, and both are the same artifact. cef.config.ts ```ts import { defineAgent, defineWorkflow } from "@cef-ai/agent-sdk/config"; export default defineAgent({ id: "hiring", version: "1.0.0", agents: [ defineAgent({ id: "scorer", version: "1.0.0", entry: "./src/scorer.ts" }), defineWorkflow({ id: "cv-review", version: "1.0.0", nodes: [ { id: "start", kind: "trigger", emit: "candidate.review" }, { id: "score", kind: "agent", use: "scorer" }, ], edges: [{ from: "start", to: "score" }], }), ], }); ``` See [Workflows overview](/agents/workflows/overview/) and [Workflows in code](/agents/workflows/author-in-code/). ## LLM agents An LLM agent is defined, not programmed: you give it instructions, a model, the tools and connections it may use, and the vault scopes it declares. Create one in the Workflow Builder (Agent step → **Choose an agent…** → **New agent**) or on the **Agents** page; ROC creates it on the Agent Bridge and registers it with your Agent Service. An A2A agent you already run elsewhere joins the same way from its Agent Card. Its answer is text; a later step reads fields from it only when the instructions or the step ask for JSON. See [LLM agents](/agents/llm-agents/overview/) and [Bring your own (A2A)](/agents/llm-agents/bring-your-own/). ## Code agents A code agent is TypeScript that reacts to events. Its building block is the **engagement**: a class marked `@Engagement` whose `@OnEvent` methods each handle one event type. src/agent.ts ```ts import { Engagement, OnEvent, type Context, type Event } from "@cef-ai/agent-sdk"; @Engagement({ id: "default", goal: "Reply to a user message" }) export default class EchoAgent { @OnEvent("user_message") async onMessage(event: Event<{ text: string }>, ctx: Context) { await ctx.vault.publish("reply", { text: `you said: ${event.payload.text}` }); } } ``` * **One engagement can own a whole flow.** Several `@OnEvent` handlers on one class cover related event types. * **Handlers are stateless.** Nothing in memory survives to the next event. State lives in a [cubby](/agents/building-blocks/cubbies/) or the [Memory Bank](/vaults/memory-bank/). * **Selection.** When an agent declares several engagements, the platform picks one per Job and pins it. `@Condition` (a CEL expression) filters, `@Priority(n)` picks the tier (lower wins), and `@Weight` splits within a tier. `@Limit` caps how often one is chosen, and `@Params` sets its parameter values. * **Everything goes through `ctx`:** `ctx.models`, `ctx.cubby`, `ctx.memory`, `ctx.vault` (publish events, read and write objects), `ctx.settings`, `ctx.params`, `ctx.self`, `ctx.close`. See [Write an agent](/agents/code-agents/overview/) and [Runs and events](/agents/how-agents-run/). ## How to choose | You want | Build | | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | | A process people can read and change: steps, approvals, connector actions, model calls | A **workflow** | | Judgement a prompt can describe: classify, review, draft, decide with tools | An **LLM agent**, used from a workflow’s Agent step | | Custom logic a prompt cannot express: parsing, loops over APIs, a state machine | A **code agent**, used from a workflow’s Agent step or on its own | | To reuse an agent you already run elsewhere, in any language | An **LLM agent** you [bring over A2A](/agents/llm-agents/bring-your-own/) | A common shape is all three: a workflow orchestrates, and its Agent steps call LLM agents and code agents for the parts that need them. ## Related * [Next: Runs and events](/agents/how-agents-run/) * [Workflows](/agents/workflows/overview/) * [LLM agents](/agents/llm-agents/overview/) * [Code agents](/agents/code-agents/overview/) * [Your Agent Service](/get-started/agent-service/) # Connect to a vault > Connect an agent or workflow to a vault so it can run on that vault's data: from ROC or with the vault SDK, reconnect to pick up new code, and set spend limits. Deploying makes a version available. Nothing runs on a vault’s data until the vault owner **connects** the agent. A connection is a signed, scoped, revocable agreement between the vault and the agent: the owner signs it with their wallet, and the vault records it with the settings they chose and the code they consented to. ## What a connection holds | Part | Meaning | | ---------- | ------------------------------------------------------------------------------------------------------------------------ | | Scopes | The vault scopes the agent may read and publish in. Each must be in the manifest’s `requiredScopes` (default `default`). | | Settings | Values for the agent’s declared `settings`, validated against its schema. | | Bundle pin | The bundle of the version the owner consented to. Dispatch runs exactly that code. | | Ceiling | Optional spend limits. See [Spend limits](#spend-limits). | The connection also provisions the agent’s cubbies in the vault. ## Connect from ROC | Where | How | | -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Workflow Builder** | **Deploy** connects the workflow, and any agent it calls, to your vault. Your wallet may ask for your passkey. | | **Sandbox** | Pick the agent and a vault; the pill reads **agent not connected** with a **Connect** button. When connected, it reads **agent connected**, with a control to disconnect. | | A widget | `WidgetRuntime.connectAgent()` connects the widget’s agent for the signed-in reader. | ## Reconnect to pick up new code A connection pins the bundle the owner consented to. Pushing and deploying a new version does **not** change the code a connected vault runs; the vault keeps running the pinned bundle until the owner reconnects. Routing, params, and the engagement list do follow the deployment; only the code is pinned. When the code has changed, ROC shows a prompt: “This agent’s code changed since you connected it (was …, now …). Reconnect to the new version?” with **Reconnect** and **Not now**. **Reconnect** signs the consent for the new bundle. LLM agents have no bundle and pin nothing. ## Connect with the vault SDK ```ts import { VaultSDK, CereWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: vaultApiUrl, garEndpoint: garUrl, // required for agents.connect signer: await CereWallet.fromMnemonic(mnemonic), // or KeypairWallet.fromSeed(seed) }); const vault = await sdk.vault.ensure(); const connection = await vault.agents.connect({ agentId: ":", scope: "default", // or scopes: ["default", "research"] settings: { apiKey: "…" }, ceiling: { gpuUnits: 500, a2aTokens: 200000 }, bundleCid: reviewedBundleCid, // the code the owner reviewed }); ``` | Input | Required | Meaning | | ------------------- | ----------- | --------------------------------------------------------------------------------------------------------- | | `agentId` | yes | `:`. The pubkey is taken from the prefix unless you pass `agentServicePubkey`. | | `scope` or `scopes` | one of them | The scopes to connect into. | | `settings` | no | Values for the agent’s settings schema. | | `ceiling` | no | Spend limits, written together with the connection. | | `bundleCid` | no | The bundle the owner consented to. Required when reconnecting to a different bundle. | `connect` signs one consent agreement covering every requested scope, submits it to the agreement registry, and creates the connection. The vault reads the manifest from the registry itself; you never send it. ### Reconnect after a new version ```ts import { BundleChangedError, ReconsentRequiredError } from "@cef-ai/vault-sdk"; try { await vault.agents.connect({ agentId, scope: "default" }); } catch (e) { if (e instanceof ReconsentRequiredError) { // Show the owner the new code (e.currentCid), then: await vault.agents.connect({ agentId, scope: "default", bundleCid: e.currentCid }); } else if (e instanceof BundleChangedError) { // A new version landed after the owner reviewed; show e.currentCid and ask again. } else throw e; } ``` | Error | Code | Meaning | | ------------------------ | -------------------- | ------------------------------------------------------------------------------------------------------------------ | | `ReconsentRequiredError` | `RECONSENT_REQUIRED` | The reconnect would run different code than the owner consented to, and no `bundleCid` was given. Nothing changed. | | `BundleChangedError` | `BUNDLE_CHANGED` | The `bundleCid` you named is no longer the agent’s current bundle. Nothing was stored. | ### Manage a connection ```ts const all = await vault.agents.list(); const one = await vault.agents.get(agentId); // status, scopes, bundle, ceiling await one.update({ apiKey: "…" }); // new settings await one.setCeiling({ gpuUnits: 1000 }); // new limits await one.disconnect(); ``` Disconnecting keeps the cubby data. ## Spend limits A connection can carry two limits. Blank or `0` means no limit. | ROC label | SDK field | Counts | | ------------------------------------- | ----------- | ----------------------------------------------------------------------------------------- | | **Compute limit — metered GPU units** | `gpuUnits` | GPU units the platform measured. Decimals allowed. | | **Token limit — agent-reported** | `a2aTokens` | Tokens the agent reports about itself. Whole numbers. A stop signal, not a measured cost. | Once the connection’s total reaches a limit, new tasks for that agent in that vault are refused. Spend is counted when a task finishes, so a limit stops the next run after it is crossed, not a run in progress. In ROC, set limits in the deployment dialog’s **Execution rules** section: pick the vault, enter the limits, and click **Save limits**. Limits belong to the connection, not to the deployment revision, so the agent must already be connected to that vault. ## Errors | Code | Cause | Fix | | ---------------------------- | ------------------------------------------------------------------------------------------ | ------------------------------------------------------------------ | | `MANIFEST_INVALID` | A requested scope is not in `requiredScopes`, or the agent id does not match the manifest. | Connect into a declared scope, or push a version that declares it. | | `MANIFEST_NOT_FOUND` | No manifest for that agent id in the registry. | Check the id and that the agent was pushed. | | `SETTINGS_SCHEMA_MISMATCH` | Settings do not satisfy the schema. | Fix the values. | | `AGENT_ALREADY_CONNECTED` | A connection already exists. | Use the existing connection. | | `GAR_MISSING`, `GAR_EXPIRED` | No valid consent agreement. | Connect again to sign a new one. | | `CUBBY_PROVISION_FAILED` | A cubby migration failed. | Fix the migration in a new version. | All codes: [Errors](/reference/errors/). ## Related * [Next: Vaults](/vaults/overview/) * [Connections and consent](/vaults/connections-and-consent/) * [Push and deploy](/agents/ship/push-and-deploy/) * [Vault SDK reference](/reference/vault-sdk/) * [Work with the vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) # Push and deploy > Ship an agent or workflow: build it, push a version to your Agent Service's bucket, and deploy it with deployment records that choose which version runs for which audience, from the CLI or ROC. This group covers getting any agent to run on a customer’s data: push and deploy a version, then have the vault owner connect it. Shipping has three steps, and they are the same for code agents and workflows: | Step | What it does | CLI | ROC | | ---------- | ------------------------------------------------------------------------------------------ | ------------ | ---------------------------------------------------------------------------- | | **Build** | Bundles the code and writes `manifest.json`. | `cef build` | The Workflow Builder builds from the canvas. | | **Push** | Uploads a version to the Agent Service’s DDC bucket, the registry. Versions are immutable. | `cef push` | Part of **Deploy** in the Workflow Builder. | | **Deploy** | Applies deployment records: which version runs, for which vaults and events. | `cef deploy` | Agent page → **Overview** → deployments; **Deploy** in the Workflow Builder. | Then a vault owner must **connect** the agent before it runs on their data. See [Connect to a vault](/agents/ship/connect-to-a-vault/). ## Before you start From ROC, open your Agent Service → **⋯** → **Settings** → **General** and copy: * **Agent-service public key**: pass it as `--as-pubkey`. * **Registry bucket ID**: pass it as `--bucket`. On the **Access** tab, generate two tokens: | Token | Used by | Environment variable | | -------------------- | ----------------------------------------------- | ---------------------- | | **DDC access token** | `cef push`, `cef widget push`, `cef cubby push` | `CEF_DDC_ACCESS_TOKEN` | | **CLI access token** | `cef deploy`, `cef publish` | `CEF_ACCESS_TOKEN` | Tokens are shown once. Members of an Agent Service get their DDC token through an invite; see [Team](/get-started/team/). ## Environments Every command takes `--env dev|stage|prod` (or `$CEF_ENV`). The default is `dev`. One flag selects the whole environment: the DDC network for push (`dev` = devnet, `stage` = testnet, `prod` = mainnet), the platform API for deploy, and the endpoints baked into widgets. Push and deploy to the same environment. ## Build ```bash cef build ``` Writes `dist//bundle.js`, `manifest.json`, and `widgets//` for every agent in `cef.config.ts`. A workflow written with `defineWorkflow` builds the same way; its graph and the workflow runner become the bundle. Build errors are listed in [Write an agent](/agents/code-agents/overview/#what-cef-build-rejects). ## Push ```bash export CEF_DDC_ACCESS_TOKEN=… cef push --env dev --bucket --as-pubkey ``` `cef push`: * uploads `bundle.js` and each widget directory, content-addressed; * writes `manifest.json` under `agents///` in the bucket and moves `latest` to it; * stamps the agent id `:` and the environment’s endpoints into the manifest and widgets; * declares any cubby the manifest names that the bucket does not have yet. A push body is capped at 64 MiB. Ship large assets as vault objects, not in the bundle. Before uploading anything, `cef push` checks the access token against the bucket’s on-chain owner. A token that the owner did not issue is refused with the owner’s address and the fix. See [Team](/get-started/team/#publishing-as-a-member). Bump `version` in `cef.config.ts` for each push you want to keep. An A2A agent you run elsewhere is registered with `--kind external`; see [Bring your own (A2A)](/agents/llm-agents/bring-your-own/). ## Deploy ```bash export CEF_ACCESS_TOKEN=… cef deploy --env dev --as-pubkey ``` `cef deploy` reads every record in `deployments/` and applies them as one set, replacing what is live. The folder is the desired state: delete a file and the next deploy removes that record. Each apply creates a revision you can roll back to in ROC. ### Deployment records One record per file; the filename is the record name (`^[a-z0-9][a-z0-9_-]{0,62}$`). `.json` and `.jsonc` are read, comments allowed. deployments/default.jsonc ```jsonc { "priority": 99, "targeting": "", // empty: the default record, matches everything "version": "latest", "weight": 1 } ``` deployments/sandbox-canary.jsonc ```jsonc { "priority": 10, "targeting": "vault.scope == 'sandbox'", "version": "0.2.0", "weight": 1, "params": { "temperature": 0.1 } } ``` | Field | Required | Rule | | ------------- | -------- | --------------------------------------------------------------------------------------------------------------------------- | | `priority` | yes | Among matching records, the lowest priority wins. | | `targeting` | yes | A CEL expression over `vault` (`id`, `scope`), `event` (`type`), and `connection.settings`. `""` marks the default record. | | `version` | yes | A pushed semver, or `"latest"`. | | `weight` | no | A positive integer, default `1`. Records in the same priority tier split traffic by weight: `90` and `10` is a 90/10 split. | | `enabled` | no | `false` excludes the record. | | `params` | no | Overrides the manifest’s param defaults. | | `engagements` | no | Per-engagement overrides: `{ "": { priority, weight, limit: { n, per }, enabled, params } }`. | | `expiresAt` | no | RFC 3339. Once past, the record is ignored. | The set must contain **exactly one** default record (empty `targeting`). A Job resolves its version when it is created and keeps it until it ends, so `"latest"` affects new Jobs only. | Flag | Default | Meaning | | ---------------------- | ---------------- | ---------------------------------------------------------------------------- | | `--version ` | — | Override `version` in every record without editing files. | | `--deployments ` | `deployments` | A folder, or one file holding `{ "deployments": [...] }` or a single record. | | `--dry-run` | — | Print the assembled set; apply nothing. | | `--note`, `--author` | git `user.email` | Recorded on the revision. | ### Rollout patterns | Goal | Records | | --------------------- | ------------------------------------------------------------------ | | Ship to everyone | One default record, `version` pinned or `"latest"`. | | Canary to an audience | Add a record with targeting and a lower priority than the default. | | Percentage rollout | Two records in one priority tier with weights such as `90` / `10`. | | Temporary experiment | Add `expiresAt`. | ### In ROC Open **Agents**, pick the agent, and use the deployments manager on **Overview**: **New deployment**, **Edit deployment**, **Delete deployment**, and **History** with **Roll back**. The form has **Route** (Name, Audience — targeting (CEL), Version, Priority, Weight), **Parameters**, **Engagements**, and **Execution rules**. Saving creates a new revision. ## Workflows A workflow is an agent, so the same steps apply. * **In code:** `defineWorkflow` in `cef.config.ts`, then `cef build`, `cef push`, `cef deploy`. See [Workflows in code](/agents/workflows/author-in-code/). * **In the Workflow Builder:** **Deploy {version}** writes the canvas as a new version, makes it live, and connects it to your vault. The tooltip reads “Write this canvas as a new version and make it live”. **Run** stays disabled until the workflow is deployed. ## Listing `cef publish` sends the agent’s card to the marketplace listing, and the agent page’s **Pricing** tab publishes its pricing. Neither is needed to deploy, connect, or run the agent. ## Related * [Next: Connect to a vault](/agents/ship/connect-to-a-vault/) * [Write an agent](/agents/code-agents/overview/) * [Team](/get-started/team/) * [CLI reference](/reference/cli/) # Workflows in code > Declare a workflow with defineWorkflow: project layout, nodes and edges, typed ids, models and cubbies, what cef build produces, and what round-trips with the Workflow Builder. `defineWorkflow` declares a workflow in TypeScript. It returns an ordinary agent config, so a workflow is built, pushed, versioned, deployed, and connected exactly like a code agent. What makes it a workflow is that its code is the platform’s workflow runner and its behavior is the graph you declare. Use code when you want the graph reviewed in pull requests, tested in CI, and shared between environments. The graph is the same document the Builder produces. ## Project layout ```plaintext ticket-triage/ ├── cef.config.ts ← export default defineWorkflow({ … }) ├── src/ │ └── prompts.ts ← prompts, SQL, and expressions as constants ├── cubbies/ │ └── triage/ │ └── 001-init.sql ← one directory per cubby, numbered migrations ├── deployments/ │ └── default.jsonc ← which version is live ├── test/ │ └── triage.test.ts ├── package.json └── tsconfig.json ``` Dependencies: ```bash pnpm add @cef-ai/agent-sdk@^5.8.0 pnpm add -D @cef-ai/cli@^2.8.0 @cef-ai/testing@^3.3.5 typescript vitest ``` The workflow runner ships inside `@cef-ai/agent-sdk`; you do not install or point at it. ## A complete workflow cef.config.ts ```ts import { defineWorkflow } from "@cef-ai/agent-sdk/config"; import { CLASSIFY, SCHEMA } from "./src/prompts.js"; const X = (col: number) => col * 320; export default defineWorkflow({ id: "ticket-triage", version: "0.2.0", goal: "Classify a support ticket, escalate urgent ones, and record the result", models: { llm: "https://cdn.example.com/1234/models/my-llm/1.0.0/model.json", }, cubbies: [{ alias: "triage", migrations: "./cubbies/triage" }], nodes: [ { id: "ticket", kind: "trigger", label: "New ticket", position: { x: X(0), y: 0 }, params: { sample: JSON.stringify({ ticketId: "T-1", text: "I was charged twice" }) }, }, { id: "classify", kind: "model", label: "Classify", position: { x: X(1), y: 0 }, params: { alias: "llm", input: { messages: [{ role: "user", content: CLASSIFY }], max_tokens: 128, response_format: { type: "json_schema", schema: SCHEMA }, }, into: "classification", }, }, { id: "parse", kind: "transform", label: "Read the classification", position: { x: X(2), y: 0 }, params: { expr: "({ ...item, ...JSON.parse(item.classification.text) })" }, }, { id: "route", kind: "branch", label: "Urgent?", position: { x: X(3), y: 0 } }, { id: "escalate", kind: "human", label: "Escalate", question: "=Urgent {{ $json.category }} ticket {{ $json.ticketId }}. Take it?", position: { x: X(4), y: -160 }, }, { id: "save", kind: "cubbyExec", label: "Record the result", position: { x: X(5), y: 0 }, params: { alias: "triage", sql: "INSERT OR REPLACE INTO triage_tickets (run_id, ticket_id, category, priority) " + "VALUES ('{{ $runId }}', ?, ?, ?)", args: ["={{ $json.ticketId }}", "={{ $json.category }}", "={{ $json.priority }}"], }, }, { id: "result", kind: "output", label: "Result", position: { x: X(6), y: 0 }, params: { result: { category: "={{ $json.category }}", priority: "={{ $json.priority }}" }, resultTypes: { category: "string", priority: "string" }, }, }, ], edges: [ { from: "ticket", to: "classify" }, { from: "classify", to: "parse" }, { from: "parse", to: "route" }, { from: "route", to: "escalate", when: { field: "priority", op: "eq", value: "urgent" } }, { from: "route", to: "save" }, { from: "escalate", to: "save" }, { from: "save", to: "result" }, ], }); ``` src/prompts.ts ```ts export const CLASSIFY = "=You triage support tickets. Classify the ticket below.\n\n{{ $json.text }}"; export const SCHEMA = { type: "object", properties: { category: { type: "string", enum: ["billing", "bug", "question"] }, priority: { type: "string", enum: ["low", "normal", "urgent"] }, }, required: ["category", "priority"], }; ``` ## `defineWorkflow` fields | Field | Required | What it is | | ------------- | -------- | ----------------------------------------------------------------------------------------------------------------------- | | `id` | yes | The workflow’s id. The agent id is `:`. | | `version` | yes | Semver. Bump it for every push. | | `nodes` | yes | The steps. See [Steps](/agents/workflows/steps/). | | `edges` | yes | `{ from, to, when?, loop? }`. | | `goal` | | One line on what the workflow does. Also the card description when `card` is omitted. | | `models` | | Alias → `model.json` URL for every model step. | | `cubbies` | | `{ alias, migrations }` for cubbies the workflow declares. See [Cubbies](/agents/building-blocks/cubbies/#declare-one). | | `schedules` | | Cron triggers. See [Triggers](/agents/workflows/triggers/#schedule). | | `card` | | `{ name, description }`. Defaults to the id and goal. | | `idleTimeout` | | How long a run’s Job stays open with nothing happening. Default `"30m"`. | | `runner` | | Path to a different runner build. Leave it unset. | ### Nodes Each node is `{ id, kind, label?, position?, params?, use?, emit?, question? }`. `use` and `emit` are for agent steps, `question` for human steps. Give every node a `position`. The runner ignores it, but the Builder draws a code-authored workflow with your positions, left to right if you lay them out that way. ### Edges are typed against node ids `defineWorkflow` infers the node ids as literal types, so an edge naming a step that does not exist is a compile error: ```plaintext error TS2820: Type '"reslt"' is not assignable to type '"ticket" | "classify" | "result"'. Did you mean '"result"'? ``` Declare `nodes` inline in the call (or `as const`) to keep this check. ### Keep long text out of the graph Prompts, SQL, and expressions read better as constants in `src/` and imported, as above. The graph keeps their values; nothing else changes. ### Scope A workflow is connected on the scopes it requires: `default` unless you say otherwise. To run in another scope, override `requiredScopes`: ```ts const workflow = defineWorkflow({ /* … */ }); export default { ...workflow, requiredScopes: ["support"] }; ``` ## What `cef build` checks and produces `cef build` loads `cef.config.ts`, runs the same validation the runner runs before a run, and refuses the build on any fatal problem, with every problem listed: * duplicate step ids, unknown kinds, edges to unknown steps, no trigger; * a step nothing leads to; * an agent step without `use`; * a model step whose alias is not in `models`, or a `models` value that is not a `...//models///model.json` URL; * a schedule whose `id` is not a trigger with `mode: "schedule"`; * templates without the leading `=`, cycles without a loop edge, invalid human-step params, invalid regions. It writes `dist//`: | File | What it is | | --------------- | -------------------------------------------------------------------------------------------------------------------------------- | | `bundle.js` | The workflow runner, with your model references baked in. | | `manifest.json` | The agent manifest. The graph is in `params.graph` as its default value. The runner’s own `runs` cubby is declared beside yours. | Then ship it like any agent: ```bash cef push --bucket --as-pubkey cef cubby push --bucket # when you added or changed a migration cef deploy ``` See [Push and deploy](/agents/ship/push-and-deploy/). After the first deploy, open the workflow in ROC and press **Run** on its **Executions** tab: ROC connects it to the vault first and asks you to approve the connection. See [Connect to a vault](/agents/ship/connect-to-a-vault/). ## Code and the Builder The graph is one document in two editors. **From the Builder to code.** On the canvas side rail, **Export as code** produces a `cef.config.ts` with `defineWorkflow`: id, version, goal, card, every node with its params and position (snapped to a 16 px grid), every edge, and schedules. Save it beside a `package.json` and push it with `cef push`. **From code to the Builder.** A workflow pushed with `cef push` appears in the Agent Service’s **Workflows** list and opens in the Builder read-only, marked **from the repo · read-only**: you edit it in the repo and push. Its steps are drawn at your `position` values. You can run it, read its runs, and run evaluations from ROC. What travels in each direction: | | Builder → code (Export) | Code → Builder | | --------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- | | Steps, params, edges, conditions, loops | yes | yes | | Positions and labels | yes | yes | | Schedules | yes | yes | | Builder-only step forms (Filter, Edit fields in Fields mode, Memory, Cubby) | exported as the kinds they compile to (`branch`, `transform`, `remember`/`recall`/`relate`, `cubbyQuery`/`cubbyExec`) | drawn as those kinds | | Webhook URLs and keys | no: installed by Deploy in the Builder | no | | Connection access for connector steps | no: written by Deploy in the Builder | grant it through the vault API, see [Connectors](/vaults/connectors/#access-is-part-of-the-deploy) | | Pinned widget | no: part of the canvas | no | | Model declarations (`models`) | no | yes | Versions differ too: the Builder’s Deploy picks the next patch version itself, while in code you set `version`. ## Test it Run the graph against the real runner in memory with `@cef-ai/testing`, mocking the agents it calls. See [Test a workflow](/agents/workflows/test/). ## Related * [Next: Test a workflow](/agents/workflows/test/) * [Steps](/agents/workflows/steps/) * [Push and deploy](/agents/ship/push-and-deploy/) * [CLI reference](/reference/cli/) * [Agent SDK reference](/reference/agent-sdk/) # Best practices > Rules for workflows that stay testable and safe to run twice: graph shape, step and payload budgets, cubby migrations, prompts and data, model output, idempotency, versioning, and fixtures. Each rule says what to do, why, and how to check it. Most checks are a line in your local tests; see [Test a workflow](/agents/workflows/test/) for the harness. ## Graph shape **Prefer a linear graph.** Branch only where the work really differs. *Why:* every branch doubles the paths you have to test and the runs you have to read. *Check:* one local test per path; if you cannot name each path’s test, the graph has too many. **Give every step a `position`.** *Why:* the Workflow Builder places a step that has one exactly there; without it the builder derives a layout, and a graph authored in code becomes hard to read next to its runs. *Check:* `for (const n of graph.nodes) expect(n.position).toBeDefined()`. **Gate branches on fields that exist before the branch, not on fields only a `transform` expression creates.** *Why:* a transform’s output is whatever its expression returns; the builder cannot see those fields, so it cannot show or check the gate, and a typo in the field name sends every run down the default edge. *Check:* a local test per gate value, asserting which step ran. **Join fan-out explicitly.** When several branches must all finish before a step, put a `join` step there. *Why:* a step with several inbound edges cannot tell branches that arrive together from alternative arms of a branch; only a `join` waits for all of them. *Check:* a test that the step after the join runs once, with every branch’s output merged. **Bound every loop.** A loop edge declares `max` (1 to 100) and the `counter` field; set `exhausted` to the step that handles running out. *Why:* without `exhausted`, a spent loop fails the run. *Check:* a test whose mock never succeeds, asserting the run reaches the `exhausted` step. ## Step budgets **Set a budget for steps, model calls, and agent calls per run, and assert it.** *Why:* each extra call adds latency and cost to every run, and a refactor that adds one is easy to miss in review. *Check:* `expect(graph.nodes.length).toBeLessThanOrEqual(N)`; count `createModelMock().calls` and each mock agent’s calls per run. In experiments, set the dataset’s **Max steps** and **Max tokens** so drift shows as limit breaches. ## Payload budgets **Keep the carried item small.** Recommended: at most 256 KB at any step. Pass references (a cubby row id, an object path) instead of embedding documents. *Why:* every step’s input and output is carried forward and recorded for the run. Hard limits apply further along: a published event body over 1 MiB is refused, and a step that waits (an agent, a person) cannot hold more than 1 MiB of input. See [Limits](/reference/limits/). *Check:* in a test with the largest realistic input, assert `Buffer.byteLength(JSON.stringify(x))` on the start payload, the Result, and each published payload. **Declare the payload of every publish step.** Set `params.payload` to an object whose keys are the fields to send. *Why:* without it the step broadcasts the whole carried item: every intermediate result, the raw input, and anything a model returned. *Check:* `expect(Object.keys(published).sort()).toEqual([...])` for each published event. **Cap what you store.** Truncate free text before it goes into a cubby column, and assert the cap. *Why:* a model answer or a pasted document can be arbitrarily long, and a widget that reads the row pays for it. *Check:* a test with oversized input asserting the stored column length. ## Cubby migrations **Never edit a migration that has shipped.** Add a new numbered file instead. *Why:* the platform tracks applied migrations by number only. A cubby that already applied `002` never runs your edited `002`, so the vaults that matter keep the old schema while fresh test vaults get the new one. *Check:* in review, the diff of `cubbies/` only adds files. **Create new tables only in new migration files, and prefix their names** with the workflow or feature (`triage_results`, not `results`). *Why:* in each vault, a cubby alias is one database shared by every agent and workflow in the Agent Service that names it. An unprefixed name collides with another workflow’s table, or with one a step created at run time. *Check:* `grep -h "CREATE TABLE" cubbies/**/*.sql` shows only prefixed names. **Bind values; never put run data in SQL text.** Use `?` in `sql` and the values in `args`. `{{ $runId }}` is the only template allowed in SQL. *Why:* run data comes from outside; interpolating it is SQL injection. The runner fails a Cubby step whose SQL contains `{{ $json.… }}`. *Check:* the runner enforces it; a local test of the step proves the binding works. **Key every write on the run.** Use `run_id` as (part of) the primary key and `INSERT OR IGNORE` or an upsert. *Why:* delivery is at-least-once; a retried step must not write a second row. *Check:* the “same start twice” test in [Test a workflow](/agents/workflows/test/#the-test). ## Prompts and data **Keep business-specific text out of prompts.** Put rules, labels, thresholds, and per-customer wording in data: a cubby table, the deployment’s params, or the start payload. *Why:* a prompt is code; changing it means a new version and a new experiment. Data changes without a deploy, and the same workflow serves every customer. *Check:* grep your prompts for customer names, product names, and numbers; each one is a candidate for data. **Use synthetic data in prompts, fixtures, and examples.** Never real customer content. ## Model output **Treat every model answer as untrusted input.** Validate it in a `transform` or `code` step right after the model step. *Why:* a model can return a label you did not list, a string where you wanted a number, or nothing. **Downgrade invalid output deterministically, or park the run for a person.** Map an unknown label to a fallback (`"other"`, confidence `0`), or route the run to a human step. Do not loop a model until it complies. *Why:* a deterministic fallback keeps runs reproducible and the Result well-typed; a repair loop multiplies cost and still fails sometimes. *Check:* a local test that feeds garbage through `createModelMock()` and asserts the Result. **Declare a typed Result.** `resultTypes` fails the run at the Result step when a field changes type. See [Results](/agents/workflows/evaluate/results/). ## Idempotency **Expect every event to arrive more than once.** Delivery is at-least-once. | Situation | What the platform does | What you do | | ------------------------------------------------------ | ------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | | A `workflow.start` repeated in the same context | One run per context: the repeat starts nothing | Use a stable context per logical request | | A webhook call retried with the same `Idempotency-Key` | Same event id and context, so the same run; a waiting retry gets the recorded result once the run ended | Always send an `Idempotency-Key` (at most 255 bytes) from callers that retry | | A webhook call without `Idempotency-Key` | Every call starts a new run | | | A step retried after a transient failure | The step runs again | Key cubby writes on `run_id`; make external actions idempotent | | **Re-run** in ROC | Starts a new run from the top: real agents, real tokens, a new Job | Use it to reproduce, not to “resume” | *Check:* publish the same start twice in a local test and assert one row and one model call. ## Versioning **Bump the version on every change you push, and evaluate before you switch the live version.** *Why:* an experiment pins a version that is not live for its own runs only, so you can judge a candidate on real models without exposing it. *Check:* run the dataset with `cef eval run --workflow-version ` and read the verdict against the baseline. **Move the baseline when the live version changes.** Make the live version’s experiment the baseline, so the next candidate is judged against what users actually get. **Update the dataset in the same change as the Result.** A field you rename or retype scores n/a until the cases follow. ## Fixtures **Build fixtures from small factory functions with synthetic values.** One base case, and overrides per test (`triggerFor({ category: "bug" })`). *Why:* each test then states only what makes it different, and a reviewer can see why it exists. **Make each negative fixture fail for one reason.** A fixture that is wrong in two ways cannot tell you which check caught it. *Check:* remove the guard the test is for and watch the test go red. A test that stays green without its guard tests nothing. ## Related * [Next: LLM agents](/agents/llm-agents/overview/) * [Test a workflow](/agents/workflows/test/) * [Limits](/reference/limits/) * [Monitor runs](/agents/workflows/monitor-runs/#why-a-run-failed-parked-or-stalled) * [Cubby schema and migrations](/agents/building-blocks/cubby-schema/) # Datasets > Create a dataset of cases for a workflow in ROC or with cef eval, capture cases from real runs, and understand how cases are revised, versioned, and stored. A dataset is the set of cases one workflow is judged on. Each case is a start payload and the Result that payload should produce. ## Create a dataset **In ROC.** Open the workflow, select **Evaluations**, then **New dataset** (or **Create empty dataset** on a workflow with none). Give it a **Name**, an optional **Description**, and optional limits: **Max duration (s)**, **Max tokens**, **Max steps**, **Max cost ($)**. The sidebar **Evaluations** page has the same **New dataset** button with a **Workflow** picker. **With the CLI.** Run it in the workflow’s project folder; `--workflow` defaults to the workflow your `cef.config.ts` (or `dist/`) declares. ```sh cef eval create triage --description "support tickets by category" --max-duration-ms 30000 ``` | Flag | Meaning | | ----------------------- | -------------------------------- | | `--description ` | What the dataset covers | | `--max-duration-ms ` | Warn when a run takes longer | | `--max-tokens ` | Warn when a run uses more tokens | | `--max-steps ` | Warn when a run takes more steps | | `--max-cost ` | Warn when a run costs more | Dataset names match `^[a-z0-9][a-z0-9-]{0,62}$` and are unique per workflow. Every `cef eval` command needs [store credentials](/agents/workflows/evaluate/experiments/#credentials). ![The Dataset view: cases with their inputs and expected results](/shots/evals-dataset.png) ## Case shape ```json { "id": "double-charge", "input": { "ticketId": "T-1001", "text": "I was charged twice for March" }, "expected": { "category": "billing", "confidence": ">= 0.8", "summary": { "$exists": true } }, "limits": { "maxDurationMs": 20000 }, "tags": ["billing"], "notes": "Seen in production twice in one week." } ``` | Field | Required | Meaning | | ---------- | -------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `id` | yes | `[A-Za-z0-9._-]`, 1–80 characters. Stays taken after the case is deleted. | | `input` | yes | The `workflow.start` payload, as an object. | | `expected` | no | Keyed by Result field; dot paths reach nested fields. A case with no expectations runs but is not scored. See [Results](/agents/workflows/evaluate/results/#matching). | | `limits` | no | `maxDurationMs`, `maxTokens`, `maxSteps`, `maxCost`. Each overrides the dataset’s own. | | `tags` | no | Free labels for filtering. | | `notes` | no | Free text. | | `source` | no | Set when the case was captured from a run: `{ context, workflowVersion?, capturedAt }`. | `input` is exactly what an experiment publishes as `workflow.start`, so it must contain everything your trigger step reads. ## Add cases by hand **In ROC**, on the **Dataset** view select **Add case**. Fill in **Case id**, **Input (workflow\.start payload)**, the **Expected result**, **Tags**, and **Notes**. Editing a case and choosing **Save revision** writes a new revision; the old one stays. **In your repo**, pull the dataset, edit the JSON, and push it back: ```sh cef eval pull triage # writes datasets/triage/ $EDITOR datasets/triage/cases/double-charge.json cef eval push triage ``` `pull` writes this layout (change it with `--dir`): ```plaintext datasets/triage/header.json the dataset header (without the baseline) datasets/triage/cases/.json each live case at its latest revision datasets/triage/.versions.json what the last pull wrote: revisions, content hashes, versions ``` A case file must be named `.json`. Commit the folder with your workflow so cases are reviewed in the same pull request as the change they test. | Command | Flag | Meaning | | ------- | -------------- | ------------------------------------------------------------------------------------------ | | `pull` | `--force` | Overwrite case files edited locally since the last pull or push | | `push` | `--cut [note]` | Cut a dataset version after pushing | | `push` | `--force` | Write over cases that changed in the bucket since the last pull | | `push` | `--prune` | Delete cases the last pull wrote that you have since removed locally | | `push` | `--restore` | Undelete or unarchive cases that are deleted or archived in the bucket but present locally | `push` writes a new revision only for a case whose content changed. Without `--force`, it refuses a case that someone edited in the bucket since your last pull, and `pull` refuses to overwrite a case file you edited locally. Neither removes a case file the last pull did not write. Both refuse a dataset folder whose `header.json` names another workflow. ## Add cases from real runs A run that went wrong in production is the best case you will get. Capture it instead of retyping it. **In ROC:** * On a run in **Executions**, select the **Add to dataset** button next to **Re-run**. Pick a **Dataset** (or type a **New dataset name**) and confirm the **Case id**. * On the **Dataset** view, select **Add from executions** and pick from the workflow’s recent runs. **In code**, `caseFromRun` from [`@cef-ai/eval`](/reference/eval/#cases-from-runs) builds the same case: ```ts import { addCase, caseFromRun, uniqueCaseId, listCaseIds } from "@cef-ai/eval"; const c = caseFromRun({ input: startPayload, // the run's workflow.start payload output: completedPayload.output, // workflow.completed output, if the run completed context: "support-4821", workflowVersion: "1.3.0", tags: ["from-production"], }); c.id = uniqueCaseId(c.id, await listCaseIds(store, "ticket-triage", "triage")); await addCase(store, "ticket-triage", "triage", c); ``` What a captured case expects: | The run’s output field is | The case expects | | ----------------------------------------------------- | ------------------------------------------- | | A label (60 characters or fewer and 8 words or fewer) | That exact value | | Free text (longer than that) | `{ "$exists": true }`: the field is present | | Not an object, or the run did not complete | No expectations: state them yourself | An inline `graph` in the start payload is dropped from the input. A captured case expects what the workflow **did**, not what it **should have done**. Review every captured expectation before you run an experiment, and fix the ones that captured the bug. ## Archive and delete | Action | Effect | | ------- | --------------------------------------------------------------------------------------------------------- | | Archive | Hides the case from new dataset versions. Its bytes are untouched; versions that pin it still load. | | Delete | Removes the case from the working set. Its revisions stay (versions may pin them) and its id stays taken. | Both are reversible. `cef eval push --restore` brings a case back if it is still in your repo folder. ## Versions You rarely cut a version by hand. When an experiment starts without a dataset version, it compares the working set (every live case at its latest revision) with the latest version: * unchanged: the experiment runs that version; * changed: a new version is cut with the note `auto: added, edited, removed`, and the experiment runs it. `cef eval push --cut "note"` cuts one explicitly. The **History** table on the **Dataset** view lists each version with **Version**, **Saved**, **Note**, **Changes**, **Cases**, and **Experiments**. A version pins each case at one revision by sha256. Revisions are never rewritten, so an experiment from months ago reloads exactly the cases it ran. ## Storage Datasets live in your Agent Service’s bucket, beside the experiments that ran them: ```plaintext workflows//datasets//header.json header (+ baselineId) workflows//datasets//cases//r.json write-once revisions workflows//datasets//cases//deleted.json delete flag workflows//datasets//archived/.json archive flag workflows//datasets//versions/v.json frozen versions workflows//experiments//… experiments and their runs ``` `` is the workflow’s id (its agent alias). ROC, the CLI, and `@cef-ai/eval` read and write the same keys. Anyone who can write a dataset can change what counts as a pass. Treat write access to a dataset like write access to the workflow it tests. ## Related * [Next: Experiments](/agents/workflows/evaluate/experiments/) * [Evaluations overview](/agents/workflows/evaluate/overview/) * [Results](/agents/workflows/evaluate/results/): how `expected` is matched * [`@cef-ai/eval` reference](/reference/eval/) # Experiments > Run a dataset against a workflow version from ROC or with cef eval run, read the verdict and per-case results, compare two experiments, and run evaluations in CI. An experiment runs one dataset version against one workflow version and records every run’s output, scores, and metrics. Run one before you switch the live version, and after any change to a prompt, a model, or a step that shapes the Result. ## Run an experiment in ROC 1. Open the workflow and select **Evaluations**. 2. Pick the dataset in the **Dataset** switcher. 3. Select **Run experiment** and fill in the dialog. | Field | Default | Meaning | | -------------------------- | ---------------------- | ------------------------------------------------------------------------------------------------------------------- | | **Workflow version** | The live version | The version to run. A version that is not live is pinned for this experiment’s runs only and unpinned when it ends. | | **Compare against** | The dataset’s baseline | The experiment to judge this one against, or **No baseline**. | | **Note** | empty | What changed in this version. Shown in the experiment list. | | **Repeats** (Advanced) | 1 | Runs per case, 1 to 20. Use 3 or more for a model-heavy workflow to see flaky cases. | | **Concurrency** (Advanced) | 4 | Runs in flight at once, 1 to 16. | | **Timeout (s)** (Advanced) | 300 | How long one run may take to end, 10 to 3600 seconds. | Before it starts, the dialog checks the cases against the Result the chosen version declares. A warning such as *11 cases expect “summary”, which this version does not return* means those checks score n/a on that version, not failures. See [Results](/agents/workflows/evaluate/results/#when-the-result-changes). ROC runs the experiment from your browser tab, in the vault of the organization that owns the Agent Service. Leaving the page does not stop it; closing or reloading the tab does (the browser asks first). An interrupted experiment stays **Running** in the list; open its menu and choose **Resume**, or **Mark cancelled**. **Stop** ends dispatching; runs already in flight finish. ## Run an experiment with the CLI ```sh cef eval run triage # latest dataset version × the live workflow version cef eval run triage --workflow-version 1.5.0 --repeats 3 --note "shorter prompt" cef eval run triage --dataset-version 2 # an older dataset version cef eval run triage --resume exp-mgfz8c1k-a1b2 # finish an interrupted experiment ``` | Flag | Default | Meaning | | ------------------------------------------- | -------------------------------------------------- | ------------------------------------------------------------- | | `--dataset-version ` | Latest, cut automatically when the cases changed | Dataset version to run | | `--workflow-version ` | The live version | A version that is not live is pinned for this experiment only | | `--repeats ` | `1` | Runs per case | | `--concurrency ` | `4` | Runs in flight at once | | `--timeout ` | `300000` | How long one run may take to end | | `--baseline ` | The dataset’s baseline | Experiment to compare against when done | | `--resume ` | | Resume an experiment that did not finish | | `--note ` | | Note recorded on the experiment | | `--vault ` | `$CEF_VAULT_ID`, else the wallet’s own vault | Vault to run in | | `--scope ` | `default` | Vault scope to publish into | | `--access-token ` | `$CEF_ACCESS_TOKEN` | CLI access token from ROC, for the orchestrator | | `--orchestrator `, `--vault-api ` | From `--env` | Endpoint overrides | | `--workflow ` | The workflow `cef.config.ts` (or `dist/`) declares | Workflow id | | `--json` | | Print JSON instead of tables | `run` prints each run as it finishes (`[3/11] double-charge#1 pass`), then the experiment summary, then the comparison with the baseline when there is one. It records the Result the version declares (from the live deployment when the version is live, otherwise from your project’s own graph of that version) and warns before starting about cases that expect fields the version no longer returns. A resume is a new attempt: runs already on file are reused, the rest run in fresh contexts. ### Credentials Every `cef eval` command opens the Agent Service’s dataset bucket; `run` also reaches the vault and the orchestrator. | Variable | Flag | Needed for | | ---------------------------------------------------------------- | --------------------------------------------- | ---------------------------------------------------------------- | | `CEF_AS_PUBKEY` | `--as-pubkey` | Every command: names the bucket | | `CEF_KEYSTORE` + `CEF_KEYSTORE_PASSWORD`, or `CEF_SECRET_PHRASE` | `--keystore`, `--password`, `--secret-phrase` | The owner wallet (ed25519) that mints the store credential | | `CEF_EVAL_KEY` + `CEF_EVAL_SECRET` | `--access-key`, `--secret` | An explicit S3 key for the owner’s wallet, instead of the wallet | | `CEF_ACCESS_TOKEN` | `--access-token` | `run`: the CLI access token from ROC | | `CEF_ENV` | `--env` | `dev`, `stage`, or `prod` (default `dev`) | The token `cef push` uses (`CEF_DDC_ACCESS_TOKEN`) cannot open dataset storage; `cef eval` says so if it is the only credential set. A minted store credential is cached for its hour in `~/.cef/eval-credentials.json`; set `CEF_EVAL_CREDENTIAL_CACHE` to another file, or to `off`. ## Read the results The **Experiments** view shows, for the selected dataset: | Element | What it tells you | | ---------------------------------------------- | ---------------------------------------------------------------------------------------------- | | **Accuracy** | Passed runs over all judged runs. Runs that errored count as failures here. | | **p50 duration**, **Tokens/run**, **Cost/run** | Median wall-clock per run and the mean cost per run, each with its change against the baseline | | **Regressions** | Cases worse than the baseline | | **History** | One metric over time, one point per experiment; click a point to open it | | **Accuracy by field** | Each Result field’s pass rate, latest experiment against the baseline | | **Verdict** | One word per experiment (below) | | Verdict | Meaning | | ------------------------------------- | --------------------------------------------------------------------------------------------------- | | **Baseline** | This is the experiment the others are compared to | | **Regression** | A case got worse, or accuracy dropped, or only duration/tokens/cost got worse | | **Improvement** | Something got better (a case, accuracy, or duration/tokens/cost beyond noise) and nothing got worse | | **Tradeoff** | Better on one axis, worse on another | | **Same** | No change beyond noise | | **No baseline** | Nothing to compare with yet | | **Not compared** | A baseline exists but the runs of one side could not be read | | **Running**, **Cancelled**, **Error** | The experiment’s state | Duration changes under 15% and token or cost changes under 10% count as noise. Accuracy has no noise band. Open an experiment to see every case: its input, expected and actual fields, each repeat’s status (pass, fail, flaky, error, missing), and **Open run**, which takes you to that run in **Executions**. ![Experiment drawer with per-case results](/shots/evals-experiment-drawer.png) ### Choose a baseline The baseline is per dataset. In the experiment’s menu, choose **Make baseline** (only a finished experiment can be one). Until you choose one, ROC compares against the first finished experiment of the live version. The CLI compares against the stored baseline, or `--baseline`. Good practice: run the dataset on the live version, make that experiment the baseline, and judge every candidate version against it. Move the baseline when you switch the live version. ## Compare two experiments **In ROC**, tick two experiments and select **Compare**, or choose **Compare with baseline** from an experiment’s menu. **With the CLI:** ```sh cef eval compare exp-mgfz8c1k-a1b2 exp-mgg0p2xq-9zt4 ``` The comparison lists accuracy, duration p50 and p95, tokens/run, cost/run, steps/run, and limit breaches for both sides with their deltas, then every run that changed (`improved`, `regressed`, `only-in-base`, `only-in-candidate`). Only cases present on both sides **at the same revision** are compared. When the dataset changed between the two experiments, the comparison says so (*Dataset changed between these experiments: 1 case added, 1 expectation edited. Compared on the 10 unchanged cases.*) and leaves the changed cases out. ## Use in CI `cef eval run` exits non-zero only when it cannot run (bad flags, credentials, an unknown dataset). A regression does not fail the command. Read the JSON and decide: ```sh cef eval run triage --workflow-version "$CANDIDATE" --repeats 3 --json > experiment.json jq -e '.comparison == null or .comparison.counts.regressed == 0' experiment.json jq -e '.experiment.summary.passRate >= 0.9' experiment.json ``` The JSON is `{ experiment, result, comparison? }`: the [experiment](/reference/eval/#experiment-and-runrecord) with its `summary`, the Result compatibility check, and the comparison with the baseline when there is one. `summary.passRate` is passed over passed + failed; runs the harness could not start (`status: "error"`) are reported in `summary.errored` and are not in it. Check `errored` too, or ROC’s **Accuracy** and your CI gate will disagree. Practical rules for CI: * Run against a candidate version that is pushed but not live; the experiment pins it for its own runs only. * Keep the CI dataset small (tens of cases) and the repeats low; run the full dataset before you switch the live version. * Run in a vault reserved for testing. An experiment’s runs are real runs. ## Related * [Next: Results](/agents/workflows/evaluate/results/) * [Datasets](/agents/workflows/evaluate/datasets/) * [Test a workflow](/agents/workflows/test/) * [CLI reference](/reference/cli/) * [`@cef-ai/eval` reference](/reference/eval/) # Evaluations > Run a workflow version over a dataset of known cases, score each run against the Result it should produce, and compare versions on accuracy, duration, tokens, and cost. This group covers evaluating a workflow: this overview, then [datasets](/agents/workflows/evaluate/datasets/) of cases, [experiments](/agents/workflows/evaluate/experiments/) that run them, and the typed [Result](/agents/workflows/evaluate/results/) they are matched against. An evaluation answers one question before you switch a workflow to a new version: does it still produce the right Result, and what does it cost? You keep a **dataset** of cases, run a workflow version over it as an **experiment**, and compare that experiment with a **baseline**. Evaluations are warn-only. Nothing in them blocks a push or a deploy; you decide what a regression means for your release. ## The parts | Term | What it is | | ------------------- | --------------------------------------------------------------------------------------------------------------------------- | | **Dataset** | A versioned set of cases for one workflow. Names are unique per workflow. | | **Case** | One `workflow.start` payload (`input`) and what the run’s Result should be (`expected`). | | **Dataset version** | A frozen cut of a dataset: every live case pinned at one revision. Cut automatically when an experiment runs changed cases. | | **Experiment** | One dataset version run against one workflow version, each case run `repeats` times. | | **Run** | One case × one repeat inside an experiment, with its output, scores, and metrics. | | **Result** | The fields a workflow’s Result step (`output` step) declares. Runs are scored against these fields and nothing else. | | **Baseline** | The experiment other experiments on a dataset are compared to. | ```plaintext dataset "triage" v3 ──┐ ├─► experiment exp-… ──► runs (case × repeat) ──► scores + metrics workflow v1.4.0 ──────┘ │ ▼ compare with the baseline experiment ``` ## How a run is scored 1. The experiment publishes the case’s `input` as a `workflow.start` event into a context of its own (`eval----a`). 2. The workflow runs as it would for any other start event: real model calls, real agent steps, real cubby writes. 3. When the run ends, its `workflow.completed` output (the declared [Result](/agents/workflows/evaluate/results/)) is matched field by field against the case’s `expected`. 4. The run’s duration, tokens, cost, and step count are read back from the platform and checked against the dataset’s limits. A case passes when every judged field passes. Limits only warn; they never fail a case. See [Results](/agents/workflows/evaluate/results/) for the matching rules. An experiment runs your workflow for real. Every run calls models, writes cubby rows, publishes events, and performs connector actions. Run experiments in a vault where that is safe, and make side-effecting steps idempotent. An experiment’s runs have ids starting with `run-eval-`, so rows you key on `{{ $runId }}` are easy to tell apart. ## Evaluations and tests | | Local tests | Evaluations | | ------- | ------------------------------------------------------ | ---------------------------------------------------------------- | | Runs | On your machine, in-process | On the platform, against a deployed version | | Models | Mocked | Real | | Checks | Exact behaviour: cubby rows, events, branches, budgets | The Result, plus duration, tokens, cost, steps | | Answers | “Does the graph do what I wrote?” | “Is this version as good as the last one, on real inputs?” | | Use | Every commit | Before switching the live version, after prompt or model changes | Use both. Tests pin down logic that must never change; a dataset catches the quality drift that only real model calls show. See [Test a workflow](/agents/workflows/test/). ## Where evaluations live **ROC.** Open a workflow and select the **Evaluations** tab. It has two views: * **Experiments**: the headline metrics (**Accuracy**, **p50 duration**, **Tokens/run**, **Cost/run**, **Regressions**), a **History** chart, **Accuracy by field**, and the experiment list with a **Verdict** per experiment. * **Dataset**: the cases, **Add from executions**, **Add case**, and the dataset’s **History** of versions. **Run experiment** starts an experiment from either view. The **Evaluations** item in the sidebar lists every dataset across the Agent Service’s workflows, with each one’s **Latest verdict** and **Accuracy**. ![Experiments view with metrics, history chart and experiment list](/shots/evals-experiments.png) **CLI.** `cef eval` reads and writes the same storage as ROC, so either sees what the other wrote: ```sh cef eval create triage cef eval pull triage # cases into ./datasets/triage/ cef eval push triage # your edits back as new revisions cef eval run triage # run against the live version cef eval compare ``` **Library.** [`@cef-ai/eval`](/reference/eval/) is the browser-safe library both ROC and the CLI use. Use it to build your own tooling over the same storage. ## Storage Datasets and experiments are JSON objects in your Agent Service’s bucket, under `workflows//`. Case revisions and dataset versions are write-once, so an old experiment always reloads exactly the cases it ran. See [Datasets](/agents/workflows/evaluate/datasets/#storage). ## Related * [Next: Datasets](/agents/workflows/evaluate/datasets/) * [Experiments](/agents/workflows/evaluate/experiments/) * [Results](/agents/workflows/evaluate/results/) * [Test a workflow](/agents/workflows/test/) * [`@cef-ai/eval` reference](/reference/eval/) # Results > Declare a typed Result on a workflow's Result step with result, resultTypes, and outcome, see what a wrong type does to a run, and learn how cases are matched against the Result. A workflow’s **Result** is the set of fields its Result step declares. Declared, the run’s final output is exactly those fields; undeclared, the run ends with the whole carried item (the trigger payload plus everything each step merged onto it). Declare a Result on every workflow you evaluate: two runs can only be compared on what they promised. The Result is also what a webhook caller waiting on the run receives, and what a dataset’s cases are matched against. ## Declare a Result The Result step is the `output` step kind. It takes three params. | Param | Type | Meaning | | ------------- | --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `result` | `{ field: value }` (or that object as JSON text) | The declared fields. Each value is resolved against the item like any mapped param (`"={{ $json.category }}"`) or is a literal. At least one field. | | `resultTypes` | `{ field: "string" \| "number" \| "boolean" \| "object" \| "array" }` | Optional. The type each field must resolve to. Every key must be a field of `result`. | | `outcome` | string | Optional. A label for how the run ended on this path (`"escalated"`), resolved like any param. Sent on `workflow.completed` beside the output. | ```ts { id: "result", kind: "output", label: "Result", position: { x: 1200, y: 0 }, params: { result: { category: "={{ $json.category }}", confidence: "={{ $json.confidence }}", model: "classifier-v2", }, resultTypes: { category: "string", confidence: "number" }, outcome: "={{ $json.category }}", }, } ``` The run ends with: ```json { "runId": "run-t-1", "output": { "category": "billing", "confidence": 0.93, "model": "classifier-v2" }, "outcome": "billing" } ``` **In the Workflow Builder**, add a **Result** step. In its Output pane, add each field with its type (leave a field’s type as *any* to skip the check). Under **Where each field comes from**, write one mapping per line (`category = ={{ $json.category }}`); a field you do not list is read from the item by its own name. **Outcome** is the optional label. ![Result step editor with typed fields, sources and outcome](/shots/result-step.png) `defineWorkflow` checks the declaration when your config loads (at `cef build`): `result` that is not an object, a `result` with no fields, an unknown type in `resultTypes`, a type for a field `result` does not declare, and an `outcome` that is not a string are all refused before anything deploys. ## What a wrong type does A field that resolves to a type other than its `resultTypes` entry **fails the run at the Result step**: * `workflow.step_failed` names the step: `result field 'confidence' expected number, got string`; * `workflow.failed` carries `result: result field 'confidence' expected number, got string`; * no `workflow.completed` is published, and in an experiment the run scores as failed. A typed field that resolves to nothing fails the same way (`got nothing`); so does `null` (`got null`). An untyped field that resolves to nothing is left out of the output rather than set to `null`. This is deliberate. A Result that silently changed type is the regression an evaluation exists to catch, and failing at the step names the cause instead of leaving a mismatch for a downstream consumer to find. If a model step can return the wrong type, normalise it in a `transform` step before the Result (see [Best practices](/agents/workflows/best-practices/#model-output)). ## Several Result steps A graph with several branches can end in several Result steps, each with its own `outcome`. The declared Result of the version is the union of their fields. A field two Result steps type differently is treated as untyped for scoring. ## Matching Each key of a case’s `expected` is one field score, addressed by dot path (`reply.language`). | Expected value | Passes when | | -------------------------------- | ----------------------------------------------------------------------------------------------------------- | | A plain value | The actual value is deeply equal | | A plain object | Every key it names matches (a subset match; other keys are ignored) | | A string such as `">= 0.8"` | The actual value is a number inside the band (`>=`, `<=`, `>`, `<`). Against a string it is plain equality. | | `{ "$eq": v }` | Deep equality | | `{ "$oneOf": [a, b] }` | Equal to one of the options | | `{ "$contains": v }` | A string containing `v`, an array with an item matching `v`, or an object matching `v` | | `{ "$regex": "^T-\\d+$" }` | A string matching the pattern | | `{ "$between": [min, max] }` | A number with min ≤ value ≤ max | | `{ "$gte": n }`, `{ "$lte": n }` | A number at or above / at or below `n` | | `{ "$approx": n, "$tol": t }` | A number within `t` of `n` | | `{ "$exists": true }` | Present and not `null` (`false`: absent or `null`) | A matcher naming several operators needs all of them (`{ "$exists": true, "$gte": 1 }`). A field the output does not have fails. A case **passes** when every judged field passes. It is **unscored** when it has no expectations. A run that failed or timed out fails. `$regex` patterns run synchronously during scoring. Scoring refuses a pattern longer than 256 characters or one that repeats a group which itself contains a quantifier (`(a+)+`), failing that field with `unsafe $regex refused: …`. This is a heuristic; keep patterns simple. Match free text loosely. A model rewording a sentence is not a regression, so expect `{ "$exists": true }` or `{ "$contains": "refund" }` for prose and exact values for labels, ids, and numbers with a band. ## Limits Limits (`maxDurationMs`, `maxTokens`, `maxSteps`, `maxCost`) are set on the dataset and overridden per case. Each one in force produces a limit score. A breach is counted in the experiment’s limit breaches and never fails a case. A metric the platform did not measure is not judged. ## When the Result changes An experiment records the Result its workflow version declares. Expectations that version cannot satisfy are scored **n/a**, not failed: | Situation | Score detail | | ----------------------------------------------------------------------------------- | ------------------------------------------------------- | | The case expects a field the version does not declare | `not in this version's Result` | | The expectation cannot match the declared type (`{ "$gte": 1 }` against a `string`) | `type changed: expected number, Result declares string` | The dataset is older than the version, which is not a regression. ROC’s **Run experiment** dialog and `cef eval run` both warn about this before the experiment starts. Update the cases (or the Result) and the n/a scores go away. When a version declares no Result at all, every expected field is judged against the whole output. ## Related * [Next: Monitor runs](/agents/workflows/monitor-runs/) * [Experiments](/agents/workflows/evaluate/experiments/) * [Datasets](/agents/workflows/evaluate/datasets/) * [Workflow steps](/agents/workflows/steps/) * [Workflows in code](/agents/workflows/author-in-code/) # Monitor runs > Follow workflow runs in ROC: the Executions tab, a run's steps and data, statuses, re-running and retrying, versions and rollback, metrics, and finding out why a run failed, parked, stalled, or never started. Every run of a workflow is recorded as a Job, with one Task per dispatched step. ROC reads them in the Workflow Builder’s **Executions** tab. ## Executions Open the workflow (Agent Service → **Workflows** → the workflow) and choose the **Executions** tab, shown as **Executions · N** once there are runs. * **Run** starts a run of the deployed version. It opens **Run this workflow**, with **What starts it** prefilled from the trigger’s example payload, and **Who takes part** when a step asks the run’s participants. The chip beside it shows where runs go: **Runs in {organization} · {scope}**. * The list shows each run with its status, the version it ran (`v0.1.4`), and where it stopped (**failed at {step}**, **at {step}**). * **Filter runs by status**: failed, running, waiting, stalled, done. Runs use the version that was live when they started, not the canvas. If the canvas has undeployed changes, **Run** warns that the run uses the live version. ![The Executions tab with a mix of done, waiting, and failed runs](/shots/executions-tab.png) ## Statuses | Status | Headline in ROC | Meaning | | --------- | --------------------------------------------------------------- | ------------------------------------------------------------------- | | running | Running … | Steps are executing. | | waiting | Waiting at … / Needs a decision at … / Asked you something at … | Parked on a person, on an agent’s question, or on an authorization. | | done | Finished | Reached its end. | | failed | Failed at {step} | A step failed; the run records the reason. | | stalled | Nothing picked up {step} | A step’s edges all had conditions and none passed. | | cancelled | Cancelled by {who}: {reason} | Closed by the person who started it. | ## A run’s detail Open a run to see the graph with each step’s state (**Success**, **Failed**, **Skipped**, **Waiting on a person**, **Waiting on an agent**, **Running**, **Not run**), and the run’s log. Select a step to see: * what it **Received** and **Produced**: the carried item before and after; * for an agent step, what was **Sent** to the agent and what it **Returned**; * for a model step, the alias and request; for a cubby step, the SQL and the rows read or changed; * the error, for a failed step. Long values are shortened in the run view: strings over 120 characters keep their start and say how long they were, and lists keep their first 20 elements. The run summary shows **Took**, **Steps**, **Slowest**, **Model calls**, **Tokens**, **Cost**, **Agents**, and **Cubby calls**. The **Summary** panel shows the run’s outcome (including the pinned widget), the run, and the events it **Announced**. ![One run: the canvas, its summary, and the execution log with step timings](/shots/run-detail.png) ### Replay The timeline under the canvas steps through the run hop by hop: **Play** / **Pause**, **Previous step**, **Next step**, **Previous hop**, **Next hop**. Replay only animates the recorded run; it starts nothing. ## Act on a run | Action | What it does | | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Answer a waiting step | The waiting step shows a card: **Approve** / **Reject** for an open gate, the step’s own answer buttons and form for an assigned step, a text box and **Send to {agent}** when an agent asked a question. **Send back to {step} with this note** returns the run to an earlier step. See [People in workflows](/agents/workflows/people/). | | **Retry “{step}”** | Resumes the same run at the failed step, with the input it had. Later steps appear in the same run. Use it after fixing the cause (an agent, a connection, a cubby schema). | | **Re-run** | Starts a new run from the top with the same trigger payload: real agents, real tokens, a new Job, billed again. | | **Add to dataset** | Saves the run’s input and result as a case for evaluations. See [Datasets](/agents/workflows/evaluate/datasets/). | ## Versions and history Each **Deploy** writes a new version and makes it live. The Deploy button shows **Deployed {version}** when the canvas is what runs. The arrow beside Deploy (**Versions and history**) opens: * **Versions**: every version, the live one marked, with **Deploy** on the others to make one live. * **Save as version {next} — don’t deploy**: record the canvas as a version without making it live. * **Deployment history…**: one entry per deploy, each holding the exact graph that was live. **Open** loads one onto the canvas as an unsaved draft; **Make live** rolls back to it without touching your canvas. * **Undeploy** and **Archive workflow…**. An archived workflow receives no events and is hidden from the dashboard until you restore it. **Discard** throws away unsaved canvas edits and reloads the deployed graph. ## Metrics The workflow’s agent page (Agent Service → **Agents** → the workflow → **Overview**) and the **Metrics · {name}** view opened from the Workflows list show, over 1 h, 24 h, or 7 d: | Panel | Shows | | ------------------------- | ----------------------------------------------------------------------------------------- | | Workflow runs | **Runs**, **Completion rate**, **Failed rate**, **Stuck**, and runs over time by outcome. | | Tokens | Tokens over time. | | Compute limit (GPU units) | Use against the limit. | The agent page also opens the workflow’s **Jobs** and **Logs**. ## Why a run failed, parked, or stalled Work from the run outward: read the run’s own record first, then the cubby rows it wrote, and reach for logs last. Start from the run’s first step that is not **Success**; its error names the step and the reason. ### Failed A **failed** run is over. Its reason is on the failed step (`workflow.step_failed`) and on the run (`workflow.failed`). Fix the cause, then **Retry** the step or start a new run. ### Parked A **waiting** run is parked: idle, holding nothing open, until the event it waits for arrives. | Parked at | Waits for | If it never arrives | | ------------------ | ----------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ | | A human step | A person’s answer | Expected. Check the step is assigned to someone who can see it. | | An agent step | The agent’s answer | The agent is not connected, not deployed, or failed without answering. Open the agent from **Agents** and check its **Jobs** and **Logs**. | | A connector action | `connector.action.succeeded` or `.failed` | Check the connection in the vault. | A workflow that parks anywhere other than a human step, or an agent that asked a question, has a bug in the graph or in the agent. ### Stalled A **stalled** run stopped because no edge out of a step matched: every condition on that step’s outgoing edges was closed for this item (*Nothing picked up {step}*). Check the field the condition tests in the step’s **Produced**, and add an edge without a condition as the default. ### The run never starts No run appears in **Executions** after the trigger fired: * The workflow, or an agent it calls, is not connected to the vault, or the consent was not approved. Press **Run** in ROC and approve the connection; see [Connect to a vault](/agents/ship/connect-to-a-vault/). * The trigger’s event type does not match the event that was sent. Check the trigger step. * A webhook call was refused (`401`, `403`, `413`, `429`). Read the webhook response; see [Triggers](/agents/workflows/triggers/#webhook-limits). ### Errors and their causes | What you see | Cause | Fix | | ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ | | Fails at the first step with a list of graph problems | The graph does not validate. | Fix the problems the Builder or `cef build` lists, deploy again. | | `model alias not declared: ''` | The model step’s alias is not in the workflow’s `models`, or does not match the alias inside `model.json`. | Declare it under the model’s own alias; see [Models](/agents/building-blocks/models/#declare-an-alias). | | `answered without ''` | The agent’s answer lacks a key the step’s answer shape declares. | Tighten the agent’s instructions, or the shape. | | A field is empty in a later step | The field was never on the item when that step ran; mappings to missing fields resolve to empty. | Open the step before and check **Produced**. Declare an answer shape for fields you depend on. | | `result field '' expected number, got string` | The Result’s type check caught a changed type. | Normalise the value in a step before the Result, or correct `resultTypes`. See [Results](/agents/workflows/evaluate/results/). | | `no edge out of '' matched` | Every condition after that step was closed for this item. | Fix the condition, or add an unconditioned default edge. | | `no such table` in a cubby step | The cubby or table was not declared, or the migration was never pushed. | Declare it and run `cef cubby push`. | | `no cubby is selected` / `no SQL is set` | A Cubby step has an empty alias or SQL. | Fill in the step. | | `SQL may not interpolate the run's data` | A cubby step’s SQL contains `{{ $json.… }}`. | Use `?` and `args`. | | `could not publish ''` | The event was refused, typically over 1 MiB. | Declare the publish step’s `payload`; send references, not content. | | `… over the 1048576-byte limit a waiting step can hold` | The item carried into an agent or human step is too large. | Shrink the item before the step. | | A memory step says the bank “accepted this write and kept nothing” | The connection does not allow that scope or privacy level. | Use a scope and privacy the vault granted. | | A connector step fails with `no_access` | The workflow has no grant for that action. | Deploy again from the Builder, or grant access; see [Connectors](/vaults/connectors/#access-is-part-of-the-deploy). | | `loop limit reached` | A loop edge ran out of rounds with no exhausted step. | Raise **Max rounds** or set **When the rounds run out**. | | `gave up after 12 exchanges without a final answer` | An agent and a person went back and forth without the agent finishing. | Make the agent ask for everything at once. | | Waiting, **Needs a decision** / **Asked you something** | A person or an agent’s question is pending. | Answer it from the run. | | Waiting, **Waiting for authorisation** | An agent needs access granted. | Grant it, then **I have granted it — run this step again**. | | Run **Done** but a widget shows nothing | The cubby step wrote to another alias or table than the widget reads, or the row’s key does not match. | Open the run’s Cubby step, query the row, compare with the widget’s query. | | Two rows for one request | The start was delivered twice into different contexts, or a write is not keyed on the run. | Send an `Idempotency-Key`; key writes on `run_id`. | | An experiment’s runs all **error** | The experiment could not start runs: wrong vault, workflow not connected, missing credentials. | Read the experiment’s error; check `--vault` and the connection. | See [Limits](/reference/limits/) for every hard limit and its error. ### Inspect cubby data Cubby rows are the durable record of what a run did; logs expire. * **From a run:** select a Cubby step and choose **Open cubby**. It opens the Agent Service’s cubby in the vault the run executed in, with a SQL runner. * **From the sidebar:** **Cubbies** lists the Agent Service’s cubby declarations; **Inspect data** opens the same inspector. Query by the run: every run’s id is `run-`, and a well-designed table keys on it. ```sql SELECT * FROM triage_results WHERE run_id = 'run-t-1'; ``` Design tables so a `SELECT` answers “what happened”: a `status` column that advances (`in_progress`, `done`, `failed`), and the reason on failure. A stuck row then tells you which stage broke long after the logs are gone. ### Read logs Open the workflow’s agent page (Agent Service → **Agents** → the workflow) and select **Logs**. Filter by **Time range** and **Level**, and narrow to one Job or Task. Steps that run code agents log under those agents; open them the same way. How agents log, and how to read logs with the Vault SDK, is on [Test and debug a code agent](/agents/code-agents/test-and-debug/#read-logs). ### Reproduce it locally Most workflow bugs reproduce in-process. Take the failing run’s start payload (open the run in **Executions**, or capture it with **Add to dataset**), and replay it through `@cef-ai/testing` with the model and agent answers the run received. Keep the test once it passes: it is the regression test for the fix. See [Test a workflow](/agents/workflows/test/#1-local-tests). ## Limits | Limit | Value | | ----------------------------- | ----------------------------------------------------------------------- | | Task log query window | 24 hours. Older logs come back empty. | | Finished Jobs and their Tasks | Kept 7 days, then deleted. Capture runs worth keeping as dataset cases. | ## Related * [Next: Best practices](/agents/workflows/best-practices/) * [Workflows](/agents/workflows/overview/) * [People in workflows](/agents/workflows/people/) * [Evaluations](/agents/workflows/evaluate/overview/) * [Limits](/reference/limits/) # Workflows > What a workflow is: an agent whose behavior is a typed graph of steps, built in the ROC Workflow Builder or in code, and run on vault data once connected. This group covers workflows end to end: the graph and its steps, authoring in the Builder or in code, and the workflow-only work of testing, evaluating, and monitoring runs, with best practices last. A **workflow** is an agent whose behavior is a graph of steps instead of handler code. It has an id and versions, lives in your Agent Service, is published and deployed like any agent, and runs on a customer’s data only after the vault owner connects it. What it adds is the graph: steps that call models and agents, ask people, read and write cubbies and the Memory Bank, act through connectors, and decide where to go next. You build a workflow in the ROC **Workflow Builder** or in TypeScript with `defineWorkflow`. Both produce the same document. ## The graph A workflow is a set of **steps** joined by **edges**. ```plaintext trigger → classify (model) → branch ─┬→ escalate (human) → reply (connector action) └→ file (cubby exec) → result (output) ``` * Every run starts at a **trigger** step: an event, a schedule, a webhook call, or a message arriving through a connector. See [Triggers](/agents/workflows/triggers/). * Each step does one thing and hands the run on along its edges. * An edge can carry a condition (`when`): it is taken only when a field of the carried item passes the test. * A step with several edges out sends the run down all of them at once. A **branch** step instead takes the first edge whose condition passes. A **join** step waits for several branches and merges them. * A cycle is allowed only through a **loop** edge, which carries a round limit. Eighteen step kinds cover models, agents, people, data, memory, connectors, control flow, and results. See [Steps](/agents/workflows/steps/). ![A ticket-triage workflow on the Builder canvas](/shots/builder-canvas.png) ## The carried item A run carries one JSON object from step to step: the **item**. The trigger’s payload is the first item. Each step reads it and passes on an updated version: | Step | What it does to the item | | ------------------------------ | ----------------------------------------------------------------------------------------- | | Agent, human, connector action | Merges the answer’s fields over the item. On a name clash the answer wins. | | Model, cubby query, recall | Adds the result under one field (`into`) and keeps everything else. | | Edit fields (transform), code | Replaces the item with what the expression or code returns. | | Join | Merges every branch’s item, in edge order, and adds each branch’s item under its step id. | | Aggregate | Adds every per-item result under one field. | A step’s parameters read the item through mappings: a string that starts with `=` and contains `{{ $json. }}`. ```ts params: { prompt: "=Summarise this ticket from {{ $json.customer }}:\n{{ $json.text }}" } ``` * A whole-value mapping such as `"={{ $json.score }}"` keeps the field’s type: a number stays a number. * Inside text, values are spliced in as strings. * A missing field resolves to empty, not to an error. * A string containing `{{ $json.… }}` without the leading `=` is refused before the run starts, because it would be sent as literal text. Beside `$json`, a mapping can read the run itself: `{{ $runId }}`, `{{ $run.initiator }}`, `{{ $run.members }}`, `{{ $run.participants }}`. See [People in workflows](/agents/workflows/people/#participants). Keep the item lean. A run holds the inputs of the steps it is waiting on in a 1 MiB budget; above that, the runner drops carried input from waiting steps, largest first. Store bulky data in a cubby or the Memory Bank and carry ids. ## Runs A **run** is one pass through the graph, started by one trigger event. Its record is a **Job**: every step that dispatches work adds a Task to it, with timings, tokens, and the step’s input and output. The run ends in one of these states: | Status | Meaning | | ----------- | ------------------------------------------------------------------------------------------- | | `running` | Steps are executing. | | `waiting` | Parked on a person, an agent’s question, or an authorization. Costs nothing while it waits. | | `done` | Reached its end. | | `failed` | A step failed. The run records which step and why. | | `stalled` | A step’s edges all had conditions and none passed. | | `cancelled` | The person who started it closed it. | Steps publish `workflow.step_ran` or `workflow.step_failed` into the run’s context as they finish, and the run ends with `workflow.completed` or `workflow.failed`. ROC’s **Executions** tab reads these. See [Monitor runs](/agents/workflows/monitor-runs/). ## Versions Every deploy writes a new version of the workflow and makes it live. A run uses the version that was live when it started and keeps that graph to the end, so editing a workflow never changes a run already in flight. You can open any earlier version, or make it live again. See [Monitor runs](/agents/workflows/monitor-runs/#versions-and-history). ## Builder or code | | Workflow Builder | `defineWorkflow` | | ---------- | -------------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | Where | ROC, Agent Service → **Workflows** | `cef.config.ts` in your repo | | Ship | **Deploy** button | `cef build`, `cef push`, `cef deploy` | | Validation | Problems shown on the canvas; Deploy refuses a graph that cannot run | `cef build` refuses a graph that cannot run, and types catch edges to unknown steps | | Review | Version history in ROC | Pull requests | | Tests | Runs and evaluations in ROC | `@cef-ai/testing` plus evaluations | The Builder’s **Export as code** produces a `defineWorkflow` file from a canvas, and a workflow pushed from code opens in the Builder read-only, with your step positions. See [Workflows in code](/agents/workflows/author-in-code/). ## Workflow or code agent Choose a **workflow** when: * the work is a sequence of model calls, agent calls, and data steps you want to see and re-run step by step; * people approve, correct, or decide in the middle; * non-developers on your team need to read or change the flow; * triggers are schedules, webhooks, or connector messages. Choose a **code agent** when: * you need long-lived session state or real-time streaming; * the logic needs network calls, timers, or libraries inside one step (a workflow’s code step is deterministic: no network, no clock, no imports); * you want one agent that a workflow calls as a step. A workflow can call your code agents and LLM agents through Agent steps. See [Agents and workflows](/agents/overview/) for how the two compare at the platform level. ## Related * [Next: Steps](/agents/workflows/steps/) * [Quickstart](/get-started/quickstart/) * [Triggers](/agents/workflows/triggers/) * [Workflows in code](/agents/workflows/author-in-code/) * [Connections and consent](/vaults/connections-and-consent/) # People in workflows > Park a run until vault members answer: human steps, assignees, closing policies, options and forms, shared documents, loop-back edges, and run participants. A **human** step parks the run and asks people a question. The run costs nothing while it waits: the Job goes idle between the question and the answer, so a workflow can wait days for a decision. When the answer closes the step, the run continues from there with the answer on the carried item. Use a human step for approvals, reviews, corrections, and anything a model should not decide alone. Branch on the answer afterwards (see [Steps](/agents/workflows/steps/#branch)). ## The simplest gate A human step with only a `question` is an open gate: anyone who can write the vault scope answers once, and the run moves on. ```ts { id: "approve-refund", kind: "human", question: "=Refund {{ $json.amount }} to {{ $json.customer }}?", label: "Approve refund", } ``` The answer adds `approved` (when the answer carries one), `text`, `output`, and `answeredBy: "human"` to the carried item. A following branch can gate on `approved`. `question` is resolved like any mapped param, so `={{ $json.field }}` works in it. Without a question, the step’s `label` is asked. ## Assigned steps Give the step any of `assignees`, `policy`, `options`, `fields`, `into`, `mode`, or `document` and it becomes an **assigned** step: only the people it names may answer, the answers are checked against the options and fields, and a policy decides when the step closes. ```ts { id: "review-reply", kind: "human", question: "Review the drafted reply before it goes to the customer.", params: { assignees: "participants-except-initiator", policy: { atLeast: 2 }, options: [ { value: "send", label: "Send it" }, { value: "rewrite", label: "Rewrite", requireNote: true }, ], fields: [{ key: "tone", type: "options", options: ["fine", "too formal", "too casual"] }], into: "review", }, } ``` ### Parameters In the Builder, add **Human approval** from **Flow**. Its panel has **Question**, then **Who is asked**, **Closes when**, **Mode**, **Answers**, **Form**, and **Save answers as**, which set the params below. | Param | Values | Default | What it does | | ----------- | -------------------------------------------------------------------------------------------------------------------------- | ------------------------------ | ------------------------------------------------------------------------------------------------------------------------ | | `assignees` | `"participants"`, `"initiator"`, `"participants-except-initiator"`, a member key, a list of member keys, or an `=` mapping | anyone who can write the scope | Who may answer. Keywords resolve against the run’s participants. A member key is `0x` + 64 hex characters. | | `policy` | `"any"`, `"all"`, `{ atLeast: n }` | `"any"` | When the step closes: the first answer, every assignee, or `n` distinct answers. `"all"` needs `assignees`. | | `options` | list of `{ value, label?, requireNote? }` | `approve`, `reject` | The choices an answer picks from. `requireNote` makes a note mandatory for that choice. | | `fields` | list of `{ key, type, label?, required?, options?, help? }` | none | A form each answer fills in. Types: `string`, `text`, `number`, `boolean`, `url`, `email`, `options`, `list`. | | `into` | a field name | none | Puts the step’s output under this field instead of merging it into the item. | | `mode` | `"single"` or `"open"` | `"single"` | `open` turns the step into a shared document people edit and save until one of them closes it. See below. | | `document` | `{ from, format? }` | none | Open mode only. `from` is the expression the first version is read from; `format` is `"markdown"` (default) or `"html"`. | The run fails at the step, with the reason, when its assignees resolve to nobody or when `{ atLeast: n }` is larger than the number of assignees. ### What the step produces When the policy is met, the step’s output is: ```json { "responses": [ { "member": "0x…", "option": "send", "note": "", "values": { "tone": "fine" }, "eventId": "…", "at": 1760000000000 }, { "member": "0x…", "option": "send", "note": "", "values": { "tone": "too formal" }, "eventId": "…", "at": 1760000004000 } ], "counts": { "send": 2 }, "assignees": ["0x…", "0x…", "0x…"], "decision": "send", "activation": 1 } ``` `decision` is the chosen option when every answer agreed, and `null` on any disagreement, so a branch gating on `decision` takes its fallthrough edge unless the answers were unanimous. Responses are in arrival order. With `into: "review"`, branch on `review.decision`. ### Answering Members answer from the run in ROC: a waiting step shows a card with its question, options, and form. An open gate offers **Approve** and **Reject**; an assigned step shows its own answer buttons and who it is still waiting on. An app answers by publishing `workflow.feedback` into the run’s context: ```json { "nodeId": "review-reply", "option": "rewrite", "note": "Drop the second paragraph.", "values": { "tone": "too formal" } } ``` The vault stamps who sent it, and the runner checks that person against the assignees. A refused answer leaves the run where it was and publishes `workflow.response_rejected` with a `reason`: | Reason | Meaning | | --------------------- | ----------------------------------------------------------- | | `not_assignee` | The sender is not one of the step’s assignees. | | `not_a_person` | The answer was published by an agent or app, not a member. | | `duplicate` | This member already answered this step. | | `unknown_option` | `option` is not one the step offers. | | `note_required` | The chosen option has `requireNote` and the note is empty. | | `invalid_field:` | A field value is missing (when required) or the wrong type. | | `stale_activation` | The answer is for an earlier pass of the step. | | `not_open` | Nothing is waiting at that step. | | `redo_target` | A send-back named a step it cannot return to. | Each accepted answer publishes `workflow.response_recorded`. ### Limits | What | Limit | | ----------------------------------- | --------------------------------------------------------------------------------- | | Note length | 20,000 characters | | One field value | 20,000 characters | | Fields per step, values per answer | 64 | | Document body (open mode) | 2,000,000 characters | | Input the run carries into the step | 1 MiB (1,048,576 bytes). A larger item fails the step; shrink it before the step. | ## Shared documents (open mode) An open step hands people a document to edit together. Its options default to `save` and `done`: `save` stores a new version and keeps the step open; any other option closes it. ```ts { id: "edit-reply", kind: "human", question: "Edit the reply, then mark it done.", params: { mode: "open", assignees: "participants", document: { from: "={{ $json.draft }}", format: "markdown" }, }, } ``` A save sends `{ "option": "save", "values": { "baseVersion": 3, "body": "…" } }`. A save based on an old version is refused with `stale_version` and the latest version, so two editors never overwrite each other silently. The step’s output adds `document: { version, path, body }` with the final text. An open step needs a `save` option and at least one other option when you declare `options` yourself. ## Send a run back An answer can carry `redoFrom: ""` to send the run back to an earlier agent, human, or connector-action step instead of moving on. The target receives the answer’s `text` as `feedback` on its item and runs again from there. The target must be upstream of the waiting step, and inside the same item region if the waiting step is in one. ## Loop-back edges An edge with `loop` is the only edge allowed to close a cycle. Use it for “revise until approved”. In the Builder, select the edge and turn on **Loop back**, then set **Max rounds**, **Counter name**, and **When the rounds run out**. ```ts edges: [ { from: "draft", to: "review-reply" }, { from: "review-reply", to: "send", when: { field: "decision", op: "eq", value: "send" } }, { from: "review-reply", to: "draft", loop: { max: 3, counter: "rounds", exhausted: "escalate" }, }, ], ``` | Field | What it does | | ----------- | ----------------------------------------------------------------------------------------------------------------- | | `max` | Most times the target may be entered, the first pass included. 1 to 100. | | `counter` | The item field that counts entries. It reads 1 on the first pass, so a prompt can say `Round {{ $json.rounds }}`. | | `exhausted` | The step the run goes to once the bound is spent. Without it, the run fails with “loop limit reached”. | Re-entering a human step starts a new activation: answers to the earlier pass are refused as `stale_activation`. A cycle with no loop edge is refused before the run starts. ## Participants A run can name who takes part. Participants are what the assignee keywords resolve against. Start a run with participants from the Run dialog in ROC, under **Who takes part** (shown when a Human approval step asks the run’s participants), or by publishing `workflow.start` with a `participants` list: ```json { "participants": [ { "member": "0x1f…", "name": "Dana", "roles": ["support-lead"] }, { "member": "0x9c…", "name": "Ravi" } ], "ticketId": "T-1042" } ``` | Rule | Limit | | ---------------------------------------- | --------------------------------------- | | Participants per run, initiator included | 2 to 64 | | `member` | `0x` + 64 hex characters, no duplicates | | `name` | up to 120 characters | The member who starts the run is the **initiator**. If they are not in the list they are added with the role `initiator`; if they are, they gain the role. A run started through a webhook takes no participants from its body. Templates can read the run’s people: | Mapping | Value | | -------------------------- | ------------------------------------- | | `={{ $run.participants }}` | The participant list. | | `={{ $run.members }}` | The participants’ member keys. | | `={{ $run.initiator }}` | The initiator’s member key, or empty. | | `={{ $run.id }}` | The run id. | ## Close a run The initiator can end a waiting or running run by publishing `workflow.close` with an optional `reason`. The run ends `cancelled` and `workflow.run_closed` is published. A close from anyone else is refused with `workflow.close_refused`. ## Related * [Next: Split and aggregate](/agents/workflows/split-and-aggregate/) * [Steps](/agents/workflows/steps/) * [Triggers](/agents/workflows/triggers/) * [Monitor runs](/agents/workflows/monitor-runs/) * [Vault members](/vaults/members/) # Split and aggregate > Run a stretch of a workflow once per item of a list with a split step, then collect every item's result with an aggregate step. A **split** step opens an item region and an **aggregate** step closes it. Every step between them runs once per item: an agent is asked once per ticket, a model is called once per chunk, a person answers once per item. The aggregate waits until its policy is met, then hands every item’s result on as one list. Use a region when each item needs more than one step, or a step that waits (an agent, a person, a connector action). For a single model call per list entry, the model step’s `each` option is simpler: see [Steps: Model](/agents/workflows/steps/#model). ## How a region runs ```plaintext trigger → split ──→ classify → reply ──→ aggregate → output (once per item, up to 8 at a time) ``` 1. The split reads a list from the carried item, for example `$json.tickets`. 2. Each list entry becomes an item. A step inside the region receives the carried item with the entry’s fields spread over it, plus `itemIndex` (0-based position in the list) and `itemKey`. An entry that is not an object arrives as `$json.item`. 3. The list field itself is removed from what each item carries, so 200 tickets are not copied into every per-item step. 4. Each item walks the region independently until it reaches the aggregate. 5. When the aggregate’s policy is met, the region closes and the run continues once, with the results under the aggregate’s `into` field. ![A split and aggregate drawn as a region frame around two per-item steps](/shots/workflow-region.png) ## Split parameters | Param | Builder label | Values | Default | What it does | | -------- | ------------------------------------------------------- | ------------------------------ | -------------- | ------------------------------------------------------------------------------------- | | `source` | Items from: **A list on the item**, then **List field** | `{ list: "" }` | required | The dot path of the list to walk, read from the carried item. | | `mode` | Items run: **One at a time** / **Several at once** | `"sequential"` or `"parallel"` | `"sequential"` | Sequential walks one item at a time. Parallel walks up to 8 items at once per region. | An empty list closes the region at once and the aggregate hands on zero results. ## Aggregate parameters | Param | Builder label | Values | Default | What it does | | -------- | --------------- | ---------------------------------- | --------- | ------------------------------------------------------------------------------------ | | `split` | Closes | a split step’s id | required | The split this aggregate closes. Each split has exactly one aggregate. | | `policy` | Carries on when | `"all"`, `"any"`, `{ atLeast: n }` | `"all"` | When the region closes: every item finished, the first item done, or `n` items done. | | `into` | Save items as | a field name | `"items"` | Where the results land on the carried item. | The aggregate writes this object under `into`: ```json { "items": [ { "itemKey": "split#1#0", "itemIndex": 0, "state": "done", "result": { "category": "billing" } }, { "itemKey": "split#1#1", "itemIndex": 1, "state": "done", "result": { "category": "bug" } } ], "reason": "complete", "count": 2, "total": 2 } ``` `result` is the carried item as it reached the aggregate. `count` is how many items finished `done`; `total` is how many items the list held. Results are in list order. When a policy closes the region early (`any`, `atLeast`), items still queued or running are listed with their state and no result. ## Example Triage a batch of tickets: one model call and one cubby write per ticket, several tickets at once. ```ts nodes: [ { id: "batch", kind: "trigger" }, { id: "each-ticket", kind: "split", params: { source: { list: "tickets" }, mode: "parallel" } }, { id: "classify", kind: "model", params: { alias: "llm", input: { messages: [{ role: "user", content: "=Classify as billing, bug or question:\n{{ $json.text }}" }], max_tokens: 16, }, into: "classification", }, }, { id: "save", kind: "cubbyExec", params: { alias: "triage", sql: "INSERT INTO triage_tickets (run_id, ticket_id, category) VALUES ('{{ $runId }}', ?, ?)", args: ["={{ $json.id }}", "={{ $json.classification.text }}"], }, }, { id: "collect", kind: "aggregate", params: { split: "each-ticket", policy: "all", into: "triaged" } }, { id: "done", kind: "output" }, ], edges: [ { from: "batch", to: "each-ticket" }, { from: "each-ticket", to: "classify" }, { from: "classify", to: "save" }, { from: "save", to: "collect" }, { from: "collect", to: "done" }, ], ``` Start it with a payload like `{ "tickets": [{ "id": "T-1", "text": "I was charged twice" }, { "id": "T-2", "text": "The export button does nothing" }] }`. ## Per-item waits Agent, connector action, and person steps inside a region wait per item. Each item’s answer is matched to that item, so answers can arrive in any order. A loop edge inside a region counts per item, which is how you retry one item without touching the others. A person’s step inside a region: * must be an assigned step (it declares `assignees`, `policy`, `options`, or another assigned-step param), so each answer names its item. See [People in workflows](/agents/workflows/people/). * is allowed only in a **sequential** region. In a parallel region every item would hold a person’s wait open at once. ## Failures A failed item fails the run. The aggregate’s `onItemError` param accepts only `"fail-run"`, which is also the default. ## Rules the graph must follow `cef build` and the runner refuse a graph that breaks any of these, before the run has any effect: | Rule | Why | | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------- | | Every split has exactly one aggregate whose `split` names it. | One region, one place where items come back. | | No split inside a region. | Nested regions are not supported. | | No `join` inside a region. | An item’s branches would join across items. | | No trigger inside a region. | It would start runs from inside one item. | | Edges enter a region only through its split and leave only through its aggregate. | An item that leaves by a side door never reaches the aggregate, so the region never closes. | | A loop edge may not cross the region boundary. | Same reason. | | The split has at least one edge out. | Otherwise no item is walked. | ## Limits | What | Limit | | ----------------------------------------- | --------------- | | Items in flight per region, parallel mode | 8 | | Items in flight, sequential mode | 1 | | Nested regions | not supported | | `onItemError` | `fail-run` only | ## Related * [Next: Workflows in code](/agents/workflows/author-in-code/) * [Steps](/agents/workflows/steps/) * [People in workflows](/agents/workflows/people/) * [Monitor runs](/agents/workflows/monitor-runs/) # Steps > Every workflow step kind: what it does, when to use it, its parameters, and a minimal example, grouped as the Workflow Builder palette groups them. A workflow is built from 18 step kinds. In the Workflow Builder you add them from the **Add node** palette, which groups them by what they do. In code each step is an object in `defineWorkflow`’s `nodes` list with an `id`, a `kind`, and `params`. Every step also takes: | Field | What it is | | ---------- | ---------------------------------------------------------- | | `id` | Unique within the workflow. Edges and run records name it. | | `label` | The name shown on the canvas and in run views. | | `position` | `{ x, y }` on the canvas. The runner ignores it. | Parameters marked as mappings accept `"={{ $json. }}"` to read the carried item. See [the carried item](/agents/workflows/overview/#the-carried-item). | Palette group | Builder step | Kind | | ------------- | ------------------------------------------------- | ------------------------------ | | Triggers | Event, Schedule, Webhook, one entry per connector | `trigger` | | Flow | Branch, Filter | `branch` | | | Join | `join` | | | Split, Aggregate | `split`, `aggregate` | | | Human approval | `human` | | Logic | Edit fields | `transform` | | | Code | `code` | | Agents | Agent | `agent` | | Models | Model | `model` | | Actions | one entry per connector | `action` | | Memory | Memory (Remember, Recall, Relate) | `remember`, `recall`, `relate` | | Cubbies | Cubby (Query, Exec) | `cubbyQuery`, `cubbyExec` | | Outputs | Event | `publish` | | | Result | `output` | ![The Add node palette with its groups open](/shots/node-palette.png) ## Triggers ### Trigger **Kind:** `trigger`. Where a run starts. A workflow needs at least one. The trigger’s output is the event’s payload, which becomes the first carried item. | Builder entry | `params.mode` | Starts a run when | | ------------------------------ | ------------- | ---------------------------------------------------------------------------------------------------- | | **Event** | none | a `workflow.start` event targeted at the workflow arrives: the **Run** button, an app, another agent | | **Schedule** | `"schedule"` | the vault’s clock reaches a cron boundary | | **Webhook** | `"webhook"` | another system calls the workflow’s URL with a key | | a connector (Slack, Telegram…) | `"connector"` | a message arrives through a vault connection | ```ts { id: "ticket", kind: "trigger", label: "New ticket" } ``` Parameters, limits, and payloads for each mode are on [Triggers](/agents/workflows/triggers/). ## Flow ### Branch **Kind:** `branch`. Takes exactly one way out: the first outgoing edge whose condition passes, in edge order. An edge without a condition always passes, so put it last as the fallthrough. In the Builder, a Branch has two outputs. The first carries the run when the condition holds; the second is the fallthrough. In code, conditions sit on the edges: ```ts edges: [ { from: "route", to: "escalate", when: { field: "priority", op: "eq", value: "urgent" } }, { from: "route", to: "queue" }, ] ``` A condition (`when`) tests one field of the item the branch received: | `op` | Builder label | Passes when the field | | ------------ | ----------------------------- | ----------------------------------------- | | `eq` | is | equals `value` | | `ne` | is not | does not equal `value` | | `gt` / `gte` | is greater than / is at least | is numerically greater / greater or equal | | `lt` / `lte` | is less than / is at most | is numerically less / less or equal | | `contains` | contains | is a string containing `value` | | `exists` | has any value | is present and not empty | `field` is a dot path (`classification.category`). A missing field fails every test except `ne`. A condition tests one field; to combine two, chain a second branch. If no edge passes, the run ends `stalled`. Any step, not only a branch, can have conditions on its edges. A non-branch step takes **every** edge whose condition passes, so two unconditional edges out of one step run both paths at once. ### Filter A Builder step that carries on only when its condition holds. It compiles to a `branch` with one conditional edge, so a run the filter stops ends `stalled`. ### Join **Kind:** `join`. Waits until every inbound edge has delivered, then continues once with the branches merged. Use it after a fan-out, for example three agents reviewing the same ticket. The merged item holds every branch’s fields, applied in inbound edge order (later edges win on a clash), plus each branch’s whole item under its step id. | Param | Builder label | Values | | --------- | ------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `require` | When a branch fails | `"all"` (default): a failed branch fails the run. `"any"`: the join continues with the branches that answered, and a failed branch arrives as `{ failed: true, nodeId, error }`. | ```ts { id: "panel", kind: "join", params: { require: "all" } } ``` A step with two inbound edges and no join runs once per arrival. Use a join whenever branches must arrive together. ### Split and Aggregate **Kinds:** `split`, `aggregate`. Run the steps between them once per entry of a list, then collect the results. See [Split and aggregate](/agents/workflows/split-and-aggregate/). ### Human approval **Kind:** `human`. Parks the run until people answer. Supports assignees, closing policies (`any`, `all`, `{ atLeast }`), answer options, forms, and shared document editing. See [People in workflows](/agents/workflows/people/). ```ts { id: "approve", kind: "human", question: "=Refund {{ $json.amount }}?" } ``` ## Logic Logic steps change the item deterministically: the same input always gives the same output, with no model call, no network, and no clock. `fetch`, `Date`, `performance`, `require`, `process`, and `globalThis` are shadowed, so using one fails the step. `item`, `JSON`, and `Math` are in scope. ### Edit fields **Kind:** `transform`. One JavaScript expression over `item`; its value becomes the new item. A non-object result is wrapped as `{ value: … }`. | Builder mode | What it does | | -------------- | ----------------------------------------------------------------------------------------------------------------------- | | **Fields** | Set fields one by one from mappings, and choose under **Other fields** whether to keep the rest of the item or drop it. | | **Expression** | Write the expression yourself (`expr`). | | **Review** | Judge the step before and send it back with an objection. See [Review a step](#review-a-step). | ```ts { id: "normalise", kind: "transform", params: { expr: "({ ...item, customer: item.customer.trim().toLowerCase(), words: item.text.split(/\\s+/).length })" }, } ``` ### Code **Kind:** `code`. A function body with statements, locals, loops, and an explicit `return` of the new item. Same scope and limits as Edit fields; use it when one expression is not enough. ```ts { id: "dedupe", kind: "code", params: { code: ` const seen = new Set(); const tickets = item.tickets.filter((t) => !seen.has(t.id) && seen.add(t.id)); console.log("kept", tickets.length); return { ...item, tickets }; `, }, } ``` * The body runs synchronously. Returning a Promise fails the step; use a model or agent step for anything that waits. * Returning nothing fails the step. * `console.log` lines (up to 100) are recorded with the step’s run record. * A runaway loop holds the step until the Job’s idle timeout. ## Agents ### Agent **Kind:** `agent`. Asks an agent of your Agent Service to do something, parks the run, and continues when the agent answers. The agent brings its own instructions, model, tools, and connections, can take several turns, and can ask a question back. | Param | Builder label | What it does | | -------- | ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `use` | Agent | The agent’s alias in your Agent Service, or a full agent id. Required. In the Builder, **Choose an agent…** picks one, **Change** picks another, and **Open** opens an LLM agent’s definition. | | `prompt` | What to ask it | The request, usually a mapping. Empty: the agent receives the item and works from its own instructions. | | `answer` | Answer it should give | JSON example of the answer’s shape, such as `{"category": "billing"}`. The step fails if the answer lacks any of its keys. | | | Ask the agent for this shape | Builder only, on by default: appends the shape to the request. | | `emit` | | Event type sent to the agent. Default `workflow.step`. | ```ts { id: "classify", kind: "agent", use: "ticket-classifier", params: { prompt: "=Classify this ticket:\n{{ $json.text }}", answer: JSON.stringify({ category: "billing", priority: "normal" }), }, } ``` The agent receives the carried item with the resolved params (including `prompt`) on top, plus `runId` and `nodeId`. **Which agents a step can call.** Any agent of your Agent Service: * an [LLM agent](/agents/llm-agents/overview/): one created in ROC with **New agent**, or an A2A agent you [bring yourself](/agents/llm-agents/bring-your-own/); * a [code agent](/agents/code-agents/overview/): the handler for the step’s event type (`workflow.step` by default) returns the answer. Every agent a workflow calls must be connected to the vault too. Deploying and running from ROC connects them with the workflow. **Reading the answer.** The answer becomes fields: * An answer that is an object is used as is. * A text answer is kept as `text` and `output`, and the JSON object in it (fenced in ` ```json `, or embedded in prose) is parsed and its fields added. The fields are merged over the carried item; on a clash the answer wins. With `answer` declared, an answer missing any declared key fails the step, naming the missing keys and what the agent produced. That keeps a wrong answer from silently taking the wrong branch three steps later. Declare `answer` for every field a later step or branch depends on; see [Its answer](/agents/llm-agents/overview/#its-answer). **When the agent asks a question.** An agent that ends its turn needing input (A2A `input-required`) does not move the run on. The run parks with the agent’s question; someone answers it from the run in ROC, and the answer goes back to the agent in the same A2A context, so it keeps what it had worked out. An agent needing authorization (`auth-required`) parks the same way until someone grants it and runs the step again. A step gives up after 12 exchanges without a final answer. **When the agent fails.** A failed agent fails the run at that step, unless the step feeds a join set to carry on with the branches that answered. See [Join](#join). ## Models ### Model **Kind:** `model`. Calls one model once and puts the answer on the item under `into`; the rest of the item is kept. No back-and-forth, no tools, no memory of its own. What models are, the catalogue, and the input fields are on [Models](/agents/building-blocks/models/). | Use a model step when | Use an agent step when | | ---------------------------------------------------------------------- | -------------------------------------------------------------------------- | | The task is one call: transcribe, embed, classify, extract, summarise. | The task needs judgement across several turns, tools, or connections. | | You want each call visible and re-runnable as its own step. | You want to reuse an agent other workflows and apps also use. | | You can write the whole request from the item. | The agent’s instructions are long or change independently of the workflow. | | Param | Builder label | Default | What it does | | ------- | ------------------ | --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `alias` | Model | required | The model alias. The workflow must declare it in `models`; `cef build` refuses a model step whose alias is not there. See [Declare an alias](/agents/building-blocks/models/#declare-an-alias). | | `input` | Request | `{}` | The request object, in the model’s own input schema. Values can be mappings, at any depth. | | `into` | Answer field | `output` | Where the answer lands. | | `each` | Call once per item | `false` | Call once per entry of a list. | | `over` | List field | the item itself | The list field to iterate when `each` is on. | ```ts { id: "classify", kind: "model", label: "Classify the ticket", params: { alias: "llm", input: { messages: [ { role: "system", content: "You triage support tickets." }, { role: "user", content: "=Classify as billing, bug or question:\n\n{{ $json.text }}" }, ], max_tokens: 64, }, into: "classification", }, } ``` A language model answers with `{ text }`, so the next step reads `$json.classification.text`. Set `max_tokens`: long answers are cut off at the model’s default budget. `input` can also be JSON text of an object (what the Builder’s **Request** editor saves), or a single mapping such as `"={{ $json.request }}"` that resolves to an object an earlier step built. Anything else fails the step before the model is called. **Structured output.** Ask a language model for JSON constrained to a schema with `response_format`: ```ts input: { messages: [{ role: "user", content: "=Classify this ticket:\n{{ $json.text }}" }], max_tokens: 128, response_format: { type: "json_schema", schema: { type: "object", properties: { category: { type: "string", enum: ["billing", "bug", "question"] }, priority: { type: "string", enum: ["low", "normal", "urgent"] }, }, required: ["category", "priority"], }, }, }, into: "classification", ``` The answer is still text. Parse it in the next step so later steps and branches can read fields: ```ts { id: "parse", kind: "transform", params: { expr: "({ ...item, ...JSON.parse(item.classification.text) })" } } ``` `response_format: { "type": "json_object" }` asks for any valid JSON without a schema. See [Structured output](/agents/code-agents/structured-output/). **Once per list entry.** With `each: true` the step calls the model once per entry of `over`, one after another, and collects the answers as a list under `into` (default `results`). Each answer carries the `itemIndex` of the entry that produced it. Each entry’s fields are spread over the item for that call, so `={{ $json.url }}` reads the current entry’s `url`; an entry that is not an object is `$json.item`. ```ts { id: "transcribe", kind: "model", params: { alias: "asr", each: true, over: "chunks", input: { audio: "={{ $json.url }}" }, into: "transcripts" }, } ``` If `over` is missing or not a list, the step fails naming the field. For several steps per entry, or waits per entry, use [Split and aggregate](/agents/workflows/split-and-aggregate/). **Failures.** A model call that fails in transport (a 5xx, a timeout, a dropped connection) is retried up to 3 attempts in total, after 0.4 s and 1.6 s. A 4xx (a bad request, an undeclared alias) fails the step at once. ![A Model step: model, request, and answer settings](/shots/model-step.png) ## Actions ### Connector action **Kind:** `action`. Does one thing in an external system through a vault **connection**: post to Slack, send an email, message a Telegram chat. The vault holds the connection’s credentials and makes the call; the workflow only names the connection, the action, and the input. Adding the step in the Builder, and the access Deploy grants, are on [Connectors](/vaults/connectors/#use-from-a-workflow). | Param | What it does | | ------------ | --------------------------------------------------------------------------------------------------------------------------- | | `connection` | The connection’s id. Required. | | `action` | The action name, such as `send_message`. Required. | | `input` | An object with one key per input field of the action. Values can be mappings; a value that resolves to nothing is left out. | ```ts { id: "notify", kind: "action", label: "Tell the support channel", params: { connection: "", action: "send_message", input: { channel: "C0123456", text: "=New {{ $json.category }} ticket {{ $json.ticketId }}: {{ $json.summary }}", }, }, } ``` | Connector | Actions | Events (for [connector triggers](/agents/workflows/triggers/#connector)) | | ------------ | ----------------------------------------------------------- | ------------------------------------------------------------------------ | | Slack | `send_message`, `send_dm`, `update_message`, `add_reaction` | `message.received` | | Telegram | `send_message` | `message.received` | | Email (SMTP) | `send_email` | none | | MCP server | discovered from the connected server | none | The exact input fields of each action are shown in the step’s panel. The vault’s connector catalogue is public: `GET /api/v1/connectors`. The step parks the run until the vault answers, like an agent step. On success, the action’s output fields are merged over the carried item. On failure, the run fails at the step with the vault’s error code and message: | Error code | Meaning | | ---------------------- | --------------------------------------------------------------------- | | `connection_not_found` | No such connection in this vault. | | `connection_revoked` | The connection was revoked. | | `scope_not_allowed` | The connection is not usable in the workflow’s scope. | | `no_access` | The workflow has no grant for this action, or is no longer connected. | | `action_unknown` | The connector has no such action. | | `invalid_input` | The input does not match the action’s fields. | | `provider_error` | The external system refused or failed. | | `timeout` | The external system did not answer in time. | A person answering the run can send it back to a connector action step to try again, for example after the vault owner reconnects an account. See [Send a run back](/agents/workflows/people/#send-a-run-back). ## Memory Memory steps read and write the vault’s [Memory Bank](/vaults/memory-bank/): typed, durable records that later runs and other workflows can read, subject to the connection’s grants. In the Builder they are one **Memory** step with an **Operation** of Remember, Recall, or Relate. Writes are checked by reading back through the same grants. If the vault’s connection does not allow the scope or privacy level, the step fails and says so instead of silently writing nothing. ### Remember **Kind:** `remember`. Writes one record. Idempotent on `id`: writing the same id again updates the record instead of adding a second. | Param | Builder label | Default | What it does | | --------- | -------------- | ---------------------- | ---------------------------------------------------- | | `scope` | Scope | required | The grant scope the record lives in. | | `privacy` | Privacy | | `public`, `internal`, `private`, or `restricted`. | | `type` | Record type | `record` | The record’s type. | | `id` | One record per | `:` | The record id; a mapping makes one record per value. | | `title` | Title | the id | A mapping or text. | | `body` | | the whole item as JSON | A mapping or text. | ```ts { id: "file", kind: "remember", params: { scope: "default", privacy: "internal", type: "ticket-triage", id: "={{ $json.ticketId }}", title: "={{ $json.summary }}" }, } ``` ### Recall **Kind:** `recall`. Searches the Memory Bank and adds the matching records to the item. | Param | Builder label | Default | What it does | | -------------- | ------------- | ---------- | --------------------------------------------------------- | | `match` | Search for | required | Search text; usually a mapping. | | `limit` | Most rows | 50 | Most records returned. | | `into` | Field name | `recalled` | Where the records land; `Count` holds their number. | | `bodies` | | `false` | Also read each record’s body. | | `maxBodyChars` | | 2000 | Longest body kept per record when `bodies` is on. | ### Relate **Kind:** `relate`. Draws a typed, directed edge between two records. | Param | Builder label | Default | What it does | | ------------------ | -------------- | --------- | ------------------------------------ | | `from` | From record | required | Source record id; usually a mapping. | | `to` | To record | required | Target record id. | | `type` | Edge type | `related` | The relation’s type. | | `scope`, `privacy` | Scope, Privacy | | As for Remember. | ## Cubbies ### Cubby query and Cubby exec **Kinds:** `cubbyQuery`, `cubbyExec`. Read or write a SQL [cubby](/agents/building-blocks/cubbies/) your Agent Service declares. In the Builder both are one **Cubby** step; its **Operation** is **Query — read rows back** or **Exec — write, or change the schema**. | Kind | Does | Adds to the carried item | | ------------ | --------------------------------------------------------- | ------------------------------------------------------------------------- | | `cubbyQuery` | Runs a `SELECT` and returns rows. | `` (the rows, a list) and `Count`. `into` defaults to `rows`. | | `cubbyExec` | Runs an `INSERT`, `UPDATE`, `DELETE`, or other statement. | `cubbyChanged`: the number of rows changed. | | Param | Builder label | What it is | | ------- | ------------- | -------------------------------------------------------------------------------------------------------------- | | `alias` | Cubby | The cubby to use. Required. | | `sql` | SQL | One SQL statement, with `?` for every value. Required. | | `args` | Values | One value per `?`, in order (one per line in the Builder). Usually mappings such as `"={{ $json.ticketId }}"`. | | `into` | Rows land on | `cubbyQuery` only: the field the rows land on. | ```ts { id: "save", kind: "cubbyExec", label: "Record the triage result", params: { alias: "triage", sql: "INSERT INTO triage_tickets (run_id, ticket_id, category, ts) VALUES ('{{ $runId }}', ?, ?, ?)", args: ["={{ $json.ticketId }}", "={{ $json.category }}", "={{ $json.receivedAt }}"], }, }, { id: "history", kind: "cubbyQuery", label: "Earlier tickets from this customer", params: { alias: "triage", sql: "SELECT ticket_id, category FROM triage_tickets WHERE customer = ? ORDER BY ts DESC LIMIT 20", args: ["={{ $json.customer }}"], into: "previous", }, }, ``` After `history`, `$json.previous` holds the rows and `$json.previousCount` their number. **Values go in `args`, never in the SQL.** A `{{ $json.… }}` inside `sql` fails the step: the run’s data comes from outside (a webhook body, a Slack message) and splicing it into SQL is an injection. Put `?` in the statement and the value in `args`, where the platform binds it. An `args` value that resolves to an object is sent as its JSON text; a list of plain values is sent joined with `"; "`. **`{{ $runId }}`** is the one template allowed inside `sql`: the run’s own id. Write it inside quotes, as in the example. Store it on every row a run writes, so a [widget pinned to the workflow](/agents/building-blocks/widgets/#with-a-workflow) can select that run’s rows with `WHERE run_id = ?`. Declaring a cubby and inspecting its data are on [Cubbies](/agents/building-blocks/cubbies/); migration rules are on [Cubby schema and migrations](/agents/building-blocks/cubby-schema/). A run’s view shows the cubby calls each step made; see [Monitor runs](/agents/workflows/monitor-runs/). ## Outputs ### Event (publish) **Kind:** `publish`. Publishes an event into the run’s scope, so another workflow, agent, or app subscribed to that type receives the result. This is how workflows compose: one publishes, the other is triggered on its own terms. | Param | Builder label | Default | What it does | | --------- | ---------------------- | ----------------- | -------------------------------------------------------------------------------- | | `event` | Event type | `workflow.result` | The event type. Changing it breaks subscribers. | | `payload` | What the event carries | the whole item | A JSON example object as text. Only its keys are published, taken from the item. | ```ts { id: "announce", kind: "publish", params: { event: "ticket.triaged", payload: JSON.stringify({ ticketId: "", category: "", priority: "" }) }, } ``` The published payload also carries `runId`. Declare `payload` so a publish never broadcasts the whole item. ### Result **Kind:** `output`. Ends a branch of the run and declares what the run returns: the webhook response, and what evaluations score. | Param | Builder label | What it does | | ------------- | --------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | | `result` | Where each field comes from | `{ field: mapping }`. The run’s output becomes exactly these fields. | | `resultTypes` | | `{ field: "string" \| "number" \| "boolean" \| "object" \| "array" }`. A resolved value of another type fails the run at this step. | | `outcome` | Outcome | An optional label for how the run ended, such as `"escalated"`; can be a mapping. | ```ts { id: "result", kind: "output", params: { result: { category: "={{ $json.category }}", priority: "={{ $json.priority }}" }, resultTypes: { category: "string", priority: "string" }, outcome: "triaged", }, } ``` Without `result`, the run’s output is the carried item as it reached the step. A step with no outgoing edges also ends the run with the item it produced. ## Review a step An Agent or Edit fields step can judge an earlier Agent step and send it back to try again. Set: | Param | Builder label | Default | What it does | | ------------- | ------------------------------------ | ---------- | ----------------------------------------------------------------------------------- | | `reviews` | Reviews the step / Judges which step | | The step to judge, or `"."` for the step drawn into this one. | | `feedback` | Objection field | `feedback` | The field of this step’s output holding the objection. Empty means accepted. | | `maxAttempts` | Attempts | 3 | Total tries for the judged step, the first included. | | `accept` | | | Optional condition (`{ field, op, value }`) that accepts the answer when it passes. | When the reviewer objects, the judged step is asked again with the objection as `answer` on its input. When attempts run out, the run does not fail and does not ship the rejected answer: it waits for a person, with the last objection as the question. ```ts { id: "check", kind: "transform", params: { expr: "({ ...item, feedback: ['billing', 'bug', 'question'].includes(item.category) ? '' : 'category must be billing, bug or question' })", reviews: "classify", maxAttempts: 2, }, } ``` ## Validation The Builder’s Deploy, `cef build`, and the runner check a graph before it can run and refuse it with every problem listed: unknown kinds, duplicate ids, edges to unknown steps, a missing trigger, an agent step with no agent, a model step with no alias, a connector step with no connection or action, unmarked templates, cycles without a loop edge, invalid human-step params, and invalid regions. A step nothing leads to is reported but does not block. ## Limits | What | Limit | | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | Input the run carries into an agent step | 1 MiB (1,048,576 bytes). A larger item fails the step. | | Exchanges in one agent step (agent asks, person answers) | 12. The step then fails: “gave up after 12 exchanges without a final answer”. | | Model call attempts | 3, transport failures only, 400 ms then 1600 ms apart | | Connector action: time per attempt | 20 s. The attempt fails and is retried. Keep actions small. | | Connector action: attempts | 5, with backoff. The step then fails with `connector.action.failed`; handle the failure path in the graph. | | SQL in a cubby step | No `{{ $json.… }}` templates; only `{{ $runId }}`. Otherwise the step fails: “SQL may not interpolate the run’s data”. | ## Related * [Next: Triggers](/agents/workflows/triggers/) * [Workflows](/agents/workflows/overview/) * [LLM agents](/agents/llm-agents/overview/) * [Models](/agents/building-blocks/models/) * [Cubbies](/agents/building-blocks/cubbies/) # Test a workflow > The three-rung testing ladder for a workflow: local tests with @cef-ai/testing and vitest, dataset experiments against a deployed version, and a daily live check of the deployed workflow. Test a workflow on three rungs. Each one catches what the one below cannot. | Rung | Runs | Catches | When | | ---------------------- | ------------------------------------- | -------------------------------------------------------------------------------------------------------------------- | --------------------------------- | | 1. Local tests | In-process, models mocked | Wrong graph wiring, bad SQL, broken transforms, payloads over budget, bad model output not handled | Every commit | | 2. Dataset experiments | On the platform, real models | Quality drift: a prompt, model, or step change that makes the Result worse, slower, or costlier | Before switching the live version | | 3. Live check | The deployed workflow, a real trigger | Deployment, connection, and data problems: the trigger does not fire, a cubby is not written, a widget shows nothing | After every deploy, and daily | ## 1. Local tests `@cef-ai/testing` runs the real workflow runner in-process against an in-memory vault and SQLite cubbies built from your real migrations. You mock only what leaves the workflow: models and other agents. ```sh pnpm add -D @cef-ai/testing vitest ``` The example is a support-ticket triage workflow: a model classifies the ticket, a transform keeps only known categories, a cubby step records the triage, a publish step tells the support queue, and a Result step ends the run. ```plaintext ticket-triage/ cef.config.ts cubbies/tickets/001-triage.sql test/triage.test.ts ``` ### The workflow cef.config.ts ```ts import { fileURLToPath } from "node:url"; import { defineWorkflow } from "@cef-ai/agent-sdk/config"; const CATEGORIES = ["billing", "bug", "account", "other"]; export default defineWorkflow({ id: "ticket-triage", version: "1.0.0", models: { classifier: "https://cdn.example.com/models/classifier/model.json" }, cubbies: [{ alias: "tickets", migrations: fileURLToPath(new URL("./cubbies/tickets", import.meta.url)) }], nodes: [ { id: "start", kind: "trigger", label: "Ticket arrives", position: { x: 0, y: 0 }, params: { eventType: "workflow.start" }, }, { id: "classify", kind: "model", label: "Classify the ticket", position: { x: 240, y: 0 }, params: { alias: "classifier", input: { text: "={{ $json.text }}", labels: CATEGORIES }, into: "classification", }, }, { id: "normalize", kind: "transform", label: "Keep only a known category", position: { x: 480, y: 0 }, params: { expr: `({ ticketId: item.ticketId, category: ${JSON.stringify(CATEGORIES)}.includes(item.classification?.category) ? item.classification.category : "other", confidence: typeof item.classification?.confidence === "number" ? item.classification.confidence : 0 })`, }, }, { id: "save", kind: "cubbyExec", label: "Record the triage", position: { x: 720, y: 0 }, params: { alias: "tickets", sql: "INSERT OR IGNORE INTO triage_results (run_id, ticket_id, category, confidence) VALUES (?, ?, ?, ?)", args: ["={{ $runId }}", "={{ $json.ticketId }}", "={{ $json.category }}", "={{ $json.confidence }}"], }, }, { id: "announce", kind: "publish", label: "Tell the support queue", position: { x: 960, y: 0 }, params: { event: "ticket.triaged", payload: JSON.stringify({ ticketId: "", category: "" }), }, }, { id: "result", kind: "output", label: "Result", position: { x: 1200, y: 0 }, params: { result: { category: "={{ $json.category }}", confidence: "={{ $json.confidence }}" }, resultTypes: { category: "string", confidence: "number" }, outcome: "={{ $json.category }}", }, }, ], edges: [ { from: "start", to: "classify" }, { from: "classify", to: "normalize" }, { from: "normalize", to: "save" }, { from: "save", to: "announce" }, { from: "announce", to: "result" }, ], }); ``` ```sql -- cubbies/tickets/001-triage.sql CREATE TABLE IF NOT EXISTS triage_results ( run_id TEXT PRIMARY KEY, ticket_id TEXT NOT NULL, category TEXT NOT NULL, confidence REAL NOT NULL ); ``` ### The test The test registers the workflow runner as an agent with the workflow’s own cubbies and its compiled graph as the deployment param, exactly as the platform does. A run starts with a `workflow.start` event, the same event an experiment publishes for a case. test/triage.test.ts ```ts import { afterEach, describe, expect, it } from "vitest"; import { createModelMock, testPlatform, type ModelMockHandle, type TestPlatform } from "@cef-ai/testing"; import { WorkflowRunner, type WorkflowDoc } from "@cef-ai/agent-sdk/workflow"; import config from "../cef.config.js"; // The compiled graph, exactly as `cef build` ships it. const graph = JSON.parse((config.params!.graph as { default: string }).default) as WorkflowDoc; const LABELS = ["billing", "bug", "account", "other"]; const bytes = (v: unknown) => Buffer.byteLength(JSON.stringify(v ?? null), "utf8"); async function setup(classifier: ModelMockHandle): Promise { const p = testPlatform({ agents: { "ticket-triage": { source: WorkflowRunner, cubbies: config.cubbies, params: { graph: config.params!.graph!.default }, }, }, models: { classifier }, }); await p.vault.agents.connect({ agentId: "ticket-triage" }); return p; } const start = (p: TestPlatform, context: string, payload: Record) => p.vault.scope("default").publish({ type: "workflow.start", context, payload }); async function events(p: TestPlatform, context: string, type: string) { const page = await p.vault.scope("default").stream(context).events.list>({ types: [type] }); return page.items.map((e) => e.payload); } const triageRows = (p: TestPlatform) => p.runInCubby("ticket-triage", "tickets", (cubby) => cubby.query("SELECT * FROM triage_results ORDER BY run_id")); describe("ticket-triage", () => { let p: TestPlatform | undefined; afterEach(() => p?.dispose()); it("classifies, records one row, publishes a declared payload and ends with the Result", async () => { const classifier = createModelMock(); classifier.expect({ text: "I was charged twice", labels: LABELS }).respond({ category: "billing", confidence: 0.93 }); p = await setup(classifier); await start(p, "t-1", { ticketId: "T-1", text: "I was charged twice" }); expect(await events(p, "t-1", "workflow.completed")).toEqual([ { runId: "run-t-1", output: { category: "billing", confidence: 0.93 }, outcome: "billing" }, ]); expect(await triageRows(p)).toEqual([ { run_id: "run-t-1", ticket_id: "T-1", category: "billing", confidence: 0.93 }, ]); const [announced] = await events(p, "t-1", "ticket.triaged"); expect(announced).toEqual({ ticketId: "T-1", category: "billing", runId: "run-t-1" }); }); it("downgrades an unknown category to 'other' instead of failing", async () => { const classifier = createModelMock(); classifier.expect({ text: "hello", labels: LABELS }).respond({ category: "spam!!", confidence: "high" }); p = await setup(classifier); await start(p, "t-2", { ticketId: "T-2", text: "hello" }); const [done] = await events(p, "t-2", "workflow.completed"); expect(done!.output).toEqual({ category: "other", confidence: 0 }); }); it("writes one row when the same start is delivered twice", async () => { const classifier = createModelMock(); classifier.expect({ text: "app crashes", labels: LABELS }).respond({ category: "bug", confidence: 0.8 }); p = await setup(classifier); await start(p, "t-3", { ticketId: "T-3", text: "app crashes" }); await start(p, "t-3", { ticketId: "T-3", text: "app crashes" }); expect(await triageRows(p)).toHaveLength(1); expect(classifier.calls).toHaveLength(1); }); it("stays inside its budgets", async () => { const text = "x".repeat(20_000); const classifier = createModelMock(); classifier.expect({ text, labels: LABELS }).respond({ category: "bug", confidence: 0.5 }); p = await setup(classifier); await start(p, "t-4", { ticketId: "T-4", text }); expect(graph.nodes.length).toBeLessThanOrEqual(8); for (const n of graph.nodes) expect(n.position, `${n.id} has a position`).toBeDefined(); expect(classifier.calls).toHaveLength(1); const [done] = await events(p, "t-4", "workflow.completed"); expect(bytes(done!.output)).toBeLessThanOrEqual(1024); const [announced] = await events(p, "t-4", "ticket.triaged"); expect(bytes(announced)).toBeLessThanOrEqual(1024); }); }); ``` ```sh pnpm vitest run ``` What each test pins down: | Test | Guards | | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ | | Happy path | The model is called with the mapped input; the cubby row is exactly right; the publish step sends only its declared fields; the Result and outcome are exact | | Bad model output | Invalid output is downgraded deterministically, so the run still completes with a valid Result | | Same start twice | A redelivered start does not run the workflow twice or write two rows | | Budgets | Step count, positions, model calls per run, Result and payload sizes stay under the numbers you chose | ### What else to assert * **Failures.** For input the workflow must refuse, assert the run failed and nothing was written: read `workflow.failed` from the context, and query your cubby for zero rows. The runner’s own `runs` cubby holds one row per run with `status` and `error`: `p.runInCubby("ticket-triage", "runs", (c) => c.query("SELECT status, error FROM runs WHERE run_id = ?", ["run-t-1"]))`. * **Which step wrote what.** `p.cubbyOps()` lists every cubby statement with its `cubby`, `op`, `sql`, `params`, and the `nodeId` of the step that made it. * **Agent steps.** Register each agent a step asks as a `mockAgent` that answers `workflow.step` by publishing `agent.answered` with `{ agent, text }` (or `{ agent, error }`). See the [testing reference](/reference/testing/#mockagent). * **Publish failures.** Set `p.faults.refusePublish` to make a publish fail as the vault does for a body over 1 MiB, and assert the run fails visibly. * **Fixtures.** Keep inputs synthetic: build them with small factory functions in `test/fixtures/`, never from real customer data. Local tests do not prove a real model returns what your mock does. That is the next rung. ## 2. Dataset experiments Once the workflow is pushed, run its dataset against the version before it goes live: ```sh cef eval run triage --workflow-version 1.1.0 --repeats 3 ``` Start the dataset from the cases your local tests already use (the same start payloads), then grow it with real runs captured from **Executions** with **Add to dataset**. Judge every candidate version against a baseline experiment of the live version. See [Experiments](/agents/workflows/evaluate/experiments/). ## 3. Live check Run this checklist on each deployed workflow after every deploy, and daily on workflows people depend on. It takes a few minutes per workflow. **Prepare (once per workflow)** * [ ] Write down a synthetic **check input**: a start payload that exercises the main path and is safe to run in the vault (no real customer data, no outbound message to a real person). * [ ] Write down what a good run looks like: the expected Result, the cubby row(s) it writes (table, key, columns), the events it publishes, and what the widget shows for it. * [ ] Make sure the check input is idempotent: give it a fixed id so a repeated check overwrites or ignores, not duplicates. **Trigger** * [ ] Open the workflow in ROC and confirm the **Deployed** chip shows the version you expect. * [ ] Start a run the way real traffic does: send the trigger event, call the webhook with an `Idempotency-Key`, or wait for the schedule. Use **Run** in ROC only when the real trigger cannot be fired safely. **Watch the run in Executions** * [ ] The run appears at the top of **Executions** within a few seconds. If it does not, the trigger did not fire: see [The run never starts](/agents/workflows/monitor-runs/#the-run-never-starts). * [ ] It ends **Done**. **Failed** shows the step that failed and why. **Waiting** names the step it is parked at; that is expected only at a human step. * [ ] Every step you expect ran, in order. Open each model and agent step and check its output looks right. * [ ] **Took**, **Steps**, **Model calls**, **Tokens**, and **Cubby calls** are in line with previous runs. A sudden jump is a regression even if the run is **Done**. **Inspect cubby data** * [ ] On a Cubby step of the run, select **Open cubby** and query the row the run wrote: `SELECT * FROM triage_results WHERE run_id = 'run-'`. * [ ] Every column you expect is filled; no `NULL` where a value belongs; timestamps are this run’s. * [ ] Exactly one row per run: no duplicates from a retried step. **Verify the widget** * [ ] Open the widget where it is pinned and find this run’s record. * [ ] It shows the values in the cubby row, not a placeholder or an empty state. * [ ] Reload the page: the record is still there. **Record** * [ ] Note the run id, the version, and pass or fail. On a fail, add the run to the workflow’s dataset with **Add to dataset** once the fix is known, so an experiment catches it next time. ## Related * [Next: Evaluations](/agents/workflows/evaluate/overview/) * [Best practices](/agents/workflows/best-practices/) * [`@cef-ai/testing` reference](/reference/testing/) * [Monitor runs](/agents/workflows/monitor-runs/) # Triggers > Start workflow runs from an event, a schedule, a webhook call, or a message arriving through a vault connector, in the Builder or in defineWorkflow. A trigger step is where a run starts. The event that starts it becomes the run’s first carried item. A workflow can have several triggers, for example a schedule for the nightly batch and an Event trigger for testing by hand. Each trigger mode is told apart by how its event arrives, so they never compete. | Mode | Builder entry | Starts a run when | | --------- | ---------------------------------------- | --------------------------------------------------------- | | Event | **Event** | a `workflow.start` event targeted at the workflow arrives | | Schedule | **Schedule** | the vault’s clock reaches a cron boundary | | Webhook | **Webhook** | another system calls the workflow’s URL with a key | | Connector | the connector’s name, under **Triggers** | a message arrives through a vault connection | Every trigger fires only in vaults the workflow is connected to. See [Connect to a vault](/agents/ship/connect-to-a-vault/). ## Event The default trigger. A run starts when a `workflow.start` event targeted at the workflow is published into the scope the workflow is connected on. The event’s payload is the first item. Three things publish it: * the **Run** button on the workflow’s **Executions** tab in ROC; * an app, through the vault SDK; * another agent or workflow. ```ts // From an app, with @cef-ai/vault-sdk await vault.scope("default").publish({ type: "workflow.start", context: `ticket-${ticketId}`, // one run per context target: ":ticket-triage", payload: { ticketId, text, customer }, }); ``` The `context` names the run: a second `workflow.start` in a context that already has a run is ignored. Use a fresh context per run, or a deterministic one (such as the ticket id) to make retries safe. | Param | Builder label | What it does | | ----------- | --------------- | ------------------------------------------------------------------ | | `eventType` | Event type | Keep `workflow.start`. | | `sample` | Example payload | JSON text that prefills the Run dialog’s **What starts it** field. | ```ts { id: "ticket", kind: "trigger", params: { eventType: "workflow.start", sample: JSON.stringify({ ticketId: "T-1", text: "I was charged twice" }) } } ``` A `workflow.start` payload may also carry `participants` (see [People in workflows](/agents/workflows/people/#participants)) and `from: ""` to start the run at a chosen step instead of the trigger. ## Schedule The vault starts a run on a cron schedule, once per connected vault: a workflow connected to three vaults fires three times at each boundary, each run in its own vault. ### In the Builder Add **Schedule** from **Triggers** and choose a cadence under **Every**: Every N minutes, Every N hours, Daily, Weekly, Monthly, or Custom cron. Pick a **Timezone** and, under **Input — sent with every fire**, the fields every run receives. The panel shows the next fire time. Deploy installs the schedule. ![The Schedule trigger: cadence, timezone, and the next runs](/shots/trigger-schedule.png) ### In code A schedule is a trigger step with `mode: "schedule"` plus an entry in `schedules` whose `id` is that step’s id: ```ts defineWorkflow({ id: "ticket-digest", version: "0.1.0", nodes: [ { id: "nightly", kind: "trigger", params: { mode: "schedule" } }, // … ], edges: [/* … */], schedules: [ { id: "nightly", cron: "0 6 * * 1-5", timezone: "Europe/Amsterdam", eventType: "workflow.start", payload: { window: "24h" }, }, ], }); ``` | Field | What it is | | ----------- | ------------------------------------------------------------------------------------------------------------- | | `id` | The schedule trigger’s step id. Stable across versions: renaming it removes one schedule and adds another. | | `cron` | Five fields (minute hour day-of-month month day-of-week), or a descriptor such as `@daily`. No seconds field. | | `timezone` | IANA name. UTC when omitted. | | `eventType` | `workflow.start`. | | `payload` | Merged into every fire’s payload. | ### What a run receives The declared `payload` plus a `schedule` block: ```json { "window": "24h", "schedule": { "id": "nightly", "firedAt": "2026-10-08T04:00:00Z", "previousFiredAt": "2026-10-07T04:00:00Z", "cron": "0 6 * * 1-5", "timezone": "Europe/Amsterdam" } } ``` ### Behavior * A fire that is missed (the vault was not running at that time) is skipped, never backfilled. The schedule continues from the next boundary. * A schedule with an invalid cron or timezone does not stop the deploy; it stays silent and records the error on the schedule. * Pressing **Run** on a workflow whose only trigger is a schedule enters at the schedule trigger, so you can test it by hand. ### Schedule limits | What | Limit | | ----------- | --------------------------------------------------------------------------------------------------------------------- | | Granularity | 1 minute: a 5-field cron, or an `@hourly` / `@daily`-style descriptor. An expression with a seconds field is refused. | | Late fire | A fire more than 4 minutes late is skipped. Make each run cover “since the last run”, not “the last minute”. | ## Webhook Another system calls the workflow’s URL with a key, and each call starts a run. The caller can get an answer at once or wait for the run’s result. ### Set it up 1. Add **Webhook** from **Triggers**. Choose **Respond**: **Immediately** or **Wait for result**, and a **Timeout** (seconds) for waiting calls. 2. Deploy. The trigger’s panel shows the **Webhook URL** for the vault the workflow runs in. 3. Under **Keys**, choose **Create key**, give it a **Name** and an optional expiry. Copy the key: it is shown once. A run started with a key **runs as the member who created the key**. If that member loses write access to the workflow’s scope, calls are refused. Revoke a key from the same panel; revocation is immediate. ![The Webhook URL panel with one key and the Copy as menu](/shots/trigger-webhook.png) ### Call it ```bash curl -X POST "$WEBHOOK_URL?wait=true" \ -H "Authorization: Bearer $CEF_WEBHOOK_KEY" \ -H "Content-Type: application/json" \ -H "Idempotency-Key: ticket-T-1042" \ -d '{"ticketId": "T-1042", "text": "The export button does nothing"}' ``` The **Copy as** menu on the URL gives the same call as cURL, JavaScript, Python, a prompt for an AI agent, or a tool definition. | Request part | Rules | | ---------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | `Authorization` | `Bearer `. | | Body | A JSON object becomes the run’s item. A JSON array arrives as `{ "body": [...] }`. An empty body is `{}`. Anything else is refused. | | `?wait=true` / `?wait=false` | Overrides the trigger’s Respond setting for this call. | | `Idempotency-Key` | Optional, up to 255 bytes. Repeating a call with the same key and the same value within 5 minutes returns the same run instead of starting another. | The run’s item is the body plus a `webhook` block: `{ id, endpointId, keyId, keyName, receivedAt }`. A body field named `webhook` is overwritten. ### Responses | Situation | Status | Body | | ------------------------------- | ------ | ----------------------------------------------------------------- | | Respond immediately | 202 | `{ "status": "accepted", "runId", "context", "statusUrl" }` | | Waiting, run finished | 200 | `{ "status": "completed", "runId", "context", "output" }` | | Waiting, run failed | 200 | `{ "status": "failed", "runId", "context", "error" }` | | Waiting, run parked on a person | 202 | `{ "status": "awaiting_input", "runId", "context", "statusUrl" }` | | Waiting, timeout reached | 202 | `{ "status": "running", "runId", "context", "statusUrl" }` | Poll `statusUrl` with the same `Authorization` header until `status` is `completed` or `failed`. `output` is what the workflow’s **Result** step declares; without one, it is the item the run ended with. ### Errors Error bodies are `{ "error": { "code", "message" } }`. | Status | Code | Cause | | ------ | --------------------------- | ------------------------------------------------------------------------------------------------------ | | 400 | `BAD_REQUEST` | Body is not JSON, or an invalid `wait` or `Idempotency-Key`. | | 401 | `WEBHOOK_UNAUTHORIZED` | Unknown URL, or a missing, wrong, revoked, or expired key. Always the same answer, whatever the cause. | | 403 | `WEBHOOK_CREATOR_NO_ACCESS` | The key’s creator can no longer write the workflow’s scope. | | 413 | `PAYLOAD_TOO_LARGE` | Body over 1 MiB. | | 429 | `RATE_LIMITED` | Over the key’s rate limit. Wait `Retry-After` seconds. | | 503 | `INGEST_UNAVAILABLE` | The run could not be started. Retry. | ### Webhook limits | What | Limit | | ----------------- | ------------------------------------- | | Body | 1 MiB | | Wait timeout | default 30 s, maximum 120 s | | Rate per key | 60 calls per minute, bursts of 10 | | `Idempotency-Key` | 255 bytes; deduplicated for 5 minutes | | Key name | 100 characters | The URL and keys stay the same across redeploys as long as the trigger’s step id does not change. Removing the trigger deletes its URL and keys. Webhook URLs are installed by Deploy in the Builder from the graph’s Webhook triggers; `defineWorkflow` has no webhook declaration. ## Connector A message arriving through a vault connection (a Slack message, a Telegram message) starts a run. Connections are set up by the vault owner on the **Connectors** page; see [Connectors](/vaults/connectors/). ### In the Builder Under **Triggers**, pick the connector (for example Slack), then its event. Choose the **{System} connection** and, optionally, a filter such as a channel; leave it empty to start on every message the connection receives. Deploy grants the workflow access to that connection’s event. ### In code ```ts { id: "slack-in", kind: "trigger", params: { mode: "connector", connection: "", event: "message.received" }, } ``` `connection` and `event` are required. Several connector triggers can share one workflow; a message enters at the first trigger whose connection and event match. ### What a run receives The connector’s event fields, plus `connector`, `event`, and `identity`. For Slack `message.received`: `teamId`, `channelId`, `userId`, `text`, `ts`, `threadTs`. For Telegram: `chatId`, `chatType`, `userId`, `username`, `text`, `messageId`. `identity` describes the external sender (`{ connector, externalId, displayName?, resolution }`). It is not a vault member: do not use it to authorize anything. Each Slack thread or Telegram chat is its own context, so the first message in it starts a run. ## Related * [Next: People in workflows](/agents/workflows/people/) * [Steps](/agents/workflows/steps/) * [Connectors](/vaults/connectors/) * [Connectors](/vaults/connectors/) # Your Agent Service > Your account on Manykind — create an Agent Service in ROC, what it holds (agents, workflows, cubby schemas, widgets, datasets, members), and how its identity names every agent you publish. An **Agent Service** is your account on Manykind, the way a cloud account is: everything you build lives in one. It holds your workflows, LLM agents, and code agents, the cubby schemas, widgets, and datasets they use, and the members who work on them. Your customers’ data never lives here; it stays in their vaults. ## Create one 1. Sign in to ROC, the Manykind web console. The **Agent Services** page lists the services you can open. 2. Choose **Create Service**, enter a **Name**, and choose **Create**. 3. ROC shows each stage as it runs: **Initializing account**, **Creating bucket**, **Creating agent service**, **Authorizing compute**, and **Adding to your services**. It then opens the new service. ![The Create Agent Service dialog](/shots/create-agent-service.png) If ROC reports **Agent service created — compute not yet authorized**, the service and its bucket exist, but its agents cannot run yet. Open the service’s **⋯** menu → **Access & compute** and choose **Authorize compute**. Select a service to open its side navigation: * **Workflow Builder:** Sandbox, Workflows, Evaluations, Agents, Cubbies, Widgets, Connectors * **Resources:** Memory Bank, Models * **Organization:** Members Connectors and the Memory Bank shown under a service belong to the vault of the organization that owns it. They are customer-side resources; the service’s workflows use them once connected. ## What it holds | Part | What it is | Where you see it in ROC | | ----------------- | ------------------------------------------------------------------------------------------------------------------------ | -------------------------------------- | | **Workflows** | Agents with a typed graph of steps. See [Workflows](/agents/workflows/overview/). | **Workflows** | | **LLM agents** | Agents defined by instructions, a model, tools, and connections. See [LLM agents](/agents/llm-agents/overview/). | **Agents** | | **Code agents** | TypeScript agents. See [Code agents](/agents/code-agents/overview/). | **Agents** | | **Cubby schemas** | SQLite databases every agent of the service shares, one copy per vault. See [Cubbies](/agents/building-blocks/cubbies/). | **Cubbies** | | **Widgets** | Screens the service publishes. See [Widgets](/agents/building-blocks/widgets/). | **Widgets** | | **Datasets** | Cases for evaluating a workflow. See [Evaluations](/agents/workflows/evaluate/overview/). | **Evaluations** | | **Members** | The people who work in the service and may publish to it. See [Team](/get-started/team/). | **Members**, **Settings** → **People** | | **Identity** | A public key that names the service, and a storage bucket for the agent versions you push. | | Workflows, LLM agents, and code agents are three kinds of agent on the same level. A workflow can use an LLM agent or a code agent as a step; see [Agents and workflows](/agents/overview/). ## Identity Every agent you publish is identified by an **agent id** of the form: ```text : ``` The prefix is your Agent Service’s public key, shown in ROC; the alias is the name you give the agent. Two consequences follow: * **Siblings can address each other.** An agent’s peers live under the same service, so a sibling’s id is this agent’s own with a different alias. `ctx.self.agentId` gives a running agent its own id. * **Identity is not a lookup.** When a vault connects an agent, the platform derives the service from the id’s prefix and checks it against the manifest. A listing cannot claim to be someone else’s agent. ## What the service shares across its agents The service is the boundary for shared state: * **Cubbies are per service, not per agent.** Every agent the service publishes reads and writes the same cubby by alias, so a sibling’s conclusions are simply there. An agent of a different service cannot reach them. * **Widgets read the service’s cubbies** and the connected vault’s Memory Bank. Durable conclusions that the customer should own and see belong in the vault’s [Memory Bank](/vaults/memory-bank/), not in a cubby. ## Who works in it An Agent Service has an owner: your wallet for a personal service, or an organization’s vault for an organization service. The owner can add **members**. Membership sets what someone can do in ROC; publishing is separate, because the service’s storage bucket accepts only writes that chain back to its owner. **Invite to publish** gives a member that right. How to add people, invite them to publish, and push as a member is on [Team](/get-started/team/). ## Related * [Next: Team](/get-started/team/) * [How it fits together](/get-started/how-it-fits-together/) * [Agents and workflows](/agents/overview/) * [Push and deploy](/agents/ship/push-and-deploy/) # How it fits together > The map of Manykind — the agents you build (in your Agent Service), the vault your customer owns, the connection between them, and how one run happens end to end. Manykind has two sides and one connection between them. Everything else in these docs is a detail of one of the three. ```text DEVELOPER SIDE CUSTOMER SIDE Agent Service (your account) ┌────────────────────────────┐ ┌──────────────────────────────┐ │ ┌────────────────────────┐ │ │ Vault │ │ │ Agents │ │ connect │ │ │ │ workflows ├─┼─────────►│ scopes ── events, objects │ │ │ LLM agents │ │ (signed, │ Memory Bank │ │ │ code agents │ │ scoped, │ connectors (Slack, email…) │ │ └────────────────────────┘ │ revocable│ members │ │ cubbies widgets │ │ │ │ datasets members │ │ │ └────────────────────────────┘ └──────────────────────────────┘ │ │ └──────────── a run ◄─────────────────┘ event lands in a scope → Job → steps / handlers → models, cubbies, Memory Bank, connector actions → results back into the vault ``` ## The developer side: your agents An **[agent](/agents/overview/)** is what you build, what a customer connects, and what runs on their data. It comes in three kinds, on the same level: **workflows** (a typed graph of steps), **LLM agents** (instructions, a model, tools, and connections), and **code agents** (TypeScript). A workflow can use an LLM agent or a code agent as a step; LLM agents and code agents are peers. Your agents live in your [Agent Service](/get-started/agent-service/): your account on Manykind, the way a cloud account is. You create it once in ROC, the Manykind web console, and next to your agents it holds what they share: * **[Cubby schemas](/agents/building-blocks/cubbies/).** SQLite databases every agent of the service shares, one copy per vault. * **[Widgets](/agents/building-blocks/widgets/).** UI the service publishes. * **[Models](/agents/building-blocks/models/).** Aliases bound to models from the platform’s catalogue, called by workflows and code agents alike. * **Datasets** for [evaluations](/agents/workflows/evaluate/overview/), and the [**members**](/get-started/team/) who work in the service. Every agent you publish is identified as `:`: the service’s public key, then the agent’s name. ## The customer side: the vault A [vault](/vaults/overview/) belongs to a customer: one person, or an organization with [members](/vaults/members/). It holds: * **Scopes.** Named partitions of the vault. Events and objects live in a scope, and every connection is granted per scope. * **[Memory Bank](/vaults/memory-bank/).** The vault’s durable, privacy-classified record of what its agents and workflows learned. * **[Connectors](/vaults/connectors/).** Connections to outside systems (Slack, email, Telegram, MCP servers) whose credentials are sealed in the vault. ## The connection An agent reaches a vault only after the vault owner [connects](/vaults/connections-and-consent/) it. Connecting signs an **agreement** with the owner’s wallet that names the agent service, the vault, and the scopes. The vault checks the agreement, provisions what the agent declares, and pins the exact bundle the owner consented to run. Revoke the agreement and the access ends. ## How a run happens 1. **Build.** You build a workflow in the ROC Workflow Builder, or write an agent or workflow in code with `@cef-ai/agent-sdk`. 2. **Ship.** You publish it to your Agent Service and deploy a version. See [Push and deploy](/agents/ship/push-and-deploy/). 3. **Connect.** The vault owner connects it to one or more scopes and signs the agreement. See [Connect to a vault](/agents/ship/connect-to-a-vault/). 4. **Trigger.** An event lands in a connected scope: a person runs the workflow, your app publishes, a schedule fires, a webhook is called, a Slack message arrives through a connector, or another agent publishes. 5. **Run.** The platform finds or opens a **Job** for that agent and stream, and each event becomes a **Task** in it. A workflow walks its steps; a code agent runs the matching handler in a sandbox. See [Runs and events](/agents/how-agents-run/). 6. **Work.** The run calls [models](/agents/building-blocks/models/) by alias, reads and writes cubbies, files records into the Memory Bank, and takes connector actions, all under the connection’s grant. 7. **Results.** Output lands back in the vault: events in the scope, cubby rows, Memory Bank records. People see it in ROC, in your widgets, or in your own app through the Vault SDK. ## Where each part of these docs fits Everything belongs to one owner, and the docs follow the owners. | Part | Covers | Side | | -------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------- | | **Get started** | This map, [your Agent Service](/get-started/agent-service/) and its [team](/get-started/team/), [installing](/get-started/install/) the CLI and SDKs, and the [Quickstart](/get-started/quickstart/). | Both | | **[Agents](/agents/overview/)** | How runs happen and everything your Agent Service runs: [workflows](/agents/workflows/overview/) (with testing, evaluation, and monitoring), [LLM agents](/agents/llm-agents/overview/), [code agents](/agents/code-agents/overview/), the [building blocks](/agents/building-blocks/models/) they use, and how to [ship](/agents/ship/push-and-deploy/) them to a vault. | Developer | | **[Vaults](/vaults/overview/)** | What a customer owns: the vault, its connections, Memory Bank, connectors, and members, and how your own apps [build on a vault](/vaults/build-on-a-vault/data-onboarding/) with the Vault SDK. | Customer | | **[Reference](/reference/cli/)** | Every CLI command, SDK export, limit, and error code, and the glossary. | Both | ## Where the tools fit | Tool | What you do with it | | -------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **ROC** | Create an Agent Service; build and run workflows in the Workflow Builder; try agents against a vault in the Sandbox; browse Agents, Cubbies, Widgets, Connectors, Models, and the Memory Bank; manage Members; run Evaluations. | | **`cef` CLI** (`@cef-ai/cli`) | Scaffold, build, push, and deploy code agents and workflows; register A2A agents you run elsewhere; push service cubbies and widgets; run evaluations. See [CLI reference](/reference/cli/). | | **Agent SDK** (`@cef-ai/agent-sdk`) | Write code agents (`@Engagement`, `@OnEvent`, `ctx`) and workflows in code (`defineWorkflow`). See [Agent SDK reference](/reference/agent-sdk/). | | **Vault SDK** (`@cef-ai/vault-sdk`) | Talk to a vault from your own app or server: publish events, upload objects, connect agents, read cubbies and the Memory Bank. See [Vault SDK reference](/reference/vault-sdk/). | | **Widget runtime** (`@cef-ai/widget-runtime`) | The browser half of a widget. See [Widget runtime reference](/reference/widget-runtime/). | | **Testing and eval** (`@cef-ai/testing`, `@cef-ai/eval`) | Test agents offline; build datasets and experiments for workflows. | ## Related * [Next: Your Agent Service](/get-started/agent-service/) * [Sovereign data](/vaults/sovereign-data/) * [Agents and workflows](/agents/overview/) * [Connections and consent](/vaults/connections-and-consent/) * [Runs and events](/agents/how-agents-run/) # Install the CLI > Install the cef CLI and the @cef-ai SDK packages, pick the ones you need, and set up the credentials the CLI pushes and deploys with. You can build and run a workflow entirely in ROC, with nothing installed. Install the CLI and SDKs when you want code: code agents, workflows in code, your own app talking to a vault, tests, and evaluations. ## Prerequisites * **Node.js** and a package manager. Examples use npm; pnpm, yarn, and bun work the same way. * **TypeScript** for agents. Engagements use decorators (`@Engagement`, `@OnEvent`). ## The packages | Package | Version | Kind | Use it for | | ------------------------------------------------------ | ------- | ------- | ---------------------------------------------------------------------------------------------------------- | | [`@cef-ai/cli`](/reference/cli/) (`cef`) | 2.8.0 | dev | Scaffold, build, push, and deploy agents and workflows; push service cubbies and widgets; run evaluations. | | [`@cef-ai/agent-sdk`](/reference/agent-sdk/) | 5.8.0 | runtime | Code agents (`@Engagement`, `@OnEvent`, `ctx`), `defineAgent` and `defineWorkflow` for `cef.config.ts`. | | [`@cef-ai/vault-sdk`](/reference/vault-sdk/) | 5.5.0 | runtime | Your own apps and servers talking to a vault. Not for agent code. | | [`@cef-ai/widget-runtime`](/reference/widget-runtime/) | 2.4.1 | runtime | The browser half of a widget. The CLI vendors it into built widgets. | | [`@cef-ai/testing`](/reference/testing/) | 3.3.5 | dev | Run your agent against an in-process fake platform in your own test runner. | | [`@cef-ai/eval`](/reference/eval/) | 0.1.0 | dev | Datasets, experiments, and results for workflows. | | [`@cef-ai/eslint-plugin`](/reference/eslint-plugin/) | 1.0.0 | dev | Editor rules that match the CLI’s build checks. | ## Install for building agents and workflows in code Start a project with the scaffold: ```bash npx @cef-ai/cli@2.8.0 init my-agent ``` `cef init [dir]` writes a project with a config, an engagement, a cubby migration, a widget, a test, and AI-assistant files (`AGENTS.md`, `CLAUDE.md`). Options: `--name `, `--pm `, `--install`, `--no-git`, `--no-ai`, `-y`. Set the scaffold’s `@cef-ai/*` dependencies to the versions in the table above. Or add the packages to an existing project: ```bash npm install @cef-ai/agent-sdk@5.8.0 npm install -D @cef-ai/cli@2.8.0 @cef-ai/testing@3.3.5 @cef-ai/eslint-plugin@1.0.0 ``` The CLI is a dev dependency; run it with `npx cef …` or an npm script. ## Install for an app that talks to a vault ```bash npm install @cef-ai/vault-sdk@5.5.0 ``` The package is an ES module and runs in the browser and in Node. See [Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/). ## Install for workflow evaluations ```bash npm install -D @cef-ai/eval@0.1.0 ``` `cef eval` uses the same library. See [Evaluations](/agents/workflows/evaluate/overview/). ## Verify ```bash npx cef --help ``` A list of commands (`init`, `build`, `typegen`, `inspect`, `push`, `publish`, `deploy`, `dev`, `widget`, `cubby`, `eval`) means the CLI is installed. ## Credentials The CLI talks to two places, with two kinds of credential: | For | Credential | Set with | | ------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------- | -------------------------------------------------- | | `cef push`, `cef widget push`, `cef cubby push` (writing to your Agent Service’s storage bucket) | The bucket owner’s secret phrase, or an access token the owner granted you | `CEF_DDC_SECRET_PHRASE`, or `CEF_DDC_ACCESS_TOKEN` | | `cef deploy`, `cef publish` | A CLI access token minted in ROC | `CEF_ACCESS_TOKEN` | Each has a matching command-line flag, but a secret on the command line stays in your shell history; prefer the environment variables. `--env ` (or `CEF_ENV`) picks the network. See [Push and deploy](/agents/ship/push-and-deploy/) and the [CLI reference](/reference/cli/). ## Pin your versions Pin exact versions. Check the current one before upgrading: ```bash npm view @cef-ai/cli version ``` ## Next steps * [Quickstart](/get-started/quickstart/): build and run a workflow * [Write an agent](/agents/code-agents/overview/) * [Workflows in code](/agents/workflows/author-in-code/) ## Related * [Next: Quickstart](/get-started/quickstart/) * [CLI reference](/reference/cli/) * [Push and deploy](/agents/ship/push-and-deploy/) * [Build with AI](/reference/build-with-ai/) * [How it fits together](/get-started/how-it-fits-together/) # Quickstart > Build, deploy, and run a small support-ticket triage workflow, either in the ROC Workflow Builder or in code with defineWorkflow. Build a workflow that takes a support ticket, classifies it with AI, and returns the category and priority. Then run it on your organization’s vault and read the run step by step. A workflow is an agent: it lives in your Agent Service, is deployed like any agent, and runs on vault data once it is connected to the vault. You can build it in the ROC Workflow Builder or in TypeScript; pick a tab. Both produce the same kind of workflow. ```plaintext New ticket (trigger) → Classify (AI) → Result { category, priority } ``` ## Before you begin * A ROC account with an **Agent Service** you can publish to, and an organization vault your account can write to. See [Your Agent Service](/get-started/agent-service/). * For the **Code** tab, also Node.js 18 or later, pnpm, and the CLI credentials described in [Install the CLI](/get-started/install/): your Agent Service’s bucket id and public key, a DDC access token in `CEF_DDC_ACCESS_TOKEN`, and a CLI access token from ROC in `CEF_ACCESS_TOKEN`. - Builder In the Builder, the classifier is an agent you create in place. ### 1. Create the workflow 1. In ROC, open your Agent Service and choose **Workflows** in the side navigation. 2. Choose **New workflow**. 3. **Name**: `ticket-triage`. **What it does**: `Classify a support ticket`. 4. Under **Start from**, choose **Blank** (“A trigger and one step”). 5. Choose **Create workflow**. The canvas opens on the **Editor** tab with an **Event** trigger wired to one **Agent** step. ### 2. Describe what starts a run Select the trigger. Leave **Event type** as `workflow.start`, and set **Example payload**: ```json { "ticketId": "T-1", "text": "I was charged twice for my October invoice." } ``` The Run dialog is prefilled from this example. ### 3. Create the classifier agent 1. Select the Agent step and choose **Choose an agent…**, then **New agent**. 2. Fill the form: * **Name**: `Ticket classifier` * **What it does**: `Classifies a support ticket by category and priority.` * **Instructions**: ```plaintext You triage support tickets. Answer as JSON and nothing else: {"category": "billing" | "bug" | "question", "priority": "low" | "normal" | "urgent"} ``` * **Model**: keep the default, or pick a language model. 3. Choose **Create agent**. The step now uses it. ![The New agent form: name, what it does, instructions, model](/shots/new-llm-agent.png) 4. On the step, set **What to ask it** to `={{ $json.text }}`, and **Answer it should give** to: ```json { "category": "billing", "priority": "normal" } ``` A run fails at this step, naming the missing field, if the agent answers without either key. ### 4. Return a result 1. Choose **Add a node** on the canvas side rail, open **Outputs**, and add **Result**. 2. Wire the Agent step’s output to the Result step. 3. In **Where each field comes from**, add `category` from `={{ $json.category }}` and `priority` from `={{ $json.priority }}`. ### 5. Deploy Choose **Deploy 0.1.0** in the toolbar. Deploy publishes the workflow as an agent of your Agent Service, makes this version live, and connects the workflow and its agent to your organization’s vault. If ROC asks you to approve the connection, approve it. When it is done, the button reads **Deployed 0.1.0**. A finished workflow on the canvas looks like this one, which also branches on priority and pages Slack: ![A deployed ticket-triage workflow on the Editor tab](/shots/builder-canvas.png) ### 6. Run it 1. Open the **Executions** tab. The chip beside **Run** shows where runs go: **Runs in {organization} · {scope}**. 2. Choose **Run**. **What starts it** holds your example payload; edit the text if you like. 3. Choose **Run** in the dialog. The run appears in the list and on the canvas. Each step turns **Success** as it finishes. ![The Executions tab with done, waiting, and failed runs](/shots/executions-tab.png) ### 7. Read the run Open the run and select each step: * the trigger **Produced** your payload; * the Agent step **Received** the ticket and **Produced** `category` and `priority` merged into the item; * the Result step shows the run’s output, with `category` and `priority`. ![One run: the canvas, its summary, and the execution log with step timings](/shots/run-detail.png) The run summary shows how long it took, the tokens used, and the cost. If a step failed, its error names the reason; see [Monitor runs](/agents/workflows/monitor-runs/#why-a-run-failed-parked-or-stalled). - Code In code, the workflow calls a language model directly with a model step. ### 1. Scaffold a project ```bash pnpm dlx @cef-ai/cli@^2.8.0 init ticket-triage --yes cd ticket-triage pnpm add @cef-ai/agent-sdk@^5.8.0 pnpm add -D @cef-ai/cli@^2.8.0 @cef-ai/testing@^3.3.5 ``` `cef init` creates a project with a sample agent, a `deployments/default.jsonc` that makes the newest version live, and a `package.json`. You replace the agent with a workflow next. ### 2. Pick a model Open **Models** in ROC’s side navigation and choose a language model that is available. Note its **alias** (shown as `ctx.models.`), and the bucket, name, and version it is stored under. Its `model.json` URL has the path `//models///model.json`. ### 3. Declare the workflow Replace `cef.config.ts`: ```ts import { defineWorkflow } from "@cef-ai/agent-sdk/config"; export default defineWorkflow({ id: "ticket-triage", version: "0.1.0", goal: "Classify a support ticket", models: { // Key = the model's alias. Replace the URL with your model's. llm: "https://cdn.example.com/1234/models/my-llm/1.0.0/model.json", }, nodes: [ { id: "ticket", kind: "trigger", label: "New ticket", position: { x: 0, y: 0 }, params: { sample: JSON.stringify({ ticketId: "T-1", text: "I was charged twice for my October invoice." }), }, }, { id: "classify", kind: "model", label: "Classify", position: { x: 320, y: 0 }, params: { alias: "llm", input: { messages: [ { role: "user", content: "=Classify this support ticket. Answer as JSON: " + '{"category": "billing" | "bug" | "question", "priority": "low" | "normal" | "urgent"}\n\n' + "{{ $json.text }}", }, ], max_tokens: 64, response_format: { type: "json_object" }, }, into: "classification", }, }, { id: "parse", kind: "transform", label: "Read the answer", position: { x: 640, y: 0 }, params: { expr: "({ ...item, ...JSON.parse(item.classification.text) })" }, }, { id: "result", kind: "output", label: "Result", position: { x: 960, y: 0 }, params: { result: { category: "={{ $json.category }}", priority: "={{ $json.priority }}" }, }, }, ], edges: [ { from: "ticket", to: "classify" }, { from: "classify", to: "parse" }, { from: "parse", to: "result" }, ], }); ``` What each part does: * The **trigger** starts a run on `workflow.start`; `sample` prefills ROC’s Run dialog. * The **model** step sends the request to the model declared as `llm` and puts the answer, `{ text, usage }`, under `classification`. Strings starting with `=` are mappings: `{{ $json.text }}` reads the ticket text from the carried item. * The **transform** parses the model’s JSON text into fields. * The **output** step declares the run’s result. Delete the scaffold’s `src/agent.ts` and `test/agent.test.ts`; the workflow does not use them. ### 4. Build ```bash pnpm cef build ``` ```plaintext [ok] ticket-triage -> …/dist/ticket-triage/bundle.js (… bytes) [ok] ticket-triage -> …/dist/ticket-triage/manifest.json ``` `cef build` runs the same checks the runner runs before a run. A typo in an edge, an undeclared model alias, or a template without its leading `=` fails here, not in production. ### 5. Push and deploy ```bash pnpm cef push --bucket --as-pubkey pnpm cef deploy ``` ```plaintext ✓ Pushed :ticket-triage bundle to DDC ✓ Deployed :ticket-triage — 1 deployment(s) [default] (revision 1) ``` `cef push` uploads the workflow to your Agent Service’s bucket. `cef deploy` applies `deployments/` and makes the version live. See [Push and deploy](/agents/ship/push-and-deploy/). ### 6. Connect and run in ROC 1. In ROC, open your Agent Service → **Workflows**. `ticket-triage` is listed; open it. It is marked **from the repo · read-only**, because you edit it in code. 2. Open the **Executions** tab and choose **Run**. 3. ROC connects the workflow to your organization’s vault before the first run. Approve the connection when asked. 4. In **Run this workflow**, **What starts it** holds the sample payload. Choose **Run**. ### 7. Read the run Open the run and select each step: the model step **Produced** `classification.text`, the transform added `category` and `priority`, and the Result step shows the run’s output. If a step failed, its error names the reason; see [Monitor runs](/agents/workflows/monitor-runs/#why-a-run-failed-parked-or-stalled). ## What you built A deployed workflow, connected to your organization’s vault, that classifies a ticket and returns a typed result. Every run is recorded as a Job you can open step by step, re-run, or add to a dataset. ## Next steps * Branch on the result, ask a person to approve urgent tickets, and record results in a cubby: [Steps](/agents/workflows/steps/), [People in workflows](/agents/workflows/people/), [Cubbies](/agents/building-blocks/cubbies/). * Start runs from a schedule, a webhook, or a Slack message: [Triggers](/agents/workflows/triggers/). * Test the workflow in CI: [Test a workflow](/agents/workflows/test/). * Score it against a dataset: [Evaluations](/agents/workflows/evaluate/overview/). ## Related * [Next: Agents and workflows](/agents/overview/) * [Workflows](/agents/workflows/overview/) * [Workflows in code](/agents/workflows/author-in-code/) * [Monitor runs](/agents/workflows/monitor-runs/) * [Install the CLI](/get-started/install/) # Team > Add people to an Agent Service, invite them to publish to its storage, and push agents, widgets, and cubbies as a member from the CLI. An [Agent Service](/get-started/agent-service/) holds your agents and their building blocks; this page is about the people who work in it. A service can have members besides its owner. Membership controls what someone can do in ROC. Publishing is separate: storage writes are authorized against the bucket’s on-chain owner, so a member can push only with a token that chains back to the owner. **Invite to publish** gives a member that token. ## Two kinds of Agent Service | Service | Who manages people | Where | | --------------------------------------------------- | ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Personal** (owned by your wallet) | You | Service **⋯** → **Settings** → **People** → **Invite a member**, or **⋯** → **Invite member** | | **Organization** (owned by an organization’s vault) | The organization | People are added on the organization’s page (**Members** in the sidebar). The organization’s own wallet then grants publishing with **Invite to publish**. | ## Invite someone 1. Open the service **⋯** menu → **Invite to publish** (organization service) or **Invite member** (personal service). Only the bucket’s owner sees **Invite to publish**. 2. Fill in the dialog: | Field | Rule | | --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | **Member public key** | `0x` followed by 64 hex characters. SS58 addresses are not accepted. | | **Label (optional)** | Personal services only. | | **Role** | Personal services only: `developer-admin`, `developer-write`, `developer-read`, or `integrator`. The role sets console permissions; every member gets full bucket access. Organization invites always use `developer-write`. | | **Access for** | 1 day, 3 days, 7 days, 1 month, 1 year, or **Custom date**. | 3. Click **Send invite** (or **Invite by link**), then **Copy invite link** and send it. For an organization service, only someone already in the organization with write access can accept the invite. ## Accept an invite 1. Open the link in the browser where you use ROC. The page reads **Accept publish invite** (organization) or **Accept agent-service invite**. 2. Click **Accept invite**. It may ask for an on-chain transaction. 3. Click **Open service**. The browser keeps the invite token. You need that same browser to generate CLI tokens. ## Publishing as a member 1. In the service, open **⋯** → **Settings** → **People**. In **Push from the CLI**, choose how long the token is valid and click **Copy DDC token for the CLI**. The **DDC access token** panel on the **Access** tab produces the same kind of token for you. 2. Push: ```bash export CEF_DDC_ACCESS_TOKEN=… cef push --bucket --as-pubkey ``` The same token works for `cef widget push` and `cef cubby push`. 3. Deploy with your own **CLI access token** from **Settings** → **Access**. It needs no invite. ### “bucket belongs to other owner” A token you minted with your own wallet cannot write an organization’s bucket: the storage node accepts only tokens rooted in the bucket’s owner. `cef push` checks this before uploading and stops with: ```plaintext [cef] This DDC access token can't write bucket : it is signed by but the bucket is owned by . Ask the owner for "Invite to publish" in ROC (Settings → People), open the invite in your browser, then generate the token from Settings → Access. ``` Ask the owner for **Invite to publish**, accept it in your browser, and generate the token again. ROC refuses to mint a DDC token for someone who is neither the owner nor holding an invite in that browser. ### Publish into an organization vault A member with write access to an organization vault’s scope can publish an agent through the vault instead of writing the bucket directly: ```bash cef push --vault --vault-scope --secret-phrase "$CEF_VAULT_SECRET_PHRASE" ``` The vault checks your write access on the scope and writes its own registry bucket. The agent’s alias is bound to the scope it is first published from. Authenticate with your wallet’s phrase or with `--vault-token` (`$CEF_VAULT_TOKEN`). `--bucket`, `--access-token`, and the network flags are refused with `--vault`. ## Remove someone For an organization service, remove people on the organization’s page. People still on the service’s own member list are shown on **Settings** → **People**, each with **Remove**. ## Related * [Next: Install the CLI](/get-started/install/) * [Your Agent Service](/get-started/agent-service/) * [Vault members](/vaults/members/) * [Push and deploy](/agents/ship/push-and-deploy/) * [CLI reference](/reference/cli/#cef-push) # Agent SDK (@cef-ai/agent-sdk) > Reference for @cef-ai/agent-sdk 5.8.0: decorators, the Context surface, defineAgent, defineWorkflow, isWorkflow, and the workflow runner entry. ```bash npm install @cef-ai/agent-sdk@5.8.0 ``` | Import | Contains | Used in | | ----------------------------------- | ------------------------------------------------------------------------------ | -------------------------------------------------- | | `@cef-ai/agent-sdk` | Decorators and types (`Context`, `Event`, …). | Agent source. | | `@cef-ai/agent-sdk/config` | `defineAgent`, `defineWorkflow`, `isWorkflow`, config types. | `cef.config.ts`. | | `@cef-ai/agent-sdk/workflow` | The workflow engine: `WorkflowRunner`, graph helpers, `WORKFLOW_RUNNER_ENTRY`. | Tools that inspect or test workflows. | | `@cef-ai/agent-sdk/workflow/runner` | The runner module itself; the default entry of every workflow. | `defineWorkflow({ runner })`. | | `@cef-ai/agent-sdk/runtime` | The in-bundle implementation of `ctx`. | Wired in by `cef build`; never imported by agents. | Decorators need `"experimentalDecorators": true` in `tsconfig.json`. ## Decorators | Decorator | Target | Signature / effect | | --------------------------- | ------ | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `@Engagement({ id, goal })` | class | Names an engagement. | | `@OnEvent(type)` | method | `(event: Event

, ctx: Context) => Promise`. `type` must be a string literal. A second handler for the same type is ignored with a warning. | | `@OnStart` / `@OnStart()` | method | Runs once when the Job starts. | | `@OnClose` / `@OnClose()` | method | `(ctx: Context, reason: OnCloseReason)`. | | `@Condition(expr)` | class | CEL selection expression. Repeatable; ANDed. | | `@Priority(n)` | class | Lower value wins. Last application wins. | | `@Weight(n)` | class | Split within a priority tier. | | `@Limit(n, per)` | class | `per`: `"day"`, `"connection/day"`, `"connection/month"`. | | `@Params(values)` | class | Param values for the engagement; repeated applications merge. | `OnCloseReason` is `"revoked" | "idle_timeout" | "closed_by_agent" | "failed"`. ## `Event

` | Field | Type | Notes | | ----------- | ------------------------------- | ------------------------------ | | `type` | `string` | | | `payload` | `P` | | | `timestamp` | `string` | ISO 8601, set by the vault. | | `context` | `string` | The stream key. | | `role` | `"source" \| "user" \| "agent"` | | | `from` | `string?` | Publisher identity. | | `eventId` | `string?` | | | `parents` | `string[]?` | Events this one was caused by. | ## `Context` | Member | Type | | ------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cubby(alias, attribution?)` | `CubbyHandle`: `query(sql, params?) → Promise`, `exec(sql, params?) → Promise<{ changes, lastInsertRowid }>`. The cubby belongs to the Agent Service. `attribution.nodeId` names the workflow step making the call. | | `models` | `KnownModels & Record`. `ModelHandle`: `infer(input: I) → Promise`, `stream(input: I) → AsyncIterable` (yields the complete output once). | | `vault.publish(type, payload, opts?)` | `opts: PublishOptions` = `{ target?, title?, description?, correlation? }`. | | `vault.objects` | `upload(path, data: Uint8Array, { contentType? })`, `get(path)`, `head(path)`, `presignedUrl(path, { ttlSeconds? })`, `list({ prefix? })`. No `delete`. | | `memory` | `MemoryHandle`: `upsert(record)`, `update(id, { title?, body? })`, `setPrivacy(id, privacy)`, `delete(id)`, `relation(edge)`, `search(match, { limit? })`, `get(id)`, `neighbours(id, { limit? })`, `countByType()`. | | `self` | `{ agentId?, vaultId?, scope?, context?, jobId?, taskId? }`, frozen. | | `settings` | `Readonly>`. | | `params` | `Readonly>`. | | `close(reason?)` | `Promise`. | `MemoryRecordInput`: `{ id, type, title?, body?, scope, privacy }`. `MemoryRelationInput`: `{ in, out, type, scope, privacy }`. `MemoryPrivacy`: `"public" | "internal" | "private" | "restricted"`, required on every write. `neighbours` returns `MemoryNeighbourRow`: `{ id, type, title, scope, privacy, edgeType }`. `KnownEventTypes` and `KnownModels` are interfaces filled by `cef typegen` for typed `@OnEvent`, `vault.publish`, and `models`. ## `defineAgent(config)` Returns the config unchanged, with its literal types. Fields of `AgentConfig`: | Field | Type | Default | | -------------------- | ----------------------------------------------------------------------------------------------- | ------------------------------------------------- | | `id` | `string` | required; the alias | | `version` | `string` | required | | `alias` | `string` | `id` | | `agentServicePubkey` | `string` | — | | `source` | `string` | — | | `card` | `{ name, description, iconUrl?, capabilities? }` | — | | `entry` | `string` | one of `entry` / `engagements` | | `engagements` | `{ id, entry, goal?, condition?, priority?, weight?, limit?: { n, per }, params?, enabled? }[]` | | | `requiredScopes` | `string[]` | `["default"]` | | `idleTimeout` | duration string | `"30m"`; `"0s"` disables | | `models` | `Record` | — | | `params` | `Record` | — | | `settings` | `SettingDecl[]` | — | | `cubbies` | `CubbyDecl[]` = `{ alias, migrations? }[]` | — | | `schedules` | `ScheduleDecl[]` | — | | `widgets` | `WidgetDecl[]` | — | | `eventSchemas` | `Record` | — | | `uses` | `Record` | Peer agents, typed by `cef typegen`. | | `agents` | `AgentConfig[]` | Several agents in one config. | | `kind` | `"internal" \| "external" \| "workflow"` | A classification; dispatch does not depend on it. | | Type | Shape | | -------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ParamDecl` | `{ type: "number" \| "string" \| "boolean" \| "modelAlias", default, min?, max?, enum? }` | | `SettingDecl` | `{ key, type: "string" \| "number" \| "boolean" \| "url" \| "secret", required?, label?, description?, default? }` | | `ScheduleDecl` | `{ id, cron, timezone?, eventType, payload? }` | | `WidgetDecl` | `{ id, name?, description?, cubbyAlias?, kind?, config?, queries?, events?, dir, entry }`; a query is `{ id, label?, sql?, cubby?, tool?, limit?, timeoutMs? }` | Guides: [Write an agent](/agents/code-agents/overview/), [Build a widget](/agents/building-blocks/create-a-widget/). ## `defineWorkflow(spec)` Declares a workflow and returns an `AgentConfig`, so a workflow builds, pushes, deploys, and connects like any agent. ```ts import { defineWorkflow } from "@cef-ai/agent-sdk/config"; export default defineWorkflow({ id: "triage", version: "0.1.0", goal: "Classify incoming requests", models: { llm: "https://cdn.ddc-dragon.com//models///model.json" }, nodes: [ { id: "start", kind: "trigger" }, { id: "classify", kind: "model", params: { alias: "llm", input: { prompt: "={{ $json.text }}" } } }, { id: "done", kind: "output" }, ], edges: [ { from: "start", to: "classify" }, { from: "classify", to: "done" }, ], }); ``` | `WorkflowSpec` field | Meaning | | ------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `id`, `version` | As for an agent. | | `goal` | Description; also the default card description. | | `nodes` | `WorkflowNode[]`: `{ id, kind, use?, emit?, params?, label?, question?, position? }`. | | `edges` | `{ from, to, when?, loop? }`; `from`/`to` must be node ids (checked by the type). `when`: `{ field, op, value? }`, `op` one of `eq`, `ne`, `gt`, `gte`, `lt`, `lte`, `contains`, `exists`. `loop`: `{ max, counter, exhausted? }`. | | `models`, `cubbies`, `schedules`, `card`, `idleTimeout` | As for an agent. `idleTimeout` defaults to `"30m"`. | | `runner` | Runner entry. Defaults to `"@cef-ai/agent-sdk/workflow/runner"`. | Node kinds: `trigger`, `agent`, `branch`, `join`, `publish`, `remember`, `relate`, `recall`, `transform`, `code`, `human`, `model`, `cubbyQuery`, `cubbyExec`, `action`, `output`, `split`, `aggregate`. Step semantics: [Steps](/agents/workflows/steps/). `defineWorkflow` throws at config load (that is, at `cef build`) when two nodes share an id, a kind is unknown, an `agent` node has no `use`, a `model` node names an alias not in `models`, there is no `trigger`, an edge names an unknown node, a schedule id is not a trigger with `params.mode: "schedule"`, a non-trigger node has no incoming edge, or the runner’s own validation finds a fatal problem. The returned config uses the runner as its entry, adds a `runs` cubby for the runner’s state, and carries the graph as the `graph` param. A deployment can override that param. ## `isWorkflow(manifest)` Returns `true` when a manifest carries a non-empty `graph` param, which is what makes an agent a workflow. `WORKFLOW_GRAPH_PARAM` is `"graph"`. ## `@cef-ai/agent-sdk/workflow` | Export | Use | | ------------------------------------------------ | ----------------------------------------------------------------------------------------- | | `WorkflowRunner` | The engine class every workflow runs. | | `WORKFLOW_RUNNER_ENTRY` | `"@cef-ai/agent-sdk/workflow/runner"`. | | `validate(doc)`, `isFatal(issue)` | The runner’s graph validation. | | `readContinuation(raw)` | Which step, pass, and item an answer was for. | | `readResultSpec`, `narrowResult`, `RESULT_TYPES` | A workflow’s declared Result, used by [evaluations](/agents/workflows/evaluate/results/). | | `STEP_ASK_EVENT` | `"workflow.step"`. | ## Related * [Write an agent](/agents/code-agents/overview/) * [Workflows in code](/agents/workflows/author-in-code/) * [CLI reference](/reference/cli/) * [Testing reference](/reference/testing/) # Build with AI > Point Claude Code, Cursor, or any coding agent at Manykind's machine-readable docs so it builds with real platform APIs. These docs are published in machine-readable form, so a coding agent has accurate, current platform context instead of guessing. ## llms.txt files Three files are generated from this site on every build: | File | What it is | Use it when | | ------------------------------------ | ---------------------------------------------- | ------------------------------------------------------------------- | | [`/llms.txt`](/llms.txt) | An index of the documentation, with links. | The agent should discover what exists and fetch only what it needs. | | [`/llms-full.txt`](/llms-full.txt) | The entire documentation in one Markdown file. | The agent should load the whole platform in one fetch. | | [`/llms-small.txt`](/llms-small.txt) | A condensed version of the full file. | The agent has a smaller context budget. | They follow the [llms.txt convention](https://llmstxt.org/) and are regenerated from the published pages, so they never drift from what you read here. ## Use it in your coding agent ```text Read https://developers.cere.io/llms-full.txt for context on the Manykind platform, then help me write a workflow that … ``` Give the agent the vocabulary up front: an **Agent Service** holds your agents and workflows; a customer’s **vault** holds their data, **Memory Bank**, and **connectors**; an agent or workflow runs on a vault only after the owner **connects** it. ## Scaffolded projects include agent instructions `cef init` writes `AGENTS.md`, `CLAUDE.md`, and a `.claude/` settings folder into a new project, so a coding agent opened in it starts with the project’s conventions. Pass `--no-ai` to skip them. ## Related * [Install the CLI](/get-started/install/) * [Quickstart](/get-started/quickstart/) * [How it fits together](/get-started/how-it-fits-together/) * [Glossary](/reference/glossary/) # CLI (@cef-ai/cli) > Every cef command and flag in @cef-ai/cli 2.8.0: init, build, typegen, inspect, push, widget push, cubby push, deploy, dev, publish, eval, and test. The reference section is the lookup layer: every CLI command, SDK export, limit, error code, and term. Start with the CLI. ```bash npm install -D @cef-ai/cli@2.8.0 # or: npm i -g @cef-ai/cli cef --version ``` | Command | Does | | ------------------------------------- | -------------------------------------------------------------------- | | [`cef init`](#cef-init) | Scaffold a project. | | [`cef build`](#cef-build) | Bundle agents and write manifests. | | [`cef typegen`](#cef-typegen) | Generate types for declared models and peers. | | [`cef inspect`](#cef-inspect) | Check a built agent. | | [`cef push`](#cef-push) | Upload an agent version, or register an A2A agent you run elsewhere. | | [`cef widget push`](#cef-widget-push) | Publish a service widget. | | [`cef cubby push`](#cef-cubby-push) | Publish the service’s cubby declarations. | | [`cef deploy`](#cef-deploy) | Apply deployment records. | | [`cef dev`](#cef-dev) | Serve a widget locally. | | [`cef publish`](#cef-publish) | Send the agent card to the marketplace listing. | | [`cef eval`](#cef-eval) | Workflow evaluations. | | `cef test` | Not implemented; prints a notice. Run `vitest` directly. | Any failing command prints its error and exits with code `1`. ## Environments `--env dev|stage|prod`, or `$CEF_ENV`. Default `dev`. | `--env` | DDC network (`--preset`) | | ------- | ------------------------ | | `dev` | `DEVNET` | | `stage` | `TESTNET` | | `prod` | `MAINNET` | The environment also selects the platform API for `deploy`, the marketplace for `publish`, and the endpoints written into widgets. Specific flags (`--preset`, `--endpoint`, `--marketplace`) override single values. ## Credentials | Variable | Flag | Used by | | ------------------------- | ------------------ | ------------------------------------------------------------------------------------------------------------------ | | `CEF_DDC_ACCESS_TOKEN` | `--access-token` | `push`, `widget push`, `cubby push`: a DDC token from ROC (**Settings** → **Access** → **DDC access token**). | | `CEF_DDC_SECRET_PHRASE` | `--secret-phrase` | Same commands: the bucket owner’s sr25519 phrase, instead of a token. | | `CEF_DDC_SUBJECT_PHRASE` | `--subject-phrase` | The phrase of the key a token was issued to, when the token names one. | | `CEF_ACCESS_TOKEN` | `--access-token` | `deploy`, `publish`, `eval run`: the CLI access token from ROC (**Settings** → **Access** → **CLI access token**). | | `CEF_VAULT_TOKEN` | `--vault-token` | `push --vault`. | | `CEF_VAULT_SECRET_PHRASE` | `--secret-phrase` | `push --vault`: your wallet phrase (used as ed25519). | | `CEF_ENDPOINT` | `--endpoint` | `deploy`: platform API base URL. | | `CEF_MARKETPLACE_URL` | `--marketplace` | `publish`. | Prefer environment variables to flags: a secret on the command line stays in shell history and is visible in the process list. A push needs exactly one of a phrase or a token. ## `cef init` ```bash cef init [dir] ``` Scaffolds a hello-world agent: `cef.config.ts`, `src/agent.ts`, `migrations/`, `deployments/`, a widget, a test, and, unless `--no-ai`, `AGENTS.md`, `CLAUDE.md`, and `.claude/`. | Flag | Default | Meaning | | ------------------- | --------------------- | ------------------------------------ | | `--name ` | slug of the directory | Agent alias. | | `-y, --yes` | — | Accept all defaults. | | `--pm ` | — | `pnpm`, `npm`, `yarn`, or `bun`. | | `--install` | off | Install dependencies. | | `--no-git` | — | Skip `git init`. | | `--no-ai` | — | Skip the AI assistant files. | | `--force` | — | Scaffold into a non-empty directory. | | `--template ` | `hello-world` | Template. | ## `cef build` Bundles every agent in `cef.config.ts` and writes `dist//bundle.js`, `manifest.json`, and `widgets//` (runtime vendored and manifest injected). | Flag | Default | Meaning | | ------------------- | --------------- | ------------------------------------------------------------- | | `--config ` | `cef.config.ts` | Config file. | | `--out

` | `dist` | Output directory. | | `--env ` | `dev` | Endpoints written into widget manifests. | | `--as-pubkey ` | — | Agent Service pubkey written into widget manifests’ agent id. | Build checks are listed in [Write an agent](/agents/code-agents/overview/#what-cef-build-rejects). ## `cef typegen` Fetches each declared model’s `model.json` and each peer in `uses`, then writes `.cef/generated.d.ts` and `cef.lock.json`. | Flag | Default | Meaning | | ----------------- | --------------- | ------------ | | `--config ` | `cef.config.ts` | Config file. | ## `cef inspect` ```bash cef inspect dist/ ``` Prints each engagement and checks that the bundle exports every handler the manifest routes to. Exits with code `2` when a handler is missing, so you can run it in CI. ## `cef push` Uploads the built agent in `dist/` to a bucket you can write, or through vault-api into an organization vault, or registers an A2A agent you run elsewhere (`--kind external`). ```bash cef push --bucket --as-pubkey cef push --vault --vault-scope cef push --kind external --card --bucket --as-pubkey ``` | Flag | Applies to | Default | Meaning | | --------------------------- | --------------------- | ------------------------- | ---------------------------------------------------------------------------------- | | `--kind ` | all | `internal` | `internal` pushes the bundle in `dist/`; `external` registers an A2A agent. | | `--bucket ` | bucket push, external | — | DDC bucket id (decimal). | | `--as-pubkey ` | all | config value | Agent Service pubkey; forms the agent id `:`. | | `--agent ` | internal | the only built agent | Which `dist//` to push. | | `--out ` | internal | `dist` | Build output directory. | | `--vault ` | internal | — | Publish into this organization vault through vault-api. Needs `--vault-scope`. | | `--vault-scope ` | `--vault` | — | Scope to publish from. The alias is bound to the scope it is first published from. | | `--vault-api ` | `--vault` | environment’s vault API | vault-api base URL. | | `--vault-token ` | `--vault` | `$CEF_VAULT_TOKEN` | Wallet-api bearer token. | | `--card ` | external (required) | — | A2A Agent Card URL. A bare origin gets `/.well-known/agent-card.json`. | | `--alias ` | external | slug of the card name | Registry alias; no `:`. | | `--agent-version ` | external | `1.0.0` | Version to register. | | `--scope ` | external | `default` | Scopes the agent asks for. | | `--idle-timeout ` | external | `30m` | How long a conversation survives without events. `0` is refused. | | `--secret-phrase ` | all | env | Bucket owner’s phrase, or with `--vault` your wallet’s phrase. | | `--access-token ` | bucket push, external | `$CEF_DDC_ACCESS_TOKEN` | DDC access token. | | `--subject-phrase ` | bucket push, external | `$CEF_DDC_SUBJECT_PHRASE` | Phrase for a token that names a subject. | | `--env ` | all | `dev` | Environment. | | `--endpoint ` | bucket push, external | from `--env` | DDC blockchain endpoint. | | `--preset ` | bucket push, external | from `--env` | `MAINNET`, `TESTNET`, or `DEVNET`. | | `--cdn ` | bucket push, external | from preset | CDN endpoint. | A flag that belongs to the other kind or destination is refused by name rather than ignored. Before uploading with a token, `cef push` checks that the token chains to the bucket’s on-chain owner and that it has not expired; see [Team](/get-started/team/#bucket-belongs-to-other-owner). On success it prints the agent id, bucket, bundle CID, manifest CID, version CID, uploaded widgets, and the cubbies it declared. ## `cef widget push` Publishes a built widget subproject under the bucket’s `widgets` root. Run it in the widget’s directory after its own build. | Flag | Default | Meaning | | ------------------------------------------------------- | ------------------------ | ------------------------------------------------------------------------------- | | `--bucket ` | required | DDC bucket id. | | `--dir ` | current directory | The widget subproject. | | `--out ` | `dist` | Build output within it. | | `--widget-id ` | `cef.id` or package name | Registry id. | | `--widget-version ` | package version | Version. | | `--as-pubkey ` | — | Agent Service the widget belongs to; written into the entry as `:`. | | `--secret-phrase`, `--access-token`, `--subject-phrase` | env | Credentials. | | `--env`, `--endpoint`, `--preset`, `--cdn` | from `--env` | Network and endpoints written into the widget. | The `cef` block of `package.json` declares the widget (`id`, `name`, `entry`, `scope`, `cubbyAlias`, `queries`, `events`, `kind`, `config`). See [Build a widget](/agents/building-blocks/create-a-widget/#option-b-on-its-own). ## `cef cubby push` Publishes `cubbies//*.sql` declarations to the bucket. Creates no database: the platform applies them the first time an agent touches the cubby in a vault. | Flag | Default | Meaning | | ------------------------------------------------------- | ------------ | ----------------------------- | | `--bucket ` | required | DDC bucket id. | | `--dir ` | `./cubbies` | One subdirectory per cubby. | | `--alias ` | all | Push only this cubby. | | `--cubby-version ` | `0.1.0` | Version for each declaration. | | `--secret-phrase`, `--access-token`, `--subject-phrase` | env | Credentials. | | `--env`, `--endpoint`, `--preset`, `--cdn` | from `--env` | Network. | See [Cubby schema and migrations](/agents/building-blocks/cubby-schema/). ## `cef deploy` Applies every record in `deployments/` as one set: `PUT /api/v1/agents//deployments`. | Flag | Default | Meaning | | ------------------------ | ---------------------------------- | ------------------------------------------------------ | | `--env ` | `dev` | Environment. | | `--endpoint ` | `$CEF_ENDPOINT`, then from `--env` | Platform API. | | `--agent ` | the only built agent | Which `dist//` to read identity from. | | `--as-pubkey ` | manifest value | Agent Service pubkey. | | `--access-token ` | `$CEF_ACCESS_TOKEN` | CLI access token. | | `--version ` | — | Override `version` in every record (`latest` allowed). | | `--deployments ` | `deployments` | A folder, or one file holding a set or a record. | | `--out ` | `dist` | Build output. | | `--author ` | git `user.email` | Audit author. | | `--note ` | — | Audit note. | | `--dry-run` | — | Print the set without applying it. | | `--header ` | — | Extra request header; repeatable. | Record format: [Push and deploy](/agents/ship/push-and-deploy/#deployment-records). ## `cef dev` ```bash cef dev [widgetId] --as-pubkey ``` Serves a widget from `cef.config.ts` with the runtime injected, and reloads the browser on change. | Flag | Default | Meaning | | ------------------- | --------------------- | ------------------------------------------ | | `[widgetId]` | first declared widget | Widget to serve. | | `--config ` | `cef.config.ts` | Config file. | | `--port ` | a free port | Listen port (0–65535). | | `--host ` | `127.0.0.1` | Listen host. | | `--env ` | `dev` | Endpoints written into the manifest. | | `--as-pubkey ` | — | Needed for `connectAgent()` and `query()`. | | `--no-watch` | — | Do not watch for changes. | ## `cef publish` Sends the agent’s card (not its manifest) to the marketplace listing. Not needed to deploy, connect, or run an agent. | Flag | Default | Meaning | | ------------------------ | ----------------------------------------- | --------------------- | | `--env ` | `dev` | Environment. | | `--marketplace ` | `$CEF_MARKETPLACE_URL`, then from `--env` | Marketplace base URL. | | `--agent ` | the only built agent | Which `dist//`. | | `--out ` | `dist` | Build output. | | `--as-pubkey ` | config value | Agent Service pubkey. | | `--access-token ` | `$CEF_ACCESS_TOKEN` | CLI access token. | ## `cef eval` Workflow evaluations: datasets, experiments, comparisons. Subcommands: `datasets`, `create `, `pull `, `push `, `run `, `experiments [dataset]`, `compare `, `migrate`. Flags and usage: [Eval reference](/reference/eval/) and [Experiments](/agents/workflows/evaluate/experiments/). ## Related * [Push and deploy](/agents/ship/push-and-deploy/) * [Write an agent](/agents/code-agents/overview/) * [Bring your own (A2A)](/agents/llm-agents/bring-your-own/) * [Errors](/reference/errors/) # Errors > Error codes and failure messages from the vault API, the SDKs, the cef CLI, deployments, and runs, with what each means and what to do. Match on the `code`, not on the HTTP status. Vault API errors arrive as: ```json { "error": { "code": "MANIFEST_INVALID", "message": "…", "retryable": false } } ``` The vault SDK raises them as `VaultRequestError` with `status`, `code`, and `retryable`. Retry only when `retryable` is `true`. ## Connecting an agent | Code | Status | Meaning | Fix | | -------------------------- | ------- | ------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------ | | `MANIFEST_INVALID` | 400 | A requested scope is not in the manifest’s `requiredScopes`, the agent id does not match the manifest, or the manifest’s `kind` is invalid. | Connect into a declared scope, or push a version that declares it. | | `MANIFEST_NOT_FOUND` | 404 | No manifest for that agent id in the registry. | Check the id; push the agent. | | `SETTINGS_SCHEMA_MISMATCH` | 400 | Settings do not satisfy the agent’s settings schema. | Fix the values. | | `AGENT_ALREADY_CONNECTED` | 409 | A connection already exists. | Use it. | | `BUNDLE_CHANGED` | 409 | The `bundleCid` you named is not the agent’s current bundle. Nothing was stored. | Show the new code; connect again with its CID. | | `RECONSENT_REQUIRED` | 409 | A reconnect would run different code than the owner consented to, and no `bundleCid` was given. | Ask for consent; connect with the new `bundleCid`. | | `GAR_MISSING` | 412 | No consent agreement. | Connect again; it signs one. | | `GAR_EXPIRED` | 412 | The agreement was revoked or expired. | Connect again. | | `CUBBY_PROVISION_FAILED` | 500 | A cubby migration failed. | Fix it in a new migration file and push a new version. | | `AGENT_NOT_CONNECTED` | 404/409 | The agent has no connection in this vault. | Connect it. | | `ROLE_CONFLICT` | 409 | A connection can hold only one role (workflow agent, harvester, analyzer, embedder). | Clear the other role first. | | `SETTINGS_STALE` | 409 | The settings changed since you read them. | Re-read and retry. | See [Connect to a vault](/agents/ship/connect-to-a-vault/). ## Authentication | Code | Meaning | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------- | | `AUTH_MISSING` | No signature or credential on the request. | | `AUTH_AMBIGUOUS` | The request carries more than one auth scheme, or a principal that cannot perform the action (an agent deleting a vault object, for example). | | `INVALID_SIGNATURE`, `SIGNATURE_EXPIRED` | The request signature is wrong or too old. Check the clock. | | `AUTHTOKEN_INVALID`, `AUTHTOKEN_EXPIRED` | The bearer token is invalid or expired. Generate a new one. | | `NOT_VAULT_OWNER` | The action is owner-only. | | `SCOPE_DENIED` | The caller may not perform this action in that scope. | ## Events, objects, and registry | Code | Meaning | | ----------------------------------- | -------------------------------------------------------------------------------------------------------- | | `PAYLOAD_TOO_LARGE` | The event or object is too big. Store large data as an object and publish its path. | | `DUPLICATE_EVENT_ID` | A publish was rejected because an event with that id already exists. Safe to treat as already delivered. | | `VAULT_DISCONNECTED` | The vault’s storage credential can no longer be refreshed. | | `ALIAS_NOT_YOURS` | `cef push --vault`: the alias belongs to another publisher. | | `REGISTRY_UNAVAILABLE` | `cef push --vault`: the registry could not be written. Retryable. | | `CUBBY_UNAVAILABLE` | A cubby read or write failed. | | `SCHEDULE_NOT_FOUND` | No such schedule on the connection. | | `CONNECTION_*`, `CONNECTOR_UNKNOWN` | Connector connection errors. See [Connectors](/vaults/connectors/). | | `WEBHOOK_UNAUTHORIZED` | Webhook call with an unknown endpoint or a bad, revoked, or expired key. One answer for all, by design. | | `WEBHOOK_CREATOR_NO_ACCESS` | The key’s creator can no longer write the scope. | ## Vault SDK exceptions | Exception | Meaning | | --------------------------------------------------------------- | ------------------------------------------------------------------------------- | | `BundleChangedError`, `ReconsentRequiredError` | The two 409s above, with `currentCid`. | | `OnboardingRequiredError` | `ensure()` needs the wallet’s on-chain gateway registration first. | | `OnboardingTimeoutError` | Onboarding did not finish in time. `ensure()` is idempotent; retry. | | `VaultSignerRequiredError` | The method needs a `signer`. | | `CeilingUnknownError` | A member scope write needs the member’s privacy ceiling; pass it in `standing`. | | `vault.agents.connect requires \`garEndpoint\``/`a \`signer\`\` | Configure both on `VaultSDK`. | ## Widgets | Error | Meaning | | ------------------------------- | --------------------------------------------------------------------------------------------- | | `AgentNotConnectedError` | The reader’s vault has not connected the widget’s agent. Offer `connectAgent()`. | | `WidgetSignedOutError` | The host did not answer the identity handshake. Mount `createWidgetHost` on the framing page. | | `WidgetVaultUnreachableError` | The named vault cannot be opened by this reader. | | `WidgetWalletUnconfiguredError` | The manifest has no wallet origin. Push with `--env` so endpoints are written. | ## CLI | Message | Meaning | Fix | | ----------------------------------------------------------------------------------------------- | -------------------------------------------- | ---------------------------------------------------------------------------------------- | | `This DDC access token can't write bucket : it is signed by … but the bucket is owned by …` | The token is not rooted in the bucket owner. | Get **Invite to publish**; see [Team](/get-started/team/#bucket-belongs-to-other-owner). | | `This DDC access token expired at …` | Expired token. | Generate a new one in ROC. | | `bucket does not exist on this network` | Wrong `--bucket` or `--env`. | Check both. | | `choose a destination: --bucket … or --vault …` | No destination. | Pass one. | | `--card is required for an external agent` | `--kind external` without a card. | Pass `--card`. | | `… describes an external agent — pass --kind external` | A flag for the other kind. | Fix the flags. | | `bundle not found … run \`cef build\` first\` | Nothing built. | Run `cef build`. | | `multiple built agents in dist … pass --agent ` | Several agents built. | Pass `--agent`. | | `no CLI access token` | `deploy`/`publish` without a token. | Set `CEF_ACCESS_TOKEN`. | | `"weight" must be a positive integer (>= 1)` | Deployment weight. | Use integer shares like `90` and `10`. | | `all default records (empty targeting) must share one priority tier` | Defaults at different priorities. | Keep exactly one default record. | | `banned import ""` | Node built-in or banned package. | Use sandbox globals. | | `@OnEvent argument must be a string literal` | Computed event type. | Use a literal. | | `ctx.models. is not declared in cef.config.ts models` | Undeclared model. | Add it to `models`. | | `schedule "" publishes "", which this agent does not handle` | Schedule event unhandled. | Add an `@OnEvent` for it. | | `workflow "" is not runnable` | `defineWorkflow` validation. | Fix the listed steps. | | Exit code `2` from `cef inspect` | A routed handler is missing from the bundle. | Add the handler. | ## Deployments API | Code | Status | Meaning | | ------------------- | ------ | ----------------------------------------------------------------------------------------------------------------------------------------- | | `validation_failed` | 400 | The set broke a rule: names, weights, versions, targeting that does not compile, or not exactly one default record. The message names it. | | `revision_conflict` | 409 | Another apply landed first. Re-read and apply again. | | `apply_failed` | 500 | The platform could not store the set. | ## Runs A Task that fails carries an `error.code`: | Code | Meaning | | ----------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | `execute_failed` | The runtime call failed: the handler threw, or the call could not complete. | | `execute_timeout` | The task ran past its time budget. | | `running_lease_expired` | The task was claimed but never reported back; it is retried. | | `job_terminated` | The Job ended while the task was pending. | | `agent_invalid_result` | The handler returned a value that does not survive JSON serialization, such as a circular reference. | | `agent_threw_non_error` | The handler threw a non-`Error` value. Throw an `Error`. | | `agent_returned_undefined` | The bundle’s handler returned nothing the runtime could read. | | `gpu_units_ceiling`, `a2a_tokens_ceiling` | The connection reached its spend limit; see [Spend limits](/agents/ship/connect-to-a-vault/#spend-limits). | A Job ends with a reason: `revoked`, `idle_timeout`, `closed_by_agent`, or `failed`. `@OnClose` receives it. | Gateway code | Status | Meaning | | ----------------- | ------ | ----------------------------------- | | `NO_RUNTIMES` | 503 | No agent runtime is available. | | `FLEET_SATURATED` | 503 | Every runtime is busy. Retry later. | ## Related * [Test and debug a code agent](/agents/code-agents/test-and-debug/) * [Monitor runs](/agents/workflows/monitor-runs/#why-a-run-failed-parked-or-stalled) * [Connect to a vault](/agents/ship/connect-to-a-vault/) * [CLI reference](/reference/cli/) * [Vault SDK reference](/reference/vault-sdk/) # ESLint plugin (@cef-ai/eslint-plugin) > Reference for @cef-ai/eslint-plugin 1.0.0: the recommended config and each rule, which show cef build's checks in your editor. ```bash npm install -D @cef-ai/eslint-plugin@1.0.0 @typescript-eslint/parser ``` The plugin shows the checks `cef build` runs on agent code while you type. It needs ESLint 8 or later. ## Configure ```json { "parser": "@typescript-eslint/parser", "plugins": ["@cef-ai"], "extends": ["plugin:@cef-ai/recommended"], "rules": { "@cef-ai/cubby-declared-alias": ["error", { "aliases": ["history"] }], "@cef-ai/model-declared-alias": ["error", { "aliases": ["llm"] }] } } ``` `recommended` turns on every agent rule below as an error. ## Rules | Rule | Checks | Options | | ----------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ------------------------------- | | `on-event-literal` | `@OnEvent(…)` takes a string literal. | — | | `engagement-id-literal` | `@Engagement({ … })` has `id` and `goal` as string literals. | — | | `cubby-declared-alias` | `ctx.cubby(…)` takes a string literal naming a declared cubby. | `{ aliases: string[] }` | | `model-declared-alias` | `ctx.models.X` names a declared model. | `{ aliases: string[] }` | | `publish-declared-type` | Calls of the form `ctx.publish(type, …)` take a string literal in the declared set. It does not inspect `ctx.vault.publish`. | `{ knownEventTypes: string[] }` | | `no-banned-imports` | No Node built-ins and no sync database or network packages. | — | Without options, the alias rules check only that the argument is a literal. ### Banned imports Node built-ins, with or without `node:`: `fs`, `fs/promises`, `path`, `net`, `tls`, `dgram`, `dns`, `http`, `https`, `http2`, `child_process`, `worker_threads`, `cluster`, `os`, `process`, `readline`, `repl`, `vm`, `stream`, `zlib`, `buffer`, `crypto`. Packages: `pg`, `pg-native`, `mysql`, `mysql2`, `mongodb`, `redis`, `ioredis`, `sqlite3`, `better-sqlite3`, `ws`, `node-fetch`, `axios`. Use the sandbox globals instead: `fetch` for HTTP, `globalThis.crypto` for WebCrypto, and `ctx.cubby` for storage. ## Editor and build | Check | Editor (plugin) | `cef build` | | ---------------------------------------- | ---------------------------- | ----------------------------------------- | | Non-literal `@OnEvent` | error | error | | Undeclared model alias | error, when `aliases` is set | error | | Banned import | error | error | | Cubby alias not in the agent’s `cubbies` | error, when `aliases` is set | warning: the Agent Service may declare it | | Computed cubby alias | error | warning | ## Related * [Write an agent](/agents/code-agents/overview/#what-cef-build-rejects) * [Agent SDK reference](/reference/agent-sdk/) * [CLI reference](/reference/cli/) # @cef-ai/eval > API reference for @cef-ai/eval: the browser-safe library behind ROC Evaluations and cef eval, with dataset storage, case capture, scoring, experiments, and comparison. `@cef-ai/eval` is the library ROC’s **Evaluations** tab and `cef eval` are built on. Use it to script datasets and experiments, or to build your own evaluation tooling over the same storage. ```sh pnpm add @cef-ai/eval ``` | | | | ------------ | -------------------------------------- | | Version | `0.1.0` | | Dependencies | None. Browser-safe: no Node built-ins. | | Module | ESM | The library never talks to the network itself. You inject two things: an [`EvalStore`](#evalstore) for storage and a [`WorkflowTarget`](#workflowtarget) for starting runs. For concepts, see the [Evaluations overview](/agents/workflows/evaluate/overview/). ## Types ### Case ```ts interface Case { id: string; // [A-Za-z0-9._-]{1,80} input: Record; // the workflow.start payload expected?: Record; // keyed by Result field; dot paths allowed limits?: Limits; tags?: string[]; notes?: string; source?: { context: string; workflowVersion?: string; capturedAt: string }; } interface Limits { maxDurationMs?: number; maxTokens?: number; maxSteps?: number; maxCost?: number } type Matcher = | { $eq: Json } | { $oneOf: Json[] } | { $contains: Json } | { $regex: string } | { $between: [number, number] } | { $gte: number } | { $lte: number } | { $exists: boolean } | { $approx: number; $tol: number }; ``` Matching rules: [Results](/agents/workflows/evaluate/results/#matching). ### Header and VersionManifest ```ts interface Header { dataset: string; workflowId: string; description?: string; limits?: Limits; createdAt: string; baselineId?: string; } interface VersionManifest { version: number; header: Header; createdAt: string; note?: string; cases: Array<{ id: string; rev: number; sha256: string }>; } ``` ### Experiment and RunRecord ```ts interface Experiment { id: string; // exp--<4 chars> dataset: string; datasetVersion: number; workflowId: string; workflowVersion: string; live?: boolean; repeats: number; resultFields?: ResultField[]; attempt?: number; startedAt: string; finishedAt?: string; status: "running" | "done" | "cancelled" | "error"; origin: "roc" | "cli"; createdBy?: string; baselineId?: string; note?: string; error?: string; summary?: Summary; } interface RunRecord { caseId: string; caseRev: number; repeat: number; attempt?: number; context: string; runId: string; status: "done" | "failed" | "timeout" | "error"; output: Json | null; outcome?: string; error?: string; startedAt: string; endedAt: string | null; metrics: Metrics; scores: Score[]; pass: boolean | null; // null: no expectations, or the harness errored } interface Metrics { durationMs: number | null; inputTokens: number; outputTokens: number; totalTokens: number; cost: number; steps: number; modelCalls: number; cpuUnits: number; gpuUnits: number; } interface Score { name: string; // field path, or limit. value: number | string | boolean | null; pass: boolean | null; // null: not judged scorer: string; // "eval/field@1" | "eval/limit@1" source: "eval" | "annotation"; detail?: string; } ``` `RunRecord.status`: `done` the workflow completed; `failed` it announced a failure; `timeout` it did not end in time; `error` the harness could not run it. ### Summary ```ts interface Summary { cases: number; runs: number; passed: number; failed: number; errored: number; unscored: number; passRate: number | null; // passed / (passed + failed) durationMs: { p50: number | null; p95: number | null; mean: number | null }; tokens: { total: number; meanPerRun: number }; cost: { total: number; meanPerRun: number }; steps: { meanPerRun: number }; limitBreaches: number; fields: Record; } ``` ## EvalStore ```ts interface EvalStore { list(prefix: string): Promise; // every key under prefix, sorted get(key: string): Promise; // null when absent; other failures throw put(key: string, body: string): Promise; putIfAbsent(key: string, body: string): Promise; // true when this call wrote readonly atomicity?: "conditional-write" | "single-threaded"; listPrefixes?(prefix: string): Promise; // optional: immediate child "directories" } ``` | Export | Meaning | | ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | | `memoryStore(seed?)` | An in-memory store for tests and dry runs, with `dump()`. | | `storeConformance(makeStore)` | A list of `{ name, run }` checks to run against your own implementation (for example, one per `it`). `putIfAbsent` must be a real compare-and-set. | ## Datasets and cases Every dataset function takes `(store, workflowId, dataset, …)`. | Function | Returns | Meaning | | ---------------------------------------------------------------------- | --------------------------- | ---------------------------------------------------------------- | | `createDataset(store, { dataset, workflowId, description?, limits? })` | `Header` | Refuses a name the workflow already has. | | `getHeader` / `requireHeader` | `Header \| null` / `Header` | | | `updateHeader(store, header)` | `Header` | | | `setBaseline(store, wf, dataset, expId \| null)` | `Header` | Set or clear the baseline. | | `listWorkflows(store)` | `string[]` | Workflows with anything stored. | | `listDatasets(store, wf)` | `DatasetSummary[]` | `{ dataset, cases, latestVersion, experiments, baselineId?, … }` | | `addCase(store, wf, dataset, c)` | `{ rev }` | Refuses an id that exists. | | `saveCase(store, wf, dataset, c)` | `{ rev, changed, created }` | First revision, next revision, or nothing when identical. | | `getCase(store, wf, dataset, id)` | `{ case, rev } \| null` | Latest revision of a live case. | | `getCaseRevision`, `listCaseRevisions` | | One revision; all revision numbers. | | `listCaseIds`, `listArchivedCaseIds`, `listDeletedCaseIds` | `string[]` | | | `archiveCase` / `unarchiveCase` | | Hide from new versions; bytes untouched. | | `deleteCase` / `undeleteCase` | `{ pinnedBy }` / | Tombstone; revisions and the id stay. | | `loadWorkingSet(store, wf, dataset)` | `{ header, cases }` | Every live case at its latest revision. | ## Versions | Function | Meaning | | --------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | | `ensureDatasetVersion(store, wf, dataset)` | The latest version when the working set matches it; otherwise cuts the next with the note `auto: …`. Returns `{ version, created, diff, manifest }`. | | `cutVersion(store, wf, dataset, { note? })` | Cut a version explicitly. | | `listVersions`, `latestVersionNumber`, `getVersion` | Read versions. | | `loadVersion(store, wf, dataset, version)` | The cases a version pins, verified by sha256. | | `versionHistory(store, wf, dataset)` | Each version with what changed and how many experiments ran it. | | `diffVersionEntries(from, to)` | `{ added, removed, edited }` | ## Cases from runs ```ts caseFromRun(run: CapturedRun): Case interface CapturedRun { input: Record; // workflow.start payload; an inline `graph` is dropped output?: Json | null; // workflow.completed output context: string; workflowVersion?: string; id?: string; // default: derived from the context capturedAt?: string; tags?: string[]; notes?: string; } ``` `expected` is built from an object output with `expectationFrom`: labels exactly, free text (`isFreeText`: over 60 characters or over 8 words) as `{ $exists: true }`. `slugCaseId(text)` makes a valid id from any text; `uniqueCaseId(base, taken)` appends `-2`, `-3`, … until it is free. ## Running experiments ```ts import { runExperiment, resultFieldsFromGraph } from "@cef-ai/eval"; const exp = await runExperiment({ store, workflowId: "ticket-triage", dataset: "triage", workflowVersion: "1.4.0", live: false, // not live: target.pinVersion is required resultFields: resultFieldsFromGraph(graph), repeats: 3, target, onProgress: ({ done, total, run }) => console.log(done, total, run.caseId, run.pass), }); ``` ### RunExperimentOptions | Option | Default | Meaning | | ------------------------------------------- | ------------------------------- | -------------------------------------------------------------------------------- | | `store`, `workflowId`, `dataset` | | Required. | | `workflowVersion` | | Required. The version evaluated (recorded only, when `live`). | | `live` | | Required. `true` runs on the live deployment; `false` needs `target.pinVersion`. | | `target` | | Required. A [`WorkflowTarget`](#workflowtarget). | | `datasetVersion` | `ensureDatasetVersion` | Version to run. | | `repeats` | `1` | Runs per case. | | `concurrency` | `4` (`DEFAULT_CONCURRENCY`) | Runs in flight. | | `timeoutMs` | `300000` (`DEFAULT_TIMEOUT_MS`) | Per run. | | `resultFields` | | The version’s declared Result; expectations it rules out score n/a. | | `baselineId`, `note`, `origin`, `createdBy` | | Recorded on the experiment. | | `experimentId` | | Resume this experiment: runs on file are reused, a new attempt runs the rest. | | `signal` | | `AbortSignal`: stop dispatching; the experiment ends `cancelled`. | | `onProgress` | | `({ done, total, run, reused })` after each run. | The meta is written first (`running`), each run record as it finishes, the summary last. Each run publishes into its own context, `eval----a`; `runOne` runs and scores a single case without storage. ### WorkflowTarget ```ts interface WorkflowTarget { start(input: Record, context: string): Promise; waitForEnd(context: string, timeoutMs: number): Promise; metrics(context: string): Promise; pinVersion?(version: string, contextPrefix: string): Promise<() => Promise>; } interface RunEnd { status: "done" | "failed" | "timeout"; output?: Json; outcome?: string; error?: string; startedAt?: string; endedAt?: string; // server timestamps } ``` `start` publishes `workflow.start` with the input into the context; `waitForEnd` waits for `workflow.completed` or `workflow.failed`; `metrics` reads what the run cost; `pinVersion` routes contexts starting with the prefix to a version until the returned release function is called. ### Reading experiments | Function | Meaning | | ------------------------------------------ | -------------------------------------------------- | | `getExperiment(store, wf, expId)` | The experiment, or null. | | `listExperiments(store, wf, { dataset? })` | All experiments of a workflow, or of one dataset. | | `loadExperiment(store, wf, expId)` | `{ experiment, runs }` | | `listRunRecords`, `getRunRecord` | Run records. | | `findExperiment(store, expId)` | Find an experiment when the workflow is not known. | ## Scoring | Function | Meaning | | ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | | `scoreFields(expected, output, resultFields?)` | One `eval/field@1` score per expected field. | | `scoreLimits(limits, metrics)` | One `eval/limit@1` score per limit in force. | | `effectiveLimits(header, case)` | The header’s limits, each overridden by the case’s. | | `casePass(scores)` | True when every judged field passes; null when none were judged. Limits never count. | | `matchValue(expected, actual)` | `null` on a match, otherwise the reason. | | `readField(value, path)` | Read a dot path. | | `resultFieldsFromGraph(graph)` | The fields a graph’s Result steps declare (accepts the graph, its JSON text, or a manifest’s `params.graph`). `undefined` when none declares a Result. | | `checkResultCompatibility(resultFields, cases)` | `{ missing, typeChanged, unexpectedNew }`: which expectations a version cannot satisfy. | | `unsafeRegex(pattern)` | The reason a `$regex` pattern is refused, or null. `MAX_REGEX_LENGTH` is 256. | ## Summaries and comparison | Function | Meaning | | ------------------------------------- | ------------------------------------------------------------------------------------------------------------- | | `summarize(runs)` | A `Summary` over run records. | | `compareExperiments(base, candidate)` | Each side is `{ experiment?, runs, cases? }`. Compares only cases present on both sides at the same revision. | `ExperimentComparison`: | Field | Meaning | | ------------------- | --------------------------------------------------------------------------------------------------------------------------- | | `base`, `candidate` | `{ id, workflowVersion?, datasetVersion? }` | | `comparedOn` | Number of cases compared | | `datasetDiff` | `{ added, removed, edited }` when both sides’ cases are given | | `counts` | Runs per change: `improved`, `regressed`, `unchanged`, `only-in-base`, `only-in-candidate` | | `cases`, `runs` | Per case and per run: the change, pass rates or statuses, and deltas | | `metrics` | `{ base, candidate, delta, relative }` for pass rate, duration p50 and p95, tokens/run, cost/run, steps/run, limit breaches | A pass ranks above “not judged”, which ranks above a failure; a case’s change is decided by its mean rank over repeats. ## Storage layout helpers Key builders (`headerKey`, `caseRevKey`, `versionKey`, `experimentMetaKey`, `runKey`, …) and validators (`isWorkflowId`, `isDatasetName`, `isCaseId`, `isExperimentId` and their `assert…` forms) expose the layout under `workflows//`. `canonicalJson` and `sha256Hex` are how versions pin case content. `parseCase`, `parseHeader`, `parseExperiment`, `parseRunRecord` validate stored JSON and throw `EvalContractError`. ## Related * [Evaluations overview](/agents/workflows/evaluate/overview/) * [Datasets](/agents/workflows/evaluate/datasets/) * [Experiments](/agents/workflows/evaluate/experiments/) * [Results](/agents/workflows/evaluate/results/) * [CLI reference](/reference/cli/) # Glossary > Definitions of the core Manykind terms — Agent Service, vault, connection, workflow, agent, Memory Bank, connector, Job, scope, and more — with links to the pages that cover each. ## Agent What runs on a vault’s data. Three kinds: a [workflow](#workflow), an [LLM agent](#llm-agent), and a [code agent](#code-agent). All three publish, deploy, and connect the same way, and a workflow can use either of the other two as a step. See [Agents and workflows](/agents/overview/). ## Agent id `:`: the Agent Service’s public key, then the agent’s name. The platform derives the publisher from the prefix. ## Agent Service The developer’s account on Manykind: it holds your workflows, LLM agents, and code agents, plus the cubby schemas, widgets, datasets, and members that belong to them. Created in ROC. See [Your Agent Service](/get-started/agent-service/). ## Agreement The signed record of a vault owner’s consent to connect an Agent Service’s agent to named scopes of their vault. Held in the agreement registry; revoking it removes the authority. See [Connections and consent](/vaults/connections-and-consent/). ## Bundle An agent’s built code, stored content-addressed in the Agent Service’s bucket. A connection pins the bundle the owner consented to run. ## Code agent TypeScript engagement classes that react to events, written with `@cef-ai/agent-sdk`. See [Write an agent](/agents/code-agents/overview/). ## Connection 1. **Agent connection:** the vault’s record that an agent may run on named scopes, with its settings, pinned bundle, and status (`provisioning`, `active`, `revoking`, `revoked`). 2. **Connector connection:** one configured account of a [connector](#connector) in a vault. ## Connector A kind of outside system a vault supports: Slack, Email (SMTP), Telegram, or an MCP server. Workflows use a vault’s connections as triggers and actions. See [Connectors](/vaults/connectors/). ## Context The field on an [event](#event) that names its stream. Events sharing a context form one stream and one [Job](#job). ## Cubby A SQLite database an Agent Service declares. Every agent of the service shares it by alias, one per vault. See [Cubbies](/agents/building-blocks/cubbies/). ## Dataset, case, experiment, result Evaluation terms. A dataset is a workflow’s set of cases; a case is one input and the expected result; an experiment runs a dataset version on a workflow version; a result is what a workflow declares as its output. See [Evaluations](/agents/workflows/evaluate/overview/). ## Engagement The building block of a code agent: a class marked `@Engagement({ id, goal })` whose `@OnEvent` methods handle event types. One engagement is selected per Job and pinned. ## Event One typed message published into a vault scope, with `type`, `context`, and `payload`. The only way work enters the platform. See [Runs and events](/agents/how-agents-run/). ## Execution token The short-lived credential a running agent or workflow uses to reach the vault, minted per task and bound to the connection. Your code never handles it directly. ## Item One element of a list a workflow processes with split and aggregate steps. See [Items](/agents/workflows/split-and-aggregate/). ## Job A long-lived run, keyed by vault, agent, scope, and context. The record of a workflow run is a Job. Each event in it is a [Task](#task). ## LLM agent An agent defined by instructions, a model, tools, connections, and scopes, created in ROC and hosted on the Agent Bridge; or an A2A agent you run elsewhere, registered from its Agent Card (`cef push --kind external`). See [LLM agents](/agents/llm-agents/overview/). ## Manifest The JSON document `cef build` produces for an agent: identity, bundle, models, cubbies, settings schema, event schemas, widgets. The vault reads it at connect time. ## Manykind The platform these docs describe: Agent Services on the developer side, vaults on the customer side, and the connection between them. Its packages are published under `@cef-ai/*`, and its CLI is `cef`. See [How it fits together](/get-started/how-it-fits-together/). ## Member Someone other than the owner who may use a vault, with a role (`admin` or `member`), per-scope levels (Write, Read, Audit), and an optional privacy ceiling. See [Members](/vaults/members/). ## Memory Bank The vault’s durable graph of records and relations that its agents and workflows filed, each with a scope and a privacy class. See [Memory Bank](/vaults/memory-bank/). ## Model alias The name an agent or workflow calls a model by (`ctx.models.`), bound in config to a model in the platform’s catalogue. See [Models](/agents/building-blocks/models/). ## Privacy class How sensitive a Memory Bank record or relation is: `public`, `internal`, `private`, or `restricted`. Required on every write. ## ROC The Manykind web console: where you create Agent Services, build and run workflows, and browse connectors, models, members, and the Memory Bank. See [How it fits together](/get-started/how-it-fits-together/#where-the-tools-fit). ## Run One execution of a workflow or agent, from a trigger event to its result. Its record is a [Job](#job). ## Scope A named partition of a vault. Events, objects, and grants are per scope. Shown as **Domains** in ROC. See [Vaults](/vaults/overview/#scopes). ## Step One unit of a workflow’s graph: a trigger, a model call, an agent, a person, a cubby read, a connector action, and so on. See [Steps](/agents/workflows/steps/). ## Task One event’s worth of work inside a [Job](#job), with its own status, attempts, and logs. ## Trigger The step that starts a workflow run: an event, a schedule, a webhook, or a connector message. See [Triggers](/agents/workflows/triggers/). ## Vault The customer side: a person’s or an organization’s own store, owned by a wallet. Holds scopes, events, objects, the Memory Bank, connectors, agent connections, and members. See [Vaults](/vaults/overview/). ## Wallet The cryptographic identity that owns a vault and signs agreements and grants. ## Widget A browser screen an Agent Service publishes. It acts as the signed-in person and reads the service’s cubbies and the vault’s Memory Bank. See [Widgets](/agents/building-blocks/widgets/). ## Workflow An agent whose behaviour is a typed graph of steps, built in the ROC Workflow Builder or in code with `defineWorkflow`. See [Workflows overview](/agents/workflows/overview/). ## Related * [How it fits together](/get-started/how-it-fits-together/) * [Agents and workflows](/agents/overview/) * [Vaults](/vaults/overview/) * [Errors](/reference/errors/) # Limits > The hard limits the platform enforces on webhooks, events, schedules, connector actions, workflow steps, cubbies, logs, and evaluations, with what happens when you hit each and what to do. These limits are enforced by the platform or the SDK. The values are the defaults the platform runs with. Each limit links to the page of the feature it belongs to, which states it too. Budgets you choose for your own workflow (steps per run, payload sizes you assert in tests) are on [Best practices](/agents/workflows/best-practices/). ## Webhook triggers | Limit | Value | When you hit it | What to do | | ---------------------------------------------------------------------- | ----------------------------------------------------- | ----------------------------------------------------------------- | -------------------------------------------------------------------------------- | | [Request body](/agents/workflows/triggers/#webhook-limits) | 1 MiB | `413 PAYLOAD_TOO_LARGE`, “webhook body too large” | Send a reference (an object path, a URL, an id) and fetch the content in a step. | | [Body shape](/agents/workflows/triggers/#call-it) | A JSON object or array | `400`. An array is delivered as `{ "body": [...] }`. | Send an object. | | [Calls per key](/agents/workflows/triggers/#webhook-limits) | 60 per minute, bursts of 10 | `429 RATE_LIMITED` with a `Retry-After` header in seconds | Honour `Retry-After`. Give each caller its own key. | | [`Idempotency-Key` header](/agents/workflows/triggers/#webhook-limits) | 255 bytes | `400` “Idempotency-Key is longer than 255 bytes” | Use a short id (a UUID or your request id). | | [Waiting for the result](/agents/workflows/triggers/#webhook-limits) | 30 s when the endpoint sets no timeout; 120 s at most | `202` with `status`, `runId`, and `statusUrl`; the run carries on | Poll `statusUrl`, or respond immediately and read the result later. | ## Events and the vault | Limit | Value | When you hit it | What to do | | ------------------------------------------------------------------ | ------ | ----------------------------------------------------------------- | ----------------------------------------------------------------------------------- | | [Publish request body](/agents/how-agents-run/#limits) | 1 MiB | `413 PAYLOAD_TOO_LARGE`. A workflow’s publish step fails the run. | Keep events small; declare each publish step’s `payload`. | | [Events per publish call](/agents/how-agents-run/#limits) | 100 | `413 PAYLOAD_TOO_LARGE` | Publish in batches of 100 or fewer. | | [Event retention per stream](/agents/how-agents-run/#limits) | 7 days | Older events are dropped | Keep anything you need later in a cubby or the [Memory Bank](/vaults/memory-bank/). | | [Agent push (`cef push`) body](/agents/ship/push-and-deploy/#push) | 64 MiB | The push is refused | Ship large assets as objects, not in the bundle. | ## Schedules | Limit | Value | When you hit it | What to do | | ---------------------------------------------------------- | ------------------------------------------------------------------- | --------------------------------------------- | ---------------------------------------------------------------- | | [Granularity](/agents/workflows/triggers/#schedule-limits) | 1 minute: a 5-field cron, or `@hourly` / `@daily`-style descriptors | An expression with a seconds field is refused | Use a standard 5-field cron. | | [Late fire](/agents/workflows/triggers/#schedule-limits) | A fire more than 4 minutes late is skipped | That tick does not run | Make each run cover “since the last run”, not “the last minute”. | ## Connector actions | Limit | Value | When you hit it | What to do | | --------------------------------------------------- | --------------- | ---------------------------------------------------- | ------------------------------------- | | [Time per attempt](/agents/workflows/steps/#limits) | 20 s | The attempt fails and is retried | Keep actions small. | | [Attempts](/agents/workflows/steps/#limits) | 5, with backoff | The action step fails with `connector.action.failed` | Handle the failure path in the graph. | ## Workflow runner | Limit | Value | When you hit it | What to do | | ------------------------------------------------------------------------------------- | ------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------- | | [Input held by a waiting step](/agents/workflows/steps/#limits) | 1 MiB (1,048,576 bytes) | The step fails: “the input carried by step ‘…’ is N bytes, over the 1048576-byte limit a waiting step can hold” | Shrink the carried item before agent and human steps. | | [Items in flight per region](/agents/workflows/split-and-aggregate/#limits) | 8 in a parallel region, 1 in a sequential one | Further items wait their turn | Nothing; size your item count for the latency you need. | | [Loop bound (`loop.max`)](/agents/workflows/people/#loop-back-edges) | 1 to 100 | The graph is refused at validation | Bound loops tightly and set `exhausted`. | | [Run participants](/agents/workflows/people/#participants) | 2 to 64, including the initiator | The run is refused | Split large groups. | | [Exchanges in one step (agent asks, person answers)](/agents/workflows/steps/#limits) | 12 | The step fails: “gave up after 12 exchanges without a final answer” | Ask for everything at once. | | [Model call attempts](/agents/workflows/steps/#limits) | 3 (transport failures only, 400 ms then 1600 ms apart) | The model step fails | Treat a failed model step as a normal failure path. | | [Human step answer: note, or one value](/agents/workflows/people/#limits) | 20,000 characters each | The answer is rejected (`note_too_long`, `invalid_field:`) | Ask for less, or collect a document. | | [Human step answer: values](/agents/workflows/people/#limits) | 64 | The answer is rejected (`too_many_values`) | Split the form. | | [Document saved in a step](/agents/workflows/people/#limits) | 2,000,000 characters | The save is rejected (`invalid_field:body`) | Store large documents as objects. | | [SQL in a Cubby step](/agents/workflows/steps/#limits) | No `{{ $json.… }}` templates | The step fails: “SQL may not interpolate the run’s data” | Use `?` and `args`. | `defineWorkflow` refuses at build time: duplicate step ids, an unknown step kind, an agent step without `use`, a model step whose alias is not in `models`, a graph with no trigger, an edge to an unknown step, and a step nothing is wired into. ## Cubbies | Limit | Value | When you hit it | What to do | | ------------------------------------------------------------ | ----- | ------------------------------------------------------------------ | ---------------------------------------------------------- | | [Size of one cubby](/agents/building-blocks/cubbies/#limits) | 10 GB | Writes are refused: `413 QUOTA_EXCEEDED`, “storage quota exceeded” | Delete or archive old rows; keep large content as objects. | ## Logs and run history | Limit | Value | When you hit it | What to do | | ----------------------------------------------------------------------- | ----------- | -------------------------- | ---------------------------------------------- | | [Task log query window](/agents/workflows/monitor-runs/#limits) | 24 hours | Older logs come back empty | Record anything you need later in a cubby row. | | [Finished Jobs and their Tasks](/agents/workflows/monitor-runs/#limits) | Kept 7 days | Older runs are deleted | Capture runs worth keeping as dataset cases. | ## Evaluations | Limit | Value | When you hit it | What to do | | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------- | ------------------------------------ | | [Dataset name](/agents/workflows/evaluate/datasets/#create-a-dataset) | `^[a-z0-9][a-z0-9-]{0,62}$` | Refused | Lowercase, digits, hyphens. | | [Case id](/agents/workflows/evaluate/datasets/#case-shape) | `[A-Za-z0-9._-]`, 1 to 80 characters | Refused | | | [`$regex` pattern](/agents/workflows/evaluate/results/#limits) | 256 characters, no nested quantifiers | The field fails: `unsafe $regex refused: …` | Simplify the pattern. | | [**Run experiment** in ROC](/agents/workflows/evaluate/experiments/#run-an-experiment-in-roc) | Repeats 1 to 20, Concurrency 1 to 16, Timeout 10 to 3600 s | The dialog refuses to start | Use `cef eval run` for other values. | ## Related * [Best practices](/agents/workflows/best-practices/) * [Monitor runs](/agents/workflows/monitor-runs/) * [Triggers](/agents/workflows/triggers/) * [Errors](/reference/errors/) # @cef-ai/testing > API reference for @cef-ai/testing: testAgent for one agent, testPlatform for agents and workflows against a simulated vault, mock agents and models, fetch and object-store fakes, and fault injection. `@cef-ai/testing` runs agents and workflows in-process against a simulated platform: an in-memory vault, the job machinery, and SQLite cubbies built from your real migrations. No Docker, no GPU, no network. ```sh pnpm add -D @cef-ai/testing vitest ``` | | | | ---------- | ---------------------------------------------------------- | | Version | `3.3.5` | | Depends on | `@cef-ai/agent-sdk`, `@cef-ai/vault-sdk`, `better-sqlite3` | | Module | ESM | Exports: `testAgent`, `testPlatform`, `mockAgent`, `testWallet`, `fromMarketplace`, `createModelMock`, `fakeObjectStore`, and their types. ## Which entry point | Use | When | | ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- | | `testAgent(Agent, opts)` | Unit tests of one code agent’s handlers: dispatch an event, assert publishes and cubby rows. | | `testPlatform({ agents, models })` | A workflow (register `WorkflowRunner` as an agent), several agents that talk to each other, or app code that uses the vault. | To test a workflow, see the worked example in [Test a workflow](/agents/workflows/test/#1-local-tests). ## testAgent ```ts testAgent(source: AgentCtor | TestAgentMultiSpec, opts?: TestHarnessOptions): TestHarness ``` ```ts import { afterEach, expect, it } from "vitest"; import { testAgent, type TestHarness } from "@cef-ai/testing"; import Echo from "../src/agent.js"; let h: TestHarness | undefined; afterEach(() => h?.dispose()); it("acks a message", async () => { h = testAgent(Echo, { cubbies: [{ alias: "history", migrations: "./cubbies/history" }] }); const events = await h.dispatch({ type: "user_message", payload: { text: "hello" } }); expect(events.map((e) => e.type)).toContain("ack"); }); ``` `source` is a class with `@OnEvent` / `@OnStart` / `@OnClose` handlers, or `{ engagements: [{ id, source, goal? }], targeting? }` for several engagements. ### TestHarnessOptions | Option | Type | Meaning | | ---------- | ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------ | | `cubbies` | `Array` | Cubbies to create, each from a directory of `*.sql` migrations. An alias not listed is created empty on first use. | | `models` | `Record` | `ctx.models[alias]`. | | `settings` | `Record` | `ctx.settings`, frozen. | | `params` | `Record` | `ctx.params`, frozen. Defaults to `{}`. | | `clock` | `{ now?: number }` | Initial time in unix milliseconds. Defaults to `Date.now()`. | | `self` | `{ agentId?, vaultId?, scope?, context? }` | `ctx.self`. Defaults to a valid synthetic identity. | ### TestHarness | Member | Meaning | | ------------------------------------------------------ | ------------------------------------------------------------------------------------- | | `dispatch({ type, payload?, context?, role?, from? })` | Run one event through the agent. Resolves to the events published during this call. | | `published` | Every event published since the harness was created. | | `cubby(alias)` | A cubby handle (`query`, `exec`) for assertions. | | `runInCubby(alias, fn)` | Call `fn(handle, db)` with the cubby handle and the raw `better-sqlite3` database. | | `start()` / `close(reason?)` | Run `@OnStart` / `@OnClose`. | | `lifecycle` | `{ state: "idle" \| "started" \| "closed", terminalReason? }`. | | `engagement` | The engagement the harness selected (multi-engagement sources). | | `snapshot()` / `restore(snap)` | Capture and restore cubbies, clock, and the publish log. | | `replay(jsonlPath)` | Feed a JSONL file of events through the harness. | | `advanceTime(ms)` / `now()` | Move and read the injected clock. | | `fetchMock()` | The [fetch mock](#fetch-mock) installed as the global `fetch` during each invocation. | | `models` | The model handles on `ctx.models`. | | `dispose()` | Close SQLite handles. Idempotent; call it in `afterEach`. | Agents call the global `fetch` and log with `console.*`; there is no `ctx.fetch` or `ctx.log`. ## testPlatform ```ts testPlatform(opts: TestPlatformOptions): TestPlatform ``` ```ts import { testPlatform, createModelMock } from "@cef-ai/testing"; import { WorkflowRunner } from "@cef-ai/agent-sdk/workflow"; import config from "../cef.config.js"; const p = testPlatform({ agents: { "ticket-triage": { source: WorkflowRunner, cubbies: config.cubbies, params: { graph: config.params!.graph!.default } }, }, models: { classifier: createModelMock() }, }); await p.vault.agents.connect({ agentId: "ticket-triage" }); await p.vault.scope("default").publish({ type: "workflow.start", context: "t-1", payload: { ticketId: "T-1", text: "…" } }); ``` ### TestPlatformOptions | Option | Type | Meaning | | --------- | ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `agents` | `Record` | Required. Each entry is a class, a `mockAgent(...)`, a long form `{ source, cubbies?, settings?, params? }`, or `{ engagements, cubbies?, settings?, targeting? }`. | | `models` | `Record` | Shared by every agent’s `ctx.models`. | | `users` | `Record` | One vault per user. Omit for a single implicit user. | | `vaultId` | string | The implicit user’s id when `users` is omitted. Default `"default"`. | | `scope` | string | Default scope for publishes that omit one. | | `context` | string | Default context for publishes that omit one. Default `"default"`. | | `clock` | `{ now?: number }` | Initial time. | `params` on a registration is `ctx.params`. A workflow runner reads its graph from `params.graph`, as on the platform. Cubbies are keyed by Agent Service (the part of the agent id before `:`) and by user, as on the platform: agents of one service that declare the same alias share one database, and each user’s vault has its own. ### TestPlatform | Member | Meaning | | ------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | `vault` | The single user’s vault. Throws when two or more `users` are declared. | | `user(id).vault` | One user’s vault. | | `vault.agents.connect({ agentId })` | Connect a registered agent; `disconnect`, `get`, `list`. | | `vault.scope(name).publish({ type, context?, payload, from? })` | Publish an event; agents whose handlers match run. | | `vault.scope(name).stream(context).events.list({ types?, limit? })` | Read a context’s events: `{ items, hasMore }`. | | `vault.scope(name).subscribe(...)` | Subscribe with a handler, or iterate events. | | `orchestrator.jobs`, `.list(filter?)`, `.get(jobId)` | Read-only view of Jobs across all users. | | `runInCubby(agentId, alias, fn, userId?)` | Call `fn(handle, db)` against an agent’s cubby. | | `cubbyOps()` | Every cubby statement any agent ran, oldest first: `{ agentId, cubby, op, sql, params, nodeId? }`. `nodeId` names the workflow step that ran it. | | `objects` | The in-memory `ctx.vault.objects` store (a [`FakeObjectStore`](#fakeobjectstore)). | | `faults.refusePublish` | `(type, payload) => boolean`. Return true to make `ctx.vault.publish` fail, as the vault does for a body over 1 MiB. | | `clock` | Shared clock. | | `fetchMock()` | Shared fetch mock. | | `models` | Shared model handles. | | `dispose()` | Tear down every simulated agent. Idempotent. | ## mockAgent ```ts mockAgent(spec: MockAgentSpec): MockAgentDescriptor ``` A stand-in for an agent you do not want in the test: handlers as plain functions. | Field | Meaning | | -------------------------------------- | ---------------------------------------------------------------------------------- | | `agentId` | Required. | | `on` | `Record unknown>` | | `onStart`, `onClose` | `(ctx, reason?) => unknown` | | `publishes` | Event types it publishes (descriptive) | | `engagements`, `targeting`, `settings` | The multi-engagement form: `engagements: [{ id, goal?, on?, onStart?, onClose? }]` | A workflow’s agent step publishes `workflow.step` to the agent and waits for `agent.answered`. A mock answers like this: ```ts const reviewer = mockAgent({ agentId: "reviewer", on: { "workflow.step": async (event, ctx) => { const input = event.payload as { text: string }; await ctx.vault.publish("agent.answered", { agent: "reviewer", text: JSON.stringify({ approved: input.text.length > 10 }) }); // or, to fail the step: { agent: "reviewer", error: "upstream unavailable" } }, }, }); ``` Register it under the same id the step’s `use` names, and connect it like any agent. ## createModelMock ```ts createModelMock(): ModelMockHandle ``` | Member | Meaning | | ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `expect(input).respond(output)` | Register an answer. An `infer` call whose input is deep-equal (by `JSON.stringify`) to `input` returns `output`. A later expectation for the same input wins. | | `calls` | Every call: `{ method: "infer" \| "stream", input }`. | | `infer(input)` | Returns the matching output, or throws `[modelMock] no expectation for input …`. | | `stream(input)` | Yields each element of an array output, or the output once. | Any object with an `infer` method is accepted wherever a model handle is, so a hand-written stub works too. A workflow’s model step resolves its `input` params against the item before calling the model; register the **resolved** input. ## Fetch mock | Member | Meaning | | ------------------------------------------------------- | ---------------------------------------------------------------------- | | `when({ url, method? }).reply(status, body?, headers?)` | Stub a request. `url` is a string or RegExp; stubs are tried in order. | | `assertCalled({ url, method? })` | Throw if no recorded call matched. | | `calls` | `{ url, method }` for each call. | ## FakeObjectStore `fakeObjectStore()` returns the store behind `TestPlatform.objects`. Objects are create-only, as on the platform: a second upload to the same path fails with `409`. | Member | Meaning | | ---------------------------- | ------------------------------------------------------------ | | `text(vaultId, scope, path)` | Read an uploaded object back as text. | | `objects` | The underlying map. | | `refuseUpload` | `(path) => boolean`. Return true to fail uploads with `503`. | | `handle(vaultId, scope)` | The `ctx.vault.objects` handle for that vault and scope. | ## Other exports | Export | Meaning | | ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `testWallet(id)` | `{ publicKey: "wallet-" }` for `users[id].wallet`. Signatures are not verified. | | `fromMarketplace(url)` | Throws: not available in the simulator. Use a class or `mockAgent`. | | Types from `@cef-ai/vault-sdk` | `AgentConnection`, `EventInput`, `EventsPage`, `PublishResult`, `Scope`, `Subscription`, `VaultRecord`, and others, so fixtures type-check against the real SDK. | | Types from `@cef-ai/agent-sdk` | `Rule`, `Targeting`, `TargetingEngagement`. | ## What the simulator does not do * Verify signatures or consent: wallets are stubs. * Run real models: every model call goes to a handle you provide. * Enforce platform limits other than those you inject (`faults`, `refuseUpload`). * Talk to a live vault: everything is in-process. ## Related * [Test a workflow](/agents/workflows/test/) * [Best practices](/agents/workflows/best-practices/) * [Agent SDK reference](/reference/agent-sdk/) * [Vault SDK reference](/reference/vault-sdk/) # Vault SDK (@cef-ai/vault-sdk) > Reference for @cef-ai/vault-sdk 5.5.0: construction, auth, the Vault handle, scopes and events, objects, agents and connections, jobs, memberships, and errors. ```bash npm install @cef-ai/vault-sdk@5.5.0 ``` The vault SDK is the client for a vault: create or open it, publish and read events, store objects, connect agents, read the Memory Bank, and manage members. It runs in Node and the browser. ## Construction ```ts import { VaultSDK, KeypairWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: vaultApiUrl, garEndpoint: garUrl, signer: await KeypairWallet.fromSeed(seed), }); ``` | `VaultSDKConfig` | Meaning | | ---------------------- | ------------------------------------------------------------- | | `endpoint` | vault-api base URL. Required. | | `signer` | A `Signer`. Signs requests and consent agreements. | | `auth` | A custom `AuthProvider`; overrides `signer` for requests. | | `garEndpoint` | Agreement registry URL. Needed for `vault.agents.connect`. | | `chainUrl` | Chain RPC URL. Needed by `vault.ensure()` on the wallet path. | | `s3GatewayAuthInfoUrl` | Gateway `/auth/info` URL, for onboarding. | | `marketplaceEndpoint` | Marketplace URL for `sdk.marketplace`. | | `fetch`, `timeoutMs` | Transport overrides. | | Signer / auth | Use | | ---------------------------------------------------------- | --------------------------------------------------------------- | | `KeypairWallet.fromSeed(seed)` | Server: a raw 32-byte ed25519 seed. | | `CereWallet.fromMnemonic(…)`, `CereWallet.fromKeystore(…)` | A wallet from a mnemonic or keystore. | | `new WalletAuthProvider(signer)` | Sign each request (the default when you pass `signer`). | | `new DelegationAuthProvider(token)` | Authenticate with a delegation bearer token instead of signing. | | `new NoAuth()` | Unauthenticated calls. | ## Opening a vault | Method | Returns | | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------- | | `sdk.vault.ensure(opts?)` | Your vault, creating it on first use. `opts`: `onboard` (default `true`), `onProgress`, `chainUrl`. Idempotent. | | `sdk.vault.current()` | Your own vault. | | `sdk.vault.byId(vaultId)` | A vault by id, typically an organization’s. | | `sdk.vault.rename(vaultId, displayName)` | Sets the vault’s name. Owner only. | `ensure()` throws `OnboardingRequiredError` when the wallet still needs its on-chain gateway registration, and `OnboardingTimeoutError` when onboarding does not finish in time. `onProgress` reports `inspecting-wallet`, `funding-wallet`, `gateway-authorization-required`, and later steps. ## `Vault` | Member | Does | | ---------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------- | | `id`, `bucketId`, `walletPubkey`, `displayName`, `status` | Vault record. `isDisconnected()` is true when its storage credential can no longer be refreshed. | | `scopes.list()`, `.get(name)`, `.create(req)`, `.update(name, req)`, `.delete(name)` | Manage scopes. | | `scope(name)` | A `ScopeHandle`. | | `service(asPubKey)` | An Agent Service’s cubbies in this vault: `list()`, `cubby(alias).query/exec/provision/delete`. | | `agents.connect(input)` | Connect an agent. See [Connect to a vault](/agents/ship/connect-to-a-vault/). | | `agents.list()`, `agents.get(agentId)` | `AgentConnectionHandle`s. | | `agents.connections(agentId)`, `.setConnection(agentId, connectionId, body)`, `.removeConnection(…)` | An agent’s access to the vault’s connector connections. | | `connections` | Connector connections: `list`, `create`, `update`, `revoke`, `actions`. See [Connectors](/vaults/connectors/). | | `connectors()` | The connector catalog. | | `webhooks` | Webhook triggers of connected workflows: `list(agentId)`, `createKey`, `revokeKey`. A key’s secret is returned once. | | `memory` | Memory Bank reads: `search`, `list`, `get`, `neighbours`, `countByType`. Writes happen in agents through `ctx.memory`. | | `jobs.list({ state?, cursor?, limit? })`, `jobs.get(jobId)` | Runs. `JobHandle`: `get()`, `tasks.list/get/logs/subscribe`. | ## `ScopeHandle` | Member | Does | | ----------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `publish(input)` | Publish one event. `input`: `{ type, context, payload, role?, target?, correlationId?, metadata?, parents?, timestamp? }`. Resolves `{ eventId }`; throws when the vault rejects it. | | `subscribe(filter, handler, opts?)` | Follow one stream. `filter`: `{ context, types?, from? }`. `opts`: `intervalMs` (default 1000), `pageSize` (default 100), `onError`, `maxBackoffMs`. Returns `{ unsubscribe(), closed }`. | | `subscribeAll(filter, handler, opts?)` | Follow every stream in the scope. `filter`: `{ types? }`. `refreshIntervalMs` (default 30000) picks up new streams. | | `streams.list(opts?)`, `streams.get(context)` | Stream summaries. | | `stream(context).events.list(opts?)` | A stream’s events, paged. | | `objects.upload(path, data, opts?)` | Upload bytes. `opts.publishEvent` also publishes an event carrying `vaultPath`. | | `objects.get`, `.head`, `.presignedUrl(path, { ttlSeconds? })`, `.delete`, `.list({ prefix? })` | Object storage. | ## `AgentConnectionHandle` Fields: `vaultId`, `agentId`, `scope`, `scopes`, `version`, `asPubKey`, `cubbyAliases`, `status`, `ceiling`, `bundle` (`{ bucketId, cid }`, the pinned code), `createdAt`, `updatedAt`. | Method | Does | | --------------------------------------- | ---------------------------------------------------------------- | | `update(settings)` | Update the connection’s settings. | | `setCeiling({ gpuUnits?, a2aTokens? })` | Set spend limits; `0` or absent means no limit on that axis. | | `cubby(alias)` | `query(sql, params)` / `exec(sql, params)` on the agent’s cubby. | | `disconnect()` | Remove the connection. Cubby data stays. | ## Memberships `sdk.memberships`: `mine()` (vaults you belong to), `list(vaultId)`, `signedRoster(vaultId)`, `put(…)`, `putScopes(…)`, `remove(vaultId, memberPubkey)`, `scopes(vaultId, memberPubkey)`. `put` and `putScopes` need a `signer`: the owner signs the grant document. See [Vault members](/vaults/members/). ## Other exports | Export | Use | | --------------------------------------------------------------------------------------- | --------------------------------------------- | | `signAndSubmitAgreement`, `GarClient`, `getAsPubkey` | Sign and submit a consent agreement yourself. | | `mintDelegationTokens`, `VAULT_DELEGATION_OPERATIONS`, `REGISTRY_DELEGATION_OPERATIONS` | The owner’s delegation tokens. | | `verifySignedRoster`, `verifyGrant`, `decodeGrantDoc` | Verify a signed member roster. | | `sdk.health` | Service health. | | `sdk.marketplace.getAgent(agentId, version?)`, `.list(opts?)` | Read the marketplace listing. | Lower-level clients are under `@cef-ai/vault-sdk/internal`. ## Errors | Class | When | | --------------------------------------------------- | ---------------------------------------------------------------------------------- | | `VaultRequestError` | Any vault-api error. Fields: `status`, `code`, `retryable`. Match on `code`. | | `BundleChangedError` | `409 BUNDLE_CHANGED`. `requestedCid`, `currentCid`. | | `ReconsentRequiredError` | `409 RECONSENT_REQUIRED`. `previousCid`, `currentCid`. | | `CeilingUnknownError` | A scope write could not be signed because the member’s privacy ceiling is unknown. | | `VaultNotImplementedError` | `501`. | | `VaultSignerRequiredError` | A method that needs a `signer` was called without one. | | `OnboardingRequiredError`, `OnboardingTimeoutError` | From `ensure()`. | All codes: [Errors](/reference/errors/). ## Related * [Work with the vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) * [Connect to a vault](/agents/ship/connect-to-a-vault/) * [Event streams](/agents/code-agents/event-streams/) * [Vaults](/vaults/overview/) # Widget runtime (@cef-ai/widget-runtime) > Reference for @cef-ai/widget-runtime 2.4.1: window.WidgetRuntime, built-in kinds, the widget manifest, sign-in, the host contract, and library exports. The widget runtime is the browser code every widget runs. It reads the widget’s manifest from `window.WidgetSandbox.manifest`, signs the reader in, exposes `window.WidgetRuntime`, and renders the built-in kinds. `cef build`, `cef widget push`, and `cef dev` inject the runtime and the manifest into your entry HTML; you do not add them yourself. `cef build` vendors the runtime version that the installed CLI resolves, and logs it. Upgrade the CLI to ship a newer runtime. ## `window.WidgetRuntime` | Method | Returns | Does | | -------------------------------------------- | ---------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- | | `query(ref, params?)` | `Promise` | Runs a declared query by id. `params` bind to `?` in its SQL, or are the tool’s arguments for a Memory Bank query. | | `publish(type, payload, context?, options?)` | `Promise<{ eventId }>` | Publishes an event into the widget’s scope as the reader. `options.target`: `:`, the one agent the event is for. | | `subscribe(context, opts, onEvents)` | `() => void` (stop) | Follows one stream; `onEvents` receives each poll’s new events as an ordered batch. | | `identity()` | `Promise` | The reader’s identity. | | `connect()` | `Promise` | Signs the reader in. | | `connectAgent()` | `Promise` | Connects the widget’s agent to the reader’s vault. | | `agentStatus()` | `Promise<'connected' \| 'disconnected'>` | Whether the agent is connected. | | `onIdentityChange(cb)` | `() => void` (unsubscribe) | Called when the identity changes. | | Type | Shape | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------- | | `QueryResult` | `{ columns: string[]; rows: unknown[][]; meta: { rowCount: number } }` | | `Identity` | `{ publicKey, status: 'anon' \| 'connecting' \| 'ready', context? }`. `context` is `{ runId?, recordId?, vaultId? }` when the host or link named one. | | `WidgetEvent

` | `{ eventId, type, context, payload, from?, timestamp }` | | `subscribe` options | `types?`, `intervalMs` (default 2500), `maxBackoffMs` (default 30000), `onError` (default `console.warn`) | `subscribe` delivers events already in the stream in its first batch, and each event once. It stops when you call the returned function or when the reader becomes `anon`. ## Errors | Error | When | | ------------------------------- | ------------------------------------------------------------------------------------------------------- | | `AgentNotConnectedError` | A read for a reader whose vault has not connected the agent. Has `agentId`. Offer `connectAgent()`. | | `WidgetSignedOutError` | A framed widget’s host did not answer the identity handshake within 8 seconds, or answered malformed. | | `WidgetVaultUnreachableError` | The host or link named a vault the widget cannot open. The runtime does not fall back to another vault. | | `WidgetWalletUnconfiguredError` | A standalone widget’s manifest has no wallet origin. | ## Kinds Set `kind` and `config` in the widget declaration; the runtime renders into `#app`. | Kind | `config` | | -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `record` | `{ query, fields: [{ label, column, format? }], empty? }` | | `list` | `{ query, item: { title, subtitle?, meta? }, empty?, limit? }` | | `dashboard` | `{ panels: [{ title, query, render: 'metric' \| 'bar' \| 'table' \| 'list', value?, label?, columns?, item? }] }` | | `submit` | Form + audio capture → upload → publish → poll. See [Onboard data](/agents/building-blocks/onboard-data-with-a-widget/). | | `conversation` | Turn-based audio: `startEvent?`, `turnEvent`, `confirmEvent?`, `turnsQuery`, `roleColumn`, `textColumn`, `turnIdxColumn`, `statusQuery?`, `confirmWhen?`, `doneWhen?`, `poll`, `audio`. | | `composite` | `{ groups, panels, defaultPanel, gating: { statusQuery, unlockWhen, firstRunPanel, hero? }, menu? }`. Each panel holds any other kind, or `{ kind: 'custom', renderer, query? }` drawn by a function registered with `registerCompositePanel`. | | `custom` | Your own page. | Field formats: `text`, `multiline`, `date`, `reltime`, `number`, written as `"column:format"`. ## Manifest The manifest `cef` writes into the entry HTML (`WidgetManifest`): | Field | Meaning | | ------------------ | ---------------------------------------------------------------------------------------------------------------------- | | `schemaVersion` | `1`. | | `widgetId`, `name` | Identity. | | `agentId` | `:`. Empty until `--as-pubkey` is given. | | `scope` | The vault scope the widget reads and publishes in. | | `cubbyAlias` | Default cubby for `sql` queries. | | `queries` | `{ id, label, sql?, cubby?, tool?, limit?, timeoutMs? }[]`; `tool` is `search`, `get`, `neighbours`, or `countByType`. | | `events` | `{ type, schemaRef? }[]`. | | `kind`, `config` | Built-in kind. | | `wallet` | `{ appId, env, origin }`. `origin` is the Manykind passkey wallet used for standalone sign-in. | | `endpoints` | `vault`, `gar`, `marketplace`, `s3GatewayAuthInfo`, `rpc`, from `--env`. | ## Sign-in | Opened | Identity source | | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | In a frame | The host, over the `postMessage` handshake below. A framed widget never falls back to the standalone wallet. | | Top-level | The Manykind passkey wallet at `wallet.origin`. An active session resumes silently; otherwise the runtime shows a sign-in control and opens the wallet on click. Reads use a delegation, so the reader is not asked to sign each request. | A link can name what to show, in its fragment (preferred) or query: `vaultId` or `vault`, `runId` or `run`, `recordId` or `record`. A named vault is still checked against the reader’s own access. ## Host contract To frame a widget in your own page, mount a host: ```ts import { createWidgetHost } from "@cef-ai/widget-runtime"; const dispose = createWidgetHost({ allowedOrigins: ["https://widget.example"], // required, canonical origins only, no "*" widget: document.querySelector("iframe#cef-widget")!, // optional extra check getIdentity: () => (signedIn ? { pubkey, sigType: "ed25519" } : null), sign: async (bytes) => signRaw(bytes), // raw ed25519 over the exact bytes; null = declined }); ``` | Message | Direction | Fields | | ------------------------------ | ------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | | `cef-widget:identity-request` | widget → host | `requestId`, `v` | | `cef-widget:identity-response` | host → widget | `requestId`, `v?`, `pubkey?`, `sigType?` (`ed25519` default, or `sr25519`), `token?`, `context?: { runId?, recordId?, vaultId? }`, `error?`, `errorCode?` | | `cef-widget:sign-request` | widget → host | `requestId`, `v`, `bytes: number[]` | | `cef-widget:sign-response` | host → widget | `requestId`, `v?`, `signature: number[] \| null`, `error?`, `errorCode?` | | Rule | Value | | ---------------- | -------------------------------------------------------------------------------------------------------------------------- | | Protocol version | `1` (`PROTOCOL_VERSION`). A message without `v` is v1; an unknown `v` is rejected with `errorCode: 'unsupported-version'`. | | Bytes | `number[]` of integers 0–255, at most 64 KiB (`MAX_SIGN_BYTES`). Check with `isByteArray`. | | Signing | Sign the bytes verbatim (`ed25519_signRaw`), never with a message-wrapping sign method. | | Timing | The widget retries the identity request every 400 ms for 8 s; a sign request waits up to 60 s. | | Sandboxed frames | Their origin is `"null"`; pass `allowedOrigins: ["null"]` and also set `widget`. | ## Library exports | Export | Use | | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------ | | `createWidgetHost`, `createHostBridge` | Host side of the contract. | | `IDENTITY_REQUEST_TYPE`, `SIGN_REQUEST_TYPE`, `PROTOCOL_VERSION`, `MAX_SIGN_BYTES`, `isByteArray`, `WidgetProtocolVersionError`, `WidgetMalformedMessageError` | Contract constants and checks for a hand-written host. | | `mountKind`, `renderList`, `renderRecord`, `renderDashboard`, `renderSubmit`, `renderConversation`, `renderComposite`, `registerCompositePanel` | Render the built-in kinds yourself. | | `validateConfig` | Check a kind config. | | `captureAndUploadAudio`, `createVaultObjectUploader`, `createVaultObjectStore`, `encodeWav`, `segmentPcm`, `decodeToMono16k`, `requestMicStream` | Audio capture and upload. | | `connectStandaloneWallet`, `createCefWalletGate`, `requestSignIn`, `autoSignIn` | Standalone sign-in. | | `contextFromLocation`, `currentLinkContext` | Read a link’s context. | | `@cef-ai/widget-runtime/build-tools`: `buildWidgetManifest`, `injectWidgetBridge` | What the CLI uses to build and inject the manifest. | ## Related * [Build a widget](/agents/building-blocks/create-a-widget/) * [Visualize vault data](/agents/building-blocks/visualize-vault-data/) * [Widgets](/agents/building-blocks/widgets/) * [CLI reference](/reference/cli/#cef-dev) # Build a vault integration > Build a headless server process that syncs an outside source into a customer's vault, or exports from it, on their behalf under a scoped, revocable delegation token. A **vault integration** is a server process that reads and writes a customer’s vault on their behalf, without them present to sign each request. Two shapes: * **Sync in.** Pull an outside source (a calendar, a CRM export, a fitness API) into the vault on a schedule, so agents can compute on it. * **Export out.** Read vault data and push it to another system: a warehouse, a backup, a report. ## Pick the right tool first | The outside system | Use | | ------------------------------------------ | ---------------------------------------------------------------- | | Slack, Telegram, email, or an MCP server | A vault [connector](/vaults/connectors/). No code of yours runs. | | Can call a URL when something happens | A workflow [webhook trigger](/agents/workflows/triggers/). | | Must be polled or transformed by your code | A vault integration (this page). | ## How an integration is authorized The customer grants authority once, while present; the integration then runs headless until they revoke it. | Grant | What it allows | Who issues it | | -------------------------- | -------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------- | | **Vault write delegation** | The vault service writes to the vault’s storage on the owner’s behalf. | Granted when the owner’s app calls `vault.ensure()` with a signer and `chainUrl`. | | **Delegation token** | Your integration calls the vault API as the customer, within the token’s scope, without their key. | Minted by the customer’s wallet and handed to your backend. | Your server never holds the customer’s key. Revoking the token locks the integration out. ## Part A: in the customer’s app, once ```ts import { VaultSDK, CereWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: config.vaultUrl, chainUrl: config.chainUrl, signer: await CereWallet.fromMnemonic(mnemonic), // the owner, present now }); const vault = await sdk.vault.ensure(); const scopes = await vault.scopes.list(); if (!scopes.some((s) => s.name === "calendar")) { await vault.scopes.create({ name: "calendar", displayName: "Calendar" }); } ``` Then obtain a delegation token from the customer’s wallet, and send it with `vault.id` to your backend. Store both as this customer’s integration credential. ## Part B: in the integration ```ts import { VaultSDK, DelegationAuthProvider } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: config.vaultUrl, auth: new DelegationAuthProvider(delegationToken), }); const vault = await sdk.vault.byId(vaultId); if (vault.isDisconnected()) { throw new Error("vault storage credential can no longer refresh; the owner must act"); } ``` `auth` replaces `signer`: the integration uses the token for every call. ### Sync in Publish each new record as an event in the integration’s scope. Track your source’s cursor on your side and publish only what is new: the same record published twice is two events. ```ts const calendar = vault.scope("calendar"); for (const r of await fetchNewRecords(sinceCursor)) { await calendar.publish({ type: "calendar.event", context: `calendar:${r.calendarId}`, payload: r, }); } ``` ### Export out Walk the scope’s streams and page their events: ```ts const health = vault.scope("health"); let streamCursor: string | undefined; do { const streams = await health.streams.list({ cursor: streamCursor }); for (const s of streams.items) { let cursor: string | undefined; do { const page = await health.stream(s.context).events.list({ cursor, limit: 200 }); await pushToWarehouse(page.items); cursor = page.hasMore ? page.nextCursor : undefined; } while (cursor); } streamCursor = streams.hasMore ? streams.nextCursor : undefined; } while (streamCursor); ``` To export what a connected Agent Service computed, read its cubbies with `vault.service(agentServicePubkey).cubby(alias).query(sql, params)`, or the vault’s conclusions with `vault.memory.search(...)`. ## Keep it healthy * **Treat an auth failure as consent withdrawn.** The customer revoked the token or it expired. Stop, and do not retry blindly. * **Check `vault.isDisconnected()` each run.** The storage credential cannot refresh until the owner acts in their own app. * **Make writes idempotent downstream.** Dedupe on your source cursor, or give each payload a stable id agents can upsert on. ## Related * [Next: Build a Node vault client](/vaults/build-on-a-vault/external-vault-client/) * [Onboard from your app](/vaults/build-on-a-vault/onboard-data-from-an-external-client/) * [Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) * [Connectors](/vaults/connectors/) * [Data onboarding](/vaults/build-on-a-vault/data-onboarding/) # Data onboarding > The routes by which data lands in a customer's vault — another app, a widget, your own app, a connector, a webhook, or a server integration — and why your agents and workflows do not depend on which one was used. This group covers your own apps and services working with a vault directly: the onboarding routes, the Vault SDK, and integrations and external clients. Your agents and workflows compute on data in a customer’s vault. Getting it there is **onboarding**, and there are several routes. Pick the ones that fit your product; your agent code does not change with the route. ## The routes | Route | Who writes | What lands | See | | ------------------------ | ----------------------------------------------------------------------- | ----------------------------------------- | --------------------------------------------------------------------------------------- | | **Already there** | Another app or agent the customer connected | Events and objects in a scope | | | **A widget** | Your service’s UI, acting as the signed-in person | Recordings, uploads, form submissions | [Onboard data with a widget](/agents/building-blocks/onboard-data-with-a-widget/) | | **Your own app** | A web, mobile, or server app with the Vault SDK, under the person’s key | Events and objects | [Onboard from your app](/vaults/build-on-a-vault/onboard-data-from-an-external-client/) | | **A connector** | Slack or Telegram, through a vault connection | A message event that can start a workflow | [Connectors](/vaults/connectors/) | | **A webhook** | An outside system calling a workflow’s webhook URL with a key | A trigger event | [Triggers](/agents/workflows/triggers/) | | **A server integration** | Your headless process, under a delegation token | Events and objects, on a schedule | [Build a vault integration](/vaults/build-on-a-vault/build-a-vault-integration/) | ### Already there The data may already be arriving, written by an app or agent the customer connected earlier. A health app writes heart-rate readings into the customer’s `health` scope; your coaching workflow reacts to them. You never talk to the watch and hold no integration credential. ### A widget A [widget](/agents/building-blocks/widgets/) is your service’s UI in the browser. It acts as the signed-in person, so what it writes lands under their own authority. Use it when onboarding belongs inside your product’s own screen. ### Your own app An app you already ship can write into the customer’s vault through the [Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/): publish events, upload objects. It signs with the person’s key, so the vault authorizes the write because it descends from their wallet. ### A connector or a webhook Outside systems can push straight into a vault. A Slack or Telegram message arrives through a vault [connector](/vaults/connectors/); any other system can call a workflow’s webhook. Either can start a run. ### A server integration A headless process can sync a source into the vault on the customer’s behalf, with a scoped, revocable delegation token instead of their key. ## What agents make of it Once data lands, your agents and workflows turn it into two kinds of output: * **Working state** in your service’s [cubbies](/agents/building-blocks/cubbies/). * **Durable conclusions** in the vault’s [Memory Bank](/vaults/memory-bank/), privacy-classified and kept by the customer. ## Compute does not depend on the route | Route | What your agent does | | ---------------------- | ----------------------------------------- | | Already there | Reacts to events in the scope | | A widget | Reacts to the events the widget published | | Your own app | Reacts to the events your app published | | A connector or webhook | Runs from the trigger event | | A server integration | Reacts to the events it published | The last column barely changes. The vault is the interface between onboarding and compute: an agent written for heart-rate events does not know whether they came from a watch app, a widget, or your server. Give each source its own scope and stable, namespaced event types (`health.heart_rate`, `health.steps`), so the customer can connect an agent to exactly that data and the payload shape stays a contract. ## Related * [Next: Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) * [Vaults](/vaults/overview/) * [Onboard from your app](/vaults/build-on-a-vault/onboard-data-from-an-external-client/) * [Connectors](/vaults/connectors/) * [Runs and events](/agents/how-agents-run/) # Build a Node vault client > A tutorial — build a Node.js script that opens a vault, connects an agent, sends it a message, reads its reply and cubby, and follows the run, with @cef-ai/vault-sdk. In this tutorial you build a Node.js script that talks to a vault from outside the platform. By the end it opens a vault, connects an agent, publishes a message, reads the agent’s reply, reads the agent’s cubby, and prints the run’s task logs. The agent here handles `user_message` events, writes each message to a `history` cubby with a `messages` table, and publishes a `reply` event. See [Write an agent](/agents/code-agents/overview/) to build one. Any agent works the same way; change the event types and the SQL. ## Prerequisites * Node.js with ES module support, and a project with `"type": "module"`. * A 32-byte ed25519 seed for the wallet the script acts as. * The vault API, agreement registry, and chain URLs for your environment. * The agent’s id, `:`, from ROC. ## Step 1: Install ```bash npm install @cef-ai/vault-sdk@5.5.0 ``` ## Step 2: Construct the SDK with a server signer client.ts ```ts import { VaultSDK, KeypairWallet } from "@cef-ai/vault-sdk"; const seed = Uint8Array.from(Buffer.from(process.env.SEED_HEX!, "hex")); const sdk = new VaultSDK({ endpoint: process.env.VAULT_URL!, garEndpoint: process.env.GAR_URL!, chainUrl: process.env.CHAIN_URL!, signer: await KeypairWallet.fromSeed(seed), }); ``` `KeypairWallet` signs every request directly, with no prompt. `garEndpoint` and the signer together let the script sign the connect agreement in Step 4. ## Step 3: Open the vault ```ts const vault = await sdk.vault.ensure({ name: "My Vault", onProgress: (e) => console.log("vault:", e.kind), }); console.log("vault", vault.id); ``` The first run claims the vault; later runs return it. If you see `OnboardingRequiredError`, authorize the storage gateway with `@cef-ai/account`’s `provisioning.ensure`, then run again. ## Step 4: Connect the agent ```ts import { VaultRequestError } from "@cef-ai/vault-sdk"; const agentId = process.env.AGENT_ID!; try { await vault.agents.connect({ agentId, scope: "default" }); } catch (e) { if (!(e instanceof VaultRequestError && e.code === "AGENT_ALREADY_CONNECTED")) throw e; } let connection = await vault.agents.get(agentId); while (connection.status === "provisioning") { await new Promise((r) => setTimeout(r, 1000)); connection = await vault.agents.get(agentId); } console.log("connection", connection.status); ``` `connect` signs one agreement with your key, submits it, and the vault provisions the agent’s cubbies. The connection moves from `provisioning` to `active`. A second run finds it already connected. ## Step 5: Send a message and wait for the reply ```ts const scope = vault.scope("default"); const context = `thread-${Date.now()}`; const reply = new Promise((resolve) => { const sub = scope.subscribe<{ text: string }>( { context, types: ["reply"] }, (event) => { sub.unsubscribe(); resolve(event.payload.text); }, ); }); await scope.publish({ type: "user_message", context, payload: { text: "hi" } }); console.log("agent said:", await reply); ``` Subscribe first, then publish, so the reply cannot slip past. The subscription polls the stream every second. ## Step 6: Read the agent’s cubby ```ts const asPubKey = agentId.split(":")[0]!; const { columns, rows } = await vault .service(asPubKey) .cubby("history") .query("SELECT text, ts FROM messages ORDER BY ts DESC LIMIT 5"); console.log(columns, rows); ``` A service’s cubbies are global within the vault, addressed by the service’s public key and the cubby alias. ## Step 7: Print the run’s logs ```ts const jobs = await connection.jobs.list({ limit: 1 }); const job = vault.jobs.get(jobs.items[0]!.jobId); const tasks = await job.tasks.list(); for (const t of tasks.items) { const { logs } = await job.tasks.logs(t.taskId); for (const line of logs) console.log(t.status, line.message); } ``` Each event your script published became a Task in the agent’s Job; the log lines are what the agent wrote with `console.*`. ## Run it ```bash SEED_HEX=… VAULT_URL=… GAR_URL=… CHAIN_URL=… AGENT_ID=… npx tsx client.ts ``` ## What you built A client that does, from outside the platform, what ROC does for a person: open a vault, connect an agent under a signed agreement, talk to it through events, and read what it stored and logged. ## Related * [Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) * [Onboard from your app](/vaults/build-on-a-vault/onboard-data-from-an-external-client/) * [Runs and events](/agents/how-agents-run/) * [Vault SDK reference](/reference/vault-sdk/) # Onboard from your app > Write a person's own data into their vault from a web, mobile, or server app with @cef-ai/vault-sdk, so the agents and workflows they connect can compute on it. Your app already has the data: a phone app reads a watch’s heart rate, a web app holds a person’s notes. Write it into the person’s vault, and any agent or workflow they connect to that scope can work on it. Onboarding needs no agent and no agreement: the person writes their own data into their own vault, under their own key. The running example: an iOS app writes heart-rate readings into a `health` scope. ## Before you start * `@cef-ai/vault-sdk` 5.5.0. * A `Signer` for the person: `CereWallet.fromMnemonic(...)` or `CereWallet.fromKeystore(...)` for the wallet on the device, or `KeypairWallet.fromSeed(...)` for a server acting as the person. * Your vault API URL and chain URL. ## 1. Construct the client with the person’s signer ```ts import { VaultSDK, CereWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: config.vaultUrl, chainUrl: config.chainUrl, signer: await CereWallet.fromMnemonic(mnemonic), }); ``` Every request is now signed with the person’s key, so the vault authorizes it because it descends from their wallet, not from a secret you hold. ## 2. Open their vault ```ts const vault = await sdk.vault.ensure({ name: "My Vault" }); ``` On first use this claims and provisions the vault; afterwards it returns the existing one. If it throws `OnboardingRequiredError`, run `@cef-ai/account`’s `provisioning.ensure` and retry. See [Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/#2-open-a-vault). ## 3. Give the data its own scope A scope of its own lets the person connect a fitness agent to exactly this data and nothing else. ```ts async function ensureScope(name: string, displayName: string) { const existing = await vault.scopes.list(); if (!existing.some((s) => s.name === name)) { await vault.scopes.create({ name, displayName }); } } await ensureScope("health", "Health"); ``` ## 4. Publish the readings Each reading is one event. Use one `context` per metric so each history is its own stream, and stable, namespaced types: agents subscribe on them, and the payload shape is the contract their handlers read. ```ts const health = vault.scope("health"); for (const s of await readWatchSamples()) { await health.publish({ type: "health.heart_rate", context: "heart-rate", payload: { bpm: s.bpm, measuredAt: s.timestamp }, }); } ``` `publish` sends one event and throws if the vault rejects it, so a resolved call means the reading is stored. Track what you already sent on your side: the same reading published twice is two events. ## 5. Store files as objects For files (a workout export, an audio note), upload an object and announce it in the same call: ```ts await health.objects.upload(`exports/${day}.json`, bytes, { contentType: "application/json", publishEvent: { type: "health.export", context: "exports", payload: { day } }, }); ``` The event’s payload gets the object’s `vaultPath`, so an agent can fetch it. ## 6. Read back what you wrote ```ts const { items: streams } = await health.streams.list(); const page = await health.stream("heart-rate").events.list({ limit: 50 }); for (const e of page.items) console.log(e.type, e.payload); ``` Pass `page.nextCursor` as `cursor` while `page.hasMore` is true. ## 7. Let an agent work on it The data is in the vault. To have an agent or workflow act on it, the person connects it to the same scope. That signs an agreement, so the SDK also needs `garEndpoint`: ```ts await vault.agents.connect({ agentId: ":fitness-coach", scope: "health", }); ``` From here the agent reacts to every reading you publish. See [Connect to a vault](/agents/ship/connect-to-a-vault/). ## Related * [Next: Build a vault integration](/vaults/build-on-a-vault/build-a-vault-integration/) * [Data onboarding](/vaults/build-on-a-vault/data-onboarding/) * [Work with the Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) * [Vault SDK reference](/reference/vault-sdk/) # Work with the Vault SDK > Use @cef-ai/vault-sdk from a browser, mobile, or Node app to open a vault, publish events and objects, connect agents, and read cubbies, the Memory Bank, and runs. `@cef-ai/vault-sdk` is the client for talking to a vault from **outside** an agent: a web or mobile app, a console, a server job. It signs requests, handles canonicalization, and wraps the vault’s API in typed handles. It runs in the browser and in Node. Code running *inside* an agent uses `ctx` from `@cef-ai/agent-sdk` (`ctx.vault`, `ctx.cubby`, `ctx.memory`, `ctx.models`), not this SDK. ## Install ```bash npm install @cef-ai/vault-sdk@5.5.0 ``` ## 1. Construct the client ```ts import { VaultSDK, CereWallet } from "@cef-ai/vault-sdk"; const sdk = new VaultSDK({ endpoint: config.vaultUrl, // vault API garEndpoint: config.garUrl, // agreement registry; needed for agents.connect chainUrl: config.chainUrl, // needed for vault.ensure() with a signer signer: await CereWallet.fromMnemonic(mnemonic), }); ``` | Field | Required | Purpose | | ---------------------- | ---------------------------------- | ---------------------------------------------------------------------------------------------------------- | | `endpoint` | yes | Vault API base URL. | | `signer` | one of `signer` / `auth` | Signs every request with the person’s key. | | `auth` | one of `signer` / `auth` | A custom auth provider, e.g. `new DelegationAuthProvider(token)`. Overrides `signer`. | | `garEndpoint` | for `agents.connect` | Agreement registry the connect flow submits the signed agreement to. | | `chainUrl` | for `vault.ensure()` with a signer | Used to grant the vault service write delegation before the vault is claimed. Can also be passed per call. | | `s3GatewayAuthInfoUrl` | | Storage gateway the onboarding check talks to. | | `marketplaceEndpoint` | | Marketplace API for `sdk.marketplace` reads. | | `fetch`, `timeoutMs` | | Custom fetch; per-request timeout (default 30 s). | ### Signers | Environment | Signer | | -------------------------------------------- | ------------------------------------------------- | | Browser or mobile, from a mnemonic | `await CereWallet.fromMnemonic(mnemonic)` | | From a JSON keystore | `await CereWallet.fromKeystore(json, passphrase)` | | Server, from a raw 32-byte ed25519 seed | `await KeypairWallet.fromSeed(seed)` | | Acting on someone’s behalf without their key | `auth: new DelegationAuthProvider(token)` | Any object implementing the `Signer` interface (`type`, `address`, `publicKey`, `isReady()`, `sign(bytes)`) works. ## 2. Open a vault ```ts const vault = await sdk.vault.ensure({ name: "My Vault", onProgress: (e) => console.log(e.kind), }); ``` `ensure()` is idempotent: it returns the signer’s vault, claiming it on first use. With a signer it checks onboarding status, grants the vault service write delegation (this needs `chainUrl`), then claims the vault. | `onProgress` kind | Meaning | | -------------------------------- | ----------------------------------------------------------------------------------------------------------- | | `inspecting-wallet` | Onboarding status check in flight. | | `funding-wallet` | Development networks only: the wallet is being funded. | | `gateway-authorization-required` | The wallet still needs its storage-gateway authorization. `ensure()` throws `OnboardingRequiredError` next. | | `authorizing-vault` | Granting the vault service write delegation. | | `provisioning-vault` | Claiming the vault. | On `OnboardingRequiredError`, authorize the gateway with `@cef-ai/account`’s `provisioning.ensure({ signer, chainUrl, gateway })`, then call `ensure()` again. Other ways to open a vault: | Call | Returns | | ------------------------- | ------------------------------------------------ | | `sdk.vault.current()` | Your own vault, without onboarding. | | `sdk.vault.byId(vaultId)` | A vault you are a [member](/vaults/members/) of. | ## 3. Scopes ```ts await vault.scopes.create({ name: "health", displayName: "Health" }); const scopes = await vault.scopes.list(); ``` `vault.scopes` also has `get(name)`, `update(name, patch)`, and `delete(name)`. See [Vaults](/vaults/overview/#scopes) for naming rules. ## 4. Publish and follow events ```ts const scope = vault.scope("default"); const { eventId } = await scope.publish({ type: "user_message", context: "thread-123", payload: { text: "hi" }, }); const sub = scope.subscribe( { context: "thread-123", types: ["reply"] }, (event) => console.log(event.type, event.payload), ); // later sub.unsubscribe(); await sub.closed; ``` * `publish` sends one event and throws if the vault rejects it. `role` defaults to `user`; `timestamp` defaults to now. * `subscribe` polls (default every 1000 ms, 100 events per page), de-duplicates by event id, and reports poll errors to `onError`. Set `maxBackoffMs` to back off while polls fail. * `subscribeAll({ types }, handler)` follows every stream in the scope and picks up new ones (every 30 s by default; `refreshIntervalMs`). * `scope.streams.list()` and `scope.stream(context).events.list({ cursor, limit })` page history. To start a workflow from your app, publish `workflow.start` with `target: ":"`. ## 5. Store objects ```ts const res = await scope.objects.upload("notes/2026-10-08.txt", bytes, { contentType: "text/plain", }); const data = await scope.objects.get(res.vaultPath); ``` `scope.objects` also has `head`, `presignedUrl(path, { ttlSeconds })`, `list({ prefix })`, and `delete`. Pass `publishEvent: { type, context, payload }` to `upload` to announce the object in the same call; the event’s payload gets `vaultPath`. ## 6. Connect an agent ```ts const connection = await vault.agents.connect({ agentId: ":", scope: "default", // or scopes: ["support", "sales"] settings: { language: "en" }, }); console.log(connection.status); // "provisioning" → "active" ``` `connect` builds one agreement for the scope set, signs it, submits it to the agreement registry, and creates the connection. It needs `garEndpoint` and a `signer`. Optional fields: `ceiling` (`{ gpuUnits?, a2aTokens? }`) and `bundleCid` (the bundle the person reviewed). See [Connections and consent](/vaults/connections-and-consent/). | Call | Does | | -------------------------------------------------- | ------------------------------------------ | | `vault.agents.list()`, `vault.agents.get(agentId)` | Read connections. | | `connection.update(settings)` | Change the settings values. | | `connection.setCeiling({ gpuUnits, a2aTokens })` | Change the spend ceiling; `{}` clears it. | | `connection.disconnect()` | Remove the connection; cubby data is kept. | ## 7. Read cubbies A service’s cubbies are global within the vault. Address them by the service’s public key: ```ts const { columns, rows } = await vault .service(agentServicePubkey) .cubby("history") .query("SELECT text, ts FROM messages ORDER BY ts DESC LIMIT ?", [20]); ``` `.exec(sql, params)` writes. `vault.service(pubkey).list()` lists the service’s cubbies. ## 8. Read the Memory Bank ```ts const page = await vault.memory.search("refund approval", { scope: "finance", limit: 20 }); ``` Also `list(type?, opts)`, `get(id)`, `neighbours(id, opts)`, and `countByType(opts)`. Results are `{ columns, rows, cursor? }`; pass `cursor` back for the next page. Writes are agent-only. See [Memory Bank](/vaults/memory-bank/). ## 9. Follow runs ```ts const jobs = await vault.jobs.list({ limit: 20 }); const job = vault.jobs.get(jobs.items[0].jobId); const tasks = await job.tasks.list(); const logs = await job.tasks.logs(tasks.items[0].taskId, { limit: 100 }); ``` `connection.jobs.list()` narrows to one agent. `job.tasks.subscribe({ since }, handler)` follows new tasks. ## Handle errors by code ```ts import { VaultRequestError, BundleChangedError } from "@cef-ai/vault-sdk"; try { await vault.agents.connect({ agentId, scope: "default" }); } catch (e) { if (e instanceof BundleChangedError) { // show the person the current bundle and ask again } else if (e instanceof VaultRequestError && e.code === "AGENT_ALREADY_CONNECTED") { // reuse the existing connection } else { throw e; } } ``` | Error | Meaning | | ---------------------------------------------- | ---------------------------------------------------------------- | | `VaultRequestError` | Any error from the vault; read `.code`, `.status`, `.retryable`. | | `BundleChangedError`, `ReconsentRequiredError` | The connect’s bundle consent does not match. | | `VaultSignerRequiredError` | The call needs a signer and none was configured. | | `VaultNotImplementedError` | The vault answered 501. | | `OnboardingRequiredError` | The wallet needs gateway authorization before `ensure()`. | | `OnboardingTimeoutError` | Onboarding did not finish in time. | ## Related * [Next: Onboard from your app](/vaults/build-on-a-vault/onboard-data-from-an-external-client/) * [Vault SDK reference](/reference/vault-sdk/) * [Build a vault integration](/vaults/build-on-a-vault/build-a-vault-integration/) * [Build a Node vault client](/vaults/build-on-a-vault/external-vault-client/) * [Connect to a vault](/agents/ship/connect-to-a-vault/) # Connections and consent > How an agent or workflow gets access to a vault — the agreement the owner signs, what the platform checks at connect time, the bundle a connection pins, and how access ends. An agent or workflow runs on a vault’s data only after the vault owner **connects** it. Connecting does two things: the owner’s wallet signs an **agreement**, and the vault creates a **connection** that records what was granted. Consent is what the owner gives; the agreement is the signed record that proves it. ## What the owner signs The agreement is held in the agreement registry, not in the vault. It binds: | Bound | Form | | ------------------ | ------------------------------------------------ | | The owner’s wallet | the public key that signs | | Your Agent Service | `agentServicePubkey`, the prefix of the agent id | | The vault | the vault id | | The scopes | the set of scopes the connection may use | One agreement covers the whole scope set for one (Agent Service, owner) pair. Because the agreement is the authority, you cannot widen your own access; you can only ask the owner to sign again. ## Where connecting happens * **In ROC.** Running a workflow from the Workflow Builder connects its agents to the vault first. The first run asks for a passkey to sign the agreement. * **From your own app.** `vault.agents.connect({ agentId, scope })` in the Vault SDK builds the agreement, signs it with the configured signer, submits it, and creates the connection in one call. The full walkthrough is [Connect to a vault](/agents/ship/connect-to-a-vault/). ## What the vault checks A valid signature is necessary but not sufficient. At connect time the vault: 1. Verifies the agreement is present and not revoked or expired. 2. Fetches your agent’s manifest from your Agent Service’s registry by agent id. The manifest, not any listing, is the authority on what the agent declares. 3. Checks that the agent id’s prefix matches the manifest’s Agent Service. 4. Checks every requested scope is one the manifest asks for (`default` when it declares none). 5. Validates the owner’s `settings` values against the manifest’s settings schema. 6. Provisions the cubbies the manifest declares, running their migrations. Only then does the connection become `active`. ## What a connection records | Field | Meaning | | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `agentId` | `:` | | `scope` / `scopes` | The scopes it may use. | | `status` | `provisioning` → `active` → `revoking` → `revoked`. | | `settings` | The owner’s values for the manifest’s settings schema. | | `bundle` | The exact bundle the owner consented to run (see below). | | `ceiling` | Optional spend ceiling: `gpuUnits` (platform-metered) and `a2aTokens` (self-reported by an LLM agent). The platform refuses new work once a ceiling is reached. | ### The bundle is pinned A connection runs the bundle the owner consented to, not whatever you published last. Deploying a new version does **not** change what a connected vault runs. To move a vault to new code, reconnect with the new bundle’s content id (`bundleCid`): | Code | When | | -------------------- | ------------------------------------------------------------------------------------------------------------------------ | | `BUNDLE_CHANGED` | The `bundleCid` sent no longer matches the agent’s current manifest: something was published between review and connect. | | `RECONSENT_REQUIRED` | A reconnect to a different bundle was sent without `bundleCid`. | ### Workflows and connectors A workflow’s access to the vault’s [connectors](/vaults/connectors/) is part of its consent. When you deploy a workflow, ROC grants it exactly the connector actions and events its graph uses, per connection, and removes access the graph no longer uses. ## How a request proves authority Every call into a vault traces back to a signature: | Proof | Used by | How | | -------------------- | -------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | | **Signed request** | Apps acting with the owner’s or a member’s key | The body is canonicalized and signed; sent as `X-Public-Key` / `X-Signature`. | | **Delegation token** | Clients acting on someone’s behalf without the key | `Authorization: Bearer `, scoped and revocable. | | **Execution token** | A running agent or workflow | Minted per task, short-lived, carrying the connection’s authority for that vault, scope, and run. Your code never sees a credential; `ctx` uses it. | ## Ending access | Action | Effect | | ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- | | **Disconnect** (`connection.disconnect()`) | Deletes the connection. Cubby data is kept, so a later reconnect resumes. | | **Revoke the agreement** | Removes the root of authority. The vault checks the registry for every active connection and tears down any whose agreement is revoked or expired. | Agreement reads are cached briefly, so revocation takes effect within that cache window rather than instantly. What the agent wrote stays in the vault. ## Errors Match on the error `code`, not the HTTP status. | Code | Meaning | | -------------------------------------- | -------------------------------------------------- | | `AUTH_MISSING` | No usable signature or token on the request. | | `AUTH_AMBIGUOUS` | Conflicting auth schemes on one request. | | `AGENT_ALREADY_CONNECTED` | A live connection already exists for this agent. | | `MANIFEST_NOT_FOUND` | No manifest for this agent id in the registry. | | `CUBBY_PROVISION_FAILED` | A declared cubby’s migration failed. | | `BUNDLE_CHANGED`, `RECONSENT_REQUIRED` | See [The bundle is pinned](#the-bundle-is-pinned). | ## Related * [Next: Memory Bank](/vaults/memory-bank/) * [Connect to a vault](/agents/ship/connect-to-a-vault/) * [Sovereign data](/vaults/sovereign-data/) * [Vaults](/vaults/overview/) * [Members](/vaults/members/) # Connectors > Vault-level connections to outside systems — Slack, email, Telegram, and MCP servers — how to add one in ROC, how its secrets are held, and how workflows use it as a trigger or an action. A **connector** links a vault to an outside system. It is set up once per vault, its secrets are sealed in the vault, and workflows the vault connects can then start on that system’s messages and act on it. | Term | Means | | -------------- | ----------------------------------------------------------------------------------------------------- | | **Connector** | A kind of outside system the vault supports, e.g. Slack. | | **Connection** | One configured account of a connector in a vault, holding its sealed credentials. | | **Access** | One agent’s permission to use one connection: which actions it may take and which events it receives. | Connections belong to the customer, not to your Agent Service. Your workflow uses a connection the vault owner set up; it never sees the secret. ## Available connectors | Connector | Credential | Starts a run on | Actions | | ---------------- | ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------- | | **Slack** | Bot token and signing secret from your own Slack app | **Message received**: a person posted in a channel the app is in, or mentioned it. Filter by channel. | **Send message**, **Send direct message**, **Update message**, **Add reaction** | | **Email (SMTP)** | SMTP host, port, `starttls` or `tls`, username, password | | **Send email** | | **Telegram** | Bot token from @BotFather | **Message received**: a message to your bot. Filter by chat. | **Send message** | | **MCP server** | Server URL and a static key (optional for a public server) | | The server’s own tools, discovered per connection | Every connector describes its settings, secrets, actions, and events. Read the catalogue with `vault.connectors()`. ## Add a connection in ROC Only an admin of the vault can add, change, or revoke connections. 1. Open **Connectors** and choose **Add connection**. 2. Pick the system. 3. Fill in **Name**, **Usable in** (the scopes workflows may use it from), **Settings**, and **Secrets**. Follow **Before you start** for the steps on the outside system. 4. For Slack, copy the **Request URL** shown after the connection is added into your Slack app’s Event Subscriptions. Telegram registers its webhook for you. ![The Add a connection dialog for Slack](/shots/connectors-add.png) Secrets are sealed in the vault and never shown again: not to you, not to agents. To change one, use **Update secrets**. To remove a connection and every agent’s access to it, open it and choose **Revoke this connection**. ## Use from a workflow In the Workflow Builder, a connector appears where its step goes: * **Under Triggers**, when it can start runs. Pick the system, then the event, then the connection. A run starts each time a matching message arrives. See [Triggers](/agents/workflows/triggers/). * **Under Actions**, when it can act. See below. ### Add a connector action step 1. Open **Add node** → **Actions** and pick the connector, then the action. A connector whose actions come from the connected server (an MCP server) offers **Choose an action** instead. 2. Choose the **{System} connection**. Only connections usable in the workflow’s scope are offered. If none is, add one (or add the scope to one) on the **Connectors** page and press **Refresh**. 3. Fill the **Input** fields: type a value, or click a box and then a field of the incoming item to map it. 4. Deploy. The step’s parameters, each connector’s actions, its error codes, and its limits are on [Steps: Connector action](/agents/workflows/steps/#connector-action). ### Access is part of the deploy When you deploy a workflow, ROC grants it exactly the actions and events its graph uses on each connection, in the workflow’s own scope, and removes access to connections it no longer uses. Editing a graph to reach a new connection takes effect only when you deploy it. A workflow deployed from code needs the same grant. The vault owner (or a member with write on the workflow’s scope) grants it through the vault API: ```plaintext PUT /api/v1/vaults/:vaultId/agents/:agentId/connections/:connectionId { "actions": ["send_message"], "events": [], "filter": {} } ``` The workflow must already be connected to the vault. Sending empty `actions` and `events` removes the grant. `vault.agents.setConnection` below does the same from the Vault SDK. ## From code The Vault SDK exposes the same model, for consoles and admin tools: ```ts const catalogue = await vault.connectors(); const slack = await vault.connections.create({ connector: "slack", name: "Support workspace", scopes: ["support"], config: { defaultChannel: "C0123456789" }, secrets: { botToken: "xoxb-…", signingSecret: "…" }, }); await vault.agents.setConnection(agentId, slack.id, { actions: ["send_message"], events: ["message.received"], }); ``` | Call | Does | | --------------------------------------------------------------------------------- | ------------------------------------------------------------ | | `vault.connections.list()` | The vault’s connections. Non-admins get no `config`. | | `vault.connections.create(body)` | Add one; `secrets` go in and never come back out. | | `vault.connections.update(id, body)` | Omitted keys stay; omitted `secrets` keeps the sealed ones. | | `vault.connections.revoke(id)` | Revoke it and every agent’s access to it. | | `vault.connections.actions(id)` | The actions a connection offers now (an MCP server’s tools). | | `vault.agents.connections(agentId)` | What one agent may do with each connection. | | `vault.agents.setConnection(agentId, connectionId, { actions, events, filter? })` | Set that access. Both lists empty removes it. | | `vault.agents.removeConnection(agentId, connectionId)` | Remove that access. | ## Related * [Next: Members](/vaults/members/) * [Triggers](/agents/workflows/triggers/) * [Steps: Connector action](/agents/workflows/steps/#connector-action) * [Connections and consent](/vaults/connections-and-consent/) # Data and trust > Where every kind of data lives on Manykind, who owns it, and what proves each boundary crossing — so you know where to put things and what your code can and cannot reach. You do not need to know how the platform is built to ship on it. You do need a model of where your code runs, where data lives, and what has to be true for the two to meet. This page is that model. ## Data classes | Class | Examples | Where it lives | Owner | | --------------------- | ----------------------------------------------------- | --------------------------------------------------------------------------------- | ---------------------------------------------------- | | **Vault data** | Events and objects in a scope | The customer’s vault, backed by a storage bucket their wallet owns | The vault owner | | **Memory Bank** | Records and relations, each privacy-classified | The customer’s vault | The vault owner | | **Connector secrets** | Bot tokens, SMTP passwords, API keys | Sealed in the customer’s vault; never returned, not to people and not to agents | The vault owner | | **Cubbies** | Your service’s working tables | One database per vault, service, and alias | Stored in the vault; schema declared by your service | | **Agent code** | Bundles, manifests, widgets | Your Agent Service’s storage bucket, content-addressed; each version is immutable | You | | **Agreements** | Signed consent between a vault owner and your service | The agreement registry | Signed by the vault owner | | **Run records** | Jobs, Tasks, task logs, usage | The platform | Attributed to the run’s agent and initiator | Connection settings are not kept on run records. The platform reads them from the vault when a Task runs and hands them to your code as `ctx.settings`. **Where to put something:** * Working state your service needs next time: a [cubby](/agents/building-blocks/cubbies/). * A conclusion the customer should keep and see, with a visibility class: the [Memory Bank](/vaults/memory-bank/). * A file: an object in a scope (`ctx.vault.objects`, or `scope.objects` in the Vault SDK). * A credential for an outside system: a vault [connector](/vaults/connectors/), never your bundle. Nothing survives in an isolate’s memory between events. ## How data reaches your code 1. **You publish** an agent or workflow to your Agent Service’s bucket and deploy a version. 2. **The owner connects** it, signing an agreement for named scopes. The vault pins the bundle they consented to. 3. **Data arrives** in a scope: from your app, a widget, a connector, a schedule, a webhook, another agent, or a person. See [Data onboarding](/vaults/build-on-a-vault/data-onboarding/). 4. **An event dispatches** to every active connection on that scope, opening or advancing a Job. 5. **Your code runs** with an execution token scoped to that vault, scope, and run. It reaches the vault only through `ctx` (or the workflow runner’s steps). 6. **Results land** back in the vault. ## What proves each crossing | Crossing | Proof | | --------------------------------- | ------------------------------------------------------------------------------------------------------------- | | Owner or member → vault | Wallet signature over the canonicalized request (`X-Public-Key`, `X-Signature`), or a scoped delegation token | | Owner → agreement registry | Wallet signature on the agreement | | You → your Agent Service’s bucket | The bucket owner’s key, or an access token the owner granted | | Running agent → vault | A per-task execution token, short-lived and bound to the connection | | Running agent → outside system | A connector action, using the secret sealed in the vault | Your code never signs as the customer. It runs inside a Task the platform already authorized against the agreement, which is why nothing in `ctx` asks for a credential, and why nothing in your code can widen what a run may touch. ## The invariant **The customer’s data and consent do not depend on your app.** The bytes live in storage the customer’s wallet owns; the agreements are records their wallet signed. Switching agents, or switching away from yours, is a **permission change, not a data migration**: nothing has to move, because your service was never the source of truth for the customer’s data. ## Related * [Next: Data onboarding](/vaults/build-on-a-vault/data-onboarding/) * [Sovereign data](/vaults/sovereign-data/) * [Connections and consent](/vaults/connections-and-consent/) * [Runs and events](/agents/how-agents-run/) * [Vaults](/vaults/overview/) # Members > Who can do what in a vault — the owner, members and their roles, per-scope grants, and privacy ceilings. A vault has one **owner** and any number of **members**. A vault with members is an organization’s vault; everything below applies to it. ## Owner The owner is the wallet that owns the vault. The owner holds every scope by virtue of the key and signs every member’s grant. The owner is never a member row; rosters show the owner first, marked as owner. ## Members Each member has: | Field | Values | | ------------------- | ------------------------------------------------------------------------------------------------------------------------------- | | **Role** | `admin` or `member`. Admins can manage the vault’s settings, such as its [connectors](/vaults/connectors/). | | **Scope grants** | One level per scope (below). A member sees only scopes they are granted. | | **Privacy ceiling** | Optional: the most sensitive [Memory Bank](/vaults/memory-bank/) class they may read: Public, Internal, Private, or Restricted. | | **State** | `active` or `suspended`. | | **Expiry** | Optional. | ### Scope levels | Level | ROC label | Allows | | -------- | --------- | ------------------------------------------------------------------------------------------ | | `member` | Write | Read, write, and run in the scope. | | `reader` | Read | Read only. | | `audit` | Audit | Knowing the scope exists, never its contents. An audit-only member’s requests are refused. | ROC shows a vault’s scopes as **Domains**. ## Grants are signed The owner’s wallet signs each member’s grant: their role, scope levels, expiry, and privacy ceiling. A party that does not trust the vault service can fetch the signed roster and verify every grant against the owner’s key. ## In ROC Open **Members** under Organization. The **Members** tab lists people, their roles, and their access; the **Domains** tab lists scopes and who can reach each. ## From code ```ts const mine = await sdk.memberships.mine(); // vaults I am a member of const org = await sdk.vault.byId(mine[0].vaultId); // open one const roster = await sdk.memberships.list(org.id); // its members // Owner only: add a member and grant them scopes. await sdk.memberships.put(org.id, memberPubkey, "member"); await sdk.memberships.putScopes(org.id, memberPubkey, { sales: "member", support: "reader" }); ``` `put` and `putScopes` sign the grant, so they need the owner’s `signer` on the `VaultSDK`. `sdk.vault.current()` is always your own vault; `mine()` does not include it. ## Related * [Next: Data and trust](/vaults/data-and-trust/) * [Vaults](/vaults/overview/) * [Memory Bank](/vaults/memory-bank/) * [Connectors](/vaults/connectors/) * [Team](/get-started/team/) # Memory Bank > The vault's durable, privacy-classified record of what its agents and workflows learned — what it holds, how data gets in, how agents, workflows, apps, and LLM agents read it, and what ROC shows. The **Memory Bank** is the vault’s long-term memory: what its agents and workflows have learned, and how each run builds on the ones before it. It belongs to the vault, so it belongs to the customer. Agents come and go; the Memory Bank stays. ## What it holds A graph of **records** and **relations**. | Record field | Meaning | | --------------- | ----------------------------------------------------------------------- | | `id` | Stable id you choose. Writing the same id again updates the record. | | `type` | What kind of thing it is, e.g. `fact`, `decision`, `principle`, `note`. | | `title`, `body` | The content. | | `scope` | Where it lives: the vault scope that governs who may read it. | | `privacy` | How sensitive it is: `public`, `internal`, `private`, or `restricted`. | A **relation** is a directed, typed edge from one record (`in`) to another (`out`). It carries its own `scope` and `privacy`, because a link can be more sensitive than the records it joins. Records also move through review states (`proposed`, `accepted`, `rejected`, `superseded`). Reads return accepted records unless they ask for other states. ### Two axes of access Every read is gated twice: * **Scope** decides *where* a reader may look. A reader sees only scopes their grants reach. An ungranted scope returns no rows, never an error. * **Privacy** decides *how sensitive* a fact a reader may see. A vault member can hold a privacy ceiling (Public, Internal, Private, or Restricted). A run reads with the permissions of the person who started it. ## How data gets in Writes come from agents and workflows only. Your app reads the Memory Bank but cannot write to it. | Writer | How | | ------------ | ---------------------------------------------------------------------------------------------------------------- | | A workflow | `remember` (file a record) and `relate` (link two records) steps. | | A code agent | `ctx.memory.upsert(record)`, `.update(id, patch)`, `.setPrivacy(id, privacy)`, `.delete(id)`, `.relation(edge)`. | `privacy` is required on every write and never defaulted. Before filing anything, answer: *who reads this later, and does each row need its own visibility?* ```ts await ctx.memory.upsert({ id: `decision:${ticketId}`, type: "decision", title: "Refunds over 500 need a second approver", body: "Agreed after the March audit.", scope: "finance", privacy: "internal", }); await ctx.memory.relation({ in: `decision:${ticketId}`, out: "principle:four-eyes", type: "applies", scope: "finance", privacy: "internal", }); ``` Work in progress belongs in a [cubby](/agents/building-blocks/cubbies/). File only what the organization should keep. ## How agents and workflows read it | Reader | How | | ------------ | ----------------------------------------------------------------------------------------------------------------------------------- | | A workflow | `recall` step: search, then read each hit’s body. | | A code agent | `ctx.memory.search(match, { limit })`, `.get(id)`, `.neighbours(id, { limit })`, `.countByType()`. Rows come back as plain objects. | | A widget | A named query with `tool: "search" \| "get" \| "neighbours" \| "countByType"`. See [Widgets](/agents/building-blocks/widgets/). | | Your app | `vault.memory.search`, `.list`, `.get`, `.neighbours`, `.countByType` in the Vault SDK. | | An LLM agent | The vault’s MCP endpoint (below). | ```ts const hits = await ctx.memory.search("refund approval", { limit: 5 }); ``` ### From your app The Vault SDK returns rows exactly as the vault sends them: a `columns` list and positional `rows`, plus a `cursor` when more rows exist. ```ts const page = await vault.memory.search("refund approval", { scope: "finance", limit: 20 }); // page.columns: id, type, title, scope, privacy, rank, snippet const next = page.cursor ? await vault.memory.search("refund approval", { scope: "finance", cursor: page.cursor }) : undefined; ``` Read `neighbours` rows by position: the far record’s `type` and the edge’s `type` share a column name. ### From an LLM agent: MCP The vault serves its Memory Bank as read-only MCP tools at: ```text POST /api/v1/vaults/:vaultId/mcp ``` It is a stateless Streamable HTTP server, authenticated like any vault read. During a run, an LLM agent that receives the run’s execution token presents it, so it searches with the permissions of the person who started the run. See [Authentication](/agents/llm-agents/bring-your-own/#authentication). | Tool | Arguments | Limits | | ------------ | --------------------------------------------------------- | ---------------------------------------------------- | | `search` | `query`, optional `domains`, `types`, `statuses`, `limit` | `limit` default 10, at most 25 | | `get` | `ids`, optional `body_chars`, `offset` | 1 to 10 ids; `body_chars` default 4000, at most 8000 | | `neighbours` | `id`, optional `edge_type`, `direction`, `limit` | `limit` default 10, at most 25 | Search matches words of three or more letters by stem, ranks records that match more and rarer words first, and returns ids, titles, and a snippet; call `get` for full text. Tool failures come back as MCP tool errors the model can act on; authentication failures stay HTTP 401 or 403. ## In ROC The **Memory Bank** page shows the vault behind the selected Agent Service, and only what your grants reach. | Tab | Shows | | ------------ | ------------------------------------------------------------------------------------------------------------------------------------ | | **Overview** | Ask a question and get an answer drawn from the records only; knowledge and run counts, knowledge by type, recently changed records. | | **Journeys** | Runs for one domain and subject, step by step, with key takeaways; compare a run with others. | | **Browse** | Search and open records; filter to **Needs review**; **Show run entries** to include the runs’ own bookkeeping. | | **Outcomes** | What a workflow wrote and what a person decided afterwards. | ![Memory Bank Overview tab with the Ask box and knowledge counts](/shots/memory-bank-overview.png) ## Related * [Next: Connectors](/vaults/connectors/) * [Vaults](/vaults/overview/) * [Members](/vaults/members/) * [Cubbies](/agents/building-blocks/cubbies/) * [Workflow steps](/agents/workflows/steps/) # Vaults > The customer side of Manykind — what a vault is, who owns it, what it holds, and how scopes partition it. This group covers the customer side: the vault, sovereign data, connections and consent, the Memory Bank, connectors, members, and how your own apps build on a vault. A **vault** is a customer’s own store for their own data. Your agents and workflows work inside it; your apps read and write it through the Vault SDK. The customer owns it, not you and not the platform. ## Who owns a vault A vault is owned by a **wallet**. One wallet owns exactly one vault, and claiming it is idempotent: `sdk.vault.ensure()` returns the caller’s vault, creating it on first use. A vault can belong to: * **A person.** One owner, no other members. * **An organization.** The owner adds [members](/vaults/members/) and grants each of them access to scopes. A vault becomes an organization’s vault by having a second member; its name is the vault’s display name. ## What a vault holds | Part | What it is | | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------- | | **Scopes** | Named partitions. Events, objects, and connections are all per scope. | | **Events** | Typed messages published into a scope, grouped into streams by `context`. See [Runs and events](/agents/how-agents-run/). | | **Objects** | Files stored in a scope at a path you choose. | | **[Memory Bank](/vaults/memory-bank/)** | Records and relations its agents and workflows filed, each with a privacy class. | | **[Connectors](/vaults/connectors/)** | Configured accounts on outside systems, with their secrets sealed in the vault. | | **Agent connections** | The agents and workflows the owner [connected](/vaults/connections-and-consent/), with their agreements, settings, and pinned bundles. | | **Cubbies** | The databases of each connected Agent Service. See [Cubbies](/agents/building-blocks/cubbies/). | | **[Members](/vaults/members/)** | Who else may use the vault, and at what level. | The vault’s storage is a bucket the owner’s wallet controls. ## Scopes A scope is the unit every grant is made in. A customer can connect a fitness agent to `health` without exposing `finance`. | Rule | Value | | --------------- | --------------------------------------------------------------------------------------------------- | | Name pattern | `^[a-z][a-z0-9-]*$`, at most 63 characters | | Always present | `default` | | Reserved | `default`, `shared` | | Optional fields | `displayName`, `streamLabel`, `metadata` (a free-form map; updates shallow-merge per top-level key) | In ROC, a vault’s scopes are shown as **Domains** on the Members page. ```ts await vault.scopes.create({ name: "health", displayName: "Health" }); const scopes = await vault.scopes.list(); ``` ## Credential status A vault’s storage credential refreshes in the background. `vault.status` is `active`, `refreshing`, or `disconnected`; `vault.isDisconnected()` is true when the credential can no longer be refreshed and the owner has to act. ## Related * [Next: Sovereign data](/vaults/sovereign-data/) * [Memory Bank](/vaults/memory-bank/) * [Members](/vaults/members/) * [Data onboarding](/vaults/build-on-a-vault/data-onboarding/) # Sovereign data > The idea Manykind is built on — the data belongs to the customer and stays in their vault, and your agents and workflows are guests that compute against it under a signed, scoped, revocable agreement. This section is the customer side: the idea it rests on, the vault and its connections, the Memory Bank, connectors, and members a customer owns, and how your own apps build on a vault. Before any API, there is one idea the platform is arranged around: **the data belongs to the customer, and it stays in their vault.** What you build does not own that data or copy it out. It is a *guest* with permission to compute against it. ## The inversion In a conventional app the user signs up, connects an account, and their records land in *your* database. From then on their data lives in your product, on your terms: **the user is the guest.** They can ask you to delete it, but they cannot check that you did, and they cannot take their history with them. Manykind turns that around. The data sits in the customer’s [vault](/vaults/overview/), backed by storage their wallet controls. Your agent or workflow travels to the data: it runs against the vault under an agreement the owner signed, touches only what that agreement covers, and leaves everything where it was when the owner withdraws it. **Now you are the guest.** | | Conventional app | Manykind | | ------------------------- | ------------------------------------------ | ------------------------------------------------------ | | Whose data it is | Yours to hold, once handed over | The customer’s, always | | Where it lives | Copied into your database | In the customer’s vault | | Who is the guest | The user, inside your product | Your agent, inside the customer’s vault | | How access starts | The user accepts terms and hands data over | The owner signs an agreement naming what you may touch | | How access ends | You are asked to delete your copy | The owner revokes the agreement | | After the customer leaves | Your copy stays with your company | Nothing to keep; it was never yours | | Who holds the authority | Your platform account | The owner’s wallet | The customer keeps control of their history. You stop running a data pipeline you never wanted: no records to migrate, no store of user data to secure, no deletion requests to service. ## What being a guest means Your agent or workflow operates under a **signed, scoped, revocable** agreement: * **Signed.** The vault owner’s wallet signs it. That signature is the authority, and nothing you hold can stand in for it. * **Scoped.** An agreement is never “access to the vault.” It names one or more **scopes**, named partitions of the vault, and your agent works only there. * **Revocable.** The owner can withdraw it at any time. What your agent wrote stays in the vault; connect a different agent later and it can pick up from there. The mechanics are in [Connections and consent](/vaults/connections-and-consent/). ## Authority roots in the owner’s wallet On conventional platforms, capability arrives as a credential *you* hold: an API key or service token. On Manykind no such credential reaches a customer’s data. You could hold every secret you own and still read none of it. Capability rests on two things that belong to the customer: * the **wallet** that owns their vault, and * an **agreement that wallet signed**, naming which Agent Service may operate in which scopes. You cannot widen your own access. You can only ask the owner to sign a broader agreement. That is what *self-sovereign* means in practice. ## The standard behind it None of this holds if it is one platform’s policy, so the ownership model is specified rather than promised. The **Self-Sovereign Context Protocol (SSCP)** makes the customer’s own data, not the agent and not the platform, the unit of portability, with the owner’s wallet as the only authority over it. It sits above agent-access protocols such as MCP and A2A, and beside identity protocols such as OAuth and DIDs, filling the ownership and consent layer they leave open. Manykind is the first implementation of SSCP. What you build against is stable either way: your agent connects to a vault under a signed agreement, and that contract is what SSCP standardizes. ## Everything you build works this way * **Agents and workflows** run on the platform against a vault that connected them. They keep shared working state in [cubbies](/agents/building-blocks/cubbies/) and file durable conclusions into the vault’s [Memory Bank](/vaults/memory-bank/). * **Widgets** are the UI your Agent Service ships. They run in the browser as the signed-in person, under that person’s access to the vault. * **Your own apps** (web, mobile, server) talk to a vault through the [Vault SDK](/vaults/build-on-a-vault/work-with-the-vault-sdk/) under the same rules. *Self-sovereign* here means **owned and revocable**: the customer owns their data and can revoke any agreement by signature. It does not mean a hosting operator is cryptographically unable to read the data. Ask for the narrowest scope you need, and treat a revoked agreement as a hard stop. ## Related * [Next: Connections and consent](/vaults/connections-and-consent/) * [How it fits together](/get-started/how-it-fits-together/) * [Vaults](/vaults/overview/) * [Data and trust](/vaults/data-and-trust/)