Test and debug
Test a code agent in-process first: @cef-ai/testing runs your engagement classes against an in-memory vault and SQLite cubbies built from your real migrations, with no network. Go to the platform only when the bug needs a real model, real vault data, or a real trigger, and read the agent’s logs there.
pnpm add -D @cef-ai/testing vitestThe local loop
| Loop | Command | Catches |
|---|---|---|
| Handler logic | vitest + testAgent |
Payload checks, cubby SQL, what the handler publishes, model-output handling. console.* prints to your terminal. |
| Several agents, or your app with an agent | vitest + testPlatform |
Events routed between agents, Jobs per stream, Memory Bank and object writes. |
| Build checks | cef build |
Banned imports, non-literal @OnEvent types, undeclared model aliases, unhandled schedule events. See What cef build rejects. |
| Bundle routing | cef inspect dist/<alias> |
A handler the manifest routes to that the bundle does not export. Exits with code 2, so it fits CI. |
cef test is not implemented; run vitest directly.
Test one agent with testAgent
testAgent(Agent, opts) builds a harness around one agent class. dispatch runs one event through the matching @OnEvent handler and resolves to the events the handler published. The example tests the agent from Write an agent.
import path from "node:path";import { afterEach, describe, expect, it } from "vitest";import { testAgent, type TestHarness } from "@cef-ai/testing";import HelloAgent from "../src/agent.js";
const HISTORY = path.resolve("migrations/history");
describe("hello agent", () => { let h: TestHarness | undefined; afterEach(() => h?.dispose());
it("replies and stores the message", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); const events = await h.dispatch({ type: "user_message", payload: { text: "hi" } });
expect(events.map((e) => e.payload)).toEqual([{ text: "you said: hi" }]); const rows = await h.cubby("history").query<{ text: string }>("SELECT text FROM messages"); expect(rows.map((r) => r.text)).toEqual(["hi"]); });
it("answers bad input without writing", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); await h.dispatch({ type: "user_message", payload: {} });
expect(await h.cubby("history").query("SELECT * FROM messages")).toEqual([]); });
it("runs the lifecycle hooks", async () => { h = testAgent(HelloAgent, { cubbies: [{ alias: "history", migrations: HISTORY }] }); await h.start(); await h.close("idle_timeout"); expect(h.lifecycle).toEqual({ state: "closed", terminalReason: "idle_timeout" }); });});An event type no handler declares is ignored, as on the platform. A handler that throws rejects dispatch, so assert on bad input explicitly.
Stub what leaves the agent
| Dependency | Stub it with |
|---|---|
ctx.models[alias] |
models: { llm: createModelMock() }, then llm.expect(input).respond(output). An unregistered input throws [modelMock] no expectation for input …. llm.calls lists every call. See Models. |
Global fetch |
h.fetchMock().when({ url, method? }).reply(status, body?). A request with no stub throws [fetchMock] no stub for <METHOD> <url>. assertCalled({ url }) checks a call was made. |
ctx.settings, ctx.params |
The settings and params options. Both are frozen, as on the platform. |
ctx.self |
self: { agentId, vaultId, scope, context }. The default is a valid synthetic identity, so code that derives a sibling’s id works. |
| Time | clock: { now }, then h.advanceTime(ms). |
testAgent does not back ctx.memory or ctx.vault.objects: calling them throws. Test an agent that uses either with testPlatform.
Replay a recorded session
h.replay("test/fixtures/session.jsonl") feeds a file of events, one JSON object per line with at least type, through the harness. It resolves to the events published during the replay, a row count per cubby table, and the duration. Use it to pin an agent’s behavior over a whole stream. Keep fixtures synthetic, never copied from customer data.
h.snapshot() and h.restore(snap) capture and restore cubbies, the clock, and the publish log, so several tests can start from one prepared state.
Test agents together with testPlatform
testPlatform({ agents }) runs agents in a simulated vault: a publish routes to the connected agents that handle its type and resolves once they have run, ctx.memory writes to an in-memory Memory Bank, and ctx.vault.objects to an in-memory object store (p.objects).
import path from "node:path";import { expect, it } from "vitest";import { testPlatform } from "@cef-ai/testing";import HelloAgent from "../src/agent.js";
it("replies in the same stream", async () => { const p = testPlatform({ agents: { hello: { source: HelloAgent, cubbies: [{ alias: "history", migrations: path.resolve("migrations/history") }] }, }, }); await p.vault.agents.connect({ agentId: "hello" });
await p.vault.scope("default").publish({ type: "user_message", context: "c-1", payload: { text: "hi" } });
const page = await p.vault.scope("default").stream("c-1").events.list({ types: ["reply"] }); expect(page.items.map((e) => e.payload)).toEqual([{ text: "you said: hi" }]); p.dispose();});Replace an agent you do not want in the test with mockAgent. p.orchestrator.jobs.list() shows the Jobs the publishes opened, and p.runInCubby(agentId, alias, fn) reads an agent’s cubby. Every option is in the testing reference.
The simulator does not verify signatures or consent, run real models, or enforce platform limits you do not inject. Those need the platform.
Read logs
Inside an agent, console.* is the logging API. On the platform, output is captured and shipped to the log store, not to a terminal.
logandinforecord at info level;warn,error, anddebugkeep their own levels.- Records are labelled with the agent, version, Job, and Task, so you can narrow to one run.
In ROC, open the agent from Agents and select Logs. Filter by Time range and Level, and narrow to one Job or Task.
With the Vault SDK, read one Task’s logs from a vault you own (useful in an integration test):
const page = await vault.jobs.get(jobId).tasks.logs(taskId);Pages carry { logs, total, offset, limit, hasMore }; limit defaults to 100 and is capped at 1000.
Logs are a short-term aid, not a record:
- Logs are batched before they ship, so a line can appear a moment after the code that wrote it, and a batch that cannot be delivered is dropped.
- Only the last 24 hours are queryable. Anything you must explain later belongs in a cubby row.
Common failure signals
| You see | Likely cause | Fix |
|---|---|---|
| No Job opens after you publish | The agent is not connected in that scope, or no @OnEvent handles the event’s type. An event no connection handles is dropped without an error. |
Connect the agent; check the event type against the handlers cef inspect prints. |
Connect fails with MANIFEST_INVALID |
The scope is not in the agent’s requiredScopes. |
Connect into a declared scope, or push a version that declares it. |
Connect fails with CUBBY_PROVISION_FAILED |
A cubby migration failed. | Fix it in a new migration file and push a new version. See Cubby schema. |
Task fails with execute_failed |
The handler threw, or the call could not complete. | Read the Task’s logs; validate the payload before using it. |
Task fails with execute_timeout |
The handler ran past its time budget. | Do less per event. |
Task fails with agent_threw_non_error or agent_invalid_result |
The handler threw a non-Error value, or returned a value that does not survive JSON serialization. |
Throw Errors; return plain data. |
model alias not declared: '<alias>' |
The models key does not match the alias inside the model’s model.json. |
Declare it under the model’s own alias. See Models. |
| Two rows for one event | The event was delivered more than once and the write is not idempotent. | Key rows on the event id; use ON CONFLICT or INSERT OR IGNORE. See Event streams. |
@OnClose receives idle_timeout before the work is done |
The Job had no activity for idleTimeout. |
Raise idleTimeout in cef.config.ts. |
Task fails with gpu_units_ceiling |
The connection reached its compute limit. | See Spend limits. |
| Logs are empty | The run is older than 24 hours, or the batch was dropped. | Record what you need later in a cubby row. |
Every error code is in Errors. Debugging a workflow run is on Monitor runs.