@cef-ai/testing
@cef-ai/testing runs agents and workflows in-process against a simulated platform: an in-memory vault, the job machinery, and SQLite cubbies built from your real migrations. No Docker, no GPU, no network.
pnpm add -D @cef-ai/testing vitest| Version | 3.3.5 |
| Depends on | @cef-ai/agent-sdk, @cef-ai/vault-sdk, better-sqlite3 |
| Module | ESM |
Exports: testAgent, testPlatform, mockAgent, testWallet, fromMarketplace, createModelMock, fakeObjectStore, and their types.
Which entry point
| Use | When |
|---|---|
testAgent(Agent, opts) |
Unit tests of one code agent’s handlers: dispatch an event, assert publishes and cubby rows. |
testPlatform({ agents, models }) |
A workflow (register WorkflowRunner as an agent), several agents that talk to each other, or app code that uses the vault. |
To test a workflow, see the worked example in Test a workflow.
testAgent
testAgent(source: AgentCtor | TestAgentMultiSpec, opts?: TestHarnessOptions): TestHarnessimport { afterEach, expect, it } from "vitest";import { testAgent, type TestHarness } from "@cef-ai/testing";import Echo from "../src/agent.js";
let h: TestHarness | undefined;afterEach(() => h?.dispose());
it("acks a message", async () => { h = testAgent(Echo, { cubbies: [{ alias: "history", migrations: "./cubbies/history" }] }); const events = await h.dispatch({ type: "user_message", payload: { text: "hello" } }); expect(events.map((e) => e.type)).toContain("ack");});source is a class with @OnEvent / @OnStart / @OnClose handlers, or { engagements: [{ id, source, goal? }], targeting? } for several engagements.
TestHarnessOptions
| Option | Type | Meaning |
|---|---|---|
cubbies |
Array<string | { alias, migrations? }> |
Cubbies to create, each from a directory of *.sql migrations. An alias not listed is created empty on first use. |
models |
Record<string, ModelHandle | ModelMockHandle> |
ctx.models[alias]. |
settings |
Record<string, unknown> |
ctx.settings, frozen. |
params |
Record<string, unknown> |
ctx.params, frozen. Defaults to {}. |
clock |
{ now?: number } |
Initial time in unix milliseconds. Defaults to Date.now(). |
self |
{ agentId?, vaultId?, scope?, context? } |
ctx.self. Defaults to a valid synthetic identity. |
TestHarness
| Member | Meaning |
|---|---|
dispatch({ type, payload?, context?, role?, from? }) |
Run one event through the agent. Resolves to the events published during this call. |
published |
Every event published since the harness was created. |
cubby(alias) |
A cubby handle (query, exec) for assertions. |
runInCubby(alias, fn) |
Call fn(handle, db) with the cubby handle and the raw better-sqlite3 database. |
start() / close(reason?) |
Run @OnStart / @OnClose. |
lifecycle |
{ state: "idle" | "started" | "closed", terminalReason? }. |
engagement |
The engagement the harness selected (multi-engagement sources). |
snapshot() / restore(snap) |
Capture and restore cubbies, clock, and the publish log. |
replay(jsonlPath) |
Feed a JSONL file of events through the harness. |
advanceTime(ms) / now() |
Move and read the injected clock. |
fetchMock() |
The fetch mock installed as the global fetch during each invocation. |
models |
The model handles on ctx.models. |
dispose() |
Close SQLite handles. Idempotent; call it in afterEach. |
Agents call the global fetch and log with console.*; there is no ctx.fetch or ctx.log.
testPlatform
testPlatform(opts: TestPlatformOptions): TestPlatformimport { testPlatform, createModelMock } from "@cef-ai/testing";import { WorkflowRunner } from "@cef-ai/agent-sdk/workflow";import config from "../cef.config.js";
const p = testPlatform({ agents: { "ticket-triage": { source: WorkflowRunner, cubbies: config.cubbies, params: { graph: config.params!.graph!.default } }, }, models: { classifier: createModelMock() },});await p.vault.agents.connect({ agentId: "ticket-triage" });await p.vault.scope("default").publish({ type: "workflow.start", context: "t-1", payload: { ticketId: "T-1", text: "…" } });TestPlatformOptions
| Option | Type | Meaning |
|---|---|---|
agents |
Record<agentId, AgentRegistration> |
Required. Each entry is a class, a mockAgent(...), a long form { source, cubbies?, settings?, params? }, or { engagements, cubbies?, settings?, targeting? }. |
models |
Record<string, ModelHandle | ModelMockHandle> |
Shared by every agent’s ctx.models. |
users |
Record<userId, { wallet? }> |
One vault per user. Omit for a single implicit user. |
vaultId |
string | The implicit user’s id when users is omitted. Default "default". |
scope |
string | Default scope for publishes that omit one. |
context |
string | Default context for publishes that omit one. Default "default". |
clock |
{ now?: number } |
Initial time. |
params on a registration is ctx.params. A workflow runner reads its graph from params.graph, as on the platform.
Cubbies are keyed by Agent Service (the part of the agent id before :) and by user, as on the platform: agents of one service that declare the same alias share one database, and each user’s vault has its own.
TestPlatform
| Member | Meaning |
|---|---|
vault |
The single user’s vault. Throws when two or more users are declared. |
user(id).vault |
One user’s vault. |
vault.agents.connect({ agentId }) |
Connect a registered agent; disconnect, get, list. |
vault.scope(name).publish({ type, context?, payload, from? }) |
Publish an event; agents whose handlers match run. |
vault.scope(name).stream(context).events.list({ types?, limit? }) |
Read a context’s events: { items, hasMore }. |
vault.scope(name).subscribe(...) |
Subscribe with a handler, or iterate events. |
orchestrator.jobs, .list(filter?), .get(jobId) |
Read-only view of Jobs across all users. |
runInCubby(agentId, alias, fn, userId?) |
Call fn(handle, db) against an agent’s cubby. |
cubbyOps() |
Every cubby statement any agent ran, oldest first: { agentId, cubby, op, sql, params, nodeId? }. nodeId names the workflow step that ran it. |
objects |
The in-memory ctx.vault.objects store (a FakeObjectStore). |
faults.refusePublish |
(type, payload) => boolean. Return true to make ctx.vault.publish fail, as the vault does for a body over 1 MiB. |
clock |
Shared clock. |
fetchMock() |
Shared fetch mock. |
models |
Shared model handles. |
dispose() |
Tear down every simulated agent. Idempotent. |
mockAgent
mockAgent(spec: MockAgentSpec): MockAgentDescriptorA stand-in for an agent you do not want in the test: handlers as plain functions.
| Field | Meaning |
|---|---|
agentId |
Required. |
on |
Record<eventType, (event, ctx) => unknown> |
onStart, onClose |
(ctx, reason?) => unknown |
publishes |
Event types it publishes (descriptive) |
engagements, targeting, settings |
The multi-engagement form: engagements: [{ id, goal?, on?, onStart?, onClose? }] |
A workflow’s agent step publishes workflow.step to the agent and waits for agent.answered. A mock answers like this:
const reviewer = mockAgent({ agentId: "reviewer", on: { "workflow.step": async (event, ctx) => { const input = event.payload as { text: string }; await ctx.vault.publish("agent.answered", { agent: "reviewer", text: JSON.stringify({ approved: input.text.length > 10 }) }); // or, to fail the step: { agent: "reviewer", error: "upstream unavailable" } }, },});Register it under the same id the step’s use names, and connect it like any agent.
createModelMock
createModelMock(): ModelMockHandle| Member | Meaning |
|---|---|
expect(input).respond(output) |
Register an answer. An infer call whose input is deep-equal (by JSON.stringify) to input returns output. A later expectation for the same input wins. |
calls |
Every call: { method: "infer" | "stream", input }. |
infer(input) |
Returns the matching output, or throws [modelMock] no expectation for input …. |
stream(input) |
Yields each element of an array output, or the output once. |
Any object with an infer method is accepted wherever a model handle is, so a hand-written stub works too.
A workflow’s model step resolves its input params against the item before calling the model; register the resolved input.
Fetch mock
| Member | Meaning |
|---|---|
when({ url, method? }).reply(status, body?, headers?) |
Stub a request. url is a string or RegExp; stubs are tried in order. |
assertCalled({ url, method? }) |
Throw if no recorded call matched. |
calls |
{ url, method } for each call. |
FakeObjectStore
fakeObjectStore() returns the store behind TestPlatform.objects. Objects are create-only, as on the platform: a second upload to the same path fails with 409.
| Member | Meaning |
|---|---|
text(vaultId, scope, path) |
Read an uploaded object back as text. |
objects |
The underlying map. |
refuseUpload |
(path) => boolean. Return true to fail uploads with 503. |
handle(vaultId, scope) |
The ctx.vault.objects handle for that vault and scope. |
Other exports
| Export | Meaning |
|---|---|
testWallet(id) |
{ publicKey: "wallet-<id>" } for users[id].wallet. Signatures are not verified. |
fromMarketplace(url) |
Throws: not available in the simulator. Use a class or mockAgent. |
Types from @cef-ai/vault-sdk |
AgentConnection, EventInput, EventsPage, PublishResult, Scope, Subscription, VaultRecord, and others, so fixtures type-check against the real SDK. |
Types from @cef-ai/agent-sdk |
Rule, Targeting, TargetingEngagement. |
What the simulator does not do
- Verify signatures or consent: wallets are stubs.
- Run real models: every model call goes to a handle you provide.
- Enforce platform limits other than those you inject (
faults,refuseUpload). - Talk to a live vault: everything is in-process.