Models
Building blocks are what your agents use besides their own logic: models, cubbies, and widgets. Each belongs to your Agent Service, and each page here shows how a workflow and a code agent use it.
Agents and workflows call models without holding an API key, picking a provider, or knowing where a GPU lives. You bind an alias to a model from the platform’s catalogue, call the alias, and the platform routes the request to a server that is serving that model, meters the call, and returns the output.
Where models come from
The platform serves a curated catalogue. Each model is described by a
model.json in content-addressed storage, at:
<cdn>/<bucket>/models/<name>/<version>/model.jsonThe model.json names the model, its alias, its tasks, its input and output
schema, and whether it streams. The catalogue spans text generation,
speech-to-text, embeddings, vision, and other tasks. What a model accepts and
returns is defined by its own schema, not by a fixed request shape.
Open Models in ROC’s side navigation to browse it. Each card shows the
model’s name, version, task, alias (as ctx.models.<alias>), price, and whether
it is available; search by name, alias, task, or version, and filter by
availability (Active or Pending). Open a model to see its input and
output schema and try it.
The models in your own Agent Service’s bucket are also listed by the API:
GET /api/v1/agent-services/:asPubKey/modelsDeclare an alias
A workflow and a code agent declare models the same way: a models map in
cef.config.ts, from alias to the model’s model.json URL.
import { defineAgent } from "@cef-ai/agent-sdk/config";
export default defineAgent({ id: "summarizer", version: "1.0.0", entry: "./src/agent.ts", models: { llm: "https://cdn.ddc-dragon.com/<bucket>/models/<name>/<version>/model.json", },});defineWorkflow takes the same models field.
| Rule | Detail |
|---|---|
| URL | The path must be /<bucket>/models/<name>/<version>/model.json, with an integer bucket. cef build reads the bucket, name, and version from it and refuses any other shape. |
| Key | Must equal the alias written inside that model.json, which is often not the name in its URL. A key that does not match fails at the first call with model alias not declared: '<alias>'. |
| Use | A model step or a ctx.models.<alias> reference to an alias that is not in models fails cef build. |
The alias is the stable handle; the binding lives in config. Swapping a model is a config change and a redeploy, not a code change.
In a workflow
A model step names an alias, resolves its request from the carried item,
calls the model once, and puts the answer on the item under into:
{ id: "classify", kind: "model", params: { alias: "llm", input: { prompt: "=Classify this ticket:\n{{ $json.text }}", max_tokens: 64 }, into: "classification", },}Its parameters, structured output, per-entry calls, and when to use an Agent step instead are on Steps: Model.
In a code agent
A code agent calls a declared model with ctx.models[alias].infer(input).
Type it. Run cef typegen after you declare or change models. It reads
each declared model.json and writes .cef/generated.d.ts, which types
ctx.models.<alias> from the model’s input and output schemas, and it writes
cef.lock.json.
Call it.
@OnEvent("document.added")async onDocument(event: Event<{ text: string }>, ctx: Context) { const out = (await ctx.models.llm.infer({ messages: [ { role: "system", content: "Summarize in two sentences." }, { role: "user", content: event.payload.text }, ], temperature: 0, max_tokens: 400, })) as { text: string }; await ctx.vault.publish("document.summarized", { summary: out.text });}| Method | Behavior |
|---|---|
infer(input) |
Sends input to the model and resolves with the model’s output. The input shape is the model’s own, as declared in its model.json. |
stream(input) |
Returns an async iterable. In the agent sandbox it yields the complete output once. It does not stream tokens. |
To pass a vault object to a model that fetches it, mint a short-lived URL with
ctx.vault.objects.presignedUrl(path) and send the URL.
Choose the model per deployment. Declare a modelAlias param so a
deployment can switch models without a rebuild:
models: { small: "https://cdn.ddc-dragon.com/<bucket>/models/<small>/<version>/model.json", large: "https://cdn.ddc-dragon.com/<bucket>/models/<large>/<version>/model.json",},params: { model: { type: "modelAlias", default: "small", enum: ["small", "large"] },},const alias = String(ctx.params.model);const out = await ctx.models[alias].infer({ prompt: event.payload.text, max_tokens: 400 });enum must list declared aliases. A deployment that sets a value outside
enum falls back to the default. See
Push and deploy.
Test without a cluster. Stub the alias in the test harness:
import { testAgent, createModelMock } from "@cef-ai/testing";
const h = testAgent(Summarizer, { models: { llm: { infer: async () => "A short summary.", stream: async function* () {} } },});createModelMock() returns a handle with .expect(input).respond(output) and a
.calls array for assertions. See the testing reference.
Write the input
The input is the model’s own schema; read it on the model’s page in ROC, or
rely on the types cef typegen generates. An ASR model, for example, takes an
audio URL.
The platform’s language models take this input and return { text }, plus
tool_calls when the model calls a tool and usage when the model reports it:
| Field | Default | Meaning |
|---|---|---|
messages |
— | Chat messages { role, content }. Either messages or prompt is required. |
prompt |
— | A single user prompt. |
max_tokens |
256 |
Output budget. Set it explicitly: long output, JSON in particular, is cut off at the budget. |
temperature |
0.7 |
Sampling temperature. |
top_p, top_k, stop |
— | Standard sampling controls. |
frequency_penalty, presence_penalty, repetition_penalty |
— | Repetition controls. |
response_format |
— | Constrain output: "json", { type: "json_object" }, or { type: "json_schema", schema }. See Structured output. |
tools |
— | Tool declarations, rendered into the prompt by the model’s chat template. |
image |
— | Image URL, for vision models. |
Metering
Every inference call is metered against the run that made it. Each call records its duration and, when the model reports them, input tokens, output tokens, and GPU units on the Task. GPU units count toward the connection’s Compute limit, which the vault owner can set; once it is reached, new work stops. See Spend limits.
Failures
Inference crosses the network. A model can be pending, busy, or slow.
- In a workflow, a model call that fails in transport (a 5xx, a timeout, a dropped connection) is retried up to 3 attempts in total, 0.4 s and then 1.6 s apart. A 4xx fails the step at once. Treat a failed model step as a normal failure path.
- In a code agent, a failed call rejects
infer. Catch it and degrade gracefully, and write a model’s output to a cubby as soon as you have it, so a retried Task does not pay for the same call twice.
In tests, createModelMock() from @cef-ai/testing scripts a model’s answers
for both. See Test a workflow and
Test and debug.
Limits
| Limit | Value |
|---|---|
| Model call attempts in a workflow | 3, transport failures only, 400 ms then 1600 ms apart |