Best practices
Each rule says what to do, why, and how to check it. Most checks are a line in your local tests; see Test a workflow for the harness.
Graph shape
Prefer a linear graph. Branch only where the work really differs. Why: every branch doubles the paths you have to test and the runs you have to read. Check: one local test per path; if you cannot name each path’s test, the graph has too many.
Give every step a position.
Why: the Workflow Builder places a step that has one exactly there; without it the builder derives a layout, and a graph authored in code becomes hard to read next to its runs.
Check: for (const n of graph.nodes) expect(n.position).toBeDefined().
Gate branches on fields that exist before the branch, not on fields only a transform expression creates.
Why: a transform’s output is whatever its expression returns; the builder cannot see those fields, so it cannot show or check the gate, and a typo in the field name sends every run down the default edge.
Check: a local test per gate value, asserting which step ran.
Join fan-out explicitly. When several branches must all finish before a step, put a join step there.
Why: a step with several inbound edges cannot tell branches that arrive together from alternative arms of a branch; only a join waits for all of them.
Check: a test that the step after the join runs once, with every branch’s output merged.
Bound every loop. A loop edge declares max (1 to 100) and the counter field; set exhausted to the step that handles running out.
Why: without exhausted, a spent loop fails the run.
Check: a test whose mock never succeeds, asserting the run reaches the exhausted step.
Step budgets
Set a budget for steps, model calls, and agent calls per run, and assert it.
Why: each extra call adds latency and cost to every run, and a refactor that adds one is easy to miss in review.
Check: expect(graph.nodes.length).toBeLessThanOrEqual(N); count createModelMock().calls and each mock agent’s calls per run. In experiments, set the dataset’s Max steps and Max tokens so drift shows as limit breaches.
Payload budgets
Keep the carried item small. Recommended: at most 256 KB at any step. Pass references (a cubby row id, an object path) instead of embedding documents.
Why: every step’s input and output is carried forward and recorded for the run. Hard limits apply further along: a published event body over 1 MiB is refused, and a step that waits (an agent, a person) cannot hold more than 1 MiB of input. See Limits.
Check: in a test with the largest realistic input, assert Buffer.byteLength(JSON.stringify(x)) on the start payload, the Result, and each published payload.
Declare the payload of every publish step. Set params.payload to an object whose keys are the fields to send.
Why: without it the step broadcasts the whole carried item: every intermediate result, the raw input, and anything a model returned.
Check: expect(Object.keys(published).sort()).toEqual([...]) for each published event.
Cap what you store. Truncate free text before it goes into a cubby column, and assert the cap. Why: a model answer or a pasted document can be arbitrarily long, and a widget that reads the row pays for it. Check: a test with oversized input asserting the stored column length.
Cubby migrations
Never edit a migration that has shipped. Add a new numbered file instead.
Why: the platform tracks applied migrations by number only. A cubby that already applied 002 never runs your edited 002, so the vaults that matter keep the old schema while fresh test vaults get the new one.
Check: in review, the diff of cubbies/ only adds files.
Create new tables only in new migration files, and prefix their names with the workflow or feature (triage_results, not results).
Why: in each vault, a cubby alias is one database shared by every agent and workflow in the Agent Service that names it. An unprefixed name collides with another workflow’s table, or with one a step created at run time.
Check: grep -h "CREATE TABLE" cubbies/**/*.sql shows only prefixed names.
Bind values; never put run data in SQL text. Use ? in sql and the values in args. {{ $runId }} is the only template allowed in SQL.
Why: run data comes from outside; interpolating it is SQL injection. The runner fails a Cubby step whose SQL contains {{ $json.… }}.
Check: the runner enforces it; a local test of the step proves the binding works.
Key every write on the run. Use run_id as (part of) the primary key and INSERT OR IGNORE or an upsert.
Why: delivery is at-least-once; a retried step must not write a second row.
Check: the “same start twice” test in Test a workflow.
Prompts and data
Keep business-specific text out of prompts. Put rules, labels, thresholds, and per-customer wording in data: a cubby table, the deployment’s params, or the start payload. Why: a prompt is code; changing it means a new version and a new experiment. Data changes without a deploy, and the same workflow serves every customer. Check: grep your prompts for customer names, product names, and numbers; each one is a candidate for data.
Use synthetic data in prompts, fixtures, and examples. Never real customer content.
Model output
Treat every model answer as untrusted input. Validate it in a transform or code step right after the model step.
Why: a model can return a label you did not list, a string where you wanted a number, or nothing.
Downgrade invalid output deterministically, or park the run for a person. Map an unknown label to a fallback ("other", confidence 0), or route the run to a human step. Do not loop a model until it complies.
Why: a deterministic fallback keeps runs reproducible and the Result well-typed; a repair loop multiplies cost and still fails sometimes.
Check: a local test that feeds garbage through createModelMock() and asserts the Result.
Declare a typed Result. resultTypes fails the run at the Result step when a field changes type. See Results.
Idempotency
Expect every event to arrive more than once. Delivery is at-least-once.
| Situation | What the platform does | What you do |
|---|---|---|
A workflow.start repeated in the same context |
One run per context: the repeat starts nothing | Use a stable context per logical request |
A webhook call retried with the same Idempotency-Key |
Same event id and context, so the same run; a waiting retry gets the recorded result once the run ended | Always send an Idempotency-Key (at most 255 bytes) from callers that retry |
A webhook call without Idempotency-Key |
Every call starts a new run | |
| A step retried after a transient failure | The step runs again | Key cubby writes on run_id; make external actions idempotent |
| Re-run in ROC | Starts a new run from the top: real agents, real tokens, a new Job | Use it to reproduce, not to “resume” |
Check: publish the same start twice in a local test and assert one row and one model call.
Versioning
Bump the version on every change you push, and evaluate before you switch the live version.
Why: an experiment pins a version that is not live for its own runs only, so you can judge a candidate on real models without exposing it.
Check: run the dataset with cef eval run <dataset> --workflow-version <candidate> and read the verdict against the baseline.
Move the baseline when the live version changes. Make the live version’s experiment the baseline, so the next candidate is judged against what users actually get.
Update the dataset in the same change as the Result. A field you rename or retype scores n/a until the cases follow.
Fixtures
Build fixtures from small factory functions with synthetic values. One base case, and overrides per test (triggerFor({ category: "bug" })).
Why: each test then states only what makes it different, and a reviewer can see why it exists.
Make each negative fixture fail for one reason. A fixture that is wrong in two ways cannot tell you which check caught it. Check: remove the guard the test is for and watch the test go red. A test that stays green without its guard tests nothing.