Skip to content

Results

A workflow’s Result is the set of fields its Result step declares. Declared, the run’s final output is exactly those fields; undeclared, the run ends with the whole carried item (the trigger payload plus everything each step merged onto it). Declare a Result on every workflow you evaluate: two runs can only be compared on what they promised.

The Result is also what a webhook caller waiting on the run receives, and what a dataset’s cases are matched against.

Declare a Result

The Result step is the output step kind. It takes three params.

Param Type Meaning
result { field: value } (or that object as JSON text) The declared fields. Each value is resolved against the item like any mapped param ("={{ $json.category }}") or is a literal. At least one field.
resultTypes { field: "string" | "number" | "boolean" | "object" | "array" } Optional. The type each field must resolve to. Every key must be a field of result.
outcome string Optional. A label for how the run ended on this path ("escalated"), resolved like any param. Sent on workflow.completed beside the output.
{
id: "result",
kind: "output",
label: "Result",
position: { x: 1200, y: 0 },
params: {
result: {
category: "={{ $json.category }}",
confidence: "={{ $json.confidence }}",
model: "classifier-v2",
},
resultTypes: { category: "string", confidence: "number" },
outcome: "={{ $json.category }}",
},
}

The run ends with:

{ "runId": "run-t-1", "output": { "category": "billing", "confidence": 0.93, "model": "classifier-v2" }, "outcome": "billing" }

In the Workflow Builder, add a Result step. In its Output pane, add each field with its type (leave a field’s type as any to skip the check). Under Where each field comes from, write one mapping per line (category = ={{ $json.category }}); a field you do not list is read from the item by its own name. Outcome is the optional label.

Result step editor with typed fields, sources and outcome defineWorkflow checks the declaration when your config loads (at cef build): result that is not an object, a result with no fields, an unknown type in resultTypes, a type for a field result does not declare, and an outcome that is not a string are all refused before anything deploys.

What a wrong type does

A field that resolves to a type other than its resultTypes entry fails the run at the Result step:

  • workflow.step_failed names the step: result field 'confidence' expected number, got string;
  • workflow.failed carries result: result field 'confidence' expected number, got string;
  • no workflow.completed is published, and in an experiment the run scores as failed.

A typed field that resolves to nothing fails the same way (got nothing); so does null (got null). An untyped field that resolves to nothing is left out of the output rather than set to null.

This is deliberate. A Result that silently changed type is the regression an evaluation exists to catch, and failing at the step names the cause instead of leaving a mismatch for a downstream consumer to find. If a model step can return the wrong type, normalise it in a transform step before the Result (see Best practices).

Several Result steps

A graph with several branches can end in several Result steps, each with its own outcome. The declared Result of the version is the union of their fields. A field two Result steps type differently is treated as untyped for scoring.

Matching

Each key of a case’s expected is one field score, addressed by dot path (reply.language).

Expected value Passes when
A plain value The actual value is deeply equal
A plain object Every key it names matches (a subset match; other keys are ignored)
A string such as ">= 0.8" The actual value is a number inside the band (>=, <=, >, <). Against a string it is plain equality.
{ "$eq": v } Deep equality
{ "$oneOf": [a, b] } Equal to one of the options
{ "$contains": v } A string containing v, an array with an item matching v, or an object matching v
{ "$regex": "^T-\\d+$" } A string matching the pattern
{ "$between": [min, max] } A number with min ≤ value ≤ max
{ "$gte": n }, { "$lte": n } A number at or above / at or below n
{ "$approx": n, "$tol": t } A number within t of n
{ "$exists": true } Present and not null (false: absent or null)

A matcher naming several operators needs all of them ({ "$exists": true, "$gte": 1 }). A field the output does not have fails.

A case passes when every judged field passes. It is unscored when it has no expectations. A run that failed or timed out fails.

$regex patterns run synchronously during scoring. Scoring refuses a pattern longer than 256 characters or one that repeats a group which itself contains a quantifier ((a+)+), failing that field with unsafe $regex refused: …. This is a heuristic; keep patterns simple.

Match free text loosely. A model rewording a sentence is not a regression, so expect { "$exists": true } or { "$contains": "refund" } for prose and exact values for labels, ids, and numbers with a band.

Limits

Limits (maxDurationMs, maxTokens, maxSteps, maxCost) are set on the dataset and overridden per case. Each one in force produces a limit score. A breach is counted in the experiment’s limit breaches and never fails a case. A metric the platform did not measure is not judged.

When the Result changes

An experiment records the Result its workflow version declares. Expectations that version cannot satisfy are scored n/a, not failed:

Situation Score detail
The case expects a field the version does not declare not in this version's Result
The expectation cannot match the declared type ({ "$gte": 1 } against a string) type changed: expected number, Result declares string

The dataset is older than the version, which is not a regression. ROC’s Run experiment dialog and cef eval run both warn about this before the experiment starts. Update the cases (or the Result) and the n/a scores go away. When a version declares no Result at all, every expected field is judged against the whole output.