Results
A workflow’s Result is the set of fields its Result step declares. Declared, the run’s final output is exactly those fields; undeclared, the run ends with the whole carried item (the trigger payload plus everything each step merged onto it). Declare a Result on every workflow you evaluate: two runs can only be compared on what they promised.
The Result is also what a webhook caller waiting on the run receives, and what a dataset’s cases are matched against.
Declare a Result
The Result step is the output step kind. It takes three params.
| Param | Type | Meaning |
|---|---|---|
result |
{ field: value } (or that object as JSON text) |
The declared fields. Each value is resolved against the item like any mapped param ("={{ $json.category }}") or is a literal. At least one field. |
resultTypes |
{ field: "string" | "number" | "boolean" | "object" | "array" } |
Optional. The type each field must resolve to. Every key must be a field of result. |
outcome |
string | Optional. A label for how the run ended on this path ("escalated"), resolved like any param. Sent on workflow.completed beside the output. |
{ id: "result", kind: "output", label: "Result", position: { x: 1200, y: 0 }, params: { result: { category: "={{ $json.category }}", confidence: "={{ $json.confidence }}", model: "classifier-v2", }, resultTypes: { category: "string", confidence: "number" }, outcome: "={{ $json.category }}", },}The run ends with:
{ "runId": "run-t-1", "output": { "category": "billing", "confidence": 0.93, "model": "classifier-v2" }, "outcome": "billing" }In the Workflow Builder, add a Result step. In its Output pane, add each field with its type (leave a field’s type as any to skip the check). Under Where each field comes from, write one mapping per line (category = ={{ $json.category }}); a field you do not list is read from the item by its own name. Outcome is the optional label.
defineWorkflow checks the declaration when your config loads (at cef build): result that is not an object, a result with no fields, an unknown type in resultTypes, a type for a field result does not declare, and an outcome that is not a string are all refused before anything deploys.
What a wrong type does
A field that resolves to a type other than its resultTypes entry fails the run at the Result step:
workflow.step_failednames the step:result field 'confidence' expected number, got string;workflow.failedcarriesresult: result field 'confidence' expected number, got string;- no
workflow.completedis published, and in an experiment the run scores as failed.
A typed field that resolves to nothing fails the same way (got nothing); so does null (got null). An untyped field that resolves to nothing is left out of the output rather than set to null.
This is deliberate. A Result that silently changed type is the regression an evaluation exists to catch, and failing at the step names the cause instead of leaving a mismatch for a downstream consumer to find. If a model step can return the wrong type, normalise it in a transform step before the Result (see Best practices).
Several Result steps
A graph with several branches can end in several Result steps, each with its own outcome. The declared Result of the version is the union of their fields. A field two Result steps type differently is treated as untyped for scoring.
Matching
Each key of a case’s expected is one field score, addressed by dot path (reply.language).
| Expected value | Passes when |
|---|---|
| A plain value | The actual value is deeply equal |
| A plain object | Every key it names matches (a subset match; other keys are ignored) |
A string such as ">= 0.8" |
The actual value is a number inside the band (>=, <=, >, <). Against a string it is plain equality. |
{ "$eq": v } |
Deep equality |
{ "$oneOf": [a, b] } |
Equal to one of the options |
{ "$contains": v } |
A string containing v, an array with an item matching v, or an object matching v |
{ "$regex": "^T-\\d+$" } |
A string matching the pattern |
{ "$between": [min, max] } |
A number with min ≤ value ≤ max |
{ "$gte": n }, { "$lte": n } |
A number at or above / at or below n |
{ "$approx": n, "$tol": t } |
A number within t of n |
{ "$exists": true } |
Present and not null (false: absent or null) |
A matcher naming several operators needs all of them ({ "$exists": true, "$gte": 1 }). A field the output does not have fails.
A case passes when every judged field passes. It is unscored when it has no expectations. A run that failed or timed out fails.
$regex patterns run synchronously during scoring. Scoring refuses a pattern longer than 256 characters or one that repeats a group which itself contains a quantifier ((a+)+), failing that field with unsafe $regex refused: …. This is a heuristic; keep patterns simple.
Match free text loosely. A model rewording a sentence is not a regression, so expect { "$exists": true } or { "$contains": "refund" } for prose and exact values for labels, ids, and numbers with a band.
Limits
Limits (maxDurationMs, maxTokens, maxSteps, maxCost) are set on the dataset and overridden per case. Each one in force produces a limit score. A breach is counted in the experiment’s limit breaches and never fails a case. A metric the platform did not measure is not judged.
When the Result changes
An experiment records the Result its workflow version declares. Expectations that version cannot satisfy are scored n/a, not failed:
| Situation | Score detail |
|---|---|
| The case expects a field the version does not declare | not in this version's Result |
The expectation cannot match the declared type ({ "$gte": 1 } against a string) |
type changed: expected number, Result declares string |
The dataset is older than the version, which is not a regression. ROC’s Run experiment dialog and cef eval run both warn about this before the experiment starts. Update the cases (or the Result) and the n/a scores go away. When a version declares no Result at all, every expected field is judged against the whole output.