Skip to content

@cef-ai/eval

@cef-ai/eval is the library ROC’s Evaluations tab and cef eval are built on. Use it to script datasets and experiments, or to build your own evaluation tooling over the same storage.

Terminal window
pnpm add @cef-ai/eval
Version 0.1.0
Dependencies None. Browser-safe: no Node built-ins.
Module ESM

The library never talks to the network itself. You inject two things: an EvalStore for storage and a WorkflowTarget for starting runs. For concepts, see the Evaluations overview.

Types

Case

interface Case {
id: string; // [A-Za-z0-9._-]{1,80}
input: Record<string, Json>; // the workflow.start payload
expected?: Record<string, Json | Matcher>; // keyed by Result field; dot paths allowed
limits?: Limits;
tags?: string[];
notes?: string;
source?: { context: string; workflowVersion?: string; capturedAt: string };
}
interface Limits { maxDurationMs?: number; maxTokens?: number; maxSteps?: number; maxCost?: number }
type Matcher =
| { $eq: Json } | { $oneOf: Json[] } | { $contains: Json } | { $regex: string }
| { $between: [number, number] } | { $gte: number } | { $lte: number }
| { $exists: boolean } | { $approx: number; $tol: number };

Matching rules: Results.

Header and VersionManifest

interface Header {
dataset: string; workflowId: string; description?: string; limits?: Limits;
createdAt: string; baselineId?: string;
}
interface VersionManifest {
version: number; header: Header; createdAt: string; note?: string;
cases: Array<{ id: string; rev: number; sha256: string }>;
}

Experiment and RunRecord

interface Experiment {
id: string; // exp-<base36 time>-<4 chars>
dataset: string; datasetVersion: number;
workflowId: string; workflowVersion: string; live?: boolean;
repeats: number; resultFields?: ResultField[]; attempt?: number;
startedAt: string; finishedAt?: string;
status: "running" | "done" | "cancelled" | "error";
origin: "roc" | "cli"; createdBy?: string; baselineId?: string; note?: string; error?: string;
summary?: Summary;
}
interface RunRecord {
caseId: string; caseRev: number; repeat: number; attempt?: number;
context: string; runId: string;
status: "done" | "failed" | "timeout" | "error";
output: Json | null; outcome?: string; error?: string;
startedAt: string; endedAt: string | null;
metrics: Metrics; scores: Score[];
pass: boolean | null; // null: no expectations, or the harness errored
}
interface Metrics {
durationMs: number | null; inputTokens: number; outputTokens: number; totalTokens: number;
cost: number; steps: number; modelCalls: number; cpuUnits: number; gpuUnits: number;
}
interface Score {
name: string; // field path, or limit.<name>
value: number | string | boolean | null;
pass: boolean | null; // null: not judged
scorer: string; // "eval/field@1" | "eval/limit@1"
source: "eval" | "annotation";
detail?: string;
}

RunRecord.status: done the workflow completed; failed it announced a failure; timeout it did not end in time; error the harness could not run it.

Summary

interface Summary {
cases: number; runs: number; passed: number; failed: number; errored: number; unscored: number;
passRate: number | null; // passed / (passed + failed)
durationMs: { p50: number | null; p95: number | null; mean: number | null };
tokens: { total: number; meanPerRun: number };
cost: { total: number; meanPerRun: number };
steps: { meanPerRun: number };
limitBreaches: number;
fields: Record<string, { passed: number; failed: number; notApplicable: number; rate: number | null }>;
}

EvalStore

interface EvalStore {
list(prefix: string): Promise<string[]>; // every key under prefix, sorted
get(key: string): Promise<string | null>; // null when absent; other failures throw
put(key: string, body: string): Promise<void>;
putIfAbsent(key: string, body: string): Promise<boolean>; // true when this call wrote
readonly atomicity?: "conditional-write" | "single-threaded";
listPrefixes?(prefix: string): Promise<string[]>; // optional: immediate child "directories"
}
Export Meaning
memoryStore(seed?) An in-memory store for tests and dry runs, with dump().
storeConformance(makeStore) A list of { name, run } checks to run against your own implementation (for example, one per it). putIfAbsent must be a real compare-and-set.

Datasets and cases

Every dataset function takes (store, workflowId, dataset, …).

Function Returns Meaning
createDataset(store, { dataset, workflowId, description?, limits? }) Header Refuses a name the workflow already has.
getHeader / requireHeader Header | null / Header
updateHeader(store, header) Header
setBaseline(store, wf, dataset, expId | null) Header Set or clear the baseline.
listWorkflows(store) string[] Workflows with anything stored.
listDatasets(store, wf) DatasetSummary[] { dataset, cases, latestVersion, experiments, baselineId?, … }
addCase(store, wf, dataset, c) { rev } Refuses an id that exists.
saveCase(store, wf, dataset, c) { rev, changed, created } First revision, next revision, or nothing when identical.
getCase(store, wf, dataset, id) { case, rev } | null Latest revision of a live case.
getCaseRevision, listCaseRevisions One revision; all revision numbers.
listCaseIds, listArchivedCaseIds, listDeletedCaseIds string[]
archiveCase / unarchiveCase Hide from new versions; bytes untouched.
deleteCase / undeleteCase { pinnedBy } / Tombstone; revisions and the id stay.
loadWorkingSet(store, wf, dataset) { header, cases } Every live case at its latest revision.

Versions

Function Meaning
ensureDatasetVersion(store, wf, dataset) The latest version when the working set matches it; otherwise cuts the next with the note auto: …. Returns { version, created, diff, manifest }.
cutVersion(store, wf, dataset, { note? }) Cut a version explicitly.
listVersions, latestVersionNumber, getVersion Read versions.
loadVersion(store, wf, dataset, version) The cases a version pins, verified by sha256.
versionHistory(store, wf, dataset) Each version with what changed and how many experiments ran it.
diffVersionEntries(from, to) { added, removed, edited }

Cases from runs

caseFromRun(run: CapturedRun): Case
interface CapturedRun {
input: Record<string, Json>; // workflow.start payload; an inline `graph` is dropped
output?: Json | null; // workflow.completed output
context: string;
workflowVersion?: string;
id?: string; // default: derived from the context
capturedAt?: string;
tags?: string[];
notes?: string;
}

expected is built from an object output with expectationFrom: labels exactly, free text (isFreeText: over 60 characters or over 8 words) as { $exists: true }. slugCaseId(text) makes a valid id from any text; uniqueCaseId(base, taken) appends -2, -3, … until it is free.

Running experiments

import { runExperiment, resultFieldsFromGraph } from "@cef-ai/eval";
const exp = await runExperiment({
store,
workflowId: "ticket-triage",
dataset: "triage",
workflowVersion: "1.4.0",
live: false, // not live: target.pinVersion is required
resultFields: resultFieldsFromGraph(graph),
repeats: 3,
target,
onProgress: ({ done, total, run }) => console.log(done, total, run.caseId, run.pass),
});

RunExperimentOptions

Option Default Meaning
store, workflowId, dataset Required.
workflowVersion Required. The version evaluated (recorded only, when live).
live Required. true runs on the live deployment; false needs target.pinVersion.
target Required. A WorkflowTarget.
datasetVersion ensureDatasetVersion Version to run.
repeats 1 Runs per case.
concurrency 4 (DEFAULT_CONCURRENCY) Runs in flight.
timeoutMs 300000 (DEFAULT_TIMEOUT_MS) Per run.
resultFields The version’s declared Result; expectations it rules out score n/a.
baselineId, note, origin, createdBy Recorded on the experiment.
experimentId Resume this experiment: runs on file are reused, a new attempt runs the rest.
signal AbortSignal: stop dispatching; the experiment ends cancelled.
onProgress ({ done, total, run, reused }) after each run.

The meta is written first (running), each run record as it finishes, the summary last. Each run publishes into its own context, eval-<expId>-<caseId>-<repeat>-a<attempt>; runOne runs and scores a single case without storage.

WorkflowTarget

interface WorkflowTarget {
start(input: Record<string, Json>, context: string): Promise<void>;
waitForEnd(context: string, timeoutMs: number): Promise<RunEnd>;
metrics(context: string): Promise<Metrics>;
pinVersion?(version: string, contextPrefix: string): Promise<() => Promise<void>>;
}
interface RunEnd {
status: "done" | "failed" | "timeout";
output?: Json; outcome?: string; error?: string;
startedAt?: string; endedAt?: string; // server timestamps
}

start publishes workflow.start with the input into the context; waitForEnd waits for workflow.completed or workflow.failed; metrics reads what the run cost; pinVersion routes contexts starting with the prefix to a version until the returned release function is called.

Reading experiments

Function Meaning
getExperiment(store, wf, expId) The experiment, or null.
listExperiments(store, wf, { dataset? }) All experiments of a workflow, or of one dataset.
loadExperiment(store, wf, expId) { experiment, runs }
listRunRecords, getRunRecord Run records.
findExperiment(store, expId) Find an experiment when the workflow is not known.

Scoring

Function Meaning
scoreFields(expected, output, resultFields?) One eval/field@1 score per expected field.
scoreLimits(limits, metrics) One eval/limit@1 score per limit in force.
effectiveLimits(header, case) The header’s limits, each overridden by the case’s.
casePass(scores) True when every judged field passes; null when none were judged. Limits never count.
matchValue(expected, actual) null on a match, otherwise the reason.
readField(value, path) Read a dot path.
resultFieldsFromGraph(graph) The fields a graph’s Result steps declare (accepts the graph, its JSON text, or a manifest’s params.graph). undefined when none declares a Result.
checkResultCompatibility(resultFields, cases) { missing, typeChanged, unexpectedNew }: which expectations a version cannot satisfy.
unsafeRegex(pattern) The reason a $regex pattern is refused, or null. MAX_REGEX_LENGTH is 256.

Summaries and comparison

Function Meaning
summarize(runs) A Summary over run records.
compareExperiments(base, candidate) Each side is { experiment?, runs, cases? }. Compares only cases present on both sides at the same revision.

ExperimentComparison:

Field Meaning
base, candidate { id, workflowVersion?, datasetVersion? }
comparedOn Number of cases compared
datasetDiff { added, removed, edited } when both sides’ cases are given
counts Runs per change: improved, regressed, unchanged, only-in-base, only-in-candidate
cases, runs Per case and per run: the change, pass rates or statuses, and deltas
metrics { base, candidate, delta, relative } for pass rate, duration p50 and p95, tokens/run, cost/run, steps/run, limit breaches

A pass ranks above “not judged”, which ranks above a failure; a case’s change is decided by its mean rank over repeats.

Storage layout helpers

Key builders (headerKey, caseRevKey, versionKey, experimentMetaKey, runKey, …) and validators (isWorkflowId, isDatasetName, isCaseId, isExperimentId and their assert… forms) expose the layout under workflows/<wf>/. canonicalJson and sha256Hex are how versions pin case content. parseCase, parseHeader, parseExperiment, parseRunRecord validate stored JSON and throw EvalContractError.