Files
windmill/docs/ai-agent-evals.md
T
hugocasa 9c557859c5 feat: AI agent evals: datasets, scored runs and comparison (#10633)
* feat: eval datasets and standalone runs for reusable AI agents

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: agent eval drawer with case editor, runs and capture entry points

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: document AI agent eval datasets and standalone runs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: say how many eval cases the list is not showing

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address review findings on eval datasets

- keep an edited case's conversation and tool inputs: serde(flatten) silently
  drops Box<RawValue> fields, so the update payload is spelled out
- remount the case editor per case so one case's turns cannot leak into another
- require jobs:read / flow_conversations:read on the capture endpoints, which
  UserDB does not gate by token scope
- take the dataset lock in create and update so a delete cannot be undone by a
  concurrent metadata write, and delete cases before metadata
- load more cases beyond the first page, and stop capping the agent picker
- record that the version stamp is taken at enqueue, not at resolution

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-2 review findings on eval datasets

- block operators from dataset and case writes
- pass the editor's operating workspace through the drawer and the capture
  request, instead of assuming the navigation workspace
- discard superseded case-list responses so switching datasets cannot land the
  previous dataset's cases
- reject a dataset without a case_id (or vice versa) rather than running an
  inline case under a dangling association
- run unsaved edits inline instead of silently running the stored case
- surface the API error body on a failed run
- fetch dataset metadata concurrently when listing
- $bindable() without a default on the optional open prop
- correct the permission and enqueue-time-version wording in the docs

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: run an untouched saved case by reference again

The editor writes back keys the stored case omits, so comparing the raw objects
reported every unedited case as edited: the run went inline and lost the
dataset/case stamp its history depends on. Compare a normalized form, and pin it
with a test. Also scope the history query to the drawer's workspace and drop
superseded responses.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: show a dataset's cases as a table, and fix round-4 review findings

The case list showed one case at a time with no overview. It is now a table with
the case, where it was captured from, and its last run — the last-run column is a
single jobs query on the path stamp rather than a request per row.

Review fixes in the same file:
- keep the edit baseline on the selected case rather than looking it up in the
  loaded page, so a case beyond page 1 is not treated as unedited and run stale
- release the loading state when a superseded case load returns early
- reload every loaded page after a write instead of collapsing to page 1
- last remaining 'resolved to' wording in the version tooltip

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: run a dataset as an experiment, with scorers as runnables

An experiment runs every case of a dataset against one subject and records the
exact case set it executed, so a result set stays reproducible while the dataset
keeps changing.

Each case runs as its own small flow — the agent, then a step per scorer — so a
case keeps the run stamp, history query and trajectory view a single run already
has, and scorers need no orchestration of their own. Results are read back per
step by node id rather than by walking a nested loop's status.

A scorer is any runnable taking (input, output, expected): a script, a flow, or a
reusable agent used as a judge. A judge is prompted with the case and the answer
as one JSON message; a script or flow receives them as named arguments. Scores
accept a bare number, a boolean or {score}.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: results table for an experiment, with scorer columns

One row per case: status, the agent's answer, and a column per scorer, with the
mean per scorer above the table and a link into each case's run for its
trajectory. Averages skip cases a scorer produced no number for — counting a
missing score as zero would read as a regression.

The drawer's left pane becomes Cases / Results, and Results carries the scorer
picker and Run dataset.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* feat: compare an experiment against a baseline

Per-scorer deltas on each row and on the mean, and a filter down to the rows that
regressed. Rows join by case id, so a case added after the baseline ran has no
delta instead of counting as a change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-5 review findings on experiments

- match scorers by label when diffing two experiments; joining by array position
  subtracted one scorer from another whenever the scorer sets differed
- report a row's status from the case job, not the agent step, so a case whose
  scorer failed no longer reads as a success
- delete a dataset's experiments with it: they hold copies of its cases, and a
  recreated dataset of the same path would have exposed them
- select the experiment that Run dataset just started instead of leaving the
  table on the previous one
- expected is scored now, so stop describing it as having no consumer

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-6 review findings on experiments

- hold the dataset lock across an experiment launch, so a delete landing between
  reading the cases and writing the experiment cannot recreate the deleted
  dataset's inputs
- match scorers between experiments on kind and path, not on label: labels
  default to a path's last segment, so f/a/quality and f/b/quality compared
  against each other
- average mean deltas over the cases both runs scored; comparing each run's own
  average reported a regression from a case the baseline never ran, with no
  regressed row to point at
- openapi: the row status is the job's, which is also canceled/skipped; runEval
  takes scorers; the update-case body no longer advertises source, which the
  handler deliberately ignores
- record why the experiment prefix cannot reach a sibling dataset

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-7 review findings on experiments

- release the dataset lock for the push loop and retake it for the write,
  re-checking the dataset still exists: holding it across the whole launch made
  every capture and case edit on that dataset 409 until the last job queued
- assemble experiment results with bounded concurrency; a 100-case, 3-scorer
  experiment was 400 sequential lookups, each itself several queries
- clear the baseline when it becomes the selected experiment, which was
  comparing a run against itself and reporting zero deltas
- take the header mean over the same cases as its delta while comparing, so the
  two numbers beside each other describe the same set
- a canceled or skipped case is no longer the same grey dot as a running one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-8 review findings on experiments

- verify the dataset's identity, not just its existence, before recording an
  experiment: the path can be deleted and recreated during the push loop, and
  the experiment holds copies of the old dataset's cases
- give the recording lock a longer budget than a case edit, since its jobs are
  already queued and giving up strands them, and say so when it fails
- keep score lookups sequential within a case: nesting two bounded streams
  multiplied into 32 in-flight queries against a 50-connection pool
- clear a baseline that no longer belongs to the loaded experiments, so
  switching datasets does not leave comparison mode on with nothing to compare
- keep a scorer's own mean when the baseline never ran it, instead of blanking a
  column full of numbers
- EvalCaseDraft.expected no longer claims nothing scores it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: do not trust an experiment's job ids, and require write to record one

Experiment objects live in workspace object storage, which a script can write
directly, and results are read on the unrestricted pool — so a forged experiment
naming another flow job returned output the jobs API would have refused. Only
jobs this server stamped with that experiment's id are read now.

Also from round 9:
- recording an experiment requires write on the dataset, not read: it persists
  into the dataset's namespace and its shared list
- clear the results table when the selection changes and surface a failed load,
  instead of labelling the previous experiment's numbers as the new one's
- a storage fault is no longer reported as a deleted dataset
- the lock-timeout message at the recording site no longer says to retry, which
  would run the whole dataset again on top of the jobs already queued
- ExperimentRow.status documents canceled and skipped

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: bind the experiment trust check to the requested dataset

The previous check matched jobs on the experiment id alone, which the stored
object supplies — so copying another dataset's experiment JSON under a readable
key carried its jobs' output along with it. A job is now only read if it was
stamped for this experiment *and* for the dataset the caller's read access was
checked against, and an experiment that names a different dataset is not served
from this key at all.

Also from round 10:
- add the .sqlx entry for that query; without it every SQLX_OFFLINE build failed
- serve results over GET: as POST the route-scope middleware classified a read
  as ai_evals:write, locking read-only tokens out of their own results
- clear the selected and baseline experiments synchronously when the dataset
  changes, so the previous dataset's id is not requested under the new one

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-11 review findings on experiments and scorers

- give scorers the whole case input, not just the message: an answer that came
  from attachments or a replayed conversation could not be judged on it
- accept a judge's boolean and structured {score} answers, including stringified
  ones, and pin every documented scorer shape with a test
- record an experiment for the cases that did launch when a later push fails,
  instead of leaving those jobs running with nothing to attribute them to
- do not capture a preview parent's synthetic runnable_path as a host flow; the
  saved case could not be rerun
- clear the case table before loading a dataset and surface a failed load, so a
  failure cannot leave the previous dataset's cases under the new name
- keep the results table through a refresh of the same experiment
- exclude flow-step jobs from the per-case last-run lookup
- drop case sets from the experiment list, which is only used to pick a run
- report a database failure at the recording lock as itself, not as contention

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-12 review findings on capture and run history

- load flow_node.flow for flownode parents: an agent inside a deployed branch or
  loop captured without its agent, host flow or tool bindings
- decide host_flow_path by whether the path resolves to a flow, not by job kind:
  excluding previews wholesale also dropped the flow editor's step test, whose
  path is real
- page the per-case last-run lookup by created_before until the loaded cases are
  covered; one page of 200 reported older cases as never run
- do not record an experiment when nothing launched
- only attach the case input to a job when a scorer will read it
- keep the case table through a save; only a different dataset clears it
- drop the superseded duplicate comment on the score parser

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: stop refetching run history on every case write

Reading the case list before the first await made the whole job-history query a
dependency of it, so every save, delete and Load more refetched up to 1000 job
rows and blanked the column. Read untracked instead.

- an empty Last run cell now distinguishes never-ran from not-found-within the
  page bound, which the comment already claimed and the cell did not
- reloading a dataset no longer replaces a populated table with a skeleton
- keep the score-parser comment that describes every shape it handles

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: keep eval datasets in Postgres instead of object storage

Datasets, cases and experiments become rows (`eval_dataset`, `eval_case`,
`eval_experiment`, `eval_experiment_case`) rather than objects under a
`wmill_eval_datasets/` prefix. What a run produced is still the job's:
only case inputs and an experiment's case snapshot are stored.

This removes the machinery the object store needed:

- The advisory lock and the read-modify-write of a per-dataset JSONL. A
  case is a row, so there is nothing to serialize.
- The launch-time identity check on the dataset. The foreign key makes a
  concurrent delete fail the transaction instead.
- The trust guard on an experiment's job ids, which existed because a
  script can write workspace object storage directly and could forge an
  experiment naming somebody else's job.

An experiment now chooses every job id and records itself before pushing
anything, so a launch that dies partway leaves a recorded case whose job
is missing rather than a running job nothing accounts for; cases that
never reached the queue are removed again.

Row-level security on `eval_dataset` is the authority on who may read or
write a dataset, so `extra_perms` grants work and the rule is not
mirrored in Rust. Cases and experiments carry a read policy derived from
their dataset and no write policy: they are written on the unrestricted
pool after the dataset row itself has been asked, with
`SELECT ... FOR UPDATE`, whether the caller may write it.

Cases are capped at 256 KiB each and 10 000 per dataset, refused rather
than truncated. Attachments are S3 references, not inline bytes, so a
case that approaches either cap is a mistake rather than a use case.

Evals no longer need the `parquet` feature or a configured workspace
object storage.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* style: align the eval drawer with the design system

- Scorer chips are `Badge`s rather than a hand-rolled bordered span, and
  the section header is a `Label` with its tooltip, as are the case
  editor's fields (which also gets the label colour right).
- The results table showed status as a coloured bullet, which says
  nothing to a colour-blind reader. It now carries the same icons the
  runs table uses, with the status as its accessible name.
- Feedback colours move to the `-500` shades the brand guidelines name.
- The conversation JSON error uses `TextInput`'s `error` prop for the
  border and the caption style for the message, as elsewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: author an expected answer, tags and attachments on a case

Every scorer is handed `(input, output, expected)`, but nothing could
produce an `expected` except a conversation capture: the case editor had
no field for it and a captured run left it empty. So:

- The editor gains Expected, Tags and a read-only list of the
  attachments a captured case carries. Expected is plain text, or JSON
  when the answer has structure.
- Capturing from an AI agent run keeps what that run answered, which is
  the only moment a reference answer exists for free.

The results table also laid itself out by content, so a long answer
pushed the scores — the numbers the table exists for — off the edge of
the pane. It is fixed-layout now, with the text columns bounded.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: expected is captured from a run and can be authored

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: link a saved agent when inserting an ai agent step

"AI Agent" in the step picker was a leaf that always created a blank
step, so reusing a saved agent meant inserting a blank one, opening its
step input and linking it there. It is a category now, like Flow and AI
Sandbox, listing the workspace's `ai_agent` resources next to a blank
option, filtered by the picker's own search.

A picked agent produces a step that is already linked rather than one
linked afterwards: `agent` set, no tools, and only the flow-local
`user_message`/`user_attachments` transforms. Seeding the brain keys
there would leave transforms a linked step never reads and that
`AgentResourceBar` strips on its next link change.

Each `on:new` forwarder rebuilds the insert detail field by field
instead of spreading it, so a new field is dropped unless the forwarder
names it. `agentPath` is typed on both `GraphEventHandlers.insert` and
`FlowGraphV2`'s `onInsert` so the next one to forget it fails the check.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: restore the link on cancel and simplify the agent bar

Cancel on an agent edit forked the step into a standalone copy, which is
the opposite of what the word means and needed a paragraph under the
card to explain. It discards the edits and re-links the step now,
leaving the agent untouched; diverging from an agent is Unlink's job, on
the linked card. This flow's `tool_inputs` survive the round trip as
overrides, so Cancel no longer folds them into the tools the way Unlink
does.

Linking a step to a saved agent happens in the step picker at insert
time, so the bar's own resource picker is gone and "Save as agent" is
the one action left. Its `+` button was a trap besides: it opened the
generic resource form, where an agent would have to be written as raw
JSON.

The card itself was `surface-secondary`, the sections token, so in dark
mode it was darker than the pane and read as a sunken well rather than
an elevated card. It uses `surface-tertiary` as the brand table
prescribes, its tool chips are `Badge`s, and the editing card no longer
overflows the pane and clips its own buttons. The remaining tooltip
follows the inline `Label` convention rather than sitting in a flex row
whose gap stacked on the trigger's own margin.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: rework the AI agent evals surface into one table

Evals become a single pane: a dataset of cases, one column per scorer, one
row per case, with the run being looked at chosen from the toolbar.

Runs are permanent. Running the whole dataset opens one; running a single
case records nothing at all — it is a job, and looking at what it did is
not a claim that it belongs in the history. Its result and its scores sit
over the row until they are saved as a run, which carries the cases that
were not rerun and the scoring jobs themselves, so the number that is
saved is the number that was looked at.

A scorer is a runnable: a judge agent or a script, created in one click and
edited in place. Scores carry a reason and per-assertion checks, shown on
hover with a rescore button.

What ran is always named. A run records the agent version, or — for a
configuration that is not deployed — a hash of it, so a table can say that
its numbers describe an agent that no longer exists: those rows dim and the
table offers to rerun. An agent's draft can be run directly instead of the
deployed value, and once those edits are deployed the runs that made them
are recognised as that version. A step with no agent of its own is
evaluable too, and saving it as an agent moves its history onto it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: keep an agent's in-progress edits on the agent

Editing a linked agent forks it into the step, which is what makes the
edits runnable there — but the agent is what is being edited, so that is
where the unsaved state belongs. The edit is mirrored into the agent's own
resource draft as it is made.

It then survives leaving the flow, shows the agent as drafted wherever it
appears, and is what evals run when asked to run the draft rather than what
is deployed. Deploying or cancelling clears it; opening Edit without
changing anything does not mark the agent as drafted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: shape the evals surface around a saved agent

Evals hang off an `ai_agent` resource, so the surface is now only ever about
one: the `draft` subject kind, the standalone-step subject and the move that
carried a step's history onto a newly saved agent are gone.

- A run is permanent and numbered per agent. Running a single case is a trial:
  it answers in the panel and never touches the table.
- "Run scorers only" opens a run of its own that reuses the answers of the run
  you are looking at, so a scorer added later measures what already ran without
  calling the agent again.
- A draft run whose configuration is later deployed is stamped, once, to the
  version it became, so its label stops reading `v23 + edits` forever.
- A scorer can carry a pass threshold, read off the scores already recorded.
- The table is the case, its answer and one number per scorer; datasets are
  created and edited in a drawer; a run that executed an earlier state of the
  current draft says so above the table, in one line.
- Which agent a step is, whether it is being edited, and which version it is on
  is a strip above the step's tabs, because it is true of every tab.
- Capturing a case from a step test or a conversation is dropped, and with it
  the `memory` override on a linked step that nothing set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: run past versions of an agent, and number versions per resource

The evals home becomes one table of every run of the agent, whichever dataset
each is of, with one badge per scorer. A list spanning datasets cannot hold every
dataset's scorers to look a name up, so a score carries its name and kind with
its number, and thresholds are joined in per run and column.

Run now asks what to run: the latest agent, resolved when the run executes as a
flow step does, any past version, or the unsaved edits. Pinning is a subject kind
of its own, since a linked step resolves the resource live and inlining is the
only way to run a version that is no longer current.

Scorers move into the edit-dataset drawer. The column header over a run reports
and nothing else: a run is permanent, and a control there that changed the
columns would edit the past from the one place that must not. Adding one offers
four ways rather than two, writing and reusing being different jobs, and both new
kinds open with a summary filled in.

Versions are numbered per resource. `resource_version.id` is one identity
sequence for the whole table, so an agent saved nine times read v4 ... v24, and
the gaps counted writes in workspaces the reader cannot see. The id stays how a
version is addressed; the new number is what it is called, in the resource
history drawer as well as here. It is assigned on write rather than counted on
read because trimming past the cap and clearing a history both take the oldest
rows, and counting the survivors would renumber a version a run already names.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: read the dataset a remembered selection names

Reopening the evals modal restored the last dataset from storage as a bare path,
without reading the row it names. Every "is this already the one?" test compared
against that selection, so all of them short-circuited and the dataset was never
loaded: editing it opened a drawer with no summary, no scorers and no cases.

The remembered path is now brought into context the same way any other choice is,
and the tests compare against the dataset that is loaded rather than the one that
is selected, so a selection can no longer stand for a read that did not happen.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: give dialogs a trail in their header

A dialog deep enough to navigate had nowhere to say where you were: the header
held a fixed title, and the way back was a control each body placed for itself,
somewhere in a toolbar that moves with everything else the toolbar holds. The
header is the one part of the surface that does not move, which is where the
trail belongs.

`Modal` takes an optional `trail` of levels below its title, rendered as a
breadcrumb whose ancestors are the way back. Declarative on purpose: callers of
this depth already hold the state that says where they are, so the dialog reads
it rather than owning a stack they would have to push and pop in step with it.

Escape follows the trail. Leaving a level is what someone deep in a dialog means
by it, and closing the whole surface throws away the navigating they did to get
there; at the root it closes as before. That only works if a dialog can tell it
is the surface being addressed, so `Disposable` now answers `isTopmost()` and the
dialog asks before acting: it keeps Escape for itself, so nothing else was
arbitrating between it and a drawer opened from inside it, and both were acting
on one key press.

Evals is the first caller: its runs list is the root, a run is a level in it, and
the back button that used to sit above the table is gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: portal dialogs out of wherever they were opened from

A dialog rendered in place inherits whatever the calling component happens to sit
inside. One `transform`, `filter` or `overflow` anywhere above it makes its
`fixed` positioning resolve against that ancestor instead of the viewport, and a
surface meant to cover the app is then confined to a box it never asked for: the
nav rail paints over it and its own edges are clipped.

Drawers have always portalled for this reason. Dialogs only did so when an
enclosing pane claimed them, and rendered in place otherwise, so the same screen
could show a drawer over everything and a dialog trapped behind the nav. They now
portal the same way: to the pane when one claims it, to `body` otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: make the dialog's title the first step of its trail

The trail listed levels below the title, so a dialog one level deep read
"Evals > All runs > Run 20 · v6": three steps for two places, the first two of
them the same place under different names. The title is the root, so it is the
root's own segment, and the trail a dialog is given is now the whole path with
that segment at its head.

Its height stopped moving too. A heading carries a line-height of its own, so a
header holding only an h3 stood six pixels shorter than one holding segments as
well, and the dialog's whole top edge stepped as you navigated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: sharpen the evals controls around where you are standing

Each screen now offers what belongs to it. The list starts runs; a run is a
record, so it offers only the one thing that acts on the record itself, which is
measuring the answers it already stored. Starting a fresh run from inside one
asked which agent and which dataset from the screen least about either, and
scoring an existing run was offered from the list, where there is no run to
score. Which run and what it is read against are one question asked twice, so
they sit together rather than at opposite ends of a row.

Choosing what to run is now a toggle over the two states worth naming, the draft
and the saved agent, with every earlier version one click further: running an old
version is deliberate, and a list made all three look alike. The draft is read
when the dialog opens rather than taken from the caller's polled copy, which
could be seconds behind an agent edited a moment ago and would leave the option
out exactly when it is the reason for opening the dialog.

The dataset field carries its path under it and its edit button on hover, as a
resource picker does, so the closed field says what the open list said. Edits
waiting on an agent are a "draft" here as everywhere else in Windmill, rather
than "+ edits". The dialog runs an evaluation rather than "the agent", which is
what it was already called everywhere it is recorded. An agent being edited keeps
its evals button on a line of its own, clear of the decision to save or discard.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: settle the evals controls on the patterns Windmill already has

The version choice uses ToggleButtonMore, as the AI provider picker does: the two
states worth naming stay in the group, the rest are behind the overflow menu, and
the one you pick joins the group rather than appearing in a second control below
it. The deployed one says which version it resolves to.

A run offers nothing to start. Scoring an existing run again was the last thing
left there, and it was one button explaining a distinction that the run and the
dataset already make between them.

The warning that a run executed an earlier draft is about the run on screen, so
it goes when the run does rather than following you back to the list, and it sits
against the table instead of inside a frame of its own.

A dataset just created stays open for its scorers and cases: those are what a
dataset is, they can only be added to one that exists, and closing on create sent
you to find it again to add them. Scorer settings are a cog rather than a word,
now that the row holds three actions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: close the gap in the version toggle and say what naming a dataset does

The overflow trigger is not a pill, so the room it reserves showed as a gap
between it and the button before it; it is pulled in by that much. The dataset
field gets its clear button, which is also the slot the edit button is positioned
against, so the two now sit where a resource picker puts them.

Naming a new dataset said nothing about what happens next, and the drawer looked
like it was missing the rest of itself. It says so instead: a scorer and a case
both belong to a dataset, so there is nothing to attach either to until this one
exists, and creating it leaves the drawer open on them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: choose a dataset's scorers while naming it

A scorer is a reference to a runnable, not a child of the dataset, so it needs
the dataset's name but not its row. The list is collected in the drawer while the
dataset is being named and sent with the create, which already accepts one, so a
dataset arrives holding the columns that were chosen for it rather than being
made empty and then edited to hold them.

Cases stay where they were: a case *is* a row of the dataset, so there is nothing
for it to be a row of until one exists. The drawer says which of the two is which
instead of leaving the screen looking like it is missing the rest of itself.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: level the version toggle and name the dataset in its own field

The overflow trigger stands a row taller than a toggle button, so the group grew
to its height and left the sunken background showing under every pill beside it.
Every child of the group is the same height now, which is why the AI provider
picker never had the band: it sizes them all alike.

The dataset field says the summary with the path after it rather than carrying
the path on a line below. The list stacks the two, which a one-line field cannot
do, so it says both the other way round.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: tidy the evals forms and the run's own controls

Picking a scorer that exists chooses between two sources rather than showing
both: the ones already measuring something, and everything else in the workspace.
The first list says what each is called with its path under it and what it
already measures on the right, instead of three columns that were the same path
truncated three ways whenever a scorer had no name of its own.

A dataset's drawer says what it is for on the page rather than under an icon, and
its summary is sized like the field beneath it.

The run's own row lines up with the table under it, the warning above that table
is spaced off the rule rather than sitting on it, and adding a case is gone from
a run: a run is a record of cases that were answered, so curating them from it is
editing what it measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: create a dataset holding the cases written for it

Creating a dataset takes the cases to create it with, so one can be assembled in
a single act instead of made empty and then filled in. The drawer holds them
while the dataset is being named, gives them ids of its own to be edited by, and
sends them with the create.

Every case is checked before the dataset is written. `eval_case` grants users no
write, so the rows cannot be inserted in the transaction that creates the dataset
under the caller's own policies; validating first is what keeps "created holding
these cases" from becoming "created, holding some of them", and the rows that do
follow go in one transaction of their own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: name the button for what it opens, and say what each version is

Starting an evaluation asks which state of the agent and which dataset, and both
cost a provider bill, so a button that read as spending one on the way past was
lying about the click. It opens something, and says so. Running one case from the
panel keeps its own name and its play icon, because that one does run on click.

The version options say what they are rather than what they are not: what a flow
step would or would not run is a fact about somewhere else, and someone choosing
what to evaluate is not standing in it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: give the editing card two rows and mark evals as beta

At the width of a step panel the card's one row wrapped: the line naming the
agent, the line saying what saving does, and the two buttons deciding the edits'
fate all fought for it. Deciding gets a row of its own, and evals sits against the
line it is about, since evals of an agent being edited run the edits.

Evals is named wherever it is offered. It read as a word in one state of the card
and as an icon in the other, which is two things to recognise for one door.

The dialog carries a beta badge against its own name, before any level below it:
every way in lands there, so it is said once and stays put as you navigate.

The version toggle spells out which is which. Both are the agent at v2 and the
difference between them is the whole choice, so it is worth the width.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: name a new dataset, and lay the scorer's settings out like a step's inputs

A new dataset arrives called "Dataset 1", which the path follows as it follows
any summary: a dataset with none was one every table could only call by its path,
and the two seeds are what the summary rule already produces.

Scorer settings put each field's description between its label and its input,
where a step's inputs put theirs, and its inputs are the size the rest of the
drawer uses. The runnable behind the column is a link to it with its kind's icon,
since it is a resource of its own and the one thing about it these fields cannot
change. The line explaining that a pass line re-reads recorded scores went: the
threshold is a number to set, and how it is applied is not a decision being made
here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: curate a dataset in the drawer and save it in one act

The drawer holds the cases while they are edited and writes them when it is
saved: added, changed and dropped, whichever it is. Typing no longer writes, so a
set is never half saved while someone is still deciding what is in it, and Save
means the same thing whether the dataset exists yet or not.

A case panel offers reading rather than acting. Running one case now and editing
one from a run were the last two ways to change a record from the screen showing
it, and the machinery behind the first went with it. The answer is rendered as
the prose it is, under what it is: the case's result, whichever run is selected
above it.

The rest is what the run's table was doing to its own edges: a column name is
clipped to its column rather than running into the next, the table squares off
against an open panel, and that panel closes with the run it belonged to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: one border above a table, and a link to the run's job

The row above the table drew a bottom border and the table draws its own top
edge, so every table sat under two lines. The row keeps its spacing and the table
keeps its edge.

A column header no longer spins while its scores arrive: the cells under it are
where the numbers are missing, and they say so themselves. The beta badge is the
height of the word beside it rather than of the line it sits on.

A run is one flow and therefore one job, so the run says where that job is: what
it is doing, what it cost and what it logged are all there rather than
reconstructed from the table.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: stream scores as each scorer finishes, and show them per case

A scorer runs after the agent inside the case's own iteration, so its verdict can
be read as soon as its step is done. Waiting for the iteration to end held every
column of a case back until the last of them finished, which is why answers
arrived one at a time and scores all at once.

Reading a job that is still running needs one guard: a module with nothing in it
is a step that has not run, not one that produced nothing, and recording the
second makes a failure that never goes away.

The panel beside the table shows what each column made of the case and why. The
reason a judge gave was stored and never shown, which is the half of a score that
says anything. It stops repeating the question the header already asks, and a
case still running reads as waiting rather than as an answer that says "Running".

A run is a number beside a dataset, so the list puts the two together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: score a case with every scorer at once

The scorers of a case read the answer and never each other, so they ran one after
another for no reason: measuring a case now takes as long as its slowest column
rather than as long as all of them. Each is a branch of its own, kept from
failing the others, so a judge that errors costs its own column and no more.

An iteration is three steps again — answer, payload, scores — rather than one per
scorer, and each branch is named for the column it produces, so the graph of a
run says which scorer did what instead of spelling out an id.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: read a judge's score out of the JSON it nearly wrote

A judge quoting the agent inside its own reason writes those quotes unescaped,
which is invalid JSON and also the most ordinary sentence for it to produce. The
whole verdict was being thrown away over it, so a column that had a number
reported having none.

The number and the reason are now read straight out of such text. Deliberately
not a second JSON parser: it finds the two keys and takes what follows, which is
what survives a quote in the middle of a sentence.

A case still running says so with a spinner rather than with the word "Running"
sitting where its answer goes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: ask a judge for a shape instead of trusting it to write one

A new judge carries an output schema, so the provider holds it to `{score,
reason}` rather than the prompt asking it to. Windmill already delivers a schema
whichever way the model takes it, a tool for Claude and Bedrock and the native
parameter elsewhere, so there is no list of models to keep here.

An agent with no runs offers its first one where the first row would be, rather
than from a toolbar above a table that has nothing in it.

Starting a run no longer picks a dataset for you. It fell back to whichever came
first, which on an agent that has never run means offering another agent's set as
though it were the obvious one; and with no dataset at all it says so and offers
the one move there is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* feat: report a column that failed throughout, and hold the run dialog

The runs overview dropped any column that produced no number, so a judge
that failed on every case of a run vanished from the row and read as a
column nobody had asked for. The aggregate now reports every column that
has cells, with the count of the ones it failed on, and the badge says
"failed" where there is nothing to average. A column with no cells at all
is still left out: that one was added after the run and has nothing to say
about it.

Creating a dataset closes the drawer rather than turning it into an edit
of what it just made: scorers and cases already ship with the create, so
there is nothing left to stay open for. Reached from the run dialog, it
gives the screen back with the new dataset selected, and the dialog keeps
the version you had already chosen.

Also:
- the case panel's job link moves to the panel's own header, where its
  scope is: the job is the whole iteration, not the answer it sat over
- one action in the scorer drawer's header, as its neighbours have. The
  reuse list picks rather than adds, and says which dataset each column
  already measures
- adding a case is the last row of the list it lands in
- the pane shows what it has read rather than an empty state it has not
  earned yet, and its rows say they open
- the linked agent card loses a border it had inside another one

* fix: keep the linked agent card's outline

The card is a thing inside the step's inputs rather than a section of
them, and the outline is what says so. Only the rule inside it goes: the
detail it separates is already set apart by being detail.

* refactor: fit the eval surface to the shipped design

* feat: give a nested dialog a back control and the runs list its own moves

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: put a dialog's description under its title

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: fold a dialog's back control into the crumb it returns to

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: edit a dataset's cases as a table rather than a list beside a form

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: edit a dataset's cases in the grid the data tables are edited in

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: edit a grid cell of prose in place, and cap a dataset at one page

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: keep the cell editor's styles beside it, not in the vendored theme

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: keep an empty cell empty and cap the editor's growth

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: name the step that assembles a run for the scorers

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: run the payload step natively, and say so when nothing serves that tag

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: report an answer as answered while its scorers are still running

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* feat: let a scorer say a case is not one it measures

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: score the answer, and leave a case with no expected answer unmeasured

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: split the evals backend into modules

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: record what a run produced so it outlives its jobs

* fix: read only the agent step's own tool jobs into the payload

* fix: pin a run's configuration and give the judge the attachments

* feat: write a dataset's cases in one transaction

* chore: refresh the sqlx cache for the eval queries

* fix: drop results a newer selection has superseded

* fix: keep a draft the agent editor never opened on

* feat: let a run record what it produced instead of waiting to be read

* fix: serialize the replacements of a dataset's cases

* fix: stop the poller from superseding a read slower than its interval

* chore: refresh the sqlx cache

* fix: keep a failed read from settling a cell as a case with no answer

* fix: hold the case grid while its save is in flight

* fix: keep a failed collect step from failing the run it recorded

* chore: refresh the sqlx cache

* fix: commit an open cell into the save that reads it

* refactor: size the eval buttons with unifiedSize

* docs: describe a run as the one flow it is

* fix: show a run's recorded rows when part of it cannot be collected

* refactor: size the remaining PR-added buttons with unifiedSize

* fix: save the dataset name that was submitted, not the one typed after

* fix: force an open cell into the save that was pressed for it

* fix: refuse to score a run whose evidence could not be read

* fix: hold one lock over a dataset's case count and its writes

* fix: keep one unreadable run from costing the whole runs list

* refactor: drop the banned bindable-default from the eval props

* fix: hold the scorer controls while the dataset is written

* fix: read only the caller's own draft of an agent

* docs: say in the contract that a run pins its configuration

* fix: say a scorer did not run rather than blaming a missing answer

* feat: resume the agent draft you already had when you press Edit

* refactor: build the trail and dataset controls from Button

* fix: clear the open-cell flag when the drawer reopens

* chore: refresh the sqlx cache

* fix: read a run's configuration and its version from one snapshot

* fix: refuse a dataset path or summary the column cannot hold

* refactor: handle the agent draft the way the resource editor does

* fix: run only a configuration the launch actually read

* docs: bound dataset path and summary where they are submitted

* fix: surface a stalled agent draft instead of claiming it is kept

* fix: stop claiming a draft holds edits a failed write never sent

* fix: word a missing score only once the run says whether the case answered

* fix: let a breadcrumb crumb shrink so its truncation applies

* docs: describe where an agent's unsaved edits live and what drops them

* fix: keep harvesting scores when the run cannot yet word a missing one

* fix: report a refused draft write the card was reading as a save

* fix: drop the refused draft write when the server copy is taken instead

* refactor: build the scorer and dataset pickers from the design system

* fix: say what removing a scorer column actually does

* fix: drop a refused draft write wherever the server copy is read

* fix: let a picker row be as tall as the two lines it holds

* docs: record what removing a scorer column does to recorded runs

* fix: send a queued draft write before reopening, and drop only what it refuses

* refactor: write the agent draft at commit points instead of mirroring keystrokes

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor: run an agent's edits from the step instead of keeping them as a draft

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: make the diff badge keyboard operable and refuse an edits run without its edits

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: drop the dataset icon from the scorer picker rows

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: size the evals buttons like the rest of windmill and call a run of edits edits

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: count a brain expression as an edit of the linked agent

* fix: cap scorers per dataset and report a launched run as launched

* fix: harvest scores in one read, refuse duplicate case ids, allow group paths

* fix: mint scorer ids server-side, save a dataset edit in one request, check attachments

* fix: write a dataset edit and its cases in one transaction

* fix: atomic dataset create/edit, reset eval pane per agent, stable pending scorer ids

* refactor: govern eval_case writes by RLS so a dataset edit is one transaction

* fix: pin launch snapshot, order case locks, cap dataset size, guard stale load

* fix: cap dataset bytes on single-case writes, reset run-dialog flag on load failure

* feat: migrate eval datasets on username change, settle unspawned cases, drop unused case endpoints

* fix: resolve scorer scripts as the caller and pin their hash; migrate scorer paths on rename

* fix: bound a failed tool call's error to the payload truncation cap

* fix: pin scorer hash as a hex string, reject missing judges, migrate eval authorship

* fix: record an out-of-range scorer result as an error, not a score

* fix: resolve judges in one caller-scoped read, pin deployed scripts, bound pass_if

* fix: settle unspawned cases only when the run completes, and their score cells too

* feat: reassign eval datasets and their path references when offboarding a user

* fix: use the regex backreference in offboarding eval path rewrites

* fix: register eval datasets in offboarding registries, keep resource-version param name

* refactor: name the resource-version path param id, since it is the row id not the version

* fix: validate dataset paths canonically, clone eval data on fork, surface eval load and launch failures

* docs: note MCP tool results are not yet surfaced to eval scorers

* fix: show the eval error state on any load failure, not only an empty dataset list

* fix: preserve eval case order across a batched save

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* docs: scope the eval launch delete-safety guarantee to the assembly window

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: only offer deployed scripts as eval scorers, drop unbuilt rescore claim

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: enforce 0-1 scorer threshold in the settings drawer and clear stale eval load errors

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: scope subject version/hash reads to the caller and keep a 0 pass threshold

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: select the saved dataset when creating or renaming from the Run dialog

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: gate eval dataset rename on path ownership, not just write access

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* fix: tolerate a malformed agent config when resolving the deployed label

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h

* refactor: trim eval code and comments, fix shared select and modal paths

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: drop the rename warning when editing an eval dataset path

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: add eval dataset delete, keep summary on partial edits, settle resultless scorer cells

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test: cover parseThreshold and subjectLabel

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix: hold dataset Save during a scorer write, derive draft_hash only from the carried draft

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-24 11:08:56 +02:00

31 KiB

AI agent evals

A reusable AI agent (docs/reusable-ai-agents.md) can be run on its own, against a curated set of cases — the inputs it is expected to keep handling.

Three words, and no fourth:

  • a case is one input the agent should handle, held in a dataset;
  • a run (stored as an experiment) is one execution of a whole dataset: a single flow job that answers every case, which is what the UI labels "Run N";
  • each case is answered as one iteration of that run.

The surface is a dialog, two screens deep: this agent's runs across every dataset it has been measured on (one row per run, a badge per scorer), and, on opening a run, one table — a row per case, a column per scorer, the cell being that scorer's verdict — with a case's detail beside it. Editing a dataset is a drawer over both. It opens from where an agent already is: the agent card at the top of an AI agent step's inputs in the flow editor, and the ai_agent row on /resources.

Evals belong to a saved agent. A dataset and its runs hang off an ai_agent resource, so they outlive the step being renamed, copied or deleted, and two runs are comparable because they name the same thing. A step whose agent is written inline has nothing to hang them on: what stands in its place is Save as reusable agent, which is the only setup evals ask for.

What runs

A run is one flow: a loop over the dataset's cases, each iteration answering its case and then scoring the answer. Pushed as a RawFlow, so the agent step is the same vehicle ModuleTest.svelte uses to test an agent step and a case exercises the production branch of ai_executor.rs rather than a parallel one.

One flow rather than one job per case, because a run outlives the tab that started it: a dataset of two hundred cases against a slow provider takes long enough that nobody watches it, and only a worker can notice that the last case finished. Scoring is therefore a step, not something the client does afterwards, and the run is one thing to watch, cancel, or point a schedule at. Each iteration is read back by node id (get_result_and_success_by_id_from_flow), so an answer is fetched without walking the loop's status.

The loop is parallel with a bounded parallelism and skip_failures: a dataset is a burst of calls to one provider, and one case failing is one cell of the run rather than the end of it.

The cases are the loop's static iterator, so they live in the flow's value, which is stored once. Passing them as an argument would put a copy of the whole dataset in every iteration's arguments.

Whichever state of the agent is chosen — what is deployed, the edits in progress, or a past version — its configuration is fixed once, when the run is opened, and inlined into the step every case runs. A linked step would resolve the resource when each case reaches it, so a deploy part-way through a run would be executed by the cases after it while every row still named the version the run started against. One run measures one configuration; the cost is that a run does not exercise the linked branch the production step takes. The edits and a past version could not be run any other way: a reference resolves to what is deployed, which is exactly what neither of them is.

A saved agent's and a past version's configuration are never taken from the request: both are read from the workspace by the path they name, and a subject carrying one is refused. The edits in progress are the one kind the request has to carry — they exist only in the editor — and the run records what it was handed, inlined into its flow and hashed, so it is reproducible and attributable to "this version plus these edits"; what the server cannot assert about them is that they derive from that version.

Each iteration is an agent step, then a payload step, then one step per scorer. The payload step exists because the flow cannot see what it needs to: the agent's own result carries the answer and every message, but each tool call's arguments, result, status, duration and schema belong to the job that ran it. The step reads them back through GET /ai_evals/run_payload, whose one argument is the iteration's own job id, and hands the scorers exactly what a scorer receives anywhere else. A scorer therefore measures the agent's latency and not its own: the payload reports the agent step's duration, never the iteration's.

The edits are the transforms as authored, expressions included. One that reads results.<step>.x or a flow_input the case does not supply resolves to nothing here, the same way it would in any run of that step outside its flow.

A linked agent is not fully self-contained: a host flow can override its tools' inputs through the step's tool_inputs. A run does not reproduce that wiring — the agent runs with its own authored defaults — so an agent whose behaviour depends on one flow's overrides is measured here without them.

A case carries no conversation: one question and the answer it should produce, so a run starts from the agent's own memory configuration and nothing is replayed into it.

Where results live

Results are jobs. A run's logs, trajectory, tool-call child jobs, permissions and retention are already v2_job / v2_job_completed and the flow status's agent_actions, and none of that is stored a second time.

What the table itself is made of is the exception: each cell's answer, its outcome and every scorer's verdict are copied into the run's own rows the first time they can be read. Jobs have their own retention, and a recorded run is meant to still read as the run it was long after the jobs that produced it are gone.

The pane shows the answer and nothing else of a job; the trajectory is the run page's, and job_id is the way there. For a recorded row the answer is read off the row rather than out of the job, because job_id is the whole iteration — the agent and then the scorers that measured it — and its result is the last scorer's verdict, not the answer.

What makes a job findable again is stamped on it at push:

  • runnable_path is the agent's own path, so the existing script_path_start job filter answers "every run of this agent" with no new state.
  • _eval in the flow's args records {subject: {kind, path, version}, dataset, experiment_id}, and every iteration inherits it, so a job opened cold from the runs page explains itself. Which case an iteration ran is in its own iter.value, which is also how a cell finds its job again. Extra flow inputs are inert — the agent step reads only user_message/user_attachments.

Versioning

subject.version is the agent's version number: how many times it has been saved. It is counted per resource rather than read off resource_version.id, which is one identity sequence for the whole table — an agent saved nine times reads v4 … v24 under it, and the gaps count writes in workspaces the reader cannot see. The id stays how a version is addressed, by the history routes and by restore; the number is what a version is called, and what runs are named and compared by.

The number is stored on the row rather than counted when read, because both ways of deleting versions take the oldest: the monitor's trim past MAX_RESOURCE_VERSIONS, and clearing a history down to its current value. Counting the survivors would renumber under either, so a run recorded against v3 would later name a different version.

For an agent run the version names the configuration the run read when it opened, which is the one every case executes. Pinning an older version is a subject kind of its own — an agent_version run says which version to read, where an agent run reads whatever is deployed at the moment it starts.

A version captures the resource, not its transitive closure. Two byte-identical versions can behave differently because a $var:/$res: they reference changed underneath them, so a recorded version is necessary for attribution but not sufficient.

Experiments

Running a dataset produces an experiment: every case executed against one subject, with a row per case. The experiment records the exact case set it ran, by value — a dataset keeps changing, and a result set that cannot say which inputs produced it is not reproducible.

What a run is called

An experiment is Run N: run_number is allocated per (dataset, agent path) when the run is opened, once, and never reused. Stored rather than counted at read time, so a run keeps the name it was given as history is pruned around it. There is no user-given label.

Numbering is per agent and not per subject kind, so runs of what is deployed and runs of the edits on top of it share one sequence: "Run 7" means one thing, and which of the two ran it is what the run says beside its number.

A run is permanent

Every run is written once and then only ever read: there is no writable experiment, no partial rerun, and no cell that can be edited after the fact. A run in which some cells came from one version and some from another would not be worth comparing, and running the dataset is the only way a run appears.

  • One experiment holds one subject, and one agent keeps one history. The experiment list is filtered to the agent the pane was opened on, across both kinds, so a dataset shared by two agents never shows one agent's runs when the other is opened.
  • A version is per cell (eval_experiment_case.subject_version), not per experiment. The subject is resolved once when the run is opened and every cell is stamped from it, so the column is uniform today; it is per cell so that a run which one day executes cell by cell can say so rather than averaging two versions silently.
  • Edits are dated by their hash, not by a version, because editing moves nothing a version could record. Each cell carries subject_draft_hash, the hash of the configuration it ran (canonicalised: key order is not meaningful and serde_json preserves insertion order).
  • An agent's unsaved edits are their own subject. Their runs are keyed under agent_draft, so a number produced by edits is never quietly read as the deployed agent's. The run dialog offers them only when it was opened from the editing card, preselected there; from anywhere else the agent is what is deployed.

The table asks what version the agent is on when it opens and whenever the tab regains focus — a small subject_state read — rather than polling for it; the results endpoint reports the same version but collects the run as it goes, so it is polled only while a run is in flight, one pass at a time. An agent saved in another tab while this one stays focused is noticed on the next focus or the next run.

A run of an older version is history and says so (Run 14 · v23 beside an agent on v24); nothing flags it, since that would flag every past run the moment anything is deployed. A run whose edits were later deployed is a run of that version, and the results endpoint recognises and restamps it (see "What a run says it ran"). The hash each run carries is recorded, not shown.

In the flow editor, editing a linked agent forks the configuration into the step and clears the link; the step is the only copy of the edits until Save changes, Cancel or Discard (docs/reusable-ai-agents.md). Evals open from the agent card in both of its states: from the editing card they run the edits as the step holds them when Run is pressed; from the linked card they run the deployed agent, the same reading a linked step makes at run time. The agent's own resource draft — the one the resource editor writes — is never read by evals. The card also names the version a run is recorded against (v24, and v24 beside an unsaved-changes badge for edits on top of it), read from the resource's newest history entry since the resource itself does not carry its version.

Scoring

Scoring is not a second act with a button of its own: Run produces an answer and then scores it. Each iteration of the run's flow scores its own answer as a step, and the numbers are harvested into rows when results are read.

A run's cells are therefore measured by the scorers as they stood when it ran, and never again. A scorer edited or added afterwards has no cell in the runs that predate it: rescoring a run in place would make a permanent run editable. The one thing read through the present is the pass line — pass_if is applied when a score is read, so moving it re-reads every run with no model call.

The columns themselves are the dataset's current scorers, so the table stays comparable across the runs it lists rather than growing a column per run. Removing a scorer therefore takes its column off the runs already recorded as well: the rows it produced are not deleted, but nothing renders them, and adding the scorer back mints a new column that fills from the next run on. The removal asks first, and says that.

Two things are deliberately absent. Rescoring stored answers under edited scorers would need a run of its own that reuses a parent run's answers and is attributed to the version that produced them, so it never reads as the agent having answered again. A result cache keyed on (agent configuration, case, scorer definition) would assert the agent is deterministic, which it is not, so it has to be an explicit choice with its own answer to what a run means when half of it was computed last week.

A score is a number, and optionally a line through it

Every scorer returns a number between 0 and 1 — both templates say so, and the mean and the pass rate read it as a fraction; a scorer returning anything outside that range has its result recorded as an error rather than counted, and a pass_if threshold is held to the same range. Pass or fail is not a second kind of score: a column carries an optional pass_if, and a case scoring at or above it counts as a pass. A boolean scorer is one that returns 0 or 1 with the line at 0.5.

A column with a threshold reports a pass rate beside its mean and marks each cell; a column without one is a plain number and is not dressed up as a verdict.

The line is deliberately outside the score's definition hash: where it sits is an interpretation of a score rather than part of producing it, so moving it re-reads every run already recorded with nothing re-run. It is set when the column is added and changed later under Scorer settings in the dataset drawer, which is also where the column is named.

The name is this dataset's own name for the scorer, seeded from the summary given when it was added. It is a copy, not a link: the script or judge agent keeps whatever it is called, so renaming a column here does not rename anything a second dataset shows. Reading it live from the runnable would cost a fetch per column and leave a column blank for anyone who cannot read what it points at.

What a scorer receives

An agent is judged on its behaviour, so the final answer is the smaller half of the evidence. Every scorer — a judge prompt or a script — is handed the same EvalRun, built from the job the run already stored:

field from
input, expected the case as the experiment recorded it
output the agent step's own result
tool_calls every message carrying an agent_action, in order, with the arguments, result, error and duration of the job that call ran
tools the tools that were called, with the schema of the script version that ran
metrics steps, duration_ms, and the provider's usage when it reported any

Tool results are truncated at 4 KiB with truncated: true, so a large one cannot swamp a judge's context, and a check that reads a truncated result can say so rather than failing on the missing tail. A tool whose schema could not be resolved carries null, and a scorer validating arguments must treat that as unchecked rather than as a failure. There is no cost field: Windmill keeps no provider price table — the script template takes a rate as an argument instead.

The kinds

Two, and both are runnables:

kind is receives
agent an ai_agent resource used as a judge the run, rendered as a message
script a workspace script run, with input, output and expected also spelled out

Keeping every scorer a runnable is what makes columns comparable: each has a path, a version, and code you can open. There is no third kind stored as configuration on the dataset — a judge's model and grading prompt live on the agent resource, so editing a judge is editing that agent, and the column is not something you edit at all. Editing a column is editing the runnable it points at, so the dataset drawer opens it in place: a script in the script editor, a judge in the resource editor.

Adding a scorer chooses the kind before the form opens: a judge is created next to the dataset from the model you pick and a grading prompt that starts at the default; a script is created from the template and opened in the editor. Both are named by a summary of what they score, which becomes the column header and, prefixed with the dataset, the path.

A reason is worth returning: it is what the cell shows on hover, together with the per-assertion checks, so a number that looks wrong can be read rather than re-derived from the trajectory.

A scorer may return a bare number, a boolean, or {score, reason, checks}; a judge's answer arrives under output, sometimes as a string holding one of those, and often as a markdown code fence around it, which is still read. comment is read as reason, so a scorer written for another platform keeps its rationale. Anything with no number in it is left empty rather than guessed at, and means skip the empty ones — a missing score counted as zero would read as a regression.

{score: null} is the one exception, and it means the scorer read the case and had nothing to measure on it: a column asking whether sources were cited has no verdict on a case with nothing to cite. The cell shows n/a and is left out of the column's mean and pass rate, which is not the same as the scorer failing — that is an error, and the column reports it as one. Written out rather than merely absent, since a scorer that returns nothing at all is a scorer that is broken.

What a run says it ran

A run is v15 when it ran the deployed agent and v15 + edits when it ran that version with undeployed changes on top.

An agent_draft run records the version it is an edit of, because "the draft" is not attributable without saying which deployed state it is a draft of. It also stops being a draft by itself: the agent is hashed as deployed, in the same shape a draft is hashed in, so a run whose configuration was later saved is recognised as the version it became. Edit, run, deploy, and the run you made reads as v16 rather than staying an edit of v15 forever.

That recognition is written, not derived. When the hashes match, the run's subject is rewritten to agent at that version, once, keeping the hash it is founded on. Deriving it on every read would make the answer expire: it would only ever mean "this ran what is deployed right now", so the next deployment would send a run that already read v16 back to v15 + edits. The write goes to the unrestricted pool alongside the scores harvested in the same read, and nothing in it comes from the caller — the hash is the proof, and a run of a configuration that was never deployed simply stays an edit.

It follows that the resolution needs someone to look: a run is stamped by the first results read after its configuration is deployed. A run whose configuration was deployed and then replaced without anyone opening the table keeps saying + edits, which is the honest answer when the only evidence is a hash that matches nothing deployed.

Reusing a scorer

The add form lists the scorers this workspace already uses, most recently edited dataset first, read out of the datasets' own scorers rather than stored anywhere new. It is filtered twice, both times by what the caller can read: the datasets are read through user_db, so a scorer only appears if the dataset carrying it does; then the runnables themselves are checked the same way, so a script or agent the caller cannot open is never suggested.

A scorer is a column

A scorer is stored on the dataset as {id, name?, pass_if?, kind, path}, with the id assigned once and never reused: on a write, an incoming id is kept only when it names a column the dataset already holds, and anything else is minted, so a column that was removed cannot come back under its old id and inherit the scores recorded against it. That id is what makes a column the same column across experiments when the scorer is renamed or its definition edited, and a delta is only ever computed between two scores carrying the same id. Two scorers pointing at the same script are two columns.

A score is keyed (experiment_id, ordinal, scorer_id), not baked into the experiment, so a frozen experiment can gain a score without becoming mutable in any way that matters: what is frozen is which runs are in it.

Each score also records the definition that produced it — the kind, the path, and the script hash or resource version that actually ran, so a path alone cannot hide an edit. When two scores of one column carry different definitions the delta is still shown, marked: hiding the number would force model calls just to see anything, and showing it unmarked would let a change of judge read as a change of agent.

The surface

Opening evals selects the dataset this agent was last worked in, remembered per agent in localStorage and only restored while it still exists and is still readable; no run is opened for you. The picker lists this agent's own datasets first and everyone else's below, sorted rather than filtered, since running one dataset against a second agent is a comparison the picker exists for.

A dataset is named the way a script is: a summary of what the cases are for, from which the path follows, prefixed with the agent so it sorts with the agent's own. With no summary the fallback is <agent>_dataset1, taking the next free number. One path segment rather than a folder under the agent, because a Windmill path is <kind>/<owner>/<name> and the picker that edits it cannot express a deeper one.

The dataset is edited in a drawer over the table: the summary and the path, the scorers, then the cases in a grid. Every way of managing a scorer is in that drawer — adding, renaming, moving its pass line, opening the runnable behind it, removing it; the column header over a run's table reports and does not edit, since a run is permanent. Creating a dataset is the same drawer with no cases yet, reached from the dataset named on a row of the runs list and from the run dialog. Renaming moves the dataset, and its cases and its runs follow through the foreign keys. The drawer edits a working copy and writes it in one request when Save is pressed — the rename, the summary and the cases together — so a rename the server refuses leaves the cases as they were, and a half-finished edit is never what the next run executes. A row's panel in the results table is read-only and shows the case as the run executed it, not as the dataset holds it now; deleting a case is in the drawer, and asks first.

A case is its message, what it expects, and nothing else; the message is what identifies it. expected is what a scorer compares an answer against: plain text, or JSON when the answer has structure.

The runs list is one row per run of this agent, newest first, whichever dataset it was of: the run's number and what executed it (v24, v24 + edits, or a pinned v18), how many cases, one badge per scorer, the dataset, and when. Each badge is the headline that column reports — a pass rate where the column has a line, the mean where it does not — read through the thresholds as they are now. A column that never scored a run reads ; a run still going spins. The badges are named and resolved server-side: a list spanning datasets cannot hold every dataset's scorers to look a column's name up, so the name and the kind ride along with the number, and the thresholds are joined in per (run, column) — one grouped query over eval_score rather than a read of each run's cells. A run whose scores are still in its flow is read out of it by the list itself, capped per call and skipped for runs already collected, so the steady state is one query.

Run asks two questions: which state of the agent (v24 (latest deployed) as it is saved when you press Run; a past version as it was then; v24 + edits (current) running the step's edits as they are when you press Run, offered and preselected only from the editing card), and which dataset, with an edit button on the row and a way to start a new one without leaving. A pinned version reproduces the configuration, not the world around it: $var: and $res: references inside it still resolve at run time. The run that was just started opens straight away.

The results table's rows are the dataset's cases, in dataset order, each carrying its result in the selected experiment when it has one, so a dataset that has never been run is not an empty table. A case the experiment ran but the dataset no longer holds keeps its row at the end: the run happened, and deleting the case does not unmake it. Each column's mean sits under its header, with its delta beside it when a baseline is selected.

Picking a baseline adds a per-scorer delta to every cell and to each column's mean, and counts the cells that regressed. Every delta names its scorer; there is no single number for a dataset, since averaging a judge with an exact match would invent one. Rows are joined by case id, so a case added after the baseline ran has no delta rather than counting as a change, and a column the baseline was never scored with reports that rather than a difference that does not exist.

Storage

Datasets, cases and experiments are rows:

table holds
eval_dataset one dataset, addressed by a workspace path, and the scorers that are its columns
eval_case one case: its inputs and the answer it was expected to produce
eval_experiment one run over a dataset, against one subject; written once, then only read
eval_experiment_case the case set it executed, the job each case became, and the version or draft hash each ran against
eval_score one scorer's verdict on one run, with the definition that produced it

An experiment records its cases by value instead of pointing at eval_case, because a dataset keeps changing and a result set that cannot say which inputs produced it is not reproducible. For the same reason case_id is a plain column rather than a foreign key: deleting a case must not rewrite the history of the runs that used it.

Deleting a dataset takes its cases, its experiments, their recorded case sets and every score with it through the foreign keys. The jobs those experiments produced are left alone — they are jobs, with their own retention.

A case is text: a message and an expected answer. Attachments are S3 references rather than inline bytes, so nothing in a case is meant to be large, and three caps keep it that way — 256 KiB per case, 16 MiB and 1 000 cases per dataset — all refused at the API rather than truncated. A run scores every case by every scorer, so a dataset also holds at most 20 scorers, refused the same way.

Permissions

A dataset is permissioned like any other path-addressed object: row-level security on eval_dataset decides who may see it (readers of its folder, u/<self>, a group, or an extra_perms grant) and who may change it. Operators cannot write at all. Recording an experiment counts as a write, since it persists into the dataset.

Cases are the contents of a dataset rather than objects in their own right. eval_case carries a read policy derived from its dataset (see_parent_dataset) and write policies that check the dataset is writableeval_dataset_writable, one function holding the same disjunction the dataset's own write policies use, so a read-only grant can list a dataset's cases but not edit them. A dataset and its cases therefore move in one user_db transaction, governed by the same policies, and a rename is checked against the destination path the same way. The experiment tables are the exception: their rows are written both by a launch (which holds dataset write) and by the harvest (which holds only read of the run it copies onto its rows), so they carry read policies only and are written on the unrestricted pool after the API has checked the right access.

Why an experiment is recorded before it is launched

Launching picks the run job's id up front, writes the experiment, its case set and a pending score per cell in one transaction, and only then queues the flow. Queueing first and recording afterwards leaves a window in which a flow is running that no experiment accounts for, that nothing will collect and that a retry would silently duplicate. In this order, a launch that dies before the push leaves an experiment naming a job that never started — a run that did not run — and a push that fails deletes it, because one failed push is the whole run.

The dataset's foreign key guards a delete that races the assembly: the transaction fails, and at that point nothing has been queued. It does not cover a delete that lands after this transaction commits and before the flow is queued, which cascades the experiment away while the run still starts.

How a cell finds its job, and its score

The flow engine mints the iteration job ids, so a case is recorded before it has one. Three things fill the gap, each copied out of the flow the first time it can be read:

  • Which iteration ran which case. The case is what the loop iterates over, so it is in the iteration's own arguments by construction: args -> 'iter' -> 'value' ->> 'case_id' matches the cell, whatever order the iterations finish in.
  • What the agent answered. The agent step's result and outcome, copied onto the cell as soon as that step is done — which is well before the iteration around it, since the scorers are still reading it.
  • What the scorers returned. Each scorer step's result is read out of the iteration's flow status into the pending row that was written for it at launch.

All three are written once, when they first become readable, and every later read is of the rows. A job that was retained away before anything read it leaves the cell saying so, rather than looking like a case still being answered.

The flow itself cannot write them: it runs on workers that know nothing about these tables. So two things call the collector. A run's flow ends with a step that calls POST /ai_evals/experiments/collect on itself, which is what records a run nobody watched finish. Reading a run collects it too, which covers the run whose flow never reached that step: one cancelled part-way, or started while nothing served the nativets tag.

That step is bookkeeping, so it is continue_on_error: a run whose every case answered and scored does not become a failed job because the call did not land.