mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-08-24 16:01:42 +00:00
81a25986f6
Evals become a single pane: a dataset of cases, one column per scorer, one row per case, with the run being looked at chosen from the toolbar. Runs are permanent. Running the whole dataset opens one; running a single case records nothing at all — it is a job, and looking at what it did is not a claim that it belongs in the history. Its result and its scores sit over the row until they are saved as a run, which carries the cases that were not rerun and the scoring jobs themselves, so the number that is saved is the number that was looked at. A scorer is a runnable: a judge agent or a script, created in one click and edited in place. Scores carry a reason and per-assertion checks, shown on hover with a rescore button. What ran is always named. A run records the agent version, or — for a configuration that is not deployed — a hash of it, so a table can say that its numbers describe an agent that no longer exists: those rows dim and the table offers to rerun. An agent's draft can be run directly instead of the deployed value, and once those edits are deployed the runs that made them are recognised as that version. A step with no agent of its own is evaluable too, and saving it as an agent moves its history onto it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>