Files
windmill/ai_evals
574775d50c fix: teach the AI the raw-app job bindings, the SDK reference and the draft/deployed split (#10754)
* feat: teach the AI the raw-app job bindings and the draft/deployed split

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: scope the raw-app deploy advice to the referenced item, and stop kind-conversion from stranding fields

The draft/deployed guidance added in the previous commit was read as "deploy the
app too": the agent asked for both the flow and the app and routed a one-item
dependency through the review-and-deploy page. Only the referenced flow or
script has to exist deployed — the preview runs the app's draft — so the prompts,
the `write_app_runnable` warning and the testing rule now say to offer that one
deploy and leave the app a draft.

`buildPersistedRunnable` spread the existing runnable when rewriting it, so
converting a path runnable to inline left `runType`/`path` behind (and the
reverse left `inlineScript`). `isRunnableByName` matches the inline branch
first, so an app "wired to a flow" silently ran stale inline code.

`test_run_app_runnable` now fills ctx-bound inputs with `$ctx:<prop>` the way
RawAppBackgroundRunner does, so a ctx argument no longer arrives missing.

The SDK-reference rationale claimed WM_TOKEN may be unset, that a missing base
URL falls back to localhost, and that a job token is scoped enough to 403 a
hand-rolled REST call. None of the three is true, and it shipped to every
write-script prompt; the text now only says the client configures itself.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address review round on the raw-app AI instructions

The eval case could pass on the exact answer it exists to reject. Every
`requiredMentionsAnyOf` alternative but one was flow-agnostic, so "the app must
be deployed" satisfied "must be deployed". All alternatives now name the flow,
and a unit test pins that the app-only phrasing fails.

`instanceLine` asserted "self-hosted Community Edition" outside the browser,
where `isCloudHosted()` reads false and the license store is unset — so every
global eval was told that regardless of what it pointed at. It is now emitted
only under BROWSER.

`assistantExpect.forbiddenMentions` defaulted a missing `assistantText` to "",
which passes every entry forever on a mode whose runner does not report it.
It now fails with that as the reason.

`buildPersistedRunnable` carried `schema` across a retarget, so a path runnable
pointed at a new flow kept the previous item's schema and `genWmillTs` typed
`backend.<key>(args)` from the wrong inputs. It survives only while kind and
path both match.

The SDK header claimed "a function that is not listed below does not exist".
`windmill-client` also exports the generated services, and the Python client
exposes `Windmill.get`/`.post`, so an endpoint without a helper had no legal
move. Each language now names its own escape hatch.

`getAppInstructions` said the attached reference carries the TypeScript SDK even
when `language: "python3"` had swapped in the Python one — on the very sentence
telling the model to make that call.

The kind-conversion comment claimed a hybrid runnable "silently runs stale
inline code". It does not: `isRunnableByName`, `isRunnableByPath`,
`convertPersistedToBackendRunnable` and `rawAppPolicy.processRunnable` all
dispatch on `type` alone. The leftovers contradict the runnable's kind rather
than override it, which is what the comment now says.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address round-2 review nits on the raw-app AI instructions

`flow is deployed` was satisfied both by "once the flow is deployed, the button
works" and by a hallucinated "done — the flow is deployed", which eval mode makes
impossible and the drafts-only judge cannot see. Every alternative now states an
outstanding obligation, and two more real phrasings ("will need to be deployed")
are accepted so a correct answer is not failed on wording.

Condenses the three comment blocks that ran past the four-line limit in
AGENTS.md, and drops two claims inside them that no longer hold: the
`testRunAppRunnable` doc said it runs a runnable the way the app's own frontend
does (it is the editor preview, which a deployed app's stored policy does not
match), and `undeployedRunnableTargets` described its argument as the write
tool's raw input when the call site passes the persisted runnable.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: report the real cause when a test run fails, and label the app-runnable card

Driving `test_run_app_runnable` in a live session surfaced two defects the
API-level check could not see.

`executeTestRun` built its failure message from `error.message`, which the
generated client leaves as the bare status text while the server's message sits
in `body`. A path runnable aimed at an undeployed flow reported "Not Found"
instead of "Not found: flow not found at name u/admin/current_time" — dropping
the one diagnostic the run exists to produce. `formatToolError`, in the same
file and written for exactly this, now does it. This also applies to
test_run_script and test_run_flow, which had the same loss.

The completion card read "Flow test completed successfully" for an app runnable,
because `contextName` doubles as the jobs-tray kind and a path runnable pointing
at a flow really does queue a flow job. A `completionName` override now names
what ran without changing the kind.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: pin the deploy expectation against wrong answers, not just correct ones

`deploying the flow` was satisfied by "done deploying the flow" — a deploy the
agent only claims to have made, which eval mode makes impossible and the
drafts-only judge cannot see. Replaced with the prospective forms, and dropped
the same reading from the workflow variant.

Three review rounds each found this same class of hole in the phrasing list, so
the list is now exercised against the wrong answers themselves rather than
eyeballed: naming the app as what needs deploying, claiming the deploy is
already done, claiming to have deployed the flow, and saying nothing about
deploying all have to fail, while four real correct phrasings have to pass. The
test reads the case out of global.yaml, so a future edit to the alternatives is
checked by it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: drop the tense-neutral deploy alternatives and cover completed claims

A gerund after a preposition carries no tense, so `before`/`after`/`by deploying
the flow` all match a deploy the agent only claims to have made ("after
deploying the flow, I clicked the button and it returns the greeting") just as
the bare gerund did. All three are gone rather than swapped for whichever reads
least badly, and the two completed-deploy phrasings are now negative fixtures.
The remaining alternatives are imperative or obligational, which a claim of
having already deployed cannot satisfy.

Condenses the two comments this list carries: the YAML block to four lines, and
the test's rationale to the durable constraint about substring matching.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: encrypt sensitive inputs when test-running an app runnable

`test_run_app_runnable` sent `force_viewer_static_fields` but not
`force_viewer_sensitive_inputs`, which every other preview path derives from
the runnable's `sensitive` user fields. That list is the only thing driving the
encryption loop in apps.rs, so testing a runnable with a sensitive input wrote
the real value into the job's args in plaintext, readable by anyone with run
access to the workspace.

Verified against a running EE instance. With the list, `api_key` is stored as
`$encrypted:mvqtSRI9…` and the sentinel appears nowhere in the job record;
without it, the sentinel is readable in run details. A non-sensitive field is
left plaintext either way.

The tool claims parity with the editor preview, so it uses that same filter
(`type == 'user' && sensitive`) and omits the field entirely when empty.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
2026-08-20 11:17:33 +02:00
..

AI Evals

Small benchmark runner for the Windmill AI generation modes:

  • cli
  • flow
  • script
  • app
  • global

The benchmark always tests the current production prompts, tools, and guidance in this checkout.

Each attempt runs:

  1. the real production path
  2. deterministic validation
  3. LLM judging

Install

cd ai_evals
bun install

Frontend modes also require frontend dependencies:

cd frontend
bun install

Commands

List model aliases:

cd ai_evals
bun run cli -- models

List cases:

cd ai_evals
bun run cli -- cases
bun run cli -- cases flow

Run benchmarks:

cd ai_evals
bun run cli -- run flow
bun run cli -- run flow flow-test4-order-processing-loop --model opus
bun run cli -- run flow flow-test0-sum-two-numbers --models haiku,opus,4o
bun run cli -- run flow flow-test0-sum-two-numbers --runs 3 --verbose
bun run cli -- run flow --record
GEMINI_API_KEY=... bun run cli -- run app app-test1-counter-create --model gemini-3-flash-preview
WMILL_AI_EVAL_BACKEND_URL=http://127.0.0.1:8000 bun run cli -- run flow --backend-validation preview
bun run cli -- run global global-test1-script-create
bun run cli -- run cli bun-hello-script

Public CLI surface:

  • models
  • cases [mode]
  • run <mode> [caseIds...]

run options:

  • --runs <n>: repeat each case n times
  • --output <path>: custom result JSON path
  • --model <alias>: choose the model under test
  • --models <a,b,c>: run the same cases sequentially against several model aliases
  • --verbose: stream assistant output for frontend runs
  • --skip-judge: skip LLM judge scoring for the run
  • --execution-only: only require the model/proxy/frontend loop to complete; skip validators, tool expectations, backend artifact validation, and judge scoring
  • --record: append a compact tracked summary line to ai_evals/history/<mode>.jsonl for full-suite runs only
  • --backend-validation <mode>: optional backend smoke validation (off or preview) for script and flow evals

Models

Use bun run cli -- models to see the current aliases.

Today:

  • haiku
  • sonnet
  • opus
  • 4o
  • gpt-5.5
  • gemini-3-flash-preview
  • gemini-3.1-pro-preview
  • deepseek-v4-flash
  • deepseek-v4-pro

Notes:

  • the command also prints accepted alias spellings such as gpt-4o, gpt-55, claude-opus-4.6, and claude-haiku-4.5
  • frontend modes (flow, script, app, global) can use Anthropic, OpenAI, Gemini, and DeepSeek-backed aliases
  • cli mode always uses the Anthropic agent SDK, so only Anthropic aliases are valid there
  • the judge model is separate and currently defaults to claude-sonnet-4-6; use --skip-judge for deterministic-only runs

Case Format

Cases live in one YAML file per mode under ai_evals/cases/.

Minimal shape:

- id: flow-test0-sum-two-numbers
  prompt: |-
    Create a flow that takes two numbers, `a` and `b`, and returns their sum.
  initial: ai_evals/fixtures/...
  expected: ai_evals/fixtures/...

Optional fields:

  • initial: starting state fixture
  • expected: expected artifact fixture
  • validate: extra deterministic validation rules
  • runtime.backendPreview: optional real backend preview config for smoke validation

For flow mode, validate can express requirements such as:

  • accepted input schema shapes
  • required results.* reference validity
  • required module/code/input characteristics

For app mode, validate can express narrow hard requirements such as:

  • required frontend file paths or backend runnable keys
  • minimum backend runnable counts
  • required backend runnable types
  • minimum datatable / datatable-table counts
  • specific required datatable tables

For global mode, validate can express draft-level requirements such as:

  • required draft type/path/language
  • required or forbidden snippets in draft values
  • required or forbidden draft counts
  • forbidden draft paths

Global initial fixtures can also seed liveEditorDrafts with type, storagePath, effectivePath, and value fields. These drafts emulate the currently open script, flow, or raw app editor so cases can test prompts that refer to "this" or the "current" item.

Global initial fixtures can seed the session's artifacts{ name, versions: [{ content, note? }], role?, approvedVersion? }, oldest version first, so the artifact starts with the history list_artifact_versions reports — and the previewTabs open in its side panel, for cases that run with runtime.sessionChat: true. A tab entry names one destination and may be the active one:

"previewTabs": [{ "artifact": { "name": "Onboarding plan", "version": 2 }, "active": true }]

page ({ href, label }) and item ({ kind, path }) tabs work the same way. Tabs are driven by the production tab model, so open_preview, get_preview_status and close_page really open, report and close them, and a version is the pin a reader chose in the artifact's version picker — which only get_preview_status reports.

Global initial fixtures can seed workspace.variables with { path, value, is_secret, description?, labels?, ws_specific? } entries so cases can read and edit variables that already exist in the workspace. The mock mirrors the real get_variable, decrypt-by-default included: a secret's value is withheld only when the caller explicitly passes decryptSecret: false, and omitting the flag returns the decrypted value, exactly as against a real backend. The chat's read path passes decryptSecret: false, so a case can verify it never invents a value it was not shown. Seed a recognizable secret (the existing fixture uses sk_live_do_not_leak_me) and assert it via valueExcludes to catch a leak.

toolExpect.toolCallArgs entries additionally support fieldMustBeAbsent: true: no recorded call to that tool may pass the field at all (an explicit null counts as passing it). Use it for partial-update tools, where supplying a field the model could not have read is itself the failure — e.g. write_variable.value on a secret variable.

Global (and flow) initial fixtures can seed workspace.datatables so the list_datatables, get_datatable_table_schema, and exec_datatable_sql tools return seeded data during evals. Each entry is { datatable_name, schemas: { <schema>: { <table>: { columns, rows? } } } }. SQL runs through a small in-memory engine (datatableSqlEngine.ts), not a real database. Writes are stateful within a case: CREATE/DROP/INSERT/UPDATE/ DELETE mutate the seeded datatable in place, so a later list_datatables, get_datatable_table_schema, SELECT, or information_schema query reflects them — this is what stops a model from looping when it re-queries to verify a write. The engine is best-effort: SELECT returns all rows of the referenced (or first) table with no WHERE filtering/projection/joins, WHERE on UPDATE/DELETE supports col = value predicates joined by AND, and anything unparseable is a no-op success. So validate datatable cases through tool-use and SQL-argument assertions (requiredToolsUsed, stringIncludesAnyOf) — not through exact returned row values. An empty/absent datatables seed makes list_datatables return [], which is what the "no datatable configured" blocking cases rely on.

Set WMILL_AI_EVAL_DISABLE_ACTIVE_EDITOR_CONTEXT=1 to run those cases with the old behavior where the live editor is only discoverable through list_workspace_items.

App fixtures can also include an optional datatables.json file at the fixture root.

For flow mode, an initial fixture can also include a benchmark workspace catalog of existing scripts and flows. That lets the real search_workspace and get_runnable_details tools discover reusable workspace runnables during evals.

If --backend-validation preview is enabled:

  • script evals run a real backend script preview in an isolated temp workspace
  • flow evals run a real backend flow preview only for cases that define runtime.backendPreview
  • flow cases with initial.workspace fixtures seed those scripts and flows into the preview workspace before preview
  • when WMILL_AI_EVAL_BACKEND_WORKSPACE is set, ai_evals creates or reuses that workspace as a dedicated test workspace, clears managed eval assets under f/evals/* before each preview run, and then reseeds the current case fixtures

Supported backend env vars:

  • WMILL_AI_EVAL_BACKEND_VALIDATION=preview
  • WMILL_AI_EVAL_BACKEND_URL=http://127.0.0.1:8000
  • WMILL_AI_EVAL_BACKEND_EMAIL=admin@windmill.dev
  • WMILL_AI_EVAL_BACKEND_PASSWORD=changeme
  • WMILL_AI_EVAL_BACKEND_WORKSPACE=integration-tests to reuse an existing workspace on CE installs with low workspace limits

Frontend modes require a reachable Windmill backend and send model requests through the workspace AI proxy at /api/w/{workspace}/ai/proxy. At startup, ai_evals checks the resolved backend URL and fails early with setup guidance if the backend cannot be reached or login fails.

For frontend modes:

  • ai_evals creates a temporary backend workspace, or creates/reuses WMILL_AI_EVAL_BACKEND_WORKSPACE when it is set
  • it upserts a provider resource under f/evals/ai/<provider>
  • frontend requests go through /api/w/{workspace}/ai/proxy

Results And Artifacts

Every run writes:

  • a summary JSON under ai_evals/results/
  • generated artifacts in a sibling directory

If --record is used, the CLI also appends one compact JSON line to:

  • ai_evals/history/flow.jsonl
  • ai_evals/history/script.jsonl
  • ai_evals/history/app.jsonl
  • ai_evals/history/global.jsonl
  • ai_evals/history/cli.jsonl

Each recorded line contains:

  • run metadata (createdAt, gitSha, mode, runModel, judgeModel)
  • suite totals (caseCount, attemptCount, passedAttempts, passRate, averageDurationMs, averagePassedDurationMs, averageJudgeScore)
  • average token usage (averageTokenUsagePerAttempt, averageTokenUsagePerPassedAttempt)
  • per-case metrics under cases[] (averageDurationMs, averagePassedDurationMs, averageJudgeScore, averageTokenUsagePerAttempt, averageTokenUsagePerPassedAttempt, pass rate)
  • failedCaseIds

The CLI headline duration and token averages use passed attempts only. All-attempt averages are still recorded to make failures auditable without letting failed attempts skew success cost comparisons.

Example:

  • summary: ai_evals/results/2026-04-09T09-40-33.051Z__flow.json
  • artifacts: ai_evals/results/2026-04-09T09-40-33.051Z__flow/

Typical artifacts by mode:

  • flow: flow.json
  • script: script.json plus the generated script file
  • app: app.json plus frontend/backend files
  • global: global-drafts.json
  • cli: assistant-output.txt, trace.json, wmill-invocations.jsonl, plus generated workspace files
  • backend-validated attempts also include backend-preview.json

Layout

  • cases/: one YAML file per mode
  • fixtures/: initial and expected fixtures
  • core/: shared loading, model resolution, validation, judging, and result writing
  • modes/: one runner per mode
  • history/: optional tracked pass-rate history written by run --record, one JSONL file per mode
  • results/: local benchmark output and artifacts

Harness unit tests run in two lanes: bun test adapters/ for plain TypeScript, and bun run test:frontend-graph for *.vitest.ts files, which exercise adapters built on frontend code (Svelte runes, SvelteKit aliases) that bun cannot load.

Notes

  • Frontend modes reuse the production frontend chat code through the Vitest bridge.
  • Global mode evaluates the production global AI tools and validates the resulting AI draft store.
  • CLI mode creates an isolated workspace, writes the current checkout guidance into it, and benchmarks the real skills / AGENTS.md flow.
  • CLI mode now also records a structured trace of invoked skills, tool calls, proposed wmill commands, and any attempted wmill executions.
  • Frontend progress streams live while the benchmark is running.
  • Deterministic validators should stay focused on real correctness constraints, not one exact implementation shape.