* feat: filter ai session tools to the user's workspace capabilities Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct and tighten comments on the session capability filter Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate session deploy tools on DisableDirectDeployment and the pipeline prompt * fix: gate create_folder on the deploy capability * refactor: assemble session prompt and tools through one seam * docs: state the capability filter as best-effort, not a guarantee * refactor: take the whole deploy gate from the shared preflight `checkDeployPermission` now evaluates `DisableDirectDeployment` and folds superadmin into the admin bypass itself, so the resolver's local composition of those two terms is redundant. Delegate outright and drop the protection-rule fetch it needed, along with the two tests that restated rule semantics the preflight's own suite now pins. The preflight's per-kind narrowing stays unused: the filter runs on tool names, before the model has named a kind, so a direct-deployment lock withholds the deploy tools for schedules and triggers too. * fix: address review findings on the session capability filter Six findings from the Claude and Codex review rounds. - `discard_local_draft` is ungated. The backend exempts discarding your OWN draft from `require_can_write_path` precisely so drafts stay cleanable after a role change; gating it stranded that cleanup. - `deploy` splits into `deploy` and `deploy_gated_kinds`, mirroring `deployPermissionForKind`. A direct-deployment lock stops only the kinds that reach `check_deploy_rules`, so schedules and triggers stay deployable and the two kind-taking deploy tools survive the lock; `create_folder` does not, folder being a gated kind. The prompt now names the lock and what it leaves deployable, instead of implying nothing can be deployed. - `COVERED_ENDPOINTS` keyed `createApp` / `updateApp`, which the MCP catalog does not expose; the app-authoring endpoints it does expose, `createAppRawSource` and `updateAppRawSource`, were uncovered and reachable through `call_api_endpoint`. - The YOLO tooltip listed tools a restricted session never ships. Both it and the token estimate now read one `shippedTools`, and `sessionAccess` is reactive so the UI follows the resolution. * fix: restore the covered API-catalog names for the raw-app endpoints `COVERED_ENDPOINTS` is matched against `EndpointTool.name`, which openapi.yaml overrides with `x-mcp-tool-name` for these two operations: `createAppRawSource` and `updateAppRawSource` are served as `createApp` and `updateApp` (`mcp/auto_generated_endpoints.rs`). Keying them by operationId left both raw-app POST endpoints discoverable and callable through the API catalog tools. Restore the exposed names and record why they differ from the operationIds. * docs: state each capability invariant once, and document the draft discard The asymmetric admin/operator precedence was restated three times in sessionAccess.ts and again in its test, the fail-open rationale twice, and the deploy split across four sites. Each now lives at the one place someone would break it, within the four-line budget, with the other sites pointing at it. Ungating discard_local_draft left it undocumented for the read-only profile, which is the profile the backend exemption exists for: the only bullet naming it sits under the draft-writing gate, beneath an opener saying no change is possible. Add the one line that profile needs. * docs: record why the session tool filter runs unconditionally The filter would strip everything from a non-GLOBAL toolset, whose names carry no policy entries. That cannot happen — `changeMode` refuses to move a session chat out of GLOBAL, and `sessionAccess` is only ever set for session chats — but the dependency was not visible at the filter itself. * fix: match the server's deploy gate exactly, never exceed it The filter must be as strict as the server and no stricter. Schedules and triggers reach no deploy rule — `check_deploy_rules` runs only from the gated kinds' handlers — so no workspace refuses `deploy_workspace_item` or `delete_workspace_item` outright, whatever refusal `checkDeployPermission` reports. Gating them on a deploy capability withheld operations the server performs. Neither tool now requires a capability. Deploying still needs a draft to deploy, so it keeps the authoring relevance; deleting a deployed item does not, so it is ungated. `deploy` returns to one capability, covering the kinds the rules gate, and `create_folder` — whose kind is one of them — is the only tool that names it. The session-state note now states which kinds a refusing workspace still accepts. * feat: gate the new app-runnable preview tool on run_preview Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: fail open when whoami resolves without a role Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: count plan-mode tools in the shipped toolset Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * docs: drop the dead capability assertion and the repeated deploy rationale Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: drop dead code and a duplicated invariant from the session filter Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: reduce SessionAccess to the capability set it is read for Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: gate session tools on permission alone, never on relevance Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * test: pin the filter to the outbound request and widen the description sweep Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: collapse SessionAccess to a capability set and merge adjacent gates Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: gate get_db_schema on run_preview, it runs a query script Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: reuse the cached workspace role and derive the exhaustiveness list from assembly Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * revert: keep tool names in descriptions that ship with them Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: let an admin who is also an operator deploy, as the server does Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: derive the deploy capability from the protection rules alone Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: stop the datatable instructions naming a tool a session may not have Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: keep the prompt and tool results honest for a profile that cannot draft Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: move tool policies onto the tools and gate kinds per handler Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * chore: tighten stale comments and name deploy in the operator prompt Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: keep the assembled tool list raw so narrowing can clone its defs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: point an operator at a workspace admin for code the role refuses Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * fix: resolve session permissions when the assistant settings modal opens Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH * refactor: return per-chat tool schemas instead of writing them to shared tools Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix: forward this through the eval tool wrapper so tools see their sent def Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test: pin identity and contents in the session filter and schema tests Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs: note why the overhead estimate skips per-chat tool schemas Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
AI Evals
Small benchmark runner for the Windmill AI generation modes:
cliflowscriptappglobal
The benchmark always tests the current production prompts, tools, and guidance in this checkout.
Each attempt runs:
- the real production path
- deterministic validation
- LLM judging
Install
cd ai_evals
bun install
Frontend modes also require frontend dependencies:
cd frontend
bun install
Commands
List model aliases:
cd ai_evals
bun run cli -- models
List cases:
cd ai_evals
bun run cli -- cases
bun run cli -- cases flow
Run benchmarks:
cd ai_evals
bun run cli -- run flow
bun run cli -- run flow flow-test4-order-processing-loop --model opus
bun run cli -- run flow flow-test0-sum-two-numbers --models haiku,opus,4o
bun run cli -- run flow flow-test0-sum-two-numbers --runs 3 --verbose
bun run cli -- run flow --record
GEMINI_API_KEY=... bun run cli -- run app app-test1-counter-create --model gemini-3-flash-preview
WMILL_AI_EVAL_BACKEND_URL=http://127.0.0.1:8000 bun run cli -- run flow --backend-validation preview
bun run cli -- run global global-test1-script-create
bun run cli -- run cli bun-hello-script
Public CLI surface:
modelscases [mode]run <mode> [caseIds...]
run options:
--runs <n>: repeat each casentimes--output <path>: custom result JSON path--model <alias>: choose the model under test--models <a,b,c>: run the same cases sequentially against several model aliases--verbose: stream assistant output for frontend runs--skip-judge: skip LLM judge scoring for the run--execution-only: only require the model/proxy/frontend loop to complete; skip validators, tool expectations, backend artifact validation, and judge scoring--record: append a compact tracked summary line toai_evals/history/<mode>.jsonlfor full-suite runs only--backend-validation <mode>: optional backend smoke validation (offorpreview) forscriptandflowevals
Models
Use bun run cli -- models to see the current aliases.
Today:
haikusonnetopus4ogpt-5.5gemini-3-flash-previewgemini-3.1-pro-previewdeepseek-v4-flashdeepseek-v4-pro
Notes:
- the command also prints accepted alias spellings such as
gpt-4o,gpt-55,claude-opus-4.6, andclaude-haiku-4.5 - frontend modes (
flow,script,app,global) can use Anthropic, OpenAI, Gemini, and DeepSeek-backed aliases climode always uses the Anthropic agent SDK, so only Anthropic aliases are valid there- the judge model is separate and currently defaults to
claude-sonnet-4-6; use--skip-judgefor deterministic-only runs
Case Format
Cases live in one YAML file per mode under ai_evals/cases/.
Minimal shape:
- id: flow-test0-sum-two-numbers
prompt: |-
Create a flow that takes two numbers, `a` and `b`, and returns their sum.
initial: ai_evals/fixtures/...
expected: ai_evals/fixtures/...
Optional fields:
initial: starting state fixtureexpected: expected artifact fixturevalidate: extra deterministic validation rulesruntime.backendPreview: optional real backend preview config for smoke validation
For flow mode, validate can express requirements such as:
- accepted input schema shapes
- required
results.*reference validity - required module/code/input characteristics
For app mode, validate can express narrow hard requirements such as:
- required frontend file paths or backend runnable keys
- minimum backend runnable counts
- required backend runnable types
- minimum datatable / datatable-table counts
- specific required datatable tables
For global mode, validate can express draft-level requirements such as:
- required draft type/path/language
- required or forbidden snippets in draft values
- required or forbidden draft counts
- forbidden draft paths
Global initial fixtures can also seed liveEditorDrafts with type,
storagePath, effectivePath, and value fields. These drafts emulate the
currently open script, flow, or raw app editor so cases can test prompts that
refer to "this" or the "current" item.
Global initial fixtures can seed the session's artifacts — { name, versions: [{ content, note? }], role?, approvedVersion? }, oldest version first, so the artifact starts with the
history list_artifact_versions reports — and the previewTabs open in its side panel, for
cases that run with runtime.sessionChat: true. A tab entry names one destination and may
be the active one:
"previewTabs": [{ "artifact": { "name": "Onboarding plan", "version": 2 }, "active": true }]
page ({ href, label }) and item ({ kind, path }) tabs work the same way. Tabs are
driven by the production tab model, so open_preview, get_preview_status and
close_page really open, report and close them, and a version is the pin a reader
chose in the artifact's version picker — which only get_preview_status reports.
Global initial fixtures can seed workspace.variables with
{ path, value, is_secret, description?, labels?, ws_specific? } entries so cases can
read and edit variables that already exist in the workspace. The mock mirrors the real
get_variable, decrypt-by-default included: a secret's value is withheld only
when the caller explicitly passes decryptSecret: false, and omitting the flag returns
the decrypted value, exactly as against a real backend. The chat's read path passes
decryptSecret: false, so a case can verify it never invents a value it was not shown.
Seed a recognizable secret (the existing fixture uses sk_live_do_not_leak_me) and
assert it via valueExcludes to catch a leak.
toolExpect.toolCallArgs entries support sharedByAtLeast: <n>: at least n recorded
calls to that tool must carry the same non-blank string in the field. Use it for calls that
have to share an identifier, like two test runs of one chat conversation.
toolExpect.toolCallArgs entries additionally support fieldMustBeAbsent: true: no
recorded call to that tool may pass the field at all (an explicit null counts as
passing it). Use it for partial-update tools, where supplying a field the model could
not have read is itself the failure — e.g. write_variable.value on a secret variable.
Global (and flow) initial fixtures can seed workspace.datatables so the
list_datatables, get_datatable_table_schema, and exec_datatable_sql tools
return seeded data during evals. Each entry is
{ datatable_name, schemas: { <schema>: { <table>: { columns, rows? } } } }.
SQL runs through a small in-memory engine (datatableSqlEngine.ts), not a real
database. Writes are stateful within a case: CREATE/DROP/INSERT/UPDATE/
DELETE mutate the seeded datatable in place, so a later list_datatables,
get_datatable_table_schema, SELECT, or information_schema query reflects them
— this is what stops a model from looping when it re-queries to verify a write.
The engine is best-effort: SELECT returns all rows of the referenced (or first)
table with no WHERE filtering/projection/joins, WHERE on UPDATE/DELETE supports
col = value predicates joined by AND, and anything unparseable is a no-op
success. So validate datatable cases through tool-use and SQL-argument assertions
(requiredToolsUsed, stringIncludesAnyOf) — not through exact returned row
values. An empty/absent datatables seed makes list_datatables return [],
which is what the "no datatable configured" blocking cases rely on.
Set WMILL_AI_EVAL_DISABLE_ACTIVE_EDITOR_CONTEXT=1 to run those cases with
the old behavior where the live editor is only discoverable through
list_workspace_items.
App fixtures can also include an optional datatables.json file at the fixture root.
For flow mode, an initial fixture can also include a benchmark workspace catalog of
existing scripts and flows. That lets the real search_workspace and
get_runnable_details tools discover reusable workspace runnables during evals.
If --backend-validation preview is enabled:
scriptevals run a real backend script preview in an isolated temp workspaceflowevals run a real backend flow preview only for cases that defineruntime.backendPreviewflowcases withinitial.workspacefixtures seed those scripts and flows into the preview workspace before preview- when
WMILL_AI_EVAL_BACKEND_WORKSPACEis set,ai_evalscreates or reuses that workspace as a dedicated test workspace, clears managed eval assets underf/evals/*before each preview run, and then reseeds the current case fixtures
Supported backend env vars:
WMILL_AI_EVAL_BACKEND_VALIDATION=previewWMILL_AI_EVAL_BACKEND_URL=http://127.0.0.1:8000WMILL_AI_EVAL_BACKEND_EMAIL=admin@windmill.devWMILL_AI_EVAL_BACKEND_PASSWORD=changemeWMILL_AI_EVAL_BACKEND_WORKSPACE=integration-teststo reuse an existing workspace on CE installs with low workspace limits
Frontend modes require a reachable Windmill backend and send model requests through the workspace AI proxy at /api/w/{workspace}/ai/proxy. At startup, ai_evals checks the resolved backend URL and fails early with setup guidance if the backend cannot be reached or login fails.
For frontend modes:
ai_evalscreates a temporary backend workspace, or creates/reusesWMILL_AI_EVAL_BACKEND_WORKSPACEwhen it is set- it upserts a provider resource under
f/evals/ai/<provider> - frontend requests go through
/api/w/{workspace}/ai/proxy
Results And Artifacts
Every run writes:
- a summary JSON under
ai_evals/results/ - generated artifacts in a sibling directory
If --record is used, the CLI also appends one compact JSON line to:
ai_evals/history/flow.jsonlai_evals/history/script.jsonlai_evals/history/app.jsonlai_evals/history/global.jsonlai_evals/history/cli.jsonl
Each recorded line contains:
- run metadata (
createdAt,gitSha,mode,runModel,judgeModel) - suite totals (
caseCount,attemptCount,passedAttempts,passRate,averageDurationMs,averagePassedDurationMs,averageJudgeScore) - average token usage (
averageTokenUsagePerAttempt,averageTokenUsagePerPassedAttempt) - per-case metrics under
cases[](averageDurationMs,averagePassedDurationMs,averageJudgeScore,averageTokenUsagePerAttempt,averageTokenUsagePerPassedAttempt, pass rate) failedCaseIds
The CLI headline duration and token averages use passed attempts only. All-attempt averages are still recorded to make failures auditable without letting failed attempts skew success cost comparisons.
Example:
- summary:
ai_evals/results/2026-04-09T09-40-33.051Z__flow.json - artifacts:
ai_evals/results/2026-04-09T09-40-33.051Z__flow/
Typical artifacts by mode:
flow:flow.jsonscript:script.jsonplus the generated script fileapp:app.jsonplus frontend/backend filesglobal:global-drafts.jsoncli:assistant-output.txt,trace.json,wmill-invocations.jsonl, plus generated workspace files- backend-validated attempts also include
backend-preview.json
Layout
cases/: one YAML file per modefixtures/: initial and expected fixturescore/: shared loading, model resolution, validation, judging, and result writingmodes/: one runner per modehistory/: optional tracked pass-rate history written byrun --record, one JSONL file per moderesults/: local benchmark output and artifacts
Harness unit tests run in two lanes: bun test adapters/ for plain TypeScript, and
bun run test:frontend-graph for *.vitest.ts files, which exercise adapters built on
frontend code (Svelte runes, SvelteKit aliases) that bun cannot load.
Notes
- Frontend modes reuse the production frontend chat code through the Vitest bridge.
- Global mode evaluates the production global AI tools and validates the resulting AI draft store.
- CLI mode creates an isolated workspace, writes the current checkout guidance into it, and benchmarks the real skills /
AGENTS.mdflow. - CLI mode now also records a structured trace of invoked skills, tool calls, proposed
wmillcommands, and any attemptedwmillexecutions. - Frontend progress streams live while the benchmark is running.
- Deterministic validators should stay focused on real correctness constraints, not one exact implementation shape.