Files
windmill/ai_evals/adapters
AlexRV12andClaude Opus 5 6cd70a8e08 feat: filter ai session tools to the user's workspace capabilities (#10719)
* feat: filter ai session tools to the user's workspace capabilities

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: correct and tighten comments on the session capability filter

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: gate session deploy tools on DisableDirectDeployment and the pipeline prompt

* fix: gate create_folder on the deploy capability

* refactor: assemble session prompt and tools through one seam

* docs: state the capability filter as best-effort, not a guarantee

* refactor: take the whole deploy gate from the shared preflight

`checkDeployPermission` now evaluates `DisableDirectDeployment` and folds superadmin
into the admin bypass itself, so the resolver's local composition of those two terms
is redundant. Delegate outright and drop the protection-rule fetch it needed, along
with the two tests that restated rule semantics the preflight's own suite now pins.

The preflight's per-kind narrowing stays unused: the filter runs on tool names, before
the model has named a kind, so a direct-deployment lock withholds the deploy tools for
schedules and triggers too.

* fix: address review findings on the session capability filter

Six findings from the Claude and Codex review rounds.

- `discard_local_draft` is ungated. The backend exempts discarding your OWN draft
  from `require_can_write_path` precisely so drafts stay cleanable after a role
  change; gating it stranded that cleanup.
- `deploy` splits into `deploy` and `deploy_gated_kinds`, mirroring
  `deployPermissionForKind`. A direct-deployment lock stops only the kinds that
  reach `check_deploy_rules`, so schedules and triggers stay deployable and the
  two kind-taking deploy tools survive the lock; `create_folder` does not, folder
  being a gated kind. The prompt now names the lock and what it leaves deployable,
  instead of implying nothing can be deployed.
- `COVERED_ENDPOINTS` keyed `createApp` / `updateApp`, which the MCP catalog does
  not expose; the app-authoring endpoints it does expose, `createAppRawSource` and
  `updateAppRawSource`, were uncovered and reachable through `call_api_endpoint`.
- The YOLO tooltip listed tools a restricted session never ships. Both it and the
  token estimate now read one `shippedTools`, and `sessionAccess` is reactive so
  the UI follows the resolution.

* fix: restore the covered API-catalog names for the raw-app endpoints

`COVERED_ENDPOINTS` is matched against `EndpointTool.name`, which openapi.yaml
overrides with `x-mcp-tool-name` for these two operations: `createAppRawSource`
and `updateAppRawSource` are served as `createApp` and `updateApp`
(`mcp/auto_generated_endpoints.rs`). Keying them by operationId left both raw-app
POST endpoints discoverable and callable through the API catalog tools. Restore
the exposed names and record why they differ from the operationIds.

* docs: state each capability invariant once, and document the draft discard

The asymmetric admin/operator precedence was restated three times in
sessionAccess.ts and again in its test, the fail-open rationale twice, and the
deploy split across four sites. Each now lives at the one place someone would
break it, within the four-line budget, with the other sites pointing at it.

Ungating discard_local_draft left it undocumented for the read-only profile,
which is the profile the backend exemption exists for: the only bullet naming it
sits under the draft-writing gate, beneath an opener saying no change is
possible. Add the one line that profile needs.

* docs: record why the session tool filter runs unconditionally

The filter would strip everything from a non-GLOBAL toolset, whose names carry no
policy entries. That cannot happen — `changeMode` refuses to move a session chat
out of GLOBAL, and `sessionAccess` is only ever set for session chats — but the
dependency was not visible at the filter itself.

* fix: match the server's deploy gate exactly, never exceed it

The filter must be as strict as the server and no stricter. Schedules and triggers
reach no deploy rule — `check_deploy_rules` runs only from the gated kinds' handlers
— so no workspace refuses `deploy_workspace_item` or `delete_workspace_item`
outright, whatever refusal `checkDeployPermission` reports. Gating them on a deploy
capability withheld operations the server performs.

Neither tool now requires a capability. Deploying still needs a draft to deploy, so
it keeps the authoring relevance; deleting a deployed item does not, so it is
ungated. `deploy` returns to one capability, covering the kinds the rules gate, and
`create_folder` — whose kind is one of them — is the only tool that names it. The
session-state note now states which kinds a refusing workspace still accepts.

* feat: gate the new app-runnable preview tool on run_preview

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: fail open when whoami resolves without a role

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: count plan-mode tools in the shipped toolset

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* docs: drop the dead capability assertion and the repeated deploy rationale

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: drop dead code and a duplicated invariant from the session filter

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: reduce SessionAccess to the capability set it is read for

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: gate session tools on permission alone, never on relevance

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* test: pin the filter to the outbound request and widen the description sweep

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: collapse SessionAccess to a capability set and merge adjacent gates

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: gate get_db_schema on run_preview, it runs a query script

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: reuse the cached workspace role and derive the exhaustiveness list from assembly

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* revert: keep tool names in descriptions that ship with them

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: let an admin who is also an operator deploy, as the server does

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: derive the deploy capability from the protection rules alone

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: stop the datatable instructions naming a tool a session may not have

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: keep the prompt and tool results honest for a profile that cannot draft

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: move tool policies onto the tools and gate kinds per handler

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* chore: tighten stale comments and name deploy in the operator prompt

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: keep the assembled tool list raw so narrowing can clone its defs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: point an operator at a workspace admin for code the role refuses

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: resolve session permissions when the assistant settings modal opens

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: return per-chat tool schemas instead of writing them to shared tools

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: forward this through the eval tool wrapper so tools see their sent def

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin identity and contents in the session filter and schema tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: note why the overhead estimate skips per-chat tool schemas

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 18:59:54 +02:00
..