feat: filter ai session tools to the user's workspace capabilities (#10719)

* feat: filter ai session tools to the user's workspace capabilities

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: correct and tighten comments on the session capability filter

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: gate session deploy tools on DisableDirectDeployment and the pipeline prompt

* fix: gate create_folder on the deploy capability

* refactor: assemble session prompt and tools through one seam

* docs: state the capability filter as best-effort, not a guarantee

* refactor: take the whole deploy gate from the shared preflight

`checkDeployPermission` now evaluates `DisableDirectDeployment` and folds superadmin
into the admin bypass itself, so the resolver's local composition of those two terms
is redundant. Delegate outright and drop the protection-rule fetch it needed, along
with the two tests that restated rule semantics the preflight's own suite now pins.

The preflight's per-kind narrowing stays unused: the filter runs on tool names, before
the model has named a kind, so a direct-deployment lock withholds the deploy tools for
schedules and triggers too.

* fix: address review findings on the session capability filter

Six findings from the Claude and Codex review rounds.

- `discard_local_draft` is ungated. The backend exempts discarding your OWN draft
  from `require_can_write_path` precisely so drafts stay cleanable after a role
  change; gating it stranded that cleanup.
- `deploy` splits into `deploy` and `deploy_gated_kinds`, mirroring
  `deployPermissionForKind`. A direct-deployment lock stops only the kinds that
  reach `check_deploy_rules`, so schedules and triggers stay deployable and the
  two kind-taking deploy tools survive the lock; `create_folder` does not, folder
  being a gated kind. The prompt now names the lock and what it leaves deployable,
  instead of implying nothing can be deployed.
- `COVERED_ENDPOINTS` keyed `createApp` / `updateApp`, which the MCP catalog does
  not expose; the app-authoring endpoints it does expose, `createAppRawSource` and
  `updateAppRawSource`, were uncovered and reachable through `call_api_endpoint`.
- The YOLO tooltip listed tools a restricted session never ships. Both it and the
  token estimate now read one `shippedTools`, and `sessionAccess` is reactive so
  the UI follows the resolution.

* fix: restore the covered API-catalog names for the raw-app endpoints

`COVERED_ENDPOINTS` is matched against `EndpointTool.name`, which openapi.yaml
overrides with `x-mcp-tool-name` for these two operations: `createAppRawSource`
and `updateAppRawSource` are served as `createApp` and `updateApp`
(`mcp/auto_generated_endpoints.rs`). Keying them by operationId left both raw-app
POST endpoints discoverable and callable through the API catalog tools. Restore
the exposed names and record why they differ from the operationIds.

* docs: state each capability invariant once, and document the draft discard

The asymmetric admin/operator precedence was restated three times in
sessionAccess.ts and again in its test, the fail-open rationale twice, and the
deploy split across four sites. Each now lives at the one place someone would
break it, within the four-line budget, with the other sites pointing at it.

Ungating discard_local_draft left it undocumented for the read-only profile,
which is the profile the backend exemption exists for: the only bullet naming it
sits under the draft-writing gate, beneath an opener saying no change is
possible. Add the one line that profile needs.

* docs: record why the session tool filter runs unconditionally

The filter would strip everything from a non-GLOBAL toolset, whose names carry no
policy entries. That cannot happen — `changeMode` refuses to move a session chat
out of GLOBAL, and `sessionAccess` is only ever set for session chats — but the
dependency was not visible at the filter itself.

* fix: match the server's deploy gate exactly, never exceed it

The filter must be as strict as the server and no stricter. Schedules and triggers
reach no deploy rule — `check_deploy_rules` runs only from the gated kinds' handlers
— so no workspace refuses `deploy_workspace_item` or `delete_workspace_item`
outright, whatever refusal `checkDeployPermission` reports. Gating them on a deploy
capability withheld operations the server performs.

Neither tool now requires a capability. Deploying still needs a draft to deploy, so
it keeps the authoring relevance; deleting a deployed item does not, so it is
ungated. `deploy` returns to one capability, covering the kinds the rules gate, and
`create_folder` — whose kind is one of them — is the only tool that names it. The
session-state note now states which kinds a refusing workspace still accepts.

* feat: gate the new app-runnable preview tool on run_preview

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: fail open when whoami resolves without a role

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: count plan-mode tools in the shipped toolset

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* docs: drop the dead capability assertion and the repeated deploy rationale

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: drop dead code and a duplicated invariant from the session filter

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: reduce SessionAccess to the capability set it is read for

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: gate session tools on permission alone, never on relevance

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* test: pin the filter to the outbound request and widen the description sweep

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: collapse SessionAccess to a capability set and merge adjacent gates

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: gate get_db_schema on run_preview, it runs a query script

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: reuse the cached workspace role and derive the exhaustiveness list from assembly

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* revert: keep tool names in descriptions that ship with them

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: let an admin who is also an operator deploy, as the server does

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: derive the deploy capability from the protection rules alone

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: stop the datatable instructions naming a tool a session may not have

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: keep the prompt and tool results honest for a profile that cannot draft

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: move tool policies onto the tools and gate kinds per handler

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* chore: tighten stale comments and name deploy in the operator prompt

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: keep the assembled tool list raw so narrowing can clone its defs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: point an operator at a workspace admin for code the role refuses

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* fix: resolve session permissions when the assistant settings modal opens

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L99mAR4LitqTcYY1Kn1ATH

* refactor: return per-chat tool schemas instead of writing them to shared tools

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix: forward this through the eval tool wrapper so tools see their sent def

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test: pin identity and contents in the session filter and schema tests

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs: note why the overhead estimate skips per-chat tool schemas

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
AlexRV12
2026-09-23 18:59:54 +02:00
committed by GitHub
co-authored by Claude Opus 5
parent 0e53d53b07
commit 6cd70a8e08
26 changed files with 1040 additions and 148 deletions
@@ -98,9 +98,11 @@ export async function runEval<THelpers, TOutput>(
// Wrap tools to intercept fn calls for tracking.
// Cast to ProductionTool since the eval Tool has a narrower toolCallbacks type
// but the actual callbacks passed at runtime will satisfy both interfaces.
// `fn` forwards `this`: the chat loop calls it on a per-iteration copy carrying the def the
// model was sent, and a tool reading `this.def` must see that one, as it does in production.
const wrappedTools = tools.map((tool) => ({
...tool,
fn: async (p: any) => {
fn: async function (this: unknown, p: any) {
toolCallsCount++;
toolsCalled.push(tool.def.function.name);
let argumentsText = "";
@@ -120,7 +122,7 @@ export async function runEval<THelpers, TOutput>(
toolName: tool.def.function.name,
argumentsText,
});
return tool.fn(p);
return tool.fn.call(this, p);
},
}));
@@ -125,6 +125,9 @@ import {
isWebSearchEnabledForProvider
} from '$lib/aiStore'
import type { WorkspaceMutationTarget } from './workspaceTools'
import { resolveSessionAccess } from './global/sessionAccess'
import type { SessionAccess } from './sessionCapabilities'
import { filterSessionTools } from './global/sessionToolset'
import {
loadWorkspaceSkills,
resolveGlobalPromptIdentity,
@@ -696,6 +699,10 @@ export class AIChatManager implements ChatViewHost {
// needs the preview pane; the global side-panel chat leaves it false. Reactive because
// `planModeAvailable` derives from it.
isSessionChat = $state(false)
// Undefined until a send or the assistant settings modal resolves it. Reactive: both
// tool views derive from it.
private sessionAccess = $state<SessionAccess | undefined>(undefined)
private sessionAccessGeneration = 0
autoAcceptEditsAvailable = $derived(supportsAutoAcceptEdits(this.mode))
autoAcceptEditsActive = $derived(
this.autoAcceptEditsAvailable &&
@@ -719,15 +726,21 @@ export class AIChatManager implements ChatViewHost {
})
/** The sources the current mode assembled, concatenated. Private, because neither view
* over it is this list: a consumer reading it would advertise a tool set no request ever
* carries. */
#assembledTools = $state<Tool<any>[]>([])
* carries. Raw: every write replaces the list, and a deep proxy would reach the tool
* defs, which the session filter copies with `structuredClone` — that throws on a proxy. */
#assembledTools = $state.raw<Tool<any>[]>([])
/** Both views narrow the same way — concatenate every source, then filter — so a tool
* reaching either one can never arrive unfiltered. A session withholds what its user's
* capabilities do not cover; elsewhere `sessionAccess` is unset and nothing is dropped. */
#shipped = (planTools: Tool<any>[]): Tool<any>[] =>
filterSessionTools([...this.#assembledTools, ...planTools], this.sessionAccess)
/** What the request carries: the assembled tools plus the plan-mode transition the current
* posture offers. Read by the request path and by the YOLO disclosure. */
tools: Tool<any>[] = $derived([...this.#assembledTools, ...this.planMode.tools])
tools: Tool<any>[] = $derived(this.#shipped(this.planMode.tools))
/** What the assistant can call in this session, posture-independent: both plan-mode
* transitions, whichever one is offered right now. A reference answer, so flipping the
* autonomy picker must not change it. */
availableTools: Tool<any>[] = $derived([...this.#assembledTools, ...this.planMode.availableTools])
availableTools: Tool<any>[] = $derived(this.#shipped(this.planMode.availableTools))
helpers = $state<any | undefined>(undefined)
scriptEditorOptions = $state<ScriptOptions | undefined>(undefined)
@@ -1427,10 +1440,12 @@ export class AIChatManager implements ChatViewHost {
typeof this.systemMessage.content === 'string'
? this.systemMessage.content.length / tokenPerCharacter
: 0
const tools = this.tools
// Counts each tool's shared `def`, not the per-chat one `schemaFor` builds at send time,
// so a flow's input schema is missed. Accepted: the estimate only stands in until the
// provider reports usage, and the compaction trigger's headroom absorbs the gap.
const toolTokens =
this.tools.length > 0
? JSON.stringify(this.tools.map((t) => t.def)).length / tokenPerCharacter
: 0
tools.length > 0 ? JSON.stringify(tools.map((t) => t.def)).length / tokenPerCharacter : 0
return systemTokens + toolTokens
}
@@ -2332,8 +2347,9 @@ export class AIChatManager implements ChatViewHost {
}
) {
if (!isAIModeVisible(mode)) return
// A session chat is GLOBAL for its whole life, and the plan gate reads that mode: moving
// it lifts the gate on a session the user still has set to Plan.
// A session chat is GLOBAL for its whole life, and two things read that mode: moving it
// lifts the plan gate on a session the user still has set to Plan, and hands the
// capability filter tools that declare no `requires`, so it fails closed to nothing.
if (this.isSessionChat && mode !== AIMode.GLOBAL) {
console.error(`Refusing to move a session chat to ${mode} mode: sessions are GLOBAL-only.`)
return
@@ -2435,6 +2451,7 @@ export class AIChatManager implements ChatViewHost {
sessionId: this.sessionId,
operatingWorkspace: this.operatingWorkspace,
artifacts: this.artifacts,
access: this.sessionAccess,
getChatId: () => this.historyManager.getCurrentChatId(),
openArtifact: this.openArtifact
}
@@ -2466,6 +2483,7 @@ export class AIChatManager implements ChatViewHost {
user: this.globalIdentity,
skills: this.globalSkills,
mcpServers: this.mcpServers,
access: this.sessionAccess,
sessionContext: this.sessionContextResolver?.(),
pipelineContext: this.pipelineAiChatHelpers?.getPipelineContext()
})
@@ -2540,6 +2558,23 @@ export class AIChatManager implements ChatViewHost {
this.systemMessage = { ...target, content: `${target.content}\n\n${section}` }
}
// Same shape as refreshGlobalSkills. Re-resolved per send rather than once for the
// session's life: the operating workspace can change between sends, and a transient
// failure resolves fail-open, so the next message re-asks rather than keeping that answer.
refreshSessionAccess = async (workspace = this.operatingWorkspace ?? '') => {
if (!this.isSessionChat || !workspace) {
this.sessionAccess = undefined
return
}
const generation = ++this.sessionAccessGeneration
const access = await resolveSessionAccess(workspace)
if (generation !== this.sessionAccessGeneration) return
this.sessionAccess = workspace === (this.operatingWorkspace ?? '') ? access : undefined
if (this.mode === AIMode.GLOBAL) {
this.configureGlobalMode()
}
}
// Rebuild the GLOBAL system message in place so an updated user instruction (persisted by
// the update_user_instructions tool) is picked up on the next chat-loop iteration, which
// re-reads this.systemMessage via a getter.
@@ -3043,13 +3078,8 @@ export class AIChatManager implements ChatViewHost {
console.error('Failed to record AI usage', e)
}
},
onBeforeIteration: async (tools, _helpers, modelProvider) => {
onBeforeIteration: async (modelProvider) => {
this.lastIterationModel = modelProvider
for (const tool of tools) {
if (tool.setSchema) {
await tool.setSchema(this.helpers)
}
}
}
})
if (this.isSessionChat && this.sessionId && result.tokenUsage.total > 0) {
@@ -3553,15 +3583,16 @@ export class AIChatManager implements ChatViewHost {
return false
}
}
// Session chats commit their workspace in beforeSend; the identity, skills and
// MCP servers must all match the committed workspace before the system prompt is
// sent. Settling them here rather than mid-turn also keeps the prompt — the
// Session chats commit their workspace in beforeSend; the identity, skills, MCP
// servers and capabilities must all match the committed workspace before the system
// prompt is sent. Settling them here rather than mid-turn also keeps the prompt — the
// cached prefix of every iteration — stable for the whole request.
if (this.mode === AIMode.GLOBAL) {
await Promise.all([
this.refreshGlobalIdentity(this.operatingWorkspace ?? ''),
this.refreshGlobalSkills(this.operatingWorkspace ?? ''),
this.refreshMcpServers(this.operatingWorkspace ?? '')
this.refreshMcpServers(this.operatingWorkspace ?? ''),
this.refreshSessionAccess(this.operatingWorkspace ?? '')
])
}
// Stop/Escape during the beforeSend pre-flight aborted this send before any
@@ -1,4 +1,4 @@
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
import { afterEach, beforeEach, describe, expect, it, onTestFinished, vi } from 'vitest'
import { writable } from 'svelte/store'
import type { FlowAIChatHelpers } from './flow/core'
import type { PipelineAIChatHelpers } from './pipeline/core'
@@ -1679,7 +1679,7 @@ describe('AIChatManager queued messages', () => {
mocks.tryGetCurrentModel.mockReturnValue(a)
mocks.runChatLoop.mockImplementation(async (config: any) => {
// an iteration starts on B...
await config.onBeforeIteration?.([], config.helpers, b)
await config.onBeforeIteration?.(b)
// ...the user switches to C while B's request is in flight...
mocks.getCurrentModel.mockReturnValue(c)
mocks.tryGetCurrentModel.mockReturnValue(c)
@@ -1792,7 +1792,7 @@ describe('AIChatManager queued messages', () => {
// mid-loop switch to a model the deny-list doesn't know...
mocks.getCurrentModel.mockReturnValue(unlistedBlind)
mocks.tryGetCurrentModel.mockReturnValue(unlistedBlind)
await config.onBeforeIteration?.([], config.helpers, unlistedBlind)
await config.onBeforeIteration?.(unlistedBlind)
// ...its request carries the images and the provider rejects them
throw new Error('400 this model does not support image input')
})
@@ -4565,6 +4565,7 @@ describe('AIChatManager tool views', () => {
vi.clearAllMocks()
mocks.getCurrentModel.mockReturnValue({ provider: 'openai', model: 'gpt-4o' })
mocks.tryGetCurrentModel.mockReturnValue({ provider: 'openai', model: 'gpt-4o' })
clearWorkspaceRoleCache()
})
// Every mode assigns the private base separately. A missed site leaves it empty on a fresh
@@ -4584,6 +4585,73 @@ describe('AIChatManager tool views', () => {
expect(manager.availableTools.length).toBeGreaterThan(0)
})
const OPERATOR_WHOAMI = {
username: 'op',
email: 'admin@test',
is_admin: false,
is_super_admin: false,
operator: true,
groups: [],
folders: [],
folders_read: []
}
// The capability filter is unit-tested against a profile handed to it directly; what is
// untested there is that the manager ever hands it one. This drives a real send and reads
// the toolset off the request, so moving the resolve out of the pre-flight, or dropping
// the rebuild when it lands, fails here rather than shipping an operator write tools.
it('withholds write and preview tools from an operator for the whole request', async () => {
onTestFinished(() => mocks.whoami.mockReset())
mocks.whoami.mockResolvedValue(OPERATOR_WHOAMI)
let sent: { tools: string[]; prompt: string; deployKinds: string[] } | undefined
mocks.runChatLoop.mockImplementation(async (config: any) => {
const deploy = config.tools.find((t: any) => t.def.function.name === 'deploy_workspace_item')
sent = {
tools: config.tools.map((t: any) => t.def.function.name),
prompt: config.systemMessage.content,
deployKinds: deploy?.def.function.parameters.properties.type.enum ?? []
}
return {
addedMessages: [],
tokenUsage: { prompt: 0, completion: 0, total: 0 },
hitMaxIterations: false
}
})
const manager = new AIChatManager()
manager.isSessionChat = true
await manager.sendRequest({ instructions: 'add a script', mode: AIMode.GLOBAL })
expect(sent?.tools).toEqual(expect.arrayContaining(['list_workspace_items', 'run_script']))
// Deploying ships — a draft can predate the role change — but only for the kinds an
// operator's token can land.
expect(sent?.deployKinds).toContain('schedule')
expect(sent?.deployKinds).not.toContain('script')
for (const withheld of ['write_script', 'test_run_script']) {
expect(sent?.tools).not.toContain(withheld)
// The prompt ships beside the tools, so the rebuild must have run after the
// profile landed — otherwise it still instructs the model to call these.
expect(sent?.prompt).not.toContain(withheld)
}
})
// The assistant settings modal resolves the profile on open, before any send, so the
// list it shows is the one the first request will carry, prompt included.
it('narrows the toolset and prompt when the profile resolves outside a send', async () => {
onTestFinished(() => mocks.whoami.mockReset())
mocks.whoami.mockResolvedValue(OPERATOR_WHOAMI)
const manager = new AIChatManager()
manager.mode = AIMode.GLOBAL
manager.isSessionChat = true
manager.configureGlobalMode()
expect(manager.availableTools.map((t) => t.def.function.name)).toContain('write_script')
await manager.refreshSessionAccess()
expect(manager.availableTools.map((t) => t.def.function.name)).not.toContain('write_script')
expect(manager.systemMessage.content).not.toContain('write_script')
})
// Which transition each posture offers is planModeController.test.ts's; what this pins is
// that only `tools` follows the picker. `isSessionChat` is what offers plan mode at all.
it('holds availableTools steady across the autonomy picker', async () => {
@@ -149,6 +149,7 @@ callers that already know which section they mean, such as the "+" menu's Manage
if (!aiChatManager.loading && !aiChatManager.sendInFlight) {
void aiChatManager.refreshGlobalSkills()
void aiChatManager.refreshMcpServers()
void aiChatManager.refreshSessionAccess()
}
}
</script>
@@ -1,5 +1,6 @@
import { z } from 'zod'
import { createToolDef, type Tool, type ToolCallbacks } from '../shared'
import { createToolDef, type ToolCallbacks } from '../shared'
import { NONE, type SessionTool } from '../sessionCapabilities'
import { artifactOverflowBytes, MAX_ARTIFACT_BYTES, normalizeChangeNote } from './artifactLimits'
import {
currentVersion,
@@ -98,8 +99,9 @@ function targetsPlan(args: unknown, h: ArtifactToolHelpers): boolean {
)
}
export const artifactTools: Tool<{}>[] = [
export const artifactTools: SessionTool<{}>[] = [
{
requires: NONE,
def: createToolDef(
createArtifactSchema,
'create_artifact',
@@ -156,6 +158,7 @@ export const artifactTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
updateArtifactSchema,
'update_artifact',
@@ -208,6 +211,7 @@ export const artifactTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
listArtifactsSchema,
'list_artifacts',
@@ -242,6 +246,7 @@ export const artifactTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
readArtifactSchema,
'read_artifact',
@@ -303,6 +308,7 @@ export const artifactTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
listArtifactVersionsSchema,
'list_artifact_versions',
@@ -569,6 +569,32 @@ const assistantTools = (...ids: string[]): ChatCompletionMessageParam => ({
function: { name: 'do_thing', arguments: '{}' }
}))
})
describe('runChatLoop per-chat tool schemas', () => {
beforeEach(() => {
vi.resetAllMocks()
mocks.resolveRequestReasoning.mockReturnValue(undefined)
mocks.getOpenAIResponsesCompletion.mockResolvedValue({})
mocks.parseOpenAIResponsesCompletion.mockResolvedValue({ shouldContinue: false, tokenUsage })
})
// The def `schemaFor` returns is the chat's own: it must be what the provider is sent and
// what the tool is dispatched with, while the shared tool object keeps its own.
it('sends and dispatches the def schemaFor returns, leaving the shared tool alone', async () => {
const base = { type: 'function', function: { name: 'dyn', parameters: {} } } as any
const perChat = {
type: 'function',
function: { name: 'dyn', parameters: { marked: 1 } }
} as any
const shared = { def: base, fn: async () => '', schemaFor: async () => perChat }
await runChatLoop({ ...createConfig({ workspace: `ws-${randomUUID()}` }), tools: [shared] })
expect(mocks.getOpenAIResponsesCompletion.mock.calls[0][2]).toEqual([perChat])
expect(mocks.parseOpenAIResponsesCompletion.mock.calls[0][4][0].def).toBe(perChat)
expect(shared.def).toBe(base)
})
})
const tool = (id: string): ChatCompletionMessageParam => ({
role: 'tool',
tool_call_id: id,
@@ -71,13 +71,8 @@ export interface ChatLoopConfig {
* lets the caller recover partial output if the loop throws or is aborted.
*/
addedMessages?: ChatCompletionMessageParam[]
/** Called before each iteration (e.g. to refresh tool schemas, or to record
* which model the iteration is about to use). */
onBeforeIteration?: (
tools: Tool<any>[],
helpers: any,
modelProvider: ReasoningProviderModel
) => Promise<void>
/** Called before each iteration (e.g. to record which model it is about to use). */
onBeforeIteration?: (modelProvider: ReasoningProviderModel) => Promise<void>
/** Fired for each completed provider response, before the loop continues. The
* loop can fail or be aborted at any iteration, so spend has to be handed over
* as it happens — a callback only at the end would discard everything the
@@ -399,7 +394,6 @@ export async function runChatLoop(config: ChatLoopConfig): Promise<ChatLoopResul
// Re-read these from config each iteration so that mode changes
// (e.g. changeModeTool in Navigator) take effect immediately.
// Callers can use JS getter properties to provide dynamic values.
const tools = config.tools
const helpers = config.helpers
const systemMessage = config.systemMessage
const modelProvider = config.modelProvider
@@ -411,8 +405,11 @@ export async function runChatLoop(config: ChatLoopConfig): Promise<ChatLoopResul
!unsupportedWebSearchCache.has(webSearchCacheKey)
if (onBeforeIteration) {
await onBeforeIteration(tools, helpers, modelProvider)
await onBeforeIteration(modelProvider)
}
const tools = await Promise.all(
config.tools.map(async (t) => (t.schemaFor ? { ...t, def: await t.schemaFor(helpers) } : t))
)
const pendingUserMessage = getPendingUserMessage?.()
@@ -6,10 +6,10 @@ import { datatableReference } from '$lib/components/dbTypes'
import {
createToolDef,
executeTestRun,
type Tool,
type ToolDisplayMessage,
type ChatJobResultFormat
} from './shared'
import { NONE, RUN_PREVIEW, type SessionTool } from './sessionCapabilities'
/**
* Workspace-scoped datatable tools, with no app whitelist and no creation policy.
@@ -246,9 +246,10 @@ export function formatChatJobCompletion(
* The unrestricted workspace datatable tools, for registration in global mode.
* Helper-free: each tool reads `workspace` directly from the tool call params.
*/
export function getDatatableTools(): Tool<{}>[] {
export function getDatatableTools(): SessionTool<{}>[] {
return [
{
requires: NONE,
def: getListDatatablesToolDef(),
planModeSafe: true,
fn: async ({ args, workspace, toolId, toolCallbacks }) => {
@@ -298,6 +299,7 @@ export function getDatatableTools(): Tool<{}>[] {
}
},
{
requires: NONE,
def: getGetDatatableTableSchemaToolDef(),
planModeSafe: true,
fn: async ({ args, workspace, toolId, toolCallbacks }) => {
@@ -337,6 +339,8 @@ export function getDatatableTools(): Tool<{}>[] {
}
},
{
// `runScript` below goes through /jobs/run/preview, which jobs.rs refuses operators.
requires: RUN_PREVIEW,
def: getExecDatatableSqlToolDef(),
requiresConfirmation: true,
confirmationMessage: 'Execute SQL on datatable',
@@ -1,4 +1,4 @@
import type { Tool } from '../shared'
import { NONE, type SessionTool } from '../sessionCapabilities'
import type { ChatCompletionTool } from 'openai/resources/index.mjs'
import { DocumentationService } from '$lib/gen'
@@ -33,7 +33,8 @@ const READ_DOCS_PAGE_TOOL: ChatCompletionTool = {
}
}
export const readDocsPageTool: Tool<{}> = {
export const readDocsPageTool: SessionTool<{}> = {
requires: NONE,
def: READ_DOCS_PAGE_TOOL,
planModeSafe: true,
fn: async ({ args, toolId, toolCallbacks }) => {
@@ -85,7 +86,8 @@ const SEARCH_DOCS_TOOL: ChatCompletionTool = {
}
}
export const searchDocsTool: Tool<{}> = {
export const searchDocsTool: SessionTool<{}> = {
requires: NONE,
def: SEARCH_DOCS_TOOL,
planModeSafe: true,
fn: async ({ args, toolId, toolCallbacks }) => {
@@ -1,6 +1,7 @@
import { z } from 'zod'
import { DataMetricService, WorkspaceService } from '$lib/gen'
import { createToolDef, type Tool } from './shared'
import { createToolDef } from './shared'
import { NONE, type SessionTool } from './sessionCapabilities'
/**
* Workspace-scoped DuckLake tools, the pipeline counterpart to `list_datatables`
@@ -69,9 +70,10 @@ const NO_DATA_METRICS_NOTE =
const DEFAULT_DATA_METRICS_LIMIT = 200
/** The workspace DuckLake tools, for registration in global mode. */
export function getDucklakeTools(): Tool<{}>[] {
export function getDucklakeTools(): SessionTool<{}>[] {
return [
{
requires: NONE,
def: listDataMetricsToolDef,
planModeSafe: true,
showDetails: true,
@@ -106,6 +108,7 @@ export function getDucklakeTools(): Tool<{}>[] {
}
},
{
requires: NONE,
def: listDucklakesToolDef,
planModeSafe: true,
fn: async ({ workspace, toolId, toolCallbacks }) => {
@@ -8,7 +8,8 @@
*/
import { z } from 'zod'
import type { ChatCompletionSystemMessageParam } from 'openai/resources/chat/completions.mjs'
import { createToolDef, type Tool } from '../shared'
import { createToolDef } from '../shared'
import { NONE, type SessionTool } from '../sessionCapabilities'
import {
readFile,
searchFilesInWorker,
@@ -95,7 +96,8 @@ const searchFilesToolDef = createToolDef(
'Search the user-attached files with a regular expression and return matching lines with their line numbers. Use this to locate content before reading a specific window with read_file.'
)
export const searchFilesTool: Tool<{}> = {
export const searchFilesTool: SessionTool<{}> = {
requires: NONE,
def: searchFilesToolDef,
planModeSafe: true,
fn: async ({ args, helpers, toolId, toolCallbacks }) => {
@@ -164,7 +166,8 @@ const readFileToolDef = createToolDef(
'Read a bounded window of lines from a user-attached file. Returns each line prefixed with its 1-based number (`<n>→<content>`) plus a pagination note. Files are not in context, so use this to inspect their contents.'
)
export const readFileTool: Tool<{}> = {
export const readFileTool: SessionTool<{}> = {
requires: NONE,
def: readFileToolDef,
planModeSafe: true,
fn: async ({ args, helpers, toolId, toolCallbacks }) => {
@@ -193,7 +196,7 @@ export const readFileTool: Tool<{}> = {
}
}
export const fileTools: Tool<{}>[] = [searchFilesTool, readFileTool]
export const fileTools: SessionTool<{}>[] = [searchFilesTool, readFileTool]
function rosterLine(f: AttachedFile): string {
// Message rows are addressed by their stable id (names may collide); session
@@ -302,7 +302,7 @@ const patchFlowJsonToolDef = createToolDef(
'patch_flow_json',
'Make a quick exact text edit in the current compact flow JSON. Prefer this for small localized changes; use set_flow_json for larger structural rewrites.'
)
// Will be overridden by setSchema
// Replaced per chat by schemaFor
const testRunFlowSchema = z.object({
args: z
.object({})
@@ -435,11 +435,10 @@ export const flowTools: Tool<FlowAIChatHelpers>[] = [
contextName: 'flow'
})
},
setSchema: async function (helpers: FlowAIChatHelpers) {
await buildSchemaForTool(this.def, async () => {
const flowInputsSchema = await helpers.getFlowInputsSchema()
return flowInputsSchema
})
schemaFor: async (helpers: FlowAIChatHelpers) => {
const def = structuredClone(testRunFlowToolDef)
await buildSchemaForTool(def, () => helpers.getFlowInputsSchema())
return def
},
requiresConfirmation: true,
confirmationMessage: 'Run a test of the current flow',
@@ -7655,9 +7655,9 @@ describe('open_page workspace gating', () => {
// A workspace the user belongs to but whose `whoami` never answers.
const FLAKY = 'flaky_ws'
const openPage = () => getGlobalTool('open_page')
const pageSchema = () => (openPage().def.function.parameters as any)?.properties?.page ?? {}
let def = openPage().def
const pageSchema = () => (def.function.parameters as any)?.properties?.page ?? {}
const advertisedPages = () => (pageSchema().enum ?? []) as string[]
const pristineDef = openPage().def
beforeEach(() => {
// Admin of the workspace being browsed, plain member of the one a session operates on.
@@ -7680,16 +7680,26 @@ describe('open_page workspace gating', () => {
superadmin.set(undefined)
whoamiByWorkspace.clear()
clearWorkspaceRoleCache()
openPage().def = pristineDef
})
// Every chat shares this tool object, so a schema written back to it would reach the
// other chats' requests.
it("never writes a chat's schema back to the shared tool", async () => {
const shared = openPage().def
const snapshot = structuredClone(shared)
def = await openPage().schemaFor!({ operatingWorkspace: SESSION })
expect(def).not.toBe(shared)
expect(openPage().def).toBe(shared)
expect(shared).toEqual(snapshot)
})
// Reading the ambient `userStore` instead offers a session the pages of the workspace
// the user happens to be browsing.
it('gates on the operating workspace, not the one userStore describes', async () => {
await openPage().setSchema?.({ operatingWorkspace: NAV })
def = await openPage().schemaFor!({ operatingWorkspace: NAV })
expect(advertisedPages()).toContain('workspace_settings')
await openPage().setSchema?.({ operatingWorkspace: SESSION })
def = await openPage().schemaFor!({ operatingWorkspace: SESSION })
expect(advertisedPages()).toContain('runs')
expect(advertisedPages()).not.toContain('workspace_settings')
await expect(
@@ -7704,7 +7714,7 @@ describe('open_page workspace gating', () => {
// Neither layer may read as a denial: the role was never established, and the model
// sees the schema before it can ever reach the handler's message.
it('advertises nothing and blames no denial when the role lookup fails', async () => {
await openPage().setSchema?.({ operatingWorkspace: FLAKY })
def = await openPage().schemaFor!({ operatingWorkspace: FLAKY })
expect(advertisedPages()).toEqual([])
expect(pageSchema().description).toContain("couldn't be checked")
const refusal = await callGlobalTool('open_page', { page: 'runs' }, toolCallbacks, {
@@ -7717,7 +7727,7 @@ describe('open_page workspace gating', () => {
// A workspace absent from `userWorkspaces` is settled, not unknown: inviting a retry
// would be false, and asking `whoami` at all only earns a 401 on every iteration.
it('reports a plain denial for a workspace the user is not a member of', async () => {
await openPage().setSchema?.({ operatingWorkspace: 'unreachable_ws' })
def = await openPage().schemaFor!({ operatingWorkspace: 'unreachable_ws' })
expect(advertisedPages()).toEqual([])
expect(pageSchema().description).not.toContain("couldn't be checked")
const refusal = await callGlobalTool('open_page', { page: 'runs' }, toolCallbacks, {
@@ -7734,7 +7744,7 @@ describe('open_page workspace gating', () => {
it('asks whoami while the workspace list is still unresolved', async () => {
usersWorkspaceStore.set(undefined)
await openPage().setSchema?.({ operatingWorkspace: SESSION })
def = await openPage().schemaFor!({ operatingWorkspace: SESSION })
expect(UserService.whoami).toHaveBeenCalledWith(expect.objectContaining({ workspace: SESSION }))
expect(advertisedPages()).not.toEqual([])
@@ -135,7 +135,6 @@ import {
type CreatedResourceTriggerKind,
type PreviewCardKind,
type RunFormDisplay,
type Tool,
type ToolCallbacks,
type ToolCodeDiff,
type ToolDisplayAction
@@ -197,6 +196,15 @@ import { UserDraftDbSyncer } from '$lib/userDraftDbSyncer.svelte'
import { invalidateWorkspaceComparison } from '$lib/workspaceComparison'
import type { UserDraftItemKind } from '$lib/gen'
import { bundleRawAppDraft } from './rawAppBundlerBridge'
import {
DEPLOY,
NONE,
RUN_PREVIEW,
WRITE_DRAFT,
type SessionAccess,
type SessionTool,
type SessionToolPolicy
} from '../sessionCapabilities'
import {
buildRunsUrl,
buildSchedulesUrl,
@@ -770,7 +778,7 @@ const deployWorkspaceItemSchema = z.object({
.boolean()
.optional()
.describe(
'Deploy even if the draft was started from an older deployed version, overwriting the version deployed since. Defaults to false; prefer calling rebase_draft first to keep the newer changes.'
'Deploy even if the draft was started from an older deployed version, overwriting the version deployed since. Defaults to false.'
)
})
@@ -930,7 +938,7 @@ const runScriptSchema = z.object({
const runScriptToolDef = createToolDef(
runScriptSchema,
'run_script',
'Run a DEPLOYED script for real, under the user\'s own permissions. Fill in every argument you can infer: the user gets an argument form prefilled with `args` and decides what runs. For a secret argument prefer `$var:<path>` naming an existing workspace variable; a literal is minted into a short-lived secret before the run, but stays in this call. A required file is the user\'s to attach, so call this even when you cannot supply one rather than asking in chat. Use only when the user names the deployed version ("the deployed X", "in production", "for real"); otherwise use test_run_script.',
'Run a DEPLOYED script for real, under the user\'s own permissions. Fill in every argument you can infer: the user gets an argument form prefilled with `args` and decides what runs. For a secret argument prefer `$var:<path>` naming an existing workspace variable; a literal is minted into a short-lived secret before the run, but stays in this call. A required file is the user\'s to attach, so call this even when you cannot supply one rather than asking in chat. Use only when the user names the deployed version ("the deployed X", "in production", "for real").',
{ strict: false }
)
@@ -969,7 +977,7 @@ const runFlowSchema = z.object({
const runFlowToolDef = createToolDef(
runFlowSchema,
'run_flow',
'Run a DEPLOYED flow for real, under the user\'s own permissions. Fill in every argument you can infer: the user gets an argument form prefilled with `args` and decides what runs. For a secret argument prefer `$var:<path>` naming an existing workspace variable; a literal is minted into a short-lived secret before the run, but stays in this call. A required file is the user\'s to attach, so call this even when you cannot supply one rather than asking in chat. Use only when the user names the deployed version ("the deployed X", "in production", "for real"); otherwise use test_run_flow.',
'Run a DEPLOYED flow for real, under the user\'s own permissions. Fill in every argument you can infer: the user gets an argument form prefilled with `args` and decides what runs. For a secret argument prefer `$var:<path>` naming an existing workspace variable; a literal is minted into a short-lived secret before the run, but stays in this call. A required file is the user\'s to attach, so call this even when you cannot supply one rather than asking in chat. Use only when the user names the deployed version ("the deployed X", "in production", "for real").',
{ strict: false }
)
@@ -1139,7 +1147,7 @@ const openPreviewSchema = z.object({
kind: z
.enum(['script', 'flow', 'raw_app', 'pipeline'])
.describe(
'Item kind to preview. Use "raw_app" for code-based apps (created via init_app). Use "pipeline" to show the data-pipeline graph for a folder — here `path` is the folder name, not an item path. The legacy drag-and-drop app builder ("app") is not previewable in the session panel — don\'t pass it.'
'Item kind to preview. Use "raw_app" for code-based apps. Use "pipeline" to show the data-pipeline graph for a folder — here `path` is the folder name, not an item path. The legacy drag-and-drop app builder ("app") is not previewable in the session panel — don\'t pass it.'
),
path: z
.string()
@@ -1272,7 +1280,11 @@ type FolderPromptContext = { folders?: string[]; foldersRead?: string[]; isAdmin
// non-exhaustive hint alongside permission-agnostic guidance (the complete set
// needs a folder-listing tool — follow-up).
// Capped so a folder-heavy workspace can't dominate the prompt.
function buildFolderGuidance(username: string, ctx?: FolderPromptContext): string {
function buildFolderGuidance(
username: string,
ctx: FolderPromptContext | undefined,
canCreateFolder: boolean
): string {
if (!ctx) return ''
const MAX = 40
const writable = ctx.folders ?? []
@@ -1288,7 +1300,7 @@ function buildFolderGuidance(username: string, ctx?: FolderPromptContext): strin
writable.length > 0
? ` Folders here include ${fmt(writable)} (you can also write to others not listed).`
: ''
return `- As a workspace admin you can write to any existing folder.${known} If the user names a folder, use it; if they explicitly ask for a new folder, create it with \`create_folder\`; otherwise ask them which folder to use rather than guessing or creating one unprompted.`
return `- As a workspace admin you can write to any existing folder.${known} If the user names a folder, use it;${canCreateFolder ? ' if they explicitly ask for a new folder, create it with `create_folder`;' : ''} otherwise ask them which folder to use rather than guessing${canCreateFolder ? ' or creating one unprompted' : ''}.`
}
// Everything below states the writable set as fact, including the empty case, so an
// unresolved role has to say nothing at all: "you have no shared folders" is a claim,
@@ -1299,11 +1311,15 @@ function buildFolderGuidance(username: string, ctx?: FolderPromptContext): strin
const lines: string[] = []
if (writable.length > 0) {
lines.push(
`- Folders you can write to in this workspace: ${fmt(writable)}. For shared/team work, pick the one whose purpose matches the request; if none clearly fits, ask which folder to use (askUserQuestion) rather than inventing a path. Use \`create_folder\` only when the user explicitly asks for a new folder.`
`- Folders you can write to in this workspace: ${fmt(writable)}. For shared/team work, pick the one whose purpose matches the request; if none clearly fits, ask which folder to use (askUserQuestion) rather than inventing a path.${canCreateFolder ? ' Use `create_folder` only when the user explicitly asks for a new folder.' : ''}`
)
} else {
lines.push(
`- You have no shared folders you can write to in this workspace, so use \`u/${username}/<name>\`. If the user explicitly asks for a shared folder, create one with \`create_folder\` (you become an owner); otherwise ask before placing shared work rather than inventing an \`f/<folder>/...\` path.`
`- You have no shared folders you can write to in this workspace, so use \`u/${username}/<name>\`. ${
canCreateFolder
? 'If the user explicitly asks for a shared folder, create one with `create_folder` (you become an owner); otherwise ask'
: 'If the user explicitly asks for a shared folder, say plainly that you cannot create one here; otherwise ask'
} before placing shared work rather than inventing an \`f/<folder>/...\` path.`
)
}
if (readOnly.length > 0) {
@@ -1319,9 +1335,19 @@ const buildGlobalSystemPrompt = (
previewTools: boolean,
folderCtx?: FolderPromptContext,
skills: AiSkillListItem[] = [],
mcpServers: McpServer[] = []
mcpServers: McpServer[] = [],
access?: SessionAccess
) => {
const folderGuidance = buildFolderGuidance(username, folderCtx)
// Each `can*` mirrors the capability the matching tools declare in `requires`, so a
// rule cannot outlive the tool it describes. An unresolved profile keeps every block.
const canWriteDraft = !access || access.has('write_draft')
const canRunPreview = !access || access.has('run_preview')
// The deploy tools take their kind as an argument, so `deploy` gates only the folder text.
const canCreateFolder = !access || access.has('deploy')
// Each gated block carries its own leading newline, so dropping one leaves no blank
// line behind and a full-access prompt is byte-for-byte the ungated text.
const when = (cond: boolean, block: string) => (cond ? block : '')
const folderGuidance = buildFolderGuidance(username, folderCtx, canCreateFolder)
const folderGuidanceBlock = folderGuidance ? `\n${folderGuidance}` : ''
// `previewTools` doubles as "this is a session chat" — sessions are the only
// chats that get the preview tool set. The alpha heads-up only makes sense
@@ -1353,48 +1379,89 @@ const buildGlobalSystemPrompt = (
The current user's workspace username is "${username}".${instanceLine}
Use tools to inspect workspace items and create per-user drafts (saved server-side, visible only to this user — not deployed) for scripts, flows, schedules, triggers, resources, variables, and raw apps.
${
canWriteDraft
? 'Use tools to inspect workspace items and create per-user drafts (saved server-side, visible only to this user — not deployed) for scripts, flows, schedules, triggers, resources, variables, and raw apps.'
: "Use tools to inspect workspace items and the workspace's run history, and to run items that are already deployed. You cannot create or edit scripts, flows, apps, schedules, triggers, resources or variables here — this user's role does not allow it — so when they ask for such a change, say plainly that you cannot make it rather than describing steps as if you had. Their role is refused scripts, flows and apps outside this chat too, so for those suggest asking a workspace admin rather than creating them in the editor."
}${when(
canWriteDraft,
`
Path conventions:
- A workspace path starts with one of two namespaces; its trailing <name> may itself contain "/", so a path has three or more segments:
- \`u/${username}/<name>\` — your personal scope. Default for ad-hoc, exploratory, or scratch work.
- \`f/<folder>/<name>\` — a shared folder scope; the <folder> must already exist (a bare \`f/<name>\` with no folder segment is INVALID and will fail).
- If the user supplies a fully qualified \`f/<folder>/...\` path, use that exact path; they have already chosen the folder. Do not ask for folder confirmation or substitute a \`u/${username}/...\` path unless a tool rejects it.
- Default a bare name with no namespace prefix (e.g. "create a flow called myflow") to \`u/${username}/<name>\`. Never invent an \`f/<folder>/...\` path for a folder that does not exist; create one with \`create_folder\` only when the user explicitly asks for a new folder.${folderGuidanceBlock}
- Default a bare name with no namespace prefix (e.g. "create a flow called myflow") to \`u/${username}/<name>\`. Never invent an \`f/<folder>/...\` path for a folder that does not exist${when(
canCreateFolder,
'; create one with `create_folder` only when the user explicitly asks for a new folder'
)}.${folderGuidanceBlock}`
)}${when(
canCreateFolder && !canWriteDraft,
'\n- You can create a shared folder with `create_folder` when the user explicitly asks for one. You cannot create anything inside it here, so do not offer to.'
)}
Rules:
- Draft tools create or update drafts only; they do not deploy or mutate deployed workspace items.
Rules:${when(
canWriteDraft,
`
- Draft tools create or update drafts only; they do not deploy or mutate deployed workspace items.`
)}
- Use list_workspace_items to find items and read_workspace_item before changing an existing item. For triggers, pass trigger_kind.
- If the user message includes an ACTIVE EDITOR section, treat it as the currently open item and use it for references like "this", "current", or "open editor".${activePreviewRule}
- If the user message includes an ACTIVE EDITOR section, treat it as the currently open item and use it for references like "this", "current", or "open editor".${activePreviewRule}${when(
canWriteDraft,
`
- Use deploy_workspace_item only after the user explicitly asks to deploy. It persists a draft to the workspace.
- To undo something you created or changed in this chat, use discard_local_draft: everything you write is a draft until it is explicitly deployed, so "delete it" / "never mind" / "remove that" about your own work means discarding the draft (it also clears the matching open editor draft). Use delete_workspace_item only to remove an item that is already deployed in the workspace; it mutates the workspace and fails if nothing is deployed at that path.
- Use diff to review changes — before deploying, or when the user asks what changed. It is read-only: without arguments it lists every draft in the workspace with its change status; with type+path it returns that item's unified diff (for multi-file apps, pass file to read one file's diff). In a fork, pass against="parent_workspace" to compare the deployed fork with its parent workspace instead. Pass search to grep changed lines across all diffs.
- To undo something you created or changed in this chat, use discard_local_draft: everything you write is a draft until it is explicitly deployed, so "delete it" / "never mind" / "remove that" about your own work means discarding the draft (it also clears the matching open editor draft). Use delete_workspace_item only to remove an item that is already deployed in the workspace; it mutates the workspace and fails if nothing is deployed at that path.`
)}${when(
!canWriteDraft,
`
- Three changes are still open to you where the server allows them: discard_local_draft drops a draft this user left behind, deploy_workspace_item deploys one when the user asks, and delete_workspace_item removes an item already deployed in the workspace. You cannot create or edit one.`
)}
- Use diff to review changes — before deploying, or when the user asks what changed. It is read-only: without arguments it lists every draft in the workspace with its change status; with type+path it returns that item's unified diff (for multi-file apps, pass file to read one file's diff). In a fork, pass against="parent_workspace" to compare the deployed fork with its parent workspace instead. Pass search to grep changed lines across all diffs.${when(
canWriteDraft,
`
- You can never read a variable's value, secret or not, so never invent one: when editing an existing variable, omit value (and is_secret) from write_variable and pass only the fields you are actually changing. The user can reveal a value in the variable editor; you cannot, so never tell them a value is unreadable in general. "$var:path/to/variable" is how a resource value references a variable — it is never a variable's own value.
- Use search_resource_types before write_resource, and get_trigger_schema before write_trigger: the trigger config fields differ per kind and are not listed in the write_trigger definition.
- When script or raw app code needs an external npm package you are not fully familiar with, use search_npm_packages to find it and get its documentation and type definitions. Link the package documentation in your answer when you rely on it.
- Hub scripts are prebuilt integrations for third-party services, hosted outside the workspace under \`hub/<version>/<app>/<name>\` paths. Use search_hub_scripts to find one before hand-writing an integration, then read_workspace_item with type "script" and the returned hub path to get its code, language, and input schema.
- Use get_db_schema with a database resource path to fetch its tables and columns before writing SQL (or a script querying that database).
- Use get_instructions before writing scripts, flows, resources, or apps. For scripts, pass the target language.
${pipelineBullet}
${when(canRunPreview, '- Use get_db_schema with a database resource path to fetch its tables and columns before writing SQL (or a script querying that database).\n')}- Use get_instructions before writing scripts, flows, resources, or apps. For scripts, pass the target language.
${pipelineBullet}`
)}${when(
canRunPreview && canWriteDraft,
`
- After creating or editing a script or flow draft, run test_run_script, test_run_flow, or test_run_step with representative args before reporting that it works. These tools prefer drafts, so testing does not require deployment.
- Do the same for a raw app: run test_run_app_runnable on each backend runnable you wrote or changed before saying the app works. A bundle that compiles proves nothing about whether the runnables run. An inline runnable executes the app's draft code; a path runnable executes the DEPLOYED script/flow it names, so a path runnable aimed at something you have not deployed fails here — that failure is the point: report it and offer to deploy that one target. The app itself does not need deploying to be tested.
- Do the same for a raw app: run test_run_app_runnable on each backend runnable you wrote or changed before saying the app works. A bundle that compiles proves nothing about whether the runnables run. An inline runnable executes the app's draft code; a path runnable executes the DEPLOYED script/flow it names, so a path runnable aimed at something you have not deployed fails here — that failure is the point: report it and offer to deploy that one target. The app itself does not need deploying to be tested.`
)}
- Use list_runs to find recent runs (optionally filtered by path, creator, label, or status), then get_run with a returned id to see what that run was called with, what it returned and what it logged — without starting a new test run.
- get_run also covers what a flow run did per step — statuses and results across the whole execution tree, subflow steps and loop iterations included — and works while the flow is still running. Pass step to read one step's result in full (capped at 12k chars).
- Use open_page to show a workspace page with filters applied — Runs, Schedules, Variables, Resources, Assets, Audit logs, or Workspace settings on a specific tab (e.g. "open the failed runs of f/foo/bar", "open the schedule for X", "open the git sync settings"). Carry over every filter the user described — Runs takes the page's whole filter set (time window, path, user, folder, label, tag, worker, trigger kind, args/result, ...), so don't drop a criterion just because it wasn't in the request's main clause. Only the pages listed for this user in the tool are available; don't offer pages that aren't listed. Don't use it as a substitute for list_runs when you just need the data yourself.
- Whenever you ask the user to perform a manual step in the UI — fill in a resource's credentials, set a secret variable's value, adjust a schedule or setting — call open_page in the same message, targeted at that item (pass open with its path to land in its editor, or the page's filters otherwise). Never just describe where to click.
- Whenever you ask the user to perform a manual step in the UI — fill in a resource's credentials, set a secret variable's value, adjust a schedule or setting — call open_page in the same message, targeted at that item (pass open with its path to land in its editor, or the page's filters otherwise). Never just describe where to click.${when(
canWriteDraft,
`
- When the user is happy with the changes and wants to review or deploy them, use open_page with page "compare" — it opens the Compare & Deploy review page.${
previewTools
? ' By default it preselects the items this chat modified; pass items ("<kind>:<path>" entries) to control the selection'
: ' Pass items ("<kind>:<path>" entries naming the items you changed) so the review is scoped to them — omitting items preselects every pending change in the workspace'
}, or mode ("draft" or "fork") to force which comparison is shown. Prefer offering this review page over calling deploy_workspace_item directly when several items changed.
- Default to test_run_script, test_run_flow, or test_run_step for any run request, an existing script included; they prefer drafts and need no deployment. Use run_script or run_flow only when the user names the deployed version ("the deployed X", "in production", "for real") — a bare "run X" is not that. For those two, read the item with read_workspace_item version: "deployed" first so the arguments match the deployed schema. test_run_script, test_run_flow, test_run_step, run_script and run_flow all show the user an argument form prefilled with what you sent, so fill in every argument you can infer rather than asking for it in chat. test_run_step's form is the step's own inputs, not the flow's.
previewTools
? ' By default it preselects the items this chat modified; pass items ("<kind>:<path>" entries) to control the selection'
: ' Pass items ("<kind>:<path>" entries naming the items you changed) so the review is scoped to them — omitting items preselects every pending change in the workspace'
}, or mode ("draft" or "fork") to force which comparison is shown. Prefer offering this review page over calling deploy_workspace_item directly when several items changed.`
)}${when(
canRunPreview,
`
- Default to test_run_script, test_run_flow, or test_run_step for any run request, an existing script included; they prefer drafts and need no deployment. Use run_script or run_flow only when the user names the deployed version ("the deployed X", "in production", "for real") — a bare "run X" is not that. For those two, read the item with read_workspace_item version: "deployed" first so the arguments match the deployed schema. test_run_script, test_run_flow, test_run_step, run_script and run_flow all show the user an argument form prefilled with what you sent, so fill in every argument you can infer rather than asking for it in chat. test_run_step's form is the step's own inputs, not the flow's.`
)}${when(
!canRunPreview,
`
- run_script and run_flow are how you run anything here: they run the DEPLOYED item under this user's own permissions, and show them an argument form prefilled with what you sent. Read the item with read_workspace_item version: "deployed" first so the arguments match the deployed schema, and fill in every argument you can infer rather than asking for it in chat.`
)}
- When a required decision is ambiguous, use askUserQuestion with two to ten clear proposed answer strings instead of guessing. The user can also type a custom answer when none of the proposed answers fit. Set multiSelect: true only when the answers can genuinely co-apply and the user may pick several (not mutually exclusive).
- When the user asks you to remember a lasting preference, always/never do something, or change/stop a behavior going forward, call update_user_instructions to persist it. It edits only the USER INSTRUCTIONS block (not WORKSPACE INSTRUCTIONS). Keep each instruction concise; do not use it for one-off requests scoped to the current task.
- Keep context targeted.${
previewTools
? `
? `${when(
canWriteDraft,
`
- After writing or substantially editing a script / flow / app draft, show it via open_preview(kind, path) so the user sees the editor and live preview right next to the chat. First check whether it is already shown: if unsure, call get_preview_status. Only call open_preview (or offer to) when no preview is open or it is showing a different item — don't re-open a preview already showing the item you just edited.
- Building a data pipeline: call open_preview(kind="pipeline", path="<folder>") as the FIRST step, before creating any node — this opens the pipeline editor the user reviews in. path is the folder, not an item; an empty or not-yet-created folder is fine (create_folder first if needed, then open it). Opening it registers build_pipeline_node / edit_pipeline_node — use ONLY those to add or change pipeline nodes, never write_script for a pipeline node — they apply directly as unsaved drafts on the canvas (no separate accept/reject step) that the user reviews and deploys. Do not write pipeline scripts without first opening the editor.
- Building a data pipeline: call open_preview(kind="pipeline", path="<folder>") as the FIRST step, before creating any node — this opens the pipeline editor the user reviews in. path is the folder, not an item; an empty ${when(canCreateFolder, 'or not-yet-created ')}folder is fine${when(canCreateFolder, ' (create_folder first if needed, then open it)')}. Opening it registers build_pipeline_node / edit_pipeline_node — use ONLY those to add or change pipeline nodes, never write_script for a pipeline node — they apply directly as unsaved drafts on the canvas (no separate accept/reject step) that the user reviews and deploys. Do not write pipeline scripts without first opening the editor.`
)}
- When debugging a running raw app, call get_app_runtime_logs to read the live preview's browser console output. It needs the raw app preview open (open_preview kind="raw_app").
- To inspect what actually rendered in a running raw app (verify an edit landed on screen, diagnose a blank/empty or wrong view, answer "what's showing"), use search_dom (regex over the live HTML) and read_dom (a line-numbered window). Pass a \`selector\` to scope to an element — prefer the selector from a DOM element chip the user attached — or omit it for the whole page. When a chip lists an \`app_path\`, pass it too so the RIGHT app is read (several previews can be open; a query without \`app_path\` hits the visible one). The DOM is read live and is never in context; no match means the element isn't rendered. Both need the raw app preview open.
- get_app_runtime_logs only shows the app's browser console. For the server-side logs of a backend runnable the app invoked (a backend.<id> call), call list_app_runs to get that run's job_id from the live preview, then get_run with it. Use this when a backend call errors or returns something unexpected.
@@ -1418,29 +1485,41 @@ Documentation:
Flows:
- read_workspace_item returns compact flow JSON. Inline script bodies appear as "inline_script.<moduleId>".
- Use read_flow_module_code and set_flow_module_code for inline script bodies.
- Use patch_flow_json for structural flow edits and write_flow for full flow rewrites.
- Use read_flow_module_code${when(canWriteDraft, ' and set_flow_module_code')} for inline script bodies.${when(
canWriteDraft,
`
- Use patch_flow_json for structural flow edits and write_flow for full flow rewrites.`
)}
Raw apps:
- The app tools below only work on raw (code) apps. \`rawApp\` says which: false is a drag-and-drop app: you can list it and read its metadata, but not read its contents, edit it or deploy it. Check it before offering to change an app.
- read_workspace_item returns app metadata only. Use read_app_file for file and inline runnable contents.
- read_workspace_item returns app metadata only. Use read_app_file for file and inline runnable contents.${when(
canWriteDraft,
`
- A draft app is reachable by nobody; deploying is what exposes its backend runnables. deploy_workspace_item says so when the deploy widens who may open the app: anonymous means anyone with the URL, without logging in; guest means anyone the instance's identity provider authenticates, member of this workspace or not. Relay that in plain words and carry on. This is disclosure, not a gate: do not stop and ask for permission, and do not refuse the deploy. You cannot change who may open an app from chat; it is set on the app's deploy settings.
- Use write_app_file, patch_app_file, and delete_app_file for frontend files.
- Use write_app_runnable and delete_app_runnable for backend runnables.
- Use init_app only after confirming framework, path, and summary with the user.
- Use deploy_workspace_item after explicit user deploy intent; raw app deploy bundles JS/CSS before saving.
- Use deploy_workspace_item after explicit user deploy intent; raw app deploy bundles JS/CSS before saving.`
)}
Data Tables:
- Datatables are workspace-scoped managed PostgreSQL databases, shared across the workspace (not owned by any single app). They must be configured by the user in their workspace settings (Workspace settings → Data Tables); they cannot be created via SQL.
- Use list_datatables to discover the available datatables and their tables. Reuse an existing table rather than creating a duplicate. If list_datatables reports none, this is a blocking prerequisite — tell the user to set up a datatable in their workspace settings and stop; do not assume a "main" datatable exists or call exec_datatable_sql.
- Use get_datatable_table_schema only when you need a table's column names/types; list_datatables is enough for table-list or availability summaries.
- Use exec_datatable_sql to explore data, run queries, mutate rows, or change schema (CREATE/ALTER/DROP). Creating a table is a normal CREATE TABLE statement — it appears in list_datatables afterward, with no registration step.${
- Use list_datatables to discover the available datatables and their tables. Reuse an existing table rather than creating a duplicate. If list_datatables reports none, this is a blocking prerequisite — tell the user to set up a datatable in their workspace settings and stop; do not assume a "main" datatable exists${when(canRunPreview, ' or call exec_datatable_sql')}.
- Use get_datatable_table_schema only when you need a table's column names/types; list_datatables is enough for table-list or availability summaries.${when(
canRunPreview,
`
- Use exec_datatable_sql to explore data, run queries, mutate rows, or change schema (CREATE/ALTER/DROP). Creating a table is a normal CREATE TABLE statement — it appears in list_datatables afterward, with no registration step.`
)}${
isCloudHosted()
? ''
: `
- A raw app may use a datatable through a role (\`data.roles\` in its raw_app.yaml). When working on such an app, pass that role to the datatable tools, and to wmill.datatable in its runnables, so you see and change only what the app itself can.`
}
- When writing runnable code (inline app runnables, scripts, flow modules) that reads or writes datatable data at runtime, it accesses a datatable via wmill.datatable(). Default to TypeScript (bun) unless the user asked for another language. Call get_instructions with subject "datatable" and language "bun" for the TypeScript SQL SDK reference (or language "python3" for Python) — it returns only that language so you get just what you need.${
}${when(
canWriteDraft,
`
- When writing runnable code (inline app runnables, scripts, flow modules) that reads or writes datatable data at runtime, it accesses a datatable via wmill.datatable(). Default to TypeScript (bun) unless the user asked for another language. Call get_instructions with subject "datatable" and language "bun" for the TypeScript SQL SDK reference (or language "python3" for Python) — it returns only that language so you get just what you need.`
)}${
skills.length > 0
? `
@@ -2337,7 +2416,7 @@ function getDatatableInstructions(language?: ScriptLang): string {
const lang = language ?? 'bun'
return `# Datatable SQL SDK reference
Datatables are workspace-scoped managed PostgreSQL databases. In chat, explore and shape them with the \`list_datatables\`, \`get_datatable_table_schema\`, and \`exec_datatable_sql\` tools. The reference below is for code you author inside runnables (inline app runnables, scripts, or flow rawscript modules) that reads or writes datatable data at runtime.
Datatables are workspace-scoped managed PostgreSQL databases. In chat, explore and shape them with the datatable tools you were given. The reference below is for code you author inside runnables (inline app runnables, scripts, or flow rawscript modules) that reads or writes datatable data at runtime.
- A runnable accesses a datatable via \`wmill.datatable()\` (the default "main") or \`wmill.datatable('<name>')\`, referencing tables as \`schema.table\`.
- Use parameterized queries (the tagged template in TypeScript, \`$1\`/\`$2\` placeholders in Python) — never interpolate untrusted values into SQL strings.
@@ -2400,12 +2479,23 @@ export type SessionPromptContext = {
/** Session-state guidance appended to the global system prompt so the model
* knows where its work lands (staged fork vs the live workspace). */
export function getSessionContextPromptSection(ctx: SessionPromptContext): string {
export function getSessionContextPromptSection(
ctx: SessionPromptContext,
access?: SessionAccess
): string {
// Concatenated onto an already capability-gated prompt, so it has to honour the same
// profile rather than assume the gating happened upstream.
const canDeploy = !access || access.has('deploy')
const canWriteDraft = !access || access.has('write_draft')
const canRunPreview = !access || access.has('run_preview')
const targets = ['reads', canWriteDraft && 'drafts', canRunPreview && 'test runs', 'deploys']
.filter(Boolean)
.join(', ')
const lines = [
'',
'',
'Session state:',
'- This chat is a Windmill AI session with its own operating workspace: every tool call (reads, drafts, test runs, deploys) targets that workspace.'
`- This chat is a Windmill AI session with its own operating workspace: every tool call (${targets}) targets that workspace.`
]
if (ctx.pendingForkOf) {
lines.push(
@@ -2432,6 +2522,13 @@ export function getSessionContextPromptSection(ctx: SessionPromptContext): strin
'- No operating workspace is set yet; the user picks one (or a new staged fork) before the first message is sent.'
)
}
// The kind enum already withholds what the rules refuse; this says why, and where the user
// promotes the rest instead. The rules gate deletes too, so it is owed to every such profile.
if (!canDeploy) {
lines.push(
"- This workspace refuses direct deployment for this user, except for schedules and triggers — those are the only kinds deploy_workspace_item and delete_workspace_item can still act on. Scripts, flows, apps, resources and variables must be promoted from the session's deploy panel (fork or pull request); do not offer to deploy or delete them directly."
)
}
return lines.join('\n')
}
@@ -2478,7 +2575,8 @@ const readSkillSchema = z.object({
.describe('The exact skill resource path as listed in the Skills section of the system prompt.')
})
export const readSkillTool: Tool<{}> = {
export const readSkillTool: SessionTool<{}> = {
requires: NONE,
def: createToolDef(
readSkillSchema,
'read_skill',
@@ -2730,7 +2828,7 @@ const RUNS_TIMEFRAME_LABELS = runsTimeframes.map((tf) => tf.label) as [string, .
// input_schema requires; a top-level oneOf would be rejected. Each per-page URL builder
// drops any key that isn't one of its page's real query params, so a field that doesn't
// apply to the chosen page is harmless. This full schema is used to PARSE tool args; the
// advertised schema (what the model sees) is narrowed per-user in `setSchema`.
// advertised schema (what the model sees) is narrowed per-user in `schemaFor`.
const openPageFullSchema = z.object({
page: z.enum(OPEN_PAGE_NAMES).describe('Which page to open'),
path: z
@@ -3177,8 +3275,9 @@ function summarizeOpenPage(url: string, page: OpenPageName): string {
return parts.length ? parts.join(', ') : `all ${OPEN_PAGE_LABELS[page].toLowerCase()}`
}
export const openPageTool: Tool<{}> = {
// The initial def assumes an untracked chat and no resolved role; setSchema below
export const openPageTool: SessionTool<{}> = {
requires: NONE,
// The initial def assumes an untracked chat and no resolved role; schemaFor below
// rebuilds it with the caller's real surface before each iteration.
def: createToolDef(
buildOpenPageDefSchema(
@@ -3197,9 +3296,9 @@ export const openPageTool: Tool<{}> = {
autoCollapseDetails: false,
// Re-narrow the advertised `page` enum to this user's permissions each iteration, so
// the model never sees (or suggests) a page the user can't reach.
setSchema: async function (helpers) {
schemaFor: async (helpers) => {
const access = await allowedOpenPages(operatingWorkspaceFromHelpers(helpers))
this.def = createToolDef(
return createToolDef(
buildOpenPageDefSchema(
access.pages,
allowedTriggerKinds(),
@@ -3282,10 +3381,11 @@ export const openPageTool: Tool<{}> = {
}
}
export const globalTools: Tool<{}>[] = [
export const globalTools: SessionTool<{}>[] = [
readSkillTool,
openPageTool,
{
requires: NONE,
def: createToolDef(
getInstructionsSchema,
'get_instructions',
@@ -3295,6 +3395,15 @@ export const globalTools: Tool<{}>[] = [
fn: async (ctx) => {
const { args, toolId, toolCallbacks } = ctx
const parsed = getInstructionsSchema.parse(args)
// Every subject is authoring guidance written around the draft tools by name, and a
// tool result is the one place the toolset filter cannot reach.
const access = (ctx.helpers as GlobalToolHelpers | undefined)?.access
if (access && !access.has('write_draft')) {
const message =
'This session cannot create or edit workspace items, so there is no authoring guidance to give. Tell the user plainly rather than describing how it would be done.'
toolCallbacks.setToolStatus(toolId, { content: 'No authoring guidance' })
return message
}
const label =
parsed.subject === 'script' && parsed.language
? `${parsed.subject} (${parsed.language})`
@@ -3308,6 +3417,7 @@ export const globalTools: Tool<{}>[] = [
searchDocsTool,
readDocsPageTool,
{
requires: NONE,
def: createToolDef(
askUserQuestionSchema,
'askUserQuestion',
@@ -3376,6 +3486,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
updateUserInstructionsSchema,
'update_user_instructions',
@@ -3449,6 +3560,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
listWorkspaceItemsSchema,
'list_workspace_items',
@@ -3513,6 +3625,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
readWorkspaceItemSchema,
'read_workspace_item',
@@ -3562,6 +3675,8 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// folders.rs `create_folder` runs `check_deploy_rules` and has no operator check.
requires: DEPLOY,
def: createToolDef(
createFolderSchema,
'create_folder',
@@ -3604,6 +3719,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(writeScriptSchema, 'write_script', 'Create or overwrite a draft script.'),
showDetails: true,
streamArguments: true,
@@ -3614,6 +3730,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(writeFlowSchema, 'write_flow', 'Create or overwrite a draft flow.'),
showDetails: true,
streamArguments: true,
@@ -3655,6 +3772,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeScheduleToolSchema,
'write_schedule',
@@ -3680,6 +3798,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeTriggerSchema,
'write_trigger',
@@ -3707,10 +3826,11 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
getTriggerSchemaSchema,
'get_trigger_schema',
'Get the configuration schema for one trigger kind. Call before write_trigger.'
'Get the configuration schema for one trigger kind — its config fields differ per kind.'
),
planModeSafe: true,
fn: async (ctx) => {
@@ -3719,15 +3839,17 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
z.object({}),
'get_schedule_schema',
"Get the shape of write_schedule's `advanced` object: retry, pausing, tags, and error-handler tuning."
"Get the shape of a schedule's `advanced` object: retry, pausing, tags, and error-handler tuning."
),
planModeSafe: true,
fn: async () => JSON.stringify(advancedScheduleShape(), null, 2)
},
{
requires: WRITE_DRAFT,
def: createToolDef(
editScriptSchema,
'edit_script',
@@ -3742,6 +3864,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
patchFlowJsonSchema,
'patch_flow_json',
@@ -3756,6 +3879,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: RUN_PREVIEW,
def: testRunScriptToolDef,
fn: async (ctx) => {
const parsed = testRunScriptSchema.parse(ctx.args)
@@ -3773,6 +3897,9 @@ export const globalTools: Tool<{}>[] = [
autoCollapseDetails: false
},
{
// Ungated, unlike the test runs: this executes the DEPLOYED item under the user's own
// permissions, the one run an operator's token allows. `run_flow` likewise.
requires: NONE,
def: runScriptToolDef,
fn: async (ctx) => {
const parsed = runScriptSchema.parse(ctx.args)
@@ -3787,6 +3914,7 @@ export const globalTools: Tool<{}>[] = [
autoCollapseDetails: false
},
{
requires: RUN_PREVIEW,
def: testRunFlowToolDef,
fn: async (ctx) => {
const parsed = testRunFlowSchema.parse(ctx.args)
@@ -3801,6 +3929,7 @@ export const globalTools: Tool<{}>[] = [
autoCollapseDetails: false
},
{
requires: NONE,
def: runFlowToolDef,
fn: async (ctx) => {
const parsed = runFlowSchema.parse(ctx.args)
@@ -3814,6 +3943,7 @@ export const globalTools: Tool<{}>[] = [
autoCollapseDetails: false
},
{
requires: RUN_PREVIEW,
def: testRunStepToolDef,
fn: async (ctx) => {
const parsed = testRunStepSchema.parse(ctx.args)
@@ -3828,6 +3958,7 @@ export const globalTools: Tool<{}>[] = [
autoCollapseDetails: false
},
{
requires: NONE,
def: createToolDef(
listRunsSchema,
'list_runs',
@@ -3857,6 +3988,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
z.object({}),
'list_workers',
@@ -3900,6 +4032,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
getRunSchema,
'get_run',
@@ -3929,6 +4062,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
cancelJobSchema,
'cancel_job',
@@ -3956,6 +4090,21 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// Ungated as a whole: the draft may predate a role change, and what the server accepts
// turns on the kind. Each kind's handler, reached through `deployDraft`'s switch:
requires: NONE,
kindRequires: {
// scripts.rs `create_script_internal`, flows.rs `create_flow`/`update_flow` and
// apps.rs `create_app_raw`/`update_app_raw` refuse operators; all run the rules.
script: ['deploy', 'manage_code'],
flow: ['deploy', 'manage_code'],
app: ['deploy', 'manage_code'],
resource: DEPLOY,
variable: DEPLOY,
// The schedule and trigger handlers check neither.
schedule: NONE,
trigger: NONE
} satisfies Record<(typeof ITEM_TYPES)[number], SessionToolPolicy>,
def: createToolDef(
deployWorkspaceItemSchema,
'deploy_workspace_item',
@@ -3972,6 +4121,8 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// Rebasing writes a fresh draft, so drafts.rs `require_can_write_path` applies.
requires: WRITE_DRAFT,
def: createToolDef(
rebaseDraftSchema,
'rebase_draft',
@@ -3986,6 +4137,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
diffSchema,
'diff',
@@ -4009,6 +4161,20 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// Ungated as a whole, like deploying. Each kind's handler, reached through
// `deleteWorkspaceItem`'s switch:
requires: NONE,
kindRequires: {
// scripts.rs `delete_script_by_path` calls `require_admin`; a non-admin archives.
script: ['admin'],
// flows.rs `delete_flow_by_path` and apps.rs `delete_app` refuse operators.
flow: ['deploy', 'manage_code'],
app: ['deploy', 'manage_code'],
resource: DEPLOY,
variable: DEPLOY,
schedule: NONE,
trigger: NONE
} satisfies Record<(typeof ITEM_TYPES)[number], SessionToolPolicy>,
def: createToolDef(
deleteWorkspaceItemSchema,
'delete_workspace_item',
@@ -4025,6 +4191,9 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// Ungated: discarding your OWN draft skips drafts.rs `require_can_write_path`, so a
// user who has LOST write access can still clean up.
requires: NONE,
def: createToolDef(
discardLocalDraftSchema,
'discard_local_draft',
@@ -4040,6 +4209,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeResourceSchema,
'write_resource',
@@ -4055,6 +4225,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeVariableSchema,
'write_variable',
@@ -4070,6 +4241,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
searchResourceTypesSchema,
'search_resource_types',
@@ -4105,6 +4277,7 @@ export const globalTools: Tool<{}>[] = [
updateEditorCache: false
}),
{
requires: NONE,
def: createToolDef(
readFlowModuleCodeSchema,
'read_flow_module_code',
@@ -4117,6 +4290,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
setFlowModuleCodeSchema,
'set_flow_module_code',
@@ -4131,6 +4305,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
initAppSchema,
'init_app',
@@ -4145,6 +4320,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
readAppFileSchema,
'read_app_file',
@@ -4157,6 +4333,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
searchAppSchema,
'search_app',
@@ -4169,6 +4346,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeAppFileSchema,
'write_app_file',
@@ -4183,6 +4361,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
deleteAppFileSchema,
'delete_app_file',
@@ -4194,6 +4373,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
patchAppFileSchema,
'patch_app_file',
@@ -4208,6 +4388,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
writeAppRunnableSchema,
'write_app_runnable',
@@ -4223,6 +4404,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: WRITE_DRAFT,
def: createToolDef(
deleteAppRunnableSchema,
'delete_app_runnable',
@@ -4234,6 +4416,9 @@ export const globalTools: Tool<{}>[] = [
}
},
{
// Reaches apps.rs `execute_component` rather than jobs.rs, but that handler refuses
// operators too once `force_viewer_static_fields` marks the call a preview.
requires: RUN_PREVIEW,
def: testRunAppRunnableToolDef,
fn: async (ctx) => {
const parsed = testRunAppRunnableSchema.parse(ctx.args)
@@ -4248,6 +4433,7 @@ export const globalTools: Tool<{}>[] = [
},
...artifactTools,
{
requires: NONE,
def: createToolDef(
openPreviewSchema,
'open_preview',
@@ -4259,6 +4445,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
getPreviewStatusSchema,
'get_preview_status',
@@ -4268,6 +4455,7 @@ export const globalTools: Tool<{}>[] = [
fn: async (ctx) => getSessionPreviewStatus(sessionIdFromCtx(ctx))
},
{
requires: NONE,
def: createToolDef(
closePageSchema,
'close_page',
@@ -4279,6 +4467,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
getRuntimeLogsSchema,
'get_app_runtime_logs',
@@ -4299,6 +4488,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
listAppRunsSchema,
'list_app_runs',
@@ -4318,6 +4508,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
searchDomSchema,
'search_dom',
@@ -4346,6 +4537,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
readDomSchema,
'read_dom',
@@ -4374,6 +4566,7 @@ export const globalTools: Tool<{}>[] = [
}
},
{
requires: NONE,
def: createToolDef(
takeScreenshotSchema,
'take_screenshot',
@@ -4454,7 +4647,7 @@ export const SESSION_PREVIEW_TOOL_NAMES = new Set([
* chat, or `globalTools` minus the preview tools for the regular global
* side-panel chat.
*/
export function globalToolsFor({ sessionPreview }: { sessionPreview: boolean }): Tool<{}>[] {
export function globalToolsFor({ sessionPreview }: { sessionPreview: boolean }): SessionTool<{}>[] {
const tools = sessionPreview
? globalTools
: globalTools.filter((t) => !SESSION_PREVIEW_TOOL_NAMES.has(t.def.function.name))
@@ -4493,6 +4686,9 @@ type WriteDraftCtx = {
export type SessionToolHelpers = { sessionId?: string }
export type GlobalToolHelpers = SessionToolHelpers & {
/** The session's capability profile. Undefined outside a session, or before the first
* send resolves it. */
access?: SessionAccess
/** Runs the flow editor mounted on `storagePath`, if one is. `memoryId` names the
* chat-mode conversation the turn belongs to. */
testActiveFlow?: (
@@ -8442,6 +8638,9 @@ export function prepareGlobalSystemMessage(
user?: GlobalPromptIdentity
skills?: AiSkillListItem[]
mcpServers?: McpServer[]
/** Capabilities of the chat's operating workspace; undefined keeps every block. Must
* match the profile the toolset was filtered with, or the prompt names withheld tools. */
access?: SessionAccess
}
): ChatCompletionSystemMessageParam {
const user = opts?.user ?? get(userStore)
@@ -8458,7 +8657,8 @@ export function prepareGlobalSystemMessage(
opts?.previewTools ?? false,
folderCtx,
opts?.skills ?? [],
opts?.mcpServers ?? []
opts?.mcpServers ?? [],
opts?.access
)
if (instructions?.workspace?.trim()) {
content = `${content}\n\nWORKSPACE INSTRUCTIONS (configured by a workspace admin, shared by everyone in this workspace — you cannot modify these):\n${instructions.workspace.trim()}`
@@ -4,7 +4,7 @@
* on another. Plan mode's decoration is not here: it is applied per request, later.
*/
import type { ChatCompletionSystemMessageParam } from 'openai/resources/index.mjs'
import type { Tool } from '../shared'
import type { SessionTool } from '../sessionCapabilities'
import {
getSessionContextPromptSection,
globalToolsFor,
@@ -29,15 +29,15 @@ export function assembleGlobalSystemMessage(
): ChatCompletionSystemMessageParam {
const systemMessage = prepareGlobalSystemMessage(instructions, opts)
if (opts.sessionContext) {
systemMessage.content += getSessionContextPromptSection(opts.sessionContext)
systemMessage.content += getSessionContextPromptSection(opts.sessionContext, opts.access)
}
if (opts.pipelineContext) {
systemMessage.content += getPipelinePromptSection(opts.pipelineContext)
systemMessage.content += getPipelinePromptSection(opts.pipelineContext, opts.access)
}
return systemMessage
}
export function assembleGlobalTools(opts: GlobalAssemblyOpts): Tool<any>[] {
export function assembleGlobalTools(opts: GlobalAssemblyOpts): SessionTool<any>[] {
return [
...globalToolsFor({ sessionPreview: opts.previewTools ?? false }),
...(opts.pipelineContext ? pipelineTools : []),
@@ -1,6 +1,7 @@
import { z } from 'zod'
import { ResourceService, type GetMcpToolsResponse } from '$lib/gen'
import { createToolDef, type Tool } from '../shared'
import { createToolDef } from '../shared'
import { NONE, type SessionTool } from '../sessionCapabilities'
import { enabledMcpPaths } from '$lib/components/mcp/enabledServers'
import { isOwnOrSharedMcpPath, MCP_LIST_PER_PAGE, mcpViewer } from '$lib/components/mcp/ownServers'
@@ -349,9 +350,10 @@ const callMcpToolSchema = z.object({
* mode and behind the user's confirmation — building both from one body keeps
* them from drifting.
*/
function createCallTool(servers: McpServer[], mode: 'read' | 'write'): Tool<{}> {
function createCallTool(servers: McpServer[], mode: 'read' | 'write'): SessionTool<{}> {
const isRead = mode === 'read'
return {
requires: NONE,
def: createToolDef(
callMcpToolSchema,
isRead ? 'call_mcp_read_tool' : 'call_mcp_write_tool',
@@ -419,12 +421,13 @@ function createCallTool(servers: McpServer[], mode: 'read' | 'write'): Tool<{}>
* are not registered at all, so a workspace without an MCP connection pays no
* per-iteration schema cost for them.
*/
export function createMcpTools(servers: McpServer[]): Tool<{}>[] {
export function createMcpTools(servers: McpServer[]): SessionTool<{}>[] {
if (servers.length === 0) return []
const serverList = servers.map((s) => s.path).join(', ')
return [
{
requires: NONE,
def: createToolDef(
searchMcpToolsSchema,
'search_mcp_tools',
@@ -0,0 +1,91 @@
import { beforeEach, describe, expect, it, vi } from 'vitest'
const { whoami, deployRules } = vi.hoisted(() => ({
whoami: vi.fn(),
deployRules: vi.fn()
}))
vi.mock('$lib/gen', () => ({ UserService: { whoami } }))
vi.mock('$lib/utils_workspace_deploy', () => ({ checkDeployRules: deployRules }))
import { resolveSessionAccess } from './sessionAccess'
import { clearWorkspaceRoleCache } from '$lib/user'
type WhoamiOverrides = { is_admin?: boolean; is_super_admin?: boolean; operator?: boolean }
function user(overrides: WhoamiOverrides) {
return {
email: 'u@windmill.dev',
username: 'u',
is_admin: false,
is_super_admin: false,
operator: false,
created_at: '',
disabled: false,
groups: [],
folders: [],
folders_read: [],
folders_owners: [],
...overrides
}
}
async function capabilitiesFor(overrides: WhoamiOverrides, workspace = 'ws') {
whoami.mockResolvedValueOnce(user(overrides))
return await resolveSessionAccess(workspace)
}
describe('resolveSessionAccess', () => {
beforeEach(() => {
vi.clearAllMocks()
deployRules.mockResolvedValue({ ok: true })
// The role memo is app-wide, so without this each case answers from the previous one.
clearWorkspaceRoleCache()
})
it('gives a developer every capability but admin', async () => {
const caps = await capabilitiesFor({})
expect([...caps].sort()).toEqual(['deploy', 'manage_code', 'run_preview', 'write_draft'])
})
// The folder, resource and variable handlers run the rules but have no operator check,
// so withholding `deploy` from an operator would be stricter than the server.
it('leaves an operator the deploy-rule capability, and nothing their token refuses', async () => {
const caps = await capabilitiesFor({ operator: true })
expect([...caps]).toEqual(['deploy'])
})
it('takes deploy from the rules alone, so a rule blocks an operator too', async () => {
deployRules.mockResolvedValue({ ok: false, refusedBy: 'DisableDirectDeployment' })
const caps = await capabilitiesFor({ operator: true })
expect([...caps]).toEqual([])
})
// Both spellings of `authed.is_admin`, which the draft path honours and the handlers
// refusing `authed.is_operator` do not — the session path never clears that flag.
it.each([{ is_admin: true }, { is_super_admin: true }])(
'lets an admin who is also an operator draft and deploy, but not preview or manage code (%o)',
async (role) => {
const caps = await capabilitiesFor({ ...role, operator: true })
expect([...caps].sort()).toEqual(['admin', 'deploy', 'write_draft'])
}
)
it('keeps drafting when a rule refuses deploy', async () => {
deployRules.mockResolvedValue({
ok: false,
reason: 'restricted to deployers',
refusedBy: 'RestrictDeployToDeployers'
})
const caps = await capabilitiesFor({})
expect(caps.has('deploy')).toBe(false)
expect(caps.has('write_draft')).toBe(true)
})
it('grants everything when the role cannot be resolved', async () => {
whoami.mockRejectedValueOnce(new Error('network'))
const access = await resolveSessionAccess('ws')
expect(access.has('write_draft')).toBe(true)
expect(access.has('deploy')).toBe(true)
})
})
@@ -0,0 +1,68 @@
import { getWorkspaceRole } from '$lib/user'
import { checkDeployRules } from '$lib/utils_workspace_deploy'
import type { SessionAccess, SessionCapability } from '../sessionCapabilities'
const ALL_CAPABILITIES: SessionCapability[] = [
'write_draft',
'run_preview',
'deploy',
'manage_code',
'admin'
]
/** Fail open, here and at every resolution failure below: blanking a toolset on a
* transient error tells a developer mid-session that they cannot author anything,
* which is worse and far less legible than the 403 they get by trying. Matches
* `checkDeployPermission`, which fails open for the same reason. */
export function fullSessionAccess(): SessionAccess {
return new Set(ALL_CAPABILITIES)
}
/** Pure, so the profiles this can produce are enumerable rather than listed by hand.
* `isAdmin` is one field because the server's `authed.is_admin` is
* `usr.is_admin || super_admin` (auth.rs), which `whoami` reports as two. */
export function capabilitiesForRole(role: {
isAdmin: boolean
operator: boolean
deployRulesPass: boolean
}): SessionAccess {
const capabilities = new Set<SessionCapability>()
// Per-capability precedence, NOT a role ladder: drafts.rs `require_can_write_path`
// returns Ok on `authed.is_admin` BEFORE its operator branch, while jobs.rs
// `run_preview_*` and the script/flow/app handlers refuse `authed.is_operator` with no
// admin escape — and on the session path that flag is never cleared for an admin.
if (role.isAdmin || !role.operator) {
capabilities.add('write_draft')
}
if (!role.operator) {
capabilities.add('run_preview')
capabilities.add('manage_code')
}
// No operator term: the rules are their own gate, and the handlers that also refuse
// operators say so through `manage_code`. Admins bypass the rules inside the check.
if (role.deployRulesPass) {
capabilities.add('deploy')
}
if (role.isAdmin) {
capabilities.add('admin')
}
return capabilities
}
export async function resolveSessionAccess(workspace: string): Promise<SessionAccess> {
// Shares the 5-minute memo with the identity the same pre-flight resolves beside this
// one, so a send costs one `whoami` rather than two, at the cost of a role changed
// elsewhere landing within that window rather than on the very next message.
// `lookup_failed` covers a rejected request AND a body too malformed to map.
const lookup = await getWorkspaceRole(workspace)
if (lookup.kind !== 'resolved') {
return fullSessionAccess()
}
const me = lookup.user
return capabilitiesForRole({
isAdmin: !!me.is_admin || !!me.is_super_admin,
operator: !!me.operator,
deployRulesPass: (await checkDeployRules(workspace, me)).ok
})
}
@@ -0,0 +1,219 @@
import { describe, expect, it, vi } from 'vitest'
// The toolset pulls in the script/flow editor tools, hence monaco. Same stand-ins as
// global/core.test.ts; nothing here executes a tool, so bare shapes are enough.
vi.mock('monaco-editor', () => ({
editor: {},
languages: {},
KeyCode: {},
Uri: { parse: (value: string) => ({ toString: () => value }) },
MarkerSeverity: { Error: 8, Warning: 4, Info: 2, Hint: 1 }
}))
vi.mock('@codingame/monaco-vscode-standalone-typescript-language-features', () => ({
getTypeScriptWorker: async () => async () => ({}),
typescriptVersion: 'test'
}))
vi.mock('@codingame/monaco-vscode-languages-service-override', () => ({ default: () => ({}) }))
vi.mock('$lib/components/vscode', () => ({}))
// `globalToolsFor` withholds take_screenshot off Blink, and node reports a `Node.js/<v>`
// user agent — so without this the description sweep below silently skips a tool that
// ships in every real Chromium session. Claim the maximal toolset instead.
vi.stubGlobal('navigator', { userAgent: 'Mozilla/5.0 Chrome/120.0.0.0' })
import { globalTools, prepareGlobalSystemMessage, type SessionPromptContext } from './core'
import { appendPlanModeInstructions } from '../planMode'
import { PlanModeController } from '../planModeController.svelte'
import { pipelineTools } from '../pipeline/core'
import { assembleGlobalSystemMessage, assembleGlobalTools } from './globalAssembly'
import { filterSessionTools, sessionToolAllowed } from './sessionToolset'
import { capabilitiesForRole, fullSessionAccess } from './sessionAccess'
import type { SessionAccess, SessionCapability, SessionTool } from '../sessionCapabilities'
const ASSEMBLY_OPTS = {
previewTools: true,
pipelineContext: { folder: 'my_pipeline', mode: 'edit', nodes: [], assets: [] },
mcpServers: [{ path: 'f/test/server' } as any]
}
/** Every tool that can reach a session toolset: what assembly builds, so a source added
* there needs no edit here, unioned with the static sources so the set cannot shrink if
* assembly stops reaching one, plus the plan tools the manager appends. */
function sessionReachableTools(): SessionTool<any>[] {
const plan = new PlanModeController({ available: true } as any).availableTools
const byName = new Map<string, SessionTool<any>>()
for (const t of [
...assembleGlobalTools(ASSEMBLY_OPTS),
...globalTools,
...pipelineTools,
...plan
]) {
byName.set(t.def.function.name, t)
}
return [...byName.values()]
}
function withheldNames(access: SessionAccess): string[] {
return sessionReachableTools()
.filter((t) => !sessionToolAllowed(t, access))
.map((t) => t.def.function.name)
}
/** The tools a session actually ships, through the same assembly production uses. */
function shippedSessionTools(access: SessionAccess) {
return filterSessionTools(assembleGlobalTools(ASSEMBLY_OPTS), access)
}
function accessWith(capabilities: SessionCapability[]): SessionAccess {
return new Set(capabilities)
}
/** The profiles a real session can hold, keyed for the test name, over every combination
* of the role facts `capabilitiesForRole` reads. */
const REACHABLE_PROFILES = [
...new Map(
[false, true].flatMap((isAdmin) =>
[false, true].flatMap((operator) =>
// `checkDeployRules` bypasses on admin, so an admin has only the one outcome.
(isAdmin ? [true] : [false, true]).map((deployRulesPass) => {
const access = capabilitiesForRole({ isAdmin, operator, deployRulesPass })
return [[...access].sort().join(',') || 'none', access] as const
})
)
)
).entries()
]
/** One per branch of getSessionContextPromptSection — each words the deploy target
* differently, so a gate fixed in one branch can still leak in another. */
const SESSION_CONTEXTS: SessionPromptContext[] = [
{ pendingForkOf: 'parent' },
{ workspaceId: 'dev', parentWorkspaceId: 'parent', isDevWorkspace: true },
{ workspaceId: 'fork', parentWorkspaceId: 'parent' },
{ workspaceId: 'fork', forkParentUnknown: true },
{ workspaceId: 'live' },
{}
]
describe('session tool policies', () => {
// Full access must be a no-op, or every existing session (and the ai_evals
// baseline measured against it) changes behaviour.
it('withholds nothing from a session with every capability', () => {
const tools = sessionReachableTools()
const filtered = filterSessionTools(tools, fullSessionAccess())
expect(filtered).toHaveLength(tools.length)
filtered.forEach((tool, i) => expect(tool).toBe(tools[i]))
})
it('withholds a tool that arrives without a policy', () => {
expect(sessionToolAllowed({ def: globalTools[0].def }, fullSessionAccess())).toBe(false)
})
it('passes the toolset through untouched when access is unresolved', () => {
const tools = globalTools.map((t) => ({ def: t.def }))
expect(filterSessionTools(tools, undefined)).toBe(tools)
})
// The kind is an argument, so its enum carries the permission, per the handler each kind
// reaches. Asserted on the shipped schema, since that is all the model sees.
it('offers the kind-taking tools only the kinds their handlers accept', () => {
const shipped = (name: string, access: SessionAccess) =>
shippedSessionTools(access).find((t) => t.def.function.name === name)!
const kindsOf = (name: string, access: SessionAccess) =>
(shipped(name, access).def.function.parameters as any).properties.type.enum
// A developer a protection rule refuses keeps the two kinds no rule reaches.
const refused = accessWith(['write_draft', 'run_preview', 'manage_code'])
expect(kindsOf('deploy_workspace_item', refused)).toEqual(['schedule', 'trigger'])
expect(kindsOf('delete_workspace_item', refused)).toEqual(['schedule', 'trigger'])
// An operator where no rule applies: only the code handlers refuse them.
const operator = accessWith(['deploy'])
expect(kindsOf('deploy_workspace_item', operator)).toEqual([
'schedule',
'trigger',
'resource',
'variable'
])
// Deleting a script is admin-only, where deploying one is not.
const developer = accessWith(['write_draft', 'run_preview', 'manage_code', 'deploy'])
expect(kindsOf('deploy_workspace_item', developer)).toContain('script')
expect(kindsOf('delete_workspace_item', developer)).not.toContain('script')
// Full access narrows nothing and ships the shared object itself, so the tool defs —
// part of every iteration's cached prefix — are byte for byte what they were.
const original = globalTools.find((t) => t.def.function.name === 'deploy_workspace_item')!
expect(shipped('deploy_workspace_item', fullSessionAccess())).toBe(original)
// And narrowing never writes through to it: the def is shared by every session, so a
// mutation here would strip one user's kinds from everyone else's schema.
expect((original.def.function.parameters as any).properties.type.enum).toContain('script')
})
// Naming a tool the model was not given produces invented calls. Asserted over the
// ASSEMBLED message: the session-state, pipeline and plan-mode sections are each
// appended by a different caller, and each can name a withheld tool.
it.each(REACHABLE_PROFILES)(
'never names a withheld tool in the assembled prompt (%s)',
(_label, access) => {
const withheld = withheldNames(access)
// The tool DEFINITIONS ship alongside the prompt, so a withheld name in a
// description is the same broken promise as one in the prompt.
const defs = JSON.stringify(shippedSessionTools(access).map((t) => t.def))
expect(withheld.filter((n) => defs.includes(n))).toEqual([])
for (const previewTools of [false, true]) {
for (const ctx of SESSION_CONTEXTS) {
const msg = assembleGlobalSystemMessage(undefined, {
previewTools,
user: { username: 'alex', folders: ['shared'], folders_read: ['shared'] },
access,
sessionContext: ctx,
pipelineContext: { folder: 'my_pipeline', mode: 'edit', nodes: [], assets: [] }
})
// Both decoration variants: the escalation one adds its own tool mentions.
for (const blocks of [0, 9]) {
const full = appendPlanModeInstructions(msg, blocks).content as string
expect(withheld.filter((n) => full.includes(n))).toEqual([])
}
}
}
}
)
// A tool result is the one place the sweep above cannot see, and `get_instructions`
// ships to every profile with guidance written around the draft tools by name.
it('does not hand authoring guidance naming withheld tools to a session that cannot draft', async () => {
const tool = globalTools.find((t) => t.def.function.name === 'get_instructions')!
const call = (access: SessionAccess) =>
tool.fn({
args: { subject: 'script', language: 'bun' },
workspace: 'ws',
helpers: { access },
toolId: 't1',
toolCallbacks: { setToolStatus: () => {} }
} as any) as Promise<string>
const readOnly = accessWith([])
const withheld = withheldNames(readOnly)
const restricted = await call(readOnly)
expect(withheld.filter((n) => restricted.includes(n))).toEqual([])
// The gate is the whole test, so pin that it is not simply refusing everyone.
expect(await call(fullSessionAccess())).toContain('write_script')
})
// A full-access profile must gate nothing at all: the text has to match the ungated
// build byte for byte, or every session's cached prefix and the ai_evals baseline
// move underneath us.
it('builds an unchanged prompt when every capability is present', () => {
const user = { username: 'alex', folders: ['shared'], folders_read: ['shared'] }
for (const previewTools of [false, true]) {
const ungated = prepareGlobalSystemMessage(undefined, { previewTools, user }).content
const full = prepareGlobalSystemMessage(undefined, {
previewTools,
user,
access: fullSessionAccess()
}).content
expect(full).toBe(ungated)
}
})
})
@@ -0,0 +1,46 @@
import type { ChatCompletionFunctionTool } from 'openai/resources/index.mjs'
import type { SessionAccess, SessionToolPolicy } from '../sessionCapabilities'
type FilterableTool = {
def: ChatCompletionFunctionTool
requires?: SessionToolPolicy
kindRequires?: Readonly<Partial<Record<string, SessionToolPolicy>>>
}
const satisfies = (policy: SessionToolPolicy, access: SessionAccess) =>
policy.every((c) => access.has(c))
/** Fails closed: every array that feeds a session is typed to declare `requires`, so a tool
* without one arrived some other way, and leaking it is the worse of the two answers. */
export function sessionToolAllowed(tool: FilterableTool, access: SessionAccess): boolean {
return tool.requires ? satisfies(tool.requires, access) : false
}
/** Filter an assembled toolset, and cut each kind-taking tool's `type` enum to the kinds
* `access` can land. `access` undefined means "not resolved yet, or not a session" — the
* toolset passes through untouched. */
export function filterSessionTools<T extends FilterableTool>(
tools: T[],
access: SessionAccess | undefined
): T[] {
if (!access) return tools
return tools.filter((t) => sessionToolAllowed(t, access)).map((t) => narrowKinds(t, access))
}
/** Returns the tool itself when every kind is allowed, so full access ships the same bytes:
* tool defs are part of every iteration's cached prefix. Builds new objects otherwise — the
* defs are module singletons shared by every session. */
function narrowKinds<T extends FilterableTool>(tool: T, access: SessionAccess): T {
const kindRequires = tool.kindRequires
if (!kindRequires) return tool
const parameters = tool.def.function.parameters as {
properties: Record<string, { enum?: string[] }>
}
const kinds = parameters.properties.type.enum ?? []
// A kind with no entry is withheld, as a tool with no `requires` is.
const allowed = kinds.filter((k) => kindRequires[k] && satisfies(kindRequires[k], access))
if (allowed.length === kinds.length) return tool
const def = structuredClone(tool.def)
;(def.function.parameters as typeof parameters).properties.type.enum = allowed
return { ...tool, def }
}
@@ -1,7 +1,14 @@
import { z } from 'zod'
import { $ScriptLang } from '$lib/gen/schemas.gen'
import type { ScriptLang } from '$lib/gen'
import { createToolDef, executeTestRun, findAndReplace, type Tool } from '../shared'
import { createToolDef, executeTestRun, findAndReplace } from '../shared'
import {
NONE,
RUN_PREVIEW,
WRITE_DRAFT,
type SessionAccess,
type SessionTool
} from '../sessionCapabilities'
import type { PipelineOutputKind } from '$lib/components/assets/AssetGraph/pipelineTemplates'
// ============================================================================
@@ -113,7 +120,7 @@ const readPipelineNodeSchema = z.object({
const readPipelineNodeToolDef = createToolDef(
readPipelineNodeSchema,
'read_pipeline_node',
'Read the full source of one pipeline node (its in-flight draft body if it has unsaved edits, otherwise the deployed body). Use before edit_pipeline_node so edits target the exact current text.'
'Read the full source of one pipeline node (its in-flight draft body if it has unsaved edits, otherwise the deployed body). Use before editing a node so edits target the exact current text.'
)
// ----------------------------------------------------------------------------
@@ -205,8 +212,9 @@ function inferredLineageNote(reads: string[], writes: string[]): string {
return ` Inferred lineage: ${parts.join('; ')}.`
}
export const pipelineTools: Tool<PipelineToolHelpers>[] = [
export const pipelineTools: SessionTool<PipelineToolHelpers>[] = [
{
requires: NONE,
def: getPipelineGraphToolDef,
planModeSafe: true,
fn: async ({ helpers, toolId, toolCallbacks }) => {
@@ -221,6 +229,7 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
}
},
{
requires: NONE,
def: readPipelineNodeToolDef,
planModeSafe: true,
fn: async ({ args, helpers, toolId, toolCallbacks }) => {
@@ -236,6 +245,7 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
}
},
{
requires: WRITE_DRAFT,
def: buildPipelineNodeToolDef,
streamArguments: true,
showDetails: true,
@@ -258,6 +268,7 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
}
},
{
requires: WRITE_DRAFT,
def: editPipelineNodeToolDef,
streamArguments: true,
showDetails: true,
@@ -286,6 +297,7 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
}
},
{
requires: WRITE_DRAFT,
def: removePipelineNodeToolDef,
fn: async ({ args, helpers, toolId, toolCallbacks }) => {
const pipeline = requirePipeline(helpers)
@@ -300,6 +312,7 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
}
},
{
requires: RUN_PREVIEW,
def: testPipelineNodeToolDef,
requiresConfirmation: true,
confirmationMessage: 'Run pipeline node',
@@ -331,8 +344,14 @@ export const pipelineTools: Tool<PipelineToolHelpers>[] = [
* /pipeline editor is open. Describes the annotation model and the direct-draft
* workflow so the model uses the pipeline tools rather than the generic
* write_script draft tools.
*
* `access` is the session's resolved capabilities (undefined outside a session, or
* before they resolve). The node-authoring guidance is dropped without `write_draft`,
* since the tools it names are withheld from that session; the annotation model stays,
* as it is what lets the model read and explain an existing pipeline.
*/
export function getPipelinePromptSection(ctx: PipelineContext): string {
export function getPipelinePromptSection(ctx: PipelineContext, access?: SessionAccess): string {
const canWriteDraft = !access || access.has('write_draft')
return `
Data Pipeline editor (ACTIVE):
@@ -342,8 +361,12 @@ Data Pipeline editor (ACTIVE):
- \`materialize\` (the managed output): a managed \`// materialize ducklake://<name>/<table>\` means the runtime writes the node's output table FOR you — write the body as a single SELECT and the runtime wraps it in the create/replace, so do NOT also write your own CREATE TABLE / INSERT. The \`dbt://\` target below is the opposite: the node writes its own DDL and none of the write strategies apply to it. IMPORTANT: a MANAGED \`// materialize\` is **DuckDB-only** and its target MUST be a DuckLake table (\`ducklake://<name>/<table>\`) — deploy rejects a \`ducklake://\` target on any other language. For a \`python3\`/\`bun\`/\`postgresql\` node writing the lake, do NOT use \`// materialize\`; write the output via the SDK instead (e.g. \`wmill.writeS3File(...)\`, a \`CREATE TABLE\` in postgresql, or \`wmill.databaseUrlFromResource\`/ducklake helpers) and let the output be inferred. Reach for \`duckdb\` when a node should materialize a DuckLake table. The one target any language BUT DBT'S OWN may declare (a dbt project's writes come from its manifest, so \`// materialize\` on a dbt script is rejected at deploy) is a WAREHOUSE RELATION: \`// materialize manual dbt://<warehouse>/<schema>/<name>\`, with \`<warehouse>\` a warehouse the workspace configures under Settings → dbt. \`manual\` is its only mode — nothing generates warehouse DDL, so the node issues its own write and the annotation records the outcome. Use it on an ingestion node a dbt project reads as a \`source\`: the declared relation and the dbt model become ONE graph node, and a downstream \`// on dbt://<warehouse>/<schema>/<name>\` fires when that node completes. Write strategy: with no option it REPLACES the whole table each run (full refresh; the only mode whose output columns may change); \`// materialize <uri> append\` INSERT-appends rows (incremental); \`// materialize <uri> key=<col>\` merges/upserts on \`<col>\`. \`// materialize manual <uri>\` opts OUT of managed writes — the script writes its own DDL and the annotation only records the output asset for lineage. \`materialize\` is paired with partitioning for incremental pipelines: a \`// partitioned <daily|hourly|weekly|monthly|dynamic>\` node runs once per partition (append/merge into a fixed-schema table), and the \`{partition}\` token — usable in any asset URI AND in the body SQL — is substituted with the current partition's IDENTITY string at run time. To filter the source to the active slice on a time grain, use the runtime-injected macro: \`WHERE wm_partition(<ts_col>) = {partition}\`. \`wm_partition(ts)\` buckets a timestamp with the exact identity format the runtime used (daily/hourly/weekly/monthly), so it always matches and you never hand-write a \`strftime\` format. Do NOT write \`= TIMESTAMP {partition}\`: the identity string is not a valid timestamp literal for hourly/weekly/monthly and errors at runtime. For \`dynamic\` partitioning the identity is your caller-supplied key (not a timestamp, no macro), so filter on it directly: \`WHERE <your_key_col> = {partition}\`. \`materialize\` is an output DECLARATION on the node — it is not a command; there is no "materialize run".
- \`measure\` / \`dimension\` (declared metrics): on a node that materializes a DuckLake table, \`// measure <name> = <aggregate> [where <predicate>]\` names the canonical way to aggregate that table (e.g. \`// measure revenue = sum(amount) where not is_refund\`), and \`// dimension <name> = <expr>\` names a way to slice it (e.g. \`// dimension region = region\`, \`// dimension month = date_trunc('month', ordered_at)\`). They execute nothing: they are catalogued at deploy so the editor and other agents can reuse the definition instead of re-deriving it and silently disagreeing. Keep the predicate in the \`where\` clause rather than folding it into the aggregate: it is rendered as \`<agg> FILTER (WHERE <pred>)\`, which is what lets two measures with different predicates sit under one GROUP BY. DuckLake-only, and only meaningful next to \`// materialize\`. Declare one when a number carries a judgement call someone else would get wrong (refunds excluded, test rows dropped, which column is the amount); do NOT blanket every table with measures, an obvious \`count(*)\` earns nothing. To USE a metric another node declares, read that node with read_pipeline_node and reuse its exact expression rather than guessing it.
- Use get_pipeline_graph to see the current nodes/assets/triggers, and read_pipeline_node before editing one.
- Every node of this pipeline lives at \`f/${ctx.folder}/<node_name>\` — \`${ctx.folder}\` is the folder name and \`f/\` is the owner prefix every workspace path carries, so write it exactly once (never \`f/f/…\`, and never a bare \`<node_name>\`).
- Every node of this pipeline lives at \`f/${ctx.folder}/<node_name>\` — \`${ctx.folder}\` is the folder name and \`f/\` is the owner prefix every workspace path carries, so write it exactly once (never \`f/f/…\`, and never a bare \`<node_name>\`).${
canWriteDraft
? `
- Build new nodes with build_pipeline_node and edit existing ones with edit_pipeline_node. These apply directly as unsaved drafts on the canvas (like the flow/script editor applies AI edits) — they DO NOT deploy. There is no separate Accept/Reject step. Prefer these over the generic write_script/edit_script draft tools while a pipeline is open.
- Reuse existing asset paths from the graph when wiring a downstream node to an upstream one (read the upstream's write asset, then \`// on\` that same URI).
- Only deploy when the user explicitly asks; the user deploys drafts from the canvas.`
: ''
}`
}
@@ -1,6 +1,7 @@
import type { ChatCompletionSystemMessageParam } from 'openai/resources/chat/completions.mjs'
import type { ArtifactVersionTarget } from '$lib/components/sessions/previewRouter'
import { createToolDef, type Tool, type ToolCallbacks } from './shared'
import { createToolDef, type ToolCallbacks } from './shared'
import { NONE, type SessionTool } from './sessionCapabilities'
import {
appendPlanModeInstructions,
derivePlanTitle,
@@ -100,7 +101,7 @@ export class PlanModeController {
/** Only the transition the current posture allows; auto-accepting exposes neither, since
* entering is the user's choice. */
get tools(): Tool<any>[] {
get tools(): SessionTool<any>[] {
if (!this.#host.available) return []
if (this.#host.autoAccepting) return []
return this.#host.active ? [this.exitTool] : [this.enterTool]
@@ -108,12 +109,16 @@ export class PlanModeController {
/** Both transitions, whichever posture is selected: what plan mode contributes to this
* chat's capabilities, rather than to the turn it is about to send. */
get availableTools(): Tool<any>[] {
get availableTools(): SessionTool<any>[] {
return this.#host.available ? [this.enterTool, this.exitTool] : []
}
// This safety tag is what keeps plan mode escapable through its handoff tool.
exitTool: Tool<any> = {
exitTool: SessionTool<any> = {
// Ungated even for a user who can change nothing: the posture is the USER's choice,
// and its instructions order this call to hand the plan over — withholding it
// strands the model in a round nothing else can end.
requires: NONE,
def: createToolDef(exitPlanModeArgs, EXIT_PLAN_MODE_TOOL, EXIT_PLAN_MODE_TOOL_DESCRIPTION),
planModeSafe: true,
requiresConfirmation: true,
@@ -162,7 +167,8 @@ export class PlanModeController {
}
}
enterTool: Tool<any> = {
enterTool: SessionTool<any> = {
requires: NONE,
def: createToolDef(enterPlanModeArgs, ENTER_PLAN_MODE_TOOL, ENTER_PLAN_MODE_TOOL_DESCRIPTION),
planModeSafe: true,
requiresConfirmation: true,
@@ -9,6 +9,7 @@ import type {
} from 'openai/resources/index.mjs'
import { type DBSchema, dbSchemas } from '$lib/stores'
import type { ContextElement } from '../context'
import { NONE, RUN_PREVIEW, type SessionTool } from '../sessionCapabilities'
import {
createSearchHubScriptsTool,
type Tool,
@@ -462,9 +463,12 @@ export const resourceTypeTool: Tool<ScriptChatHelpers> = {
// Generic DB schema tool factory shared by the script, flow and global modes
export function createDbSchemaTool<T>(
opts: { description?: string; updateEditorCache?: boolean } = {}
): Tool<T> {
): SessionTool<T> {
const { description, updateEditorCache = true } = opts
return {
// `getDbSchemas` below introspects by running a query job through /jobs/run/preview,
// which jobs.rs refuses operators — so this reads like a lookup but gates like a run.
requires: RUN_PREVIEW,
def: description
? {
...DB_SCHEMA_FUNCTION_DEF,
@@ -610,7 +614,8 @@ const SEARCH_NPM_PACKAGES_TOOL: ChatCompletionFunctionTool = {
}
// Helpers-agnostic so both script mode and global mode can offer it.
export const searchNpmPackagesTool: Tool<{}> = {
export const searchNpmPackagesTool: SessionTool<{}> = {
requires: NONE,
def: SEARCH_NPM_PACKAGES_TOOL,
planModeSafe: true,
fn: async ({ args, toolId, toolCallbacks }) => {
@@ -0,0 +1,52 @@
import type { Tool } from './shared'
// A leaf on purpose: every module defining a session-reachable tool imports these
// constants, so importing anything with side effects here drags it into all of them.
/**
* What a user may do in ONE workspace, as an AI session's toolset needs to know it.
* Best-effort, not a boundary — the token is the enforcement point, so this narrows
* what the model is offered and guarantees nothing.
*/
export type SessionCapability =
/** Passes drafts.rs `require_can_write_path` — for a draft of any kind, schedules and
* triggers included, though an operator may create those directly. */
| 'write_draft'
/** Passes the operator refusal jobs.rs `run_preview_*` makes before starting a job. */
| 'run_preview'
/** Passes `check_deploy_rules`. */
| 'deploy'
/** Passes the operator refusal the script, flow and app handlers make before creating or
* deleting one. The same condition as `run_preview` today, but another handler family's
* check, so each follows its own gate if the two ever diverge. */
| 'manage_code'
/** Passes `require_admin`. */
| 'admin'
export type SessionAccess = ReadonlySet<SessionCapability>
/**
* The capabilities without which a tool's call cannot succeed — usually because the
* backend refuses it, occasionally because only that capability produces the tool's
* input. Never a relevance judgement: a tool the server would accept ships, even when
* it is of little use to the session, because withholding it is a stricter answer than
* the one the user would get by trying.
*/
export type SessionToolPolicy = readonly SessionCapability[]
export const NONE: SessionToolPolicy = []
export const WRITE_DRAFT: SessionToolPolicy = ['write_draft']
export const RUN_PREVIEW: SessionToolPolicy = ['run_preview']
export const DEPLOY: SessionToolPolicy = ['deploy']
/**
* A tool that can reach an AI session's toolset. `requires` is mandatory, and declared
* beside the `fn` whose call is what proves it right.
*/
export type SessionTool<T> = Tool<T> & {
requires: SessionToolPolicy
/** For a tool taking the item kind as its `type` argument, where one verdict for the whole
* tool is either too strict or too loose: what each kind's handler needs. A session is
* offered only the kinds it can land, and the rest never appear in the schema it sees. */
kindRequires?: Readonly<Partial<Record<string, SessionToolPolicy>>>
}
@@ -7,6 +7,7 @@ import type { UserDraftItemKind } from '$lib/gen'
// The gate's two refusals, from a module that holds prose and one size limit: under the
// shallow-import rule below, the rest of plan mode is not reachable from here.
import { PLAN_MODE_MESSAGES } from './planModeMessages'
import { NONE } from './sessionCapabilities'
// Import-free leaf, so it satisfies the shallow-import rule below.
import {
openItemPreviewAction,
@@ -1249,7 +1250,10 @@ export interface Tool<T> {
workspace: string
helpers: T
}) => MaybePromise<ToolRejection | undefined>
setSchema?: (helpers: any) => Promise<void>
/** This chat's definition of the tool, for a schema that depends on the chat (its workspace,
* its open flow). Returns a new def: tools are module singletons shared by every chat, so
* one written back leaks into the others. */
schemaFor?: (helpers: any) => Promise<ChatCompletionFunctionTool>
/** Safe to run while plan mode is active. Absence fails closed. */
planModeSafe?: boolean
/** The arguments a plan-mode-safe tool still refuses while the posture holds — for a tool
@@ -1549,6 +1553,7 @@ export function isHubPath(path: string): boolean {
}
export const createSearchHubScriptsTool = (withContent: boolean = false) => ({
requires: NONE,
def: searchHubScriptsToolDef,
planModeSafe: true,
fn: async ({ args, toolId, toolCallbacks }) => {
+37 -15
View File
@@ -803,33 +803,25 @@ export function deployPermissionForKinds(
}
/**
* Whether the current user may deploy into `workspace`. Mirrors `check_deploy_rules` in
* The workspace's deploy-protection rules alone. Mirrors `check_deploy_rules` in
* windmill-common so the UI can disable the action with a reason instead of letting the
* click 403: `DisableDirectDeployment` is evaluated before `RestrictDeployToDeployers`, so
* the same message wins here as on the server when both block; admins and superadmins bypass
* both rules, while `wm_deployers` members bypass only the latter.
*
* The operator refusal is not part of that mirror. The server refuses operators in the item
* handlers instead, and for fewer kinds, so refusing them for everything here is deliberately
* stricter than the server rather than a faithful copy of it.
* Separate from `checkDeployPermission`, which adds a blanket operator refusal: the server
* refuses operators per kind, in the item handlers, not in these rules.
*
* Fails open on any error — the server still enforces on the actual deploy.
* Shared by the session dock and the compare page so both gate identically.
*/
export async function checkDeployPermission(
export async function checkDeployRules(
workspace: string,
/** Pre-fetched `whoami` for `workspace`, to save a round trip when the caller already has one. */
whoami?: User
/** Pre-fetched identity for `workspace`, to save a round trip. Narrowed to the fields
* read so a caller holding a `UserExt` can pass it without a cast. */
whoami?: Pick<User, 'is_admin' | 'is_super_admin' | 'username' | 'groups'>
): Promise<DeployPermission> {
try {
const me = whoami ?? (await UserService.whoami({ workspace }))
if (me.operator) {
return {
ok: false,
reason: "You're an operator in this workspace — operators can't deploy",
refusedBy: 'operator'
}
}
const userInfo = {
is_admin: !!me.is_admin,
is_super_admin: !!me.is_super_admin,
@@ -877,6 +869,36 @@ export async function checkDeployPermission(
}
}
/**
* Whether the current user may deploy into `workspace` at all: the protection rules above,
* plus a refusal of every operator.
*
* That refusal is not part of the `check_deploy_rules` mirror. The server refuses operators
* in the item handlers instead, and for fewer kinds, so refusing them for everything here is
* deliberately stricter than the server rather than a faithful copy of it. A caller mirroring
* one handler wants `checkDeployRules`.
*
* Shared by the session dock and the compare page so both gate identically.
*/
export async function checkDeployPermission(
workspace: string,
whoami?: Pick<User, 'operator' | 'is_admin' | 'is_super_admin' | 'username' | 'groups'>
): Promise<DeployPermission> {
try {
const me = whoami ?? (await UserService.whoami({ workspace }))
if (me.operator) {
return {
ok: false,
reason: "You're an operator in this workspace — operators can't deploy",
refusedBy: 'operator'
}
}
return await checkDeployRules(workspace, me)
} catch {
return { ok: true }
}
}
/**
* Whether `me` may write at `path` in `workspace` — the per-item half of the deploy gate, which
* `checkDeployPermission`'s workspace-level rules don't cover. Advisory only: it exists so the UI