Files
windmill/docs/reusable-ai-agents.md
T
a78beff743 feat: dynamic AI agent toolsets (#11050)
* feat: dynamic ai agent toolsets, and memory as a step input

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: address review round 1 on dynamic ai agent toolsets

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: tag enabled_tools and drop the memory step input

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: let an mcp server entry be named by the path the roster shows

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: keep $res: out of the tool names the enabled tools picker offers

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: name an mcp server by its bare path on the one side that can hold it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: count the enabled tool names that matched nothing instead of logging them

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: narrow an agent's roster in one pass, by whole entries

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: pin that an mcp summary is rejected against a name that is not

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: regenerate the copilot flow schema after the merge

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: shorten the enabled tools list hint

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: take enabled_tools back to a plain list of tool names

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: keep the enabled tools add-menu hint describing the unset field

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: name a websearch tool that carries no summary of its own

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix: reserve the name web search is enabled by so no tool can share it

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: spell the reserved web search name with a hyphen

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* refactor: reserve __wm_web_search as the name web search is enabled by

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: advance ee-repo-ref past the git sync ci check work

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs: shorten the enabled tools description the run form shows

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* chore: update ee-repo-ref to e4c1b794d6c5e6e390987341b2840587bbb40348

This commit updates the EE repository reference after PR #785 was merged in windmill-ee-private.

Previous ee-repo-ref: af668462f0f06b02a5f4e0c22e6156858487a518

New ee-repo-ref: e4c1b794d6c5e6e390987341b2840587bbb40348

Automated by sync-ee-ref workflow.

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Ruben Fiszel <ruben@windmill.dev>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
2026-09-15 11:34:30 +02:00

129 lines
8.4 KiB
Markdown

# Reusable AI Agents
An AI agent flow step can be saved as a **reusable agent** — a resource of the built-in
`ai_agent` resource type that bundles the agent's brain (provider/model, system prompt,
temperature, output schema, memory…) and its tool set. Other flows can link to the same
agent, and edits to the agent propagate to every linked step.
The `ai_agent` resource type is defined in the hub (windmill-integrations) and synced into
every workspace via the standard cached-resource-type sync, like other built-in types.
## Rigid linking
`FlowModuleValue::AIAgent` has an optional `agent` field holding the resource path, plus a
`tool_inputs` map (per-tool host-flow input overrides). When `agent` is set:
- The brain config and tools are resolved at runtime from the resource
(`windmill-worker/src/ai_executor.rs`): the brain is interpolated, so a nested provider `$res:`
credential resolves automatically.
- The step keeps only the flow-local inputs (`user_message`, `user_attachments`, `enabled_tools`)
in its own `input_transforms`; the brain and tools stay in the resource (read-only in the step).
`enabled_tools` says which of the roster this step may call, narrowing one use of a shared agent
without touching the agent: an absent field carries every tool, a list carries the ones it names,
and an empty list carries none.
- The agent carries its tools' default input bindings verbatim as authored (static, AI-filled,
or flow expressions), so saving round-trips losslessly. Each host flow overrides what it
needs: `tool_inputs` stores per-tool overrides (a diff from the resource tool's own
transforms) that overlay onto the matching tools at runtime. Editing on a linked step edits
the flow's use of the agent; editing in the agent editor edits the agent itself.
In the flow editor, the AI agent step's **Step Input** tab shows a single read-only card
(*linked to <path>*, with the inherited brain + tools and an explanatory tooltip) plus
*Edit*, which opens the agent editor over the flow, and *Unlink* (fork the resolved config —
including any `tool_inputs` — back into the step as a one-off).
A linked agent's tools appear as display-only graph tool nodes (clicking one selects the
agent step); below the step's inputs, each tool gets a section with the standard schema-aware
input editors (prop picker included) and a read-only view of its code — edits persist into
`tool_inputs`.
## Drafts
The agent editor edits the resource through a **per-user resource draft** (`draft` table,
`item_kind = 'resource'`), autosaved by `useAgentDraft` and deployed by the editor's own Deploy
button. It is the same draft row the generic resource editor writes and the Review & Deploy page
lists, so an agent can be deployed from any of them.
A flow does not wait for that deploy to see the draft:
- Testing the flow, or a single linked step, runs the draft. `runFlowPreview` and `ModuleTest`
substitute each linked step for the standalone step the draft would run as
(`linkedAgentDrafts.ts`): `agent` cleared, the draft's brain as static input transforms, the
draft's tools on the step, and the step's own flow-local inputs kept on top —
the same overlay order `ai_executor.rs` applies to a linked step. `tool_inputs` is untouched,
since the worker overlays it in both branches.
- The step's linked card and the graph's tool nodes show the draft, with a *Draft* badge, so the
editor describes what a test would run. Read-only surfaces (the deployed flow page, the run
viewer) stay on the deployed agent: they resolve tools through `publishLinkedAgentTools` without
the draft flag.
- Deploying the flow lists every linked agent that has a draft in the confirmation dialog, beside
the draft triggers. Deploying one writes the resource and drops the draft; leaving one out keeps
its draft untouched, and the flow runs the agent as currently deployed. That is the one place
the two kinds differ: an undeployed draft trigger is deleted, because it belongs to the flow,
while an agent draft belongs to a resource other flows also use.
Because a draft is per-user, a flow test can behave differently for two people looking at the same
flow. That is the same contract as a flow draft, and deploying the agent is what makes it shared.
Inlining has a consequence worth knowing: a preview job's `raw_flow` then carries the agent's
config, where a linked step used to carry only the path and leave the resolution to the worker. So
an agent's prompt and tool set are readable by whoever can read that preview job, which is a wider
set than whoever can read the resource when the agent sits in a more restricted folder than the
flow. No credential travels with it — the provider stays a `$res:` reference, resolved at run time
as the runner. The agent editor's own test pane has inlined the same way since drafts existed;
closing the gap would mean the preview carrying a draft *reference* the worker resolves, rather
than the config.
Sharing works through standard resource folder permissions (save agents under `f/...`).
Only the agent's brain is interpolated when the step runs. A tool's own `$res:`/`$var:` defaults are
left alone and resolved when that tool executes, so a host flow can override a default pointing at a
resource it cannot read — and an unused tool whose default is inaccessible never fails the agent.
## Resolution is live, not pinned
A linked step resolves its agent resource when the step runs — that is what makes an edit propagate
to every linked flow. It also means a run is not a snapshot: editing the agent while a flow is
in-flight affects steps that have not started yet. The same applies one level down, where the effect
is sharper: a *nested* agent tool of a linked agent runs as its own job and looks its definition up
in the resource again by tool id, so an edit landing between the LLM selecting that tool and the
tool starting can run the changed definition, or fail if the tool was removed. Pinning would require
carrying the resolved definition into the child job rather than its id. Inline (unlinked) agents are
unaffected: their tools live in the flow value, which is snapshotted with the run.
## Version history
Editing a resource appends a row to `resource_version` (all types except `state` and `cache`,
which the platform rewrites on every job), so an agent's prompt, model and tool set can be
diffed and restored from the resource editor's History drawer. Restoring writes the old value
forward as a new version rather than rewinding, keeping the history append-only.
History captures the resource, not its transitive closure. A `$var:`/`$res:` reference is stored
as the reference, so two versions can be byte-identical while the agent behaves differently
because the referenced variable changed underneath them. Anything comparing agent runs across
versions has to account for that.
An eval run records the version its agent was at when the run was enqueued, which is what makes a
result attributable to a prompt state — see `docs/ai-agent-evals.md`.
A superseded value is retained for up to 100 versions. Values written through the UI keep their
secrets in linked variables, but one pushed by `wmill` or written by `setResource` can hold an
inline credential, and overwriting it no longer removes it from the database — anyone who can
read the resource can read it from the history. Rotating such a credential therefore does not
erase the old one on its own — follow the rotation with **Clear past versions** in the History
drawer, which drops every version but the current value for that one resource. Secret *variables*
are deliberately not versioned at all for the same reason.
## Dependencies and locks
A linked step carries `tools: []`, and no dependency job ever visits the `ai_agent` resource, so the
tool scripts inside it are outside the lockfile and dependency-map pipelines: `lock_modules` has
nothing to lock on the step, and `FlowValue::traverse_leafs` sees no leaf for them. Consequences:
- Raw-script tools saved into an agent keep whatever `lock` they had on the authoring step (`null`
if that step was never deployed), and every linked flow resolves their dependencies at job time.
- Script tools referenced by path are invisible to redeploy cascades — republishing such a script
does not re-lock the flows that link the agent.
Deploying a linked flow to another workspace pulls the `ai_agent` resource in as a dependency, and
from there its provider resource and its tools' scripts, flows, MCP resources and nested agents.