mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-10-05 00:02:24 +00:00
e1e3692fbc82d50c021ef8bf8ca7019d760705f0
95
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
6e1ef93f32 |
feat: let test_run_flow name the conversation of a chat-mode test run (#11198)
* feat: let test_run_flow name the conversation of a chat-mode test run Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refuse a non-UUID conversation_id and mint chat test-run ids in one place Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ai-evals): add a chat-flow follow-up case and mock flow preview runs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: expect the conversation id argument on the manager's flow test bridge Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ai-evals): require the chat-flow follow-up runs to share one conversation id Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: name the test_run_flow argument memory_id after the run parameter Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c297ed0052 |
feat: managed memory with an inherited or custom memory id per step (#11118)
* feat: split ai agent memory into agent policy, run memory id and step history * fix: scope string memory ids to workspace and flow, keep nested tool history inputs * chore: update sqlx cache for the flow context query * docs: describe memory id scoping as collision-free rather than isolated * chore: regenerate openflow json after merging main * fix: offer no memory id for legacy manual memory, document linked history inputs * fix: seed provided messages from legacy manual memory and hide its note once set * fix: bypass memory when a provided messages expression evaluates to null * fix: require a user message when provided messages are empty * chore: keep the empty messages comment within the line width * docs: name the history inputs wherever linked steps list their flow-local inputs * docs: keep the memory storage path on one line * feat: managed memory with an inherited or custom memory id per step Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: list a custom memory id in the test run form and name where an inherited one comes from Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep memory id out of the add-field menu and drop the memory id telemetry Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep legacy auto memory without an id working after an untouched redeploy Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: rename step messages to previous_messages and address review Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: rewrap comments and docs lines lengthened by the previous_messages rename Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: read agent memory as either a legacy shape or the current one Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: name the memory setting in ignored-input notes and keep conversions honest Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep a legacy memory count unset on open and read a cleared count as off Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: address review on cleared test history and zero-count memory Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: drop flow-local keys from a linked agent resource before interpolating it Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: keep a linked resource's own inputs as fallbacks and note ignored history on image runs Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: restore the linked agent draft tests and log ignored history on every image run Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: resolve the one-of variant from the value when the selected one leaves the list Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: treat zero-count managed memory as off when enabling chat mode and shorten comments Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * fix: stop requiring user_message in the openflow agent contract when previous messages are the prompt Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
a78beff743 |
feat: dynamic AI agent toolsets (#11050)
* feat: dynamic ai agent toolsets, and memory as a step input Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review round 1 on dynamic ai agent toolsets Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: tag enabled_tools and drop the memory step input Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: let an mcp server entry be named by the path the roster shows Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep $res: out of the tool names the enabled tools picker offers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name an mcp server by its bare path on the one side that can hold it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: count the enabled tool names that matched nothing instead of logging them Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: narrow an agent's roster in one pass, by whole entries Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: pin that an mcp summary is rejected against a name that is not Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: regenerate the copilot flow schema after the merge Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: shorten the enabled tools list hint Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: take enabled_tools back to a plain list of tool names Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: keep the enabled tools add-menu hint describing the unset field Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name a websearch tool that carries no summary of its own Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reserve the name web search is enabled by so no tool can share it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: spell the reserved web search name with a hyphen Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: reserve __wm_web_search as the name web search is enabled by Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: advance ee-repo-ref past the git sync ci check work Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: shorten the enabled tools description the run form shows Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to e4c1b794d6c5e6e390987341b2840587bbb40348 This commit updates the EE repository reference after PR #785 was merged in windmill-ee-private. Previous ee-repo-ref: af668462f0f06b02a5f4e0c22e6156858487a518 New ee-repo-ref: e4c1b794d6c5e6e390987341b2840587bbb40348 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
63cb46d7bb |
feat: add a minimal skin for the approval page and slack/teams (#11061)
* feat: add an approval skin to the approval page and slack/teams messages Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * chore: point ee-repo-ref at the teams approval skin commit Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * fix: resolve the approval skin from the step awaiting approval Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * fix: shorten the slack approval message to fit the button value limit Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * fix: rename skins to detailed/minimal and keep long slack messages Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * feat: title the minimal approval page from the step and flow summaries Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * feat: let wait_for_approval set the description approvers see Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * fix: keep a finished workflow's approval description, still gated Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * fix: keep a login-required approval locked after the run moves on Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * feat: hide the windmill version on the approval page Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KK2Ye4PizykReZVuMEm3m9 * chore: update ee-repo-ref to e92abc9d1fba3ba898640df0cfeb0af8a50849b4 This commit updates the EE repository reference after PR #786 was merged in windmill-ee-private. Previous ee-repo-ref: bf1766ff49458f62d3f11746f1a06893ae3c2325 New ee-repo-ref: e92abc9d1fba3ba898640df0cfeb0af8a50849b4 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
fa73539839 |
fix(ai-chat): test_run_flow could test a different flow than the one asked (#11066)
* fix(ai-chat): test_run_flow could test a different flow than the one asked Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TmQDoFXVPwYyEN8PfV2oL7 * fix(ai-chat): prefer the flow editor stored at the path over one renamed to it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TmQDoFXVPwYyEN8PfV2oL7 * test(ai-chat): default the flow helpers factory and trim duplicated setup Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TmQDoFXVPwYyEN8PfV2oL7 * refactor(ai-chat): resolve the flow editor to run by its storage path Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TmQDoFXVPwYyEN8PfV2oL7 * refactor(ai-chat): move the editor storage path context out of sessions Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TmQDoFXVPwYyEN8PfV2oL7 --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
0b3dc3e5c9 |
fix: build the global chat's prompt identity from the operating workspace (#10793)
* fix: do not read an unloaded workspace list as a non-membership `roleForWorkspace` settled `not_a_member` from `userWorkspaces` alone. That store and `superadmin` both start undefined and load asynchronously, so an unloaded list read as an empty one: a chat operating on any workspace other than the one being browsed advertised no pages and reported an access denial. The root layout gives up after its retries, so a load that fails leaves the denial permanent, with `whoami` never attempted. Settle a non-membership only once both stores have resolved; treat unresolved as unknown and fall through to the `whoami` lookup. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: build the global chat's prompt identity from the operating workspace The prompt's path conventions and folder guidance came from the ambient `userStore`, which describes the workspace being browsed rather than the one the chat operates on. Three of those fields are per-workspace and wrong whenever the two differ: the username (that workspace's `usr` row), the writable/readable folder sets (its ACLs), and `is_admin`, which decides whether the folder list reads as exhaustive. The backend still enforces the ACLs, so the cost is prompt quality — paths the model cannot write to, and a 403 to recover from. Resolve the identity for the operating workspace and feed that to the prompt, refreshed alongside skills and MCP servers and settled in `beforeSend` so the cached system-prompt prefix stays stable for the turn. An unresolved role now leaves the folder sets undefined rather than empty, so the guidance is dropped instead of claiming there is nothing to write to, and `create_folder` credits the workspace it wrote to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: read the AI provider resource types lazily `Object.keys(AI_PROVIDERS)` at module scope made `AI_PROVIDERS` a load-time requirement for every importer of this module, the global chat included. `AIChatManager.test.ts` mocks `../lib` without it and has been unable to load since the catalog was introduced; no CI workflow runs vitest, so nothing reported it. The constant is read in two places, both inside functions, so deferring it removes the load-time dependency without changing behaviour. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
449b1a6933 |
fix: ground the chat's AI agent provider in the workspace's models (#10774)
* fix: ground the chat's AI agent provider in the workspace's models Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: never reject an AI agent model the catalog could not confirm Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: benchmark AI agent provider grounding in ai_evals Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: only reject an AI agent model an exhaustive listing rules out Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep instance-level AI settings out of the workspace provider catalog Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep untrusted model ids out of the chat's context Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: carry completeness on the model listing itself Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct two comments left behind by the catalog rework Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bound the model listing and verify the default against it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: honour a workspace default a filtered listing names Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: recognise a workspace default past the prompt's model cap Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep an aliasing provider's unlisted model ids permissive Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2b4369d7cb |
fix: reject invalid AI agent tool names when the chat writes a flow (#10756)
* fix: reject invalid AI agent tool names when the chat writes a flow Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review nits on agent tool name validation Share one AI-agent walk between the providerless-agent and invalid-tool-name collectors, drop the unused validateToolName, and list every reserved id in the tool naming rules. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: describe an agent tool's summary as the name the agent calls it by The OpenFlow schema described `AgentTool.summary` as a short description of the tool, which is the same schema the flow write tools hand the model, so it pulled against the naming rules. Narrow those rules to flowmodule tools, since websearch and mcp tool names are never regex-checked, and let `kind` take either vocabulary its callers resolve. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name-check only the agent tools whose summary the agent calls An mcp tool exposes the MCP server's own tool names and a websearch tool's summary is a plain label, so neither reaches the worker's name check. Both default to an empty summary in the editor, which the chat then refused to write back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
caa189868c |
feat(ai-sessions): add plan mode (#10057)
* feat(sessions): let an opener name the artifact version to show A tab already remembers the version a reader pinned, and re-pointing it keeps that pin. Plan mode needs the two intents that leaves out: a plan card scrolled up the transcript wants the version it proposed, and a plan going up for approval wants the current text with no pin at all. `ArtifactVersionTarget` is those two alongside the existing one: a number, `'latest'`, or omitted. Omitted still cannot double as `'latest'` — every artifact tool re-opens the document it just wrote, so taking that as a request to move would yank a reader out of the version they chose on every edit the agent makes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(copilot): add the plan-mode gate and tag plan-mode-safe tools Plan mode is a read-only posture, so something has to decide which tools it may still run. `Tool.planModeSafe` is that tag, and processToolCall fails closed on it: untagged means mutating means blocked. Deriving it from `requiresConfirmation` was not an option — unconfirmed mutating tools exist, and a posture that leaks one is not a posture. The gate runs twice per call. Before `validateBeforeConfirmation`, so a validator cannot reach out while planning; and again after the confirmation wait, because plan mode can be entered while a mutating tool's card is already pending, and that approval must not carry it through. Arguments are read one field at a time rather than through a parse of the whole call. `change_note` is optional and cosmetic, and a model that sends it as `null` would otherwise fail the object parse and take the plan down with it — the user being told there was no plan to approve, which is false. Also here, because refusing a call well needs them: a validator may now return the row the user reads and the result the model gets separately, a tool may word its own cancellation, and a tool may start work when its card appears rather than when it is approved. The gate is consulted before any of them. `shouldAutoAcceptToolConfirmations` is asked about the tool by name, because skipping the confirmation wait is itself an answer on the user's behalf and one tool must not be answered for. Deciding that without the name would put the exception out of reach of the only path that needs it. The gate stays inert until a chat supplies `isPlanModeActive`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(copilot): give a session one versioned plan document The plan the user agrees to has to survive `/clear`, so it belongs to the session rather than the conversation, and a session holds exactly one. Its id is the session's, so the primary key is the constraint — there is no second row to mint, no index to maintain and no schema change at all. Every write reads the row it is about to replace inside the transaction that replaces it. Read outside, two tabs both see version N, both stamp N+1, and the later write silently drops the earlier one's text and its snapshot; IndexedDB serialises readwrite transactions over a store, so read and write together cannot interleave. Approval takes the same route but patches only the pointer: an approval computed while another tab was revising must not carry this tab's older content back over the newer text. Approval is `approvedVersion`, a pointer at a version, never a flag. Below the current version means the newest text is a proposal the user has not agreed to; absent means nothing here was ever approved. Only exit_plan_mode can leave the pointer behind, since every write outside plan mode carries it forward — an amendment the user's posture already trusts is still the agreed plan. Declining writes nothing at all: the refused proposal stands as the newest version, with the agreed one still in history. Nor can create_artifact confer approval. It asks for no confirmation, so the model writing a plan document is not the user agreeing to one; a plan written there holds the session's slot as a draft until a decision lands on it. That is also why the approved version is exempt from pruning. A plan approved at v1 and then planned against for twenty more rounds would otherwise lose the very version that stands as agreed, and with it the card that opens it, the banner offering it back, and read_artifact at that version. It is excluded from the pruning candidates rather than added on top, so the budget is unchanged and what survives simply stops being contiguous. The write reports whether the database took it. Most callers still degrade like the reads do, but a plan cannot: returning one the database refused would let the user approve and execute against a document that disappears on reload — a refused plan write raises instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(copilot): add plan mode — the posture and its two tools enter_plan_mode asks to hold work; exit_plan_mode hands over a plan and, on approval, gives the posture back to whatever preceded it. Both carry `planModeSafe`, since a posture with no exit is a trap. Only the transition the current posture allows is offered, so there is no tool for leaving a posture the chat is not in. A planning round runs from entering plan mode to the proposal the user decides on. It remembers only the write it made, because nothing it does is undone — and that write is shared between the card's confirmation hook and the tool's `fn`, so the plan is on screen while the user is deciding whether to approve it rather than after. The round is identified by an epoch bumped on *entering*, not by the conversation. A chat rotation mid-approval must still let that approval hand the posture back; a round the user has since left and re-entered must not, or approving the old plan would drop them out of a read-only posture they just chose. Saving a proposal revises the session's plan document and creates one only when there is none — both halves in a single transaction, so a second tab proposing at the same moment revises the row this one wrote rather than racing it. Persistence failures hold the posture. Approval is reported only once both the proposal and the approval pointer are durable, so a plan the database refused cannot unblock mutating tools. The failure is reported from `fn` and no earlier: the write settles while the card is still waiting to be confirmed, and clearing that card from underneath the wait would take away the only control that resolves it. An auto-accepting posture answers for the user through one predicate, asked by every path that answers: the pending-card sweep, the confirmation itself, and the decision to skip the wait at all. enter_plan_mode never qualifies: YOLO means "stop asking and run it", and a call from a tool set snapshotted before the switch must not answer that with a read-only posture — whether its card is already pending or has yet to be registered. Plan mode lives in its own controller with a narrow view of the chat it runs in: it reads that autonomy state and asks for the two changes it can cause, rather than owning any of it. Plan mode is offered only in a session chat, and a session chat is GLOBAL for its whole life. The gate reads that mode, so `changeMode` refuses to move one out of GLOBAL rather than resting the invariant on a picker being hidden. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(copilot): surface plan mode in the chat and the artifact list Plan mode is the only posture that refuses work, so the composer says so before the user types the request it is about to turn down: the mode pill is tinted whole rather than by its icon, and the empty placeholder carries the constraint in words. Teal, not the house green — green is the transcript's success colour a few rows up, and a mode signal in it would read as "this worked" rather than "this is held". A blocked tool renders as its own lean row naming the tool, not as an error: the call did what plan mode says it should, and "why can't it edit" is answered where it is asked. A plan card names the decision — proposed, approved, or not approved — and never the button, since a Stop and a posture switch resolve it too. Its button opens the version that card proposed, so a card far up the transcript still shows the plan it put forward rather than whatever the document has become since. The artifact list and the preview header both label the plan through one badge helper, so the two cannot disagree about what counts as one: a plan the user never approved keeps the plan icon and takes the neutral badge, leaving the teal to mean exactly one thing. In the viewer, an unapproved revision says so in a bar that cannot be scrolled past, with the version the user did agree to one click away. The autonomy picker became a table with one row per posture, so adding one touches a single place instead of four parallel switch statements. A version of a plan is read against the one the user approved, not against the newest: latest is only where the model happened to stop. So the approved version is never stale — its bar is teal and points forward to the draft rather than warning about it — the version in front of it is the draft, and anything behind it is history that is neither and takes no pill at all. The list opens a plan at the approved version for the same reason, which is what lets its pill say `plan` while an unapproved draft sits at the head. One helper answers all of it, so the list and the preview header cannot drift apart on what counts as the plan. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(ai-evals): exercise plan mode end to end A case a unit test cannot stand in for: it starts in plan mode against the real gate and the real exit_plan_mode, and grades whether the model researches and hands over a usable plan instead of guessing at one. The checklist does not grade what the harness does for the model — exit_plan_mode writes the plan document itself, so "saves the plan as an artifact" would pass on any run where the tool is called at all. The eval store seeds artifacts with history and mirrors the store's own approval rules, so a rename cannot promote a proposal the user turned down. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(ai-evals): import the plan-mode messages from the module that owns them `PLAN_MODE_MESSAGES` moved to `planModeMessages.ts`; `planMode.ts` imports it without re-exporting. Under vitest, which runs the frontend adapters, the stale import resolved to `undefined` rather than failing to link, so `global-planmode1-hands-over-a-plan` threw on the approval message after the posture had already been dropped and the tool withdrawn. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(copilot): state plan mode's constraint in neutral text The composer's two-tone placeholder becomes a plain "Read-only" beside the autonomy picker, next to where YOLO puts its own warning, and a blocked call's row drops the mode colour. Teal is left marking what the posture is — the badge, the version bars, the pill — rather than every call it refuses. ContextTextarea goes back to main with the accent: `placeholderAccent` had no other consumer, and the aria-label existed only because the accent blanked the native placeholder. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(copilot): hold the plan header's verdict until the snapshot lands Opening a plan at the version its reader approved pins a version behind the head, and until that read resolves `shownVersion` is still the head — so the header wore the draft's badge and its orange "not approved" bar over the very case the pin exists to serve, then flipped. The header now says nothing while `restoringPin`, as the body already does. Judging `pinned` instead would print the approved signal over text that is still the draft, trading a true transient signal for a false one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(copilot): refuse a hand-over once plan mode has ended A response can carry two exit_plan_mode calls, and the tool list they run against is snapshotted before the first one restores the posture. The second then found the tool with plan mode already over: under YOLO every confirmation is answered for the user, so it wrote its own summary and stamped the user's approval on a plan no card had shown them. Refused in `validateBeforeConfirmation` rather than in `fn`, since `onConfirmationRequested` writes the document too. The maintenance path is untouched — a plan still gets revised outside the posture with update_artifact, which is what the tool's own description already tells the model to use. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
09c8f3b1f3 |
feat: redesign flow step, loop and branch settings panels (#10026)
* feat: responsive modal step panel for the flow editor in sessions On narrow layouts the flow editor's step-details pane opens as a modal (double-click a graph node) instead of a split pane, with a dock/float toggle. Scoped to sessions via allowModalPanel; the full-page editor is unchanged. - FlowEditor: modal/docked modes gated by mount width + allowModalPanel, small header (step-id Badge + subtle dock/close), standing double-click hint, and a per-step hint in the name tooltip - selectionManager: onSelectIntent hook so flow-level panels (settings, input, triggers…) open the modal on single click - PropPickerWrapper: collapse the prop picker until connect and animate it in via AnimatedPane (runs-page pattern), no blue connect ring in modal mode - StepInputGen: drop the TAB/Wand autocompletion button + spinner (feature still works via focus + Tab) - InputTransformForm: decouple the Help dropdown from the AI suggestion - FlowModuleHeader: move 'Save to workspace' into an ellipsis dropdown Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: loop editor rendering and nested splitpanes splitters in the sessions modal - Loop iterator/parallelism: keep the picker split pane (forceExpanded) so the editor fills its box and the picker shows; the collapse-until-connect mode stays for the step inputs - Remove the intrusive AI TAB/Wand autocompletion button from IteratorGen (generation still runs headless via focus + Tab) - Size the iterator connect plug and restyle the loop header/labels/toggles - Scope the global `.splitter-hidden` splitter-hiding rule to direct children so it no longer leaks into nested Splitpanes under the sessions preview Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: redesign flow step advanced settings as a single toggle-first column Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: taller step test pane by default and restyle advanced section titles Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: show flow run-settings params disabled when a setting is toggled off Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: single-column for-loop panel reusing the run-settings accordion Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: single-column while-loop panel reusing the run-settings accordion Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: single-column branch panels reusing the run-settings accordion Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: auto-open modal panel when creating an AI agent tool Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: redesign branch panels with card layout and shared predicate editor Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: remove per-setting status badges from flow map nodes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: sync package-lock after windmill-utils-internal bump Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style: polish prop-picker plug button and branch panel layouts * fix: persist skip-if-stopped toggles in early stop settings * fix: open the step panel modal on demand and cap its width * fix: restore graph step setting badges, strip panel header chips instead * feat: docked panel header with detach action and open-details step menu * feat: width-based panel mode on every surface with inline detach action * refactor: single source for flow step settings and their defaults * docs: pin flow editor vocabulary in CONTEXT.md * fix: open the trigger panel on double click or a specific trigger * fix: keep module pickers inside their pane and dismissable * fix: drop the misleading chevron on the MCP tool entry * fix: resolve flow approvals against the job's workspace, not the nav one * refactor: derive the approval workspace from the job, not from callers * fix: restore S3 snippets and gate params while their setting is off * fix: restore branch mock controls and address review findings * chore: drop stray debug log from the flow map item * feat: pinned output section for loop and branch panels * fix: open the panel for deliberate navigation from the flow header * perf: mount branch predicate editors on demand * fix: skip predicate picker previews the previous step's result * fix: flow-level graph nodes open their panel on a single click * fix: open the step panel for AI chat selections, not for undo * chore: drop dead console.log and duplicated modalPanel doc * fix: re-sync expression editors and scope error-handler settings * fix: match the failure module exactly and ignore unselectable nodes * fix: keep concurrency editable, honour module cache_ttl, tighten panel ids * fix: open panel from indirect selections, use presence for value-driven toggles * fix: don't open settings on error-handler delete, flush editors on unmount * fix: guard editor destroy flush, keep retry kind reachable * refactor: name the run settings panel after the domain vocabulary * fix: only write editor flushes to the step they belong to * fix: bind step panels by id so a delete can't retarget editor writes * fix: don't let the trigger picker's escape close the drawer beneath it * docs: condense two comments to the constraint they record * fix: arbitrate escape through the overlay stack instead of deferring to it * fix: key nested step blocks by identity so anchored bindings can't go stale * fix: untrack the overlay-stack push and drop the frozen branch binding * chore: state the escape rationale once, key branch lists, format * fix: let the topmost overlay own escape instead of the graph * fix: keep the dynamic-input help box out of static template fields * fix: restore the graph connect on the for-loop iterator * fix: end connect mode with the modal and keep it to docked panels * fix: never enter graph connect mode from the modal panel * fix: reveal inserted steps, restore editor pane size, unleak the drawer stack * fix: keep the enable-AI popover reachable in session panes * feat: add the connect policy and its single armed slot * refactor: one picker for every expression input * refactor: route every connect through one armed slot * fix: give every connect button the same footprint * fix: keep the connect ring from showing through the button * fix: keep flow card actions right-aligned beside the detach button * fix: give the connect ring an opaque ground to mask against * feat: dock the panel back without reopening it * feat: dock the panel from the graph control bar * style: round the graph control bar and size its glyphs * style: customize the graph controls through their supported api * style: build the graph control bar from lucide icons * fix: use the graph's tooltip component in the zoom controls * style: pad the graph controls and enlarge their glyphs * style: pad the graph controls and put dock at the bar's end * refactor: give settings rows the same popover picker as other expressions * fix: pass the wrapper's pickable properties to nested inputs * refactor: stack step settings and render every expression through the step input form * feat: split loop panels into tabs and rework the approval form * feat: anchor drawers to their host pane and give them a size floor * fix: mark the loop iterator expression as required * refactor: badge ee-only toggles instead of a warning line * fix: flag an empty loop iterator expression as an error * refactor: pick the early-stop flow status from one toggle group * fix: keep parallel loops uncapped unless a limit is opted into * fix: scope the overlay stack to its host and disarm connect on dismissal * fix: anchor the trigger picker to its host pane * feat: move diff into the menu when the top bar is narrow * fix: gate the result logs toggle to the graph popover * feat: raise the modal-panel breakpoint to 1280 * fix: anchor flow editor popovers and fullscreen to their host pane * fix: anchor overlays to their host pane and mute them when hidden * fix: portal hosted modals and menus into the pane they anchor to * fix: keep non-listening dialogs off the overlay stack * fix: drop the topmost gate from confirmation dialogs * fix: silence overlays in a collapsed preview panel * feat: rework the branch panels with tabs, reordering and add/delete * refactor: fold the detached-panel chrome into the card header * fix: give every flow panel a titled card header * fix: stop the step panel oscillating on an auto-height editor * feat: consolidate script panel actions and restore branch predicate AI * fix: restore the logs toggle on the flow result popover * fix: collapse the idle property picker in modal step panels * fix: stop the docked pane scrolling alongside its panel * fix: space the last settings row off the panel bottom * revert: always show the property picker pane in step panels * chore: keep the inline script AI button identical to main * fix: ask for AI input suggestions on click, not on hover * fix: keep graph connects armed and remount the parallelism input * style: reveal the predicate AI button on row hover * style: give branch cards a handle and delete column * refactor: arbitrate flow overlay escape through Disposable * fix: give the popover picker its results and re-narrow the EE badge * docs: correct loopSubset and guard the modal width measurement * fix: insert picked properties at the cursor in expression inputs * fix: give the expanded-subflow panel the shared header chrome * style: rename the suspend setting to Suspend until approval/resume * feat: open a step's modal when clicking the step already selected * feat: add an auto/attached/detached toggle for the step panel * refactor: pick the step panel's placement from one named menu * refactor: keep the panel-mode module's exports to what is consumed * feat: show each configured setting's value on its badge * fix: carry the suspend rename into the step settings registry * docs: name both gestures in the step explore hint * test: pin where the step panel goes for a given width and preference --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
fb82748296 |
fix: make on_behalf_of control permissions for scripts and flows (#10438)
* fix: make on_behalf_of control permissions for scripts and flows Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: inherit the recorded on-behalf-of identity when a preserving deploy omits it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep an omitted permissioned_as from re-versioning an unchanged script Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: derive the on-behalf-of principal from the email and reject mismatched pairs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop workspace deploys from carrying a source-workspace principal Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct the onBehalfOfPermissionedAs param doc Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: pin that workspace deploys never carry a source-workspace principal Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct the omitted-principal contract and refresh generated prompts Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep external-superadmin principals on email-only redeploys Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: scope the recorded principal to its workspace and prefer real accounts Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: carry the recorded principal correctly through drafts and set-permissioned-as Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: sweep draft identity pairs on email change and offboarding Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: leave group identities alone when sweeping a user's email Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: treat only g/ without an email as a group, and match the offboard preview Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop the group guard from skipping rows with no recorded principal Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: state the group guard once instead of restating it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: make the permissioned_as the only stored on-behalf-of identity Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: skip resolving the on-behalf-of address for sync clients that discard it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address the local review of the identity refactor Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve the on-behalf-of identity coherently across clones, offboarding and no-op deploys * test: pin that a fork keeps only the on-behalf-of identities that resolve in it * fix: decide a principal prefix-first everywhere and canonicalize bare addresses * fix: prefix a slash-containing address so a reader cannot take it for a group * fix: read an address as a username before the group- convention * fix: rewrite the canonical principal when an account's address moves * fix: keep the address form of a principal to accounts without a usr row * fix: reject an identity a job row cannot carry and read it uncached at dispatch * fix: count characters against the job identity width and cap the backfill * refactor: name the script/flow principal on_behalf_of, as apps do * docs: state the caller-must-authorize contract on the identity resolvers * fix: keep writing on_behalf_of_email until every worker reads the principal * fix: err high on the compatibility version and document the last resolver * fix: keep the compatibility address current through identity mutations * fix: carry the compatibility address with the principal on every copy path * chore: re-pin the EE ref to the companion branch merged with EE main * fix: key the dbt retry lookup on the stored principal * fix: keep a mixed-version address recoverable through a fork * fix: read a round-tripped address uncached so a redeploy is not rejected * fix: refuse an email change that would make a principal unenqueueable * chore: update ee-repo-ref to ac3d7d015296f041ae44ab6bc4953485f44d36e4 This commit updates the EE repository reference after PR #704 was merged in windmill-ee-private. Previous ee-repo-ref: 219b0b03905a1a0028054b3a4985724e77d09036 New ee-repo-ref: ac3d7d015296f041ae44ab6bc4953485f44d36e4 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
032300e28e |
feat: run dbt projects as a first-class Windmill runtime (#10326)
* fix: mount only the engine in the dbt jail, reject shadowed and malformed args
Review round 42.
The jail mounted the whole dbt cache directory, whose siblings of the
engine are `repos/` and `packages/` — other workspaces' private checkouts
and package trees, kept apart by cache key rather than by permissions. A
jailed project could read them. It now mounts the engine's own directory,
which the provisioner names; verified from inside the jail that `repos/`,
`packages/` and `state/` are invisible while the engine stays usable.
A `{{ placeholder }}` may no longer take the name of a run argument this
runtime defines. It was silently dropped from the signature, so a
descriptor like `value: "{{ select }}"` deployed and then could not be
run at all: the built-in `select` is an array and the interpolation needs
a scalar. Refused at parse, so the deploy says so.
A `vars` override that is not an object is refused rather than ignored.
Argument-schema validation is opt-in, so a string or an array silently
ran the descriptor's own vars — against a different schema or alias than
the caller asked for. `select` and `exclude` already refused theirs.
* feat(dbt): the project is the script's module bundle, not a git checkout
A dbt script now carries its whole dbt project as its module bundle. The
descriptor is the script content; `<script>__dbt/` holds the project verbatim,
so importing an existing project is `cp -r` plus `wmill sync push`, and the
worker materialises the bundle into the job directory instead of cloning.
Backend
- `prepare_project` writes the script's modules and requires `dbt_project.yml`
at the bundle root. `checkout`, the git-ssh command, the clone cache and the
repository resource are gone, along with `repo`, `project`, `ref` and
`git_ssh_identity` on the descriptor.
- Run identity and the package cache key take a `project_digest` (sorted SHA256
over the bundle) where the commit used to sit, so an edited project cannot
resume a previous run's `run_results.json` or reuse its `dbt_packages`.
- The per-run graph re-ingest is now gated on `vars` placeholders and `$var:`
env alone.
- `capture_dependency_job` takes the script's modules so a dependency job, which
has no generic module-writing step, materialises them itself.
- `dbt deps` caching strips the git remote from every package it cached, not
just the tree root: `packages.yml` can render a token into a `git:` URL.
- `git_clone.rs` is dropped and `ansible_executor.rs` returns to its own copy of
the clone helpers.
CLI
- `wmill sync pull` keeps a dbt script's lock beside its folder rather than
inside it, so the folder holds nothing but the project.
- Directories dbt generates (`target-path`, `packages-install-path`,
`clean-targets` and the usual defaults, read from `dbt_project.yml`) are
excluded from the bundle, from the sync diff and from staleness hashing.
- A module-only edit now pushes its parent dbt script and is reported as a
changed module rather than passing unnoticed.
* fix: keep a script's modules in the worker's file-system cache
The first fetch of a script version reads the database and carries its
modules; every later fetch imports from the worker's cache directory, whose
`RawScript::import` hard-coded `modules: None` and whose `export` never wrote
them. A worker restart therefore started running the script without its own
files, silently — for a dbt script, without its project, which fails with
"carries no project"; for any other script with a module bundle, with the
imports missing.
`modules.json` is now written on every export and required on import, so an
entry written by an older version fails to import and is refetched rather than
serving a stripped script for as long as the directory lives.
Also derives a dbt run's `project_digest` from the bundle the run actually
carries: `handle_dbt_job` was passing `None`, which collapsed every project in
a workspace onto one digest and let `dbt retry` resume a different project's
`run_results.json`.
* fix(dbt): give every phase the script's environment, bound the cache copies
`dbt deps` ran without the script's environment variables on an unsandboxed
worker, so a `packages.yml` resolving a private package URL through
`env_var()` could not see them while the package cache key was still built on
their digest. `with_invocation_env`, applied at three of the four call sites,
is folded into `dbt_command` so no phase can be added without it, and
`DBT_TARGET_PATH` is set after both environments rather than before.
The package cache copies ran through a bare `Command::output()`: the tree is
the project's, so a cancelled or timed-out job held its worker slot until `cp`
finished. Both the restore and the publish now run under the job poller like
every other phase.
* fix(dbt): only offer commands whose writes match the graph, honour packages-install-path
`dbt_command: run` is dropped from the allowed overrides. Asset dispatch fires a
script's deploy-time writes on any successful job, and `dbt run` covers models
only, so a project with seeds or snapshots notified consumers of relations the
invocation left stale. That is the same reason `test` was already excluded.
Narrowing what a run touches is `select`/`exclude`, which scope the graph too.
`dbt deps` writes to the project's `packages-install-path`, so a project that
moved it got no package cache at all: the publish found nothing at
`dbt_packages` and every job resolved its dependencies over the network again.
The path is read from `dbt_project.yml` and validated as project-relative,
since both cache copies are rooted at it.
Also states the sidecar's mutator contract at the module level: the dbt manifest
tables carry no RLS and grant `windmill_user` full access, so a user-scoped
transaction is not enforcement and every caller must have verified write access
to the script itself.
CLI: a module file is now grouped with its parent script for the push. Left in
a group of its own it got its own `alreadySynced`, so a push touching several
files of one bundle deployed the script once per file; the resulting versions
raced, and the asset graph could end up describing none of them.
* fix(dbt): seed a project for browser-created scripts, refuse a no-op retry
A dbt script created in the browser only got a descriptor, and the runtime
refuses a script whose bundle has no `dbt_project.yml`, so the advertised
Create → Deploy → Run path always failed its dependency job. New dbt scripts
now start with a project that builds: pointing `profile.resource` at a
warehouse is the one edit, and growing it is `wmill sync pull` plus a local
editor, which is where dbt development happens.
`dbt retry` builds its graph from the previous run's error, fail and skipped
nodes alone, so retrying an all-green run selected nothing and wrote nothing —
and a job that succeeds having written nothing still dispatches every
deploy-time write, waking every downstream consumer for relations no one
touched. Refused, with the reason.
CLI: a configured `target-path` or `packages-install-path` may be nested
(`build/target`), and `clean-targets` has a block form as well as an inline
one. Both are now parsed, and the exclusion compares the project-relative path
rather than the top-level segment, so a nested generated tree no longer lands
in the bundle and no longer makes a local `dbt run` look like a project change.
* fix(dbt): lock a project once, find the parent on either path separator
A dbt script's modules are its dbt project, not helper code with dependencies
of its own, so the generic per-module lock loop is skipped for it: the parent
lock already ran `dbt deps` and `dbt parse` over the whole project. Locking
each file separately re-materialised the bundle and re-invoked dbt once per
file, so a project of N files paid N project-sized passes and a large one timed
the deploy out. The 13-file fixture went from 14 relock passes to 1.
`pushParentScriptForModule` searched the raw path for `__dbt/`, so on Windows,
where the folder is spelled `__dbt\`, a module-only edit returned without
deploying its parent while the caller still recorded the file as synced. It now
goes through `getScriptBasePathFromModulePath`, which normalizes separators.
Also drops the last of the external-repository wording from the descriptor's
module docs and from the `codebase` rejection a user can hit.
* feat(dbt): infer the run form locally, keep test-only retries from cascading
`windmill-parser-wasm-yaml` 1.770.0 carries `parse_dbt`, so the browser and the
CLI derive a dbt script's run arguments from its descriptor instead of waiting
for the deploy to hand back a schema. Pins bumped in both.
`generate-metadata` was rewriting a dbt script's `lock` field on every run: a
dbt lock comes from the dependency job on a worker, so nothing generates it
locally and the resolved `!inline` reference was left inlined into the metadata
or blanked. It is restored instead, and a push straight after
`generate-metadata` is a no-op again.
A retry now needs a failed node that materialises something. `dbt retry` builds
its graph from error, fail and skipped nodes, and with `test_behavior:
after_all` a failing test is what `run_results.json` ends up describing — so the
retry reran tests, wrote nothing, succeeded, and still dispatched every
deploy-time write.
The dbt badge's destination is deterministic: writers outrank readers, and among
several writers of one relation (which the backend permits) the smallest id
wins, rather than whichever write edge arrived last.
* feat(dbt): browse the project and read a run's per-node result
Two views a dbt user expects and that the generic script surfaces do not give.
**The project.** A dbt script's editor gains a Project tab beside its
descriptor: the module bundle as the tree dbt itself expects, each file
read-only with syntax highlighting. The existing module tab strip is a flat row
built for a couple of helper files and does not survive a real project; a
13-file fixture already overflows it. Directories sort before files so it reads
like the checkout on disk, and an empty bundle explains the `cp -r` instead of
showing a blank pane.
**The run.** `DisplayResult` renders a dbt invocation's per-node breakdown above
the raw payload: totals, then a table of node, kind, target relation, rows and
time, with failures and warnings sorted first and carrying their message. The
data was already structured; it was being shown as JSON to scroll and PASS/WARN
counts to find in the log. On a failed run the same JSON rides in the error
message after the exit-status line, so it is parsed back out — that is the case
worth rendering, since the failing node is what the user came for.
* docs(dbt): say that profile.resource is what buys the asset graph
The starter descriptor described `profile.resource` as the thing rendered into
profiles.yml, with the project's own file as an equal alternative. It is not
equal: the resource PATH is the warehouse's identity in the asset graph, so a
project bringing its own profiles.yml runs fine and silently gets no assets, no
lineage and no cascade. The deploy already says so in its log; now the
descriptor a user starts from says it too, before they choose.
* fix(dbt): authorize a resource used only for asset identity, clean up after failed installs
A descriptor setting both `profile.profiles_yml` and `profile.resource` took
its connection from the project's file but returned the resource path as the
graph's warehouse identity without ever reading it. A script editor could
therefore publish `table://<any resource>/...` writes, and wake that
warehouse's subscribers, while connecting somewhere else. The resource is now
read on that path too — reading is what authorizes it — so the combination
keeps working for the case that wants it (keep your own profiles.yml, still get
lineage) and fails closed otherwise.
Provisioning cleaned up its staging directory only on the paths someone
remembered, so a run of failed or cancelled first-use installs accumulated
venvs, tarballs and installer scripts until the worker's disk was gone. All
three engines now hold their scratch paths in a guard that removes them on
drop, which is the one exit every path takes, cancellation included.
Frontend: `partial success` is dbt's word for a node that built but whose tests
failed, counted in `totals.error` and redone by a retry, so it ranks with the
failures instead of rendering green with its message hidden. And the run panel
now keys off the worker's engine discriminator rather than `{nodes, totals}`,
which is a shape an ordinary script can return. Both pinned by unit tests on
the extracted `parseDbtRun` helpers.
* feat(dbt): show a run's models on the run page
The run page is where you land on a running job, and until now it showed a dbt
run as streaming text: the per-node table only renders once the job has
produced a result, and the graph that moves per model lived on the pipeline
page you had to navigate to. A Models section now sits above the result,
scoped to the running script's own relations and its `ref()` lineage, polling
while the job is in flight so nodes move as dbt walks the DAG.
No `resolveGraph`: that merges drafts and live editor buffers into the
persisted graph, and a run page has neither.
* feat(dbt): retry failed nodes automatically, and from any worker
**Node-level retry, in the job.** `retry_failed_nodes: {attempts, delay_seconds}`
rebuilds only what a failed build left failed or skipped, before the job reports
failure. dbt confines a failure to its own subtree and `dbt retry` resumes
exactly that set, so a transient warehouse error costs those nodes rather than
the project. Doing it in-job is what keeps the state question out of it: the
previous attempt's `run_results.json` is still in the job directory, so there is
nothing to persist and no worker to land back on. This is the granularity
astronomer-cosmos gets from one Airflow task per model, without the ~6x that
per-model tasks measured.
A retry's `run_results.json` names only the nodes it redid, so it overlays the
accumulated results rather than replacing them: the job's result has to be every
node the job touched, or the nodes that succeeded before the retry settle no
materializations. Pinned by a test.
**Durable retry state.** `run_results.json` is now saved to `dbt_run_state` as
well as the worker's local cache, so an explicit `dbt_command: retry` works from
any worker of the group rather than only the one that failed. Only the results
are stored: `dbt retry` also needs `manifest.json`, roughly sixty times larger
and growing with the project (732 KB against 12 KB on the six-node fixture), but
the manifest is a pure function of the project files, vars and env, all of which
the stored identity already pins, so a worker restoring from the database
re-derives it with a `dbt parse` of about a second.
* fix(dbt): restore the sqlx cache, make retries cancellable and path-aware
**SQLx cache.** A `cargo sqlx prepare` deleted 750 entries, including the
enterprise queries CI needs under `SQLX_OFFLINE=true`, and the check that was
supposed to catch it reported zero losses because it was run from `backend/`
with a `backend/`-prefixed path, so its baseline was empty and it failed open.
All 750 are restored; the branch now adds 19 and deletes none, and
`SQLX_OFFLINE=true cargo check` passes.
**Retry backoff observes cancellation.** `canceled_by` is only written by the
job poller, which does not run between attempts, so re-reading it reported the
state as of the failed attempt and missed every cancel issued during the wait
— the whole window the check exists to cover. The wait now reads
`v2_job_queue.canceled_by` each second, and the job's deadline is honoured
before starting another dbt process.
**Retry state follows its script.** `dbt_run_state` is path-keyed like the
manifest sidecar but, unlike it, nothing regenerates it: a rename moves the row
so a resumable failure survives, while archive and delete clear it, so a script
later created at that path cannot inherit a stranger's failure and its
arguments.
**CLI.** `table` joins ducklake and s3object in the local graph's auto-trigger
kinds, matching `is_auto_trigger_kind` and the frontend's set; without it a
local graph and the generated docs omitted a cascade edge the deploy has.
* fix(dbt): carry only the project files a bundle can hold, and say what it drops
Exploring real and edge-case projects surfaced three frictions, all in the
import path a user hits first.
**A binary file broke the push, opaquely.** dbt projects carry images under
`docs/`, stray `.DS_Store` files and occasionally a parquet seed. Read as text
they become mojibake, and a NUL among them is rejected by Postgres with
`unsupported Unicode escape sequence` — which `wmill sync push` then reported as
success, exiting 0 with the script never created. Binary files are now detected
the way `git` detects them, by a NUL in the first 8000 bytes rather than by
extension, and skipped with the reason.
**The size guard the docs promised did not exist.** Now it does: 5 MB per file,
which only ever catches a committed dataset. Real dbt code is about 500 bytes
median and 1.9 KB at p90.
**Skipped files became a permanent phantom diff.** The push dropped them while
the sync diff still offered them, so every push reported changes no push could
resolve. One predicate now answers for the push, the staleness hash and the
diff.
Verified on a project with unicode filenames and content, CRLF endings, an
empty model, an ephemeral model, a disabled model, a `.md` docs block, an
extensionless README, six levels of nesting, a 7.6 MB seed and a PNG: it
pushes, round-trips byte-for-byte through pull, deploys to 7 dbt nodes and 6
`table://` assets (ephemeral and disabled correctly absent), and runs green.
* fix(dbt): resolve dbt-core against the adapter, settle partial success, unify status
**Adapters could not be provisioned.** The 1.x engine pinned `dbt-core` to a
fixed version independent of the adapter, but several adapters cap below it:
`dbt-mysql` at `~=1.7`, `dbt-oracle` and `dbt-databricks` below 1.12, and
`dbt-salesforce` has no package at all (it exists only inside Fusion). Those
projects failed at provisioning with a uv resolver dump. The install now asks
for a range and lets the adapter choose, and records what the resolver picked so
the lock pins a version that adapter can take.
The floor is the CLI this runtime invokes: resolving down to dbt-core 1.7
produced a working venv that then failed with `No such option '--target'`, which
is worse than not resolving. An adapter with no release in range now fails
naming itself and pointing at `dbt-core-2x` or `fusion`, instead of a resolver
dump. Salesforce is refused up front with the reason.
**`partial success` left a model stuck on `Running`.** It is dbt's word for a
node that built and then failed its tests, and it was already treated as a
failure when counting totals and deciding a retry — but the two sites that
settle the RELATION fell through to "says nothing", so the tailer's `Running`
was never replaced and a finished job showed a model still building. Six status
comparisons had drifted apart, two folding case and four not, while dbt-core 1.x
echoes the author's casing and 2.x uppercases; they are now one classifier.
**Agent workers.** The durable retry state and the cancellation poll both need a
database, which an agent worker reaches only through the API. The automatic node
retry is refused there rather than running a wait it could not interrupt, and
the docs say "any worker with a database connection" instead of overclaiming.
Also clears `dbt_run_state` when a path stops being a dbt script, and moves
`run_identity`'s contract onto `run_identity` from the digest helper below it.
* feat(dbt): show the transform behind a model on the run graph
The run page's graph carried a node for the script itself and drew every
relation as a bare table. Both were wrong for that page: the graph there is
already scoped to one script, so a node standing for it distinguishes nothing
(on the pipeline page it separates one project from another, which is why it
exists), and dbt's own DAG node is the model — the SQL and the relation it
writes are one thing, so a graph of relations alone leaves out what a reader
came to see.
The script node is dropped, and selecting a model now shows its SQL underneath
the canvas with its file path and materialization, read-only, the same view the
pipeline details pane gives.
* feat(dbt): move the graph with the run
The worker has always recorded a state per relation as dbt walks the DAG —
`running` when a model starts, `materialized` or `failed` when it ends — but
nothing rendered it: the graph response carries what a relation IS, not what a
particular run is doing to it, so the canvas had nothing to show and a running
job looked identical to a finished one.
`assets/run_progress/{job_id}` returns that state for one job, the run page
polls it beside the graph, and the asset node carries a spinner or its outcome.
Errors and retries need nothing extra: a failed node writes `failed`, and an
in-job retry rewrites the same row, so the node returns to `running` and on to
its new outcome by itself.
`materialized_partition` holds a relation's CURRENT state keyed by relation, so
filtering on `job_id` returns exactly what this run last touched — which is the
question a run page asks, and why a superseded older run shows nothing.
* feat(dbt): a dbt project is not a data pipeline
Deploying a dbt script marked it `auto_kind = 'pipeline'`, which enrolled it
in pipeline membership: the folder became a Pipeline entry on the home page,
the script folded into it, and `/pipeline/<folder>` opened a canvas holding
the project's whole model DAG next to the pipeline's own scripts. A folder
holding both then read as two projects in one editor, and the pipeline editor
offered to author transforms that are in fact authored in a local `dbt run`
loop and pushed as the script's bundle.
A dbt script is now never a pipeline member, and the pipeline canvas drops the
dbt script node. Its models stay, with their `ref()` lineage: the relations are
what a downstream pipeline script reads, and dropping them would break the
cascade from a dbt run — the point of giving dbt models `table://` identity.
Also drops a screenshot committed to this branch by accident.
* fix(dbt): authorize run_progress through the job, drop dbt from the local graph
`run_progress` read `materialized_partition` through `user_db` on the
assumption that RLS would scope the rows. That table has RLS disabled and no
policies, so any workspace member could pass a job id and read that run's
relation paths, row counts and error text. It now joins `v2_job`, which does
carry per-user policies, so a caller who cannot see the job sees nothing —
the same pattern `v2_job_completed` reads need. Verified as a plain member:
the old query returned 6 rows for another user's run, the new one returns 0,
while the job's owner still sees all 6.
The CLI's local graph still forced `in_pipeline` on every dbt script, so
`pipeline docs --local` and `pipeline dev` kept presenting a dbt project as a
pipeline the deploy no longer enrolls. It now skips them, matching the server.
A dbt descriptor has no asset parser locally, so nothing is lost: its models
come from the manifest the deploy derives.
Declares `run_progress` in openapi.yaml so the frontend uses the generated
client instead of a handwritten fetch; the generated `status` union also
replaces a hand-rolled string mapping.
* fix(dbt): drop the dbt node from the CLI's deployed pipeline views too
`pipeline dev` and `pipeline docs` (without `--local`) read `/assets/graph`
directly. That endpoint is asset-usage driven rather than membership driven, so
it returns a dbt script like any producer — and both commands render every
runnable, so a dbt project still showed up as a pipeline script there after the
local builder stopped emitting one.
`hideDbtRunnables` mirrors the frontend's projection of the same payload. It is
generic over the graph shape so the bounded-cascade view (`BCGraph`, a narrower
type over identical JSON) passes through without a cast.
The relations stay: they are what a downstream pipeline script reads, and the
node is what attributes them to a producer for every other consumer of the
endpoint, so the filter belongs in the views rather than the query.
* fix(dbt): narrow a selective run's cascade, settle the finished run graph
Review-round fixes.
A `select`/`exclude` run builds part of the project, but asset dispatch reads
the deploy-time write set for the whole script, so a run selecting one model
woke the subscribers of every other. Dispatch now intersects that set with the
relations the run actually recorded as materialized, scoped to dbt because it is
the only producer whose write set is decided per run. A run that recorded
nothing still dispatches everything, so an agent worker whose reconciliation
failed cascades as before. Verified both ways: `select: [extra_model]` no longer
wakes the `fct_orders` subscriber, and a full run still does.
`hideDbtRunnables` keyed its removal set on path alone while the graph keys
runnables by `(usage_kind, path)`, so a flow sharing a path with a dbt script
lost its node, edges and triggers too. Both copies now key on the pair.
The run graph never took a final reading when a job finished, so the last state
shown was whatever the tick before completion saw. Only `dbt-core-1x` streams
node events; the other engines record every relation during end-of-run
reconciliation, so their finished graph showed nothing until a reload.
`DbtNodeOutcome::Inconclusive` collapsed statuses the tally has to tell apart,
so two sites re-lowercased the status beside the classifier and `no-op` landed
in `totals.error` — a clean run reporting an error in its own result. Split into
Warn / Skipped / NoOp / Unknown so every site falls out of one match; `no-op` is
kept out of the retry set, which dbt spells as error / fail / skipped.
Also: reattach two doc comments to the items they describe, and correct the
engine-distribution table — only dbt-core-2x is baked into the images, 1.x is a
per-adapter venv provisioned on first use, and the default is compiled in rather
than an instance setting.
* fix(dbt): make the model chip inert where its project node is not on the graph
The canvas passed `onDbtSelect` unconditionally, so the chip always rendered
`cursor-pointer` and hover-highlighted — but the owner map is empty on both
graphs this feature added, since the run page carries no runnables and the
pipeline page hides the dbt node. The chip advertised a click that resolved to
nothing. It now takes its handlers only when the relation has an owner on this
graph, so it stays live on the surfaces that do show the project node.
`classify_status` and `DbtNodeOutcome` were `pub` in a private module with no
caller outside the file, unlike every neighbour.
* fix(dbt): take the cascade's write set from the run's own result
The previous narrowing read `materialized_partition`, which was wrong twice.
That table keeps one row per relation and the newest writer takes `job_id`, so
two overlapping runs over the same model erase each other's claim to it: the
earlier job would dispatch a subset of what it built, or none of it.
And an empty row set was read as "recording failed, dispatch everything" when it
is also a real answer. A `select` matching no model, or one resolving to tests
only, exits 0 having built nothing — and then woke every consumer of every model
in the project, which is the opposite of what the narrowing exists to do and is
reachable by a typo in a run argument.
The run now reports the relations it materialized in its own result, which is
immutable and per job. Absent means the producer said nothing (a job from before
the field, a non-dbt producer) and the whole deploy-time set dispatches as
before; present-and-empty means it built nothing and dispatches nothing.
Verified on all three: an unmatched selector builds nothing and wakes nobody, a
selector naming one unsubscribed model wakes nobody, and a full run wakes the
subscriber.
Also indexes `materialized_partition (workspace_id, job_id)` -- the run page
polls that shape every 2s and no existing index leads with `job_id` -- corrects
the selective-cascade section of the design doc, which still described the old
deploy-time behavior, and reattaches `buildLocalPipelineGraph`'s doc comment.
* docs(dbt): attach the CLI JSDoc to its function, correct the index rationale
The `hideDbtRunnables` JSDoc ended up documenting the type declared beneath it —
made while fixing the same mistake one function down.
The migration's comment credited the cascade with a `job_id` lookup that the
same commit replaced with a read of the job's own result. The run page's poll is
the only reader keyed on that column.
* fix(dbt): refuse graph publication for a removed script, allow test-only retries
An archived or hard-deleted script could still republish its graph: the
publication guard filtered `deleted` but not `archived`, and treated a missing
row as "nothing newer exists" rather than "nothing left to publish for". A
dependency job or dynamic run finishing after the removal put the asset,
provenance and subscription rows back with nothing left to clear them.
`dbt_command: retry` refused a run whose only failures were tests, which is
precisely what `test_behavior: after_all` produces. That restriction existed
because a successful job dispatched its whole deploy-time write set, so a
test-only retry would have woken every consumer for relations no one touched —
the cascade now dispatches what the run reports materializing, so it wakes
nobody and the restriction only blocked a legitimate retry.
* fix(dbt): gate run progress behind the job-read check, not RLS alone
The endpoint joined `v2_job` so RLS would decide visibility, which it does — but
`require_job_read_access` adds two things RLS does not: a scoped token's
`if_jobs:filter_tags` restriction, and the app-embed cutoff that stops untrusted
app JS from inheriting the viewer's broader job access. A scoped or embed token
could therefore read relation names, statuses, row counts and errors for jobs
the ordinary job endpoints deny it.
That helper is private to `windmill-api`, which depends on `windmill-api-assets`
rather than the reverse, so the endpoint moves to the job routes instead of the
check being duplicated. It is job-scoped anyway:
`/w/{ws}/assets/run_progress/{job_id}` becomes
`/w/{ws}/jobs/run_progress/{id}`, and the frontend follows the generated client.
* feat(dbt): a dbt run does not trigger downstream runs
dbt orders its own DAG, so a cascade only ever adds one thing: waking a Windmill
script that reads a mart. That edge is narrow, and only half of it can even be
expressed — nothing outside dbt can declare a `table://` write, since
`// materialize` accepts DuckLake targets only, so an ingestion script cannot
wake a dbt project.
Against that, dispatching correctly is not cheap. A run's `select` can build any
subset of the project, so the deploy-time write set is not what ran; using it
wakes consumers of relations the run never touched, and narrowing it needs a
per-job record of what was built. The per-relation state table cannot supply one
(it keeps a single row per relation stamped with the last writer), and the
result field added for it made a run's own output carry the cascade's bookkeeping.
So `asset_dispatch` returns early for `ScriptLang::Dbt`, before the producer
gate. dbt still materializes, records per-model state and publishes its graph:
models, `ref()` lineage and live run progress are unchanged, and a
`# on table://<mart>` reader still renders beside the model it reads. It simply
does not fire. Wiring it up later means deciding what a selective run should
notify, which is the actual work.
Verified: a full run of a 6-model project succeeds and starts nothing, where it
previously triggered its subscriber; the run page still reports all 6 relations
and the folder graph still carries 14 tables and 8 ref() edges.
* fix(dbt): remove the cascade surface, settle stranded models, fix nested __mod
Stopping dispatch left its surface behind. `table://` was still an auto-trigger
kind, `persist_ingest` still derived subscriptions from a manifest's reads, and
the deploy still accepted `# on table://` — so the canvas drew cascade arrows
into scripts nothing could wake. All three are gone: the kind no longer derives,
the ingest only deletes rows earlier versions wrote, and the deploy refuses the
annotation with a message saying why rather than persisting a silent no-op.
`DescriptorTriggers` went with them; every field it parsed was cascade config.
A model marked `running` by the live tailer was never settled when the run did
not finish: reconciliation only revisits nodes `run_results.json` names, and a
cancelled or timed-out run has none for the model in flight, so the finished job
showed a relation building forever. It is now settled on every exit path.
Verified by cancelling a run mid-flight: 3 models `running` before, 3 `failed`
after, none stranded.
`getScriptBasePathFromModulePath` took the first matching suffix rather than the
outermost boundary, so `proj__dbt/models/legacy__mod/a.sql` resolved to
`proj__dbt/models/legacy`. dbt owns its directory names verbatim, so a folder
ending `__mod` is legal inside a project, and a module-only sync would have
looked for a descriptor that is not there and skipped the deploy.
* fix(dbt): colour a finished run's models from its own result
`materialized_partition` keeps one row per relation stamped with whichever job
wrote it last, so reopening a run showed only the models no later run had
touched since — down to none for an old run, which reads as a broken page rather
than as stale data. Reproduced: a 6-model run reported 6 relations, then a second
run rebuilt one shared model and the first reported 5.
A finished run already carries the answer. Its result lists every node with a
status, and the graph carries each asset's dbt `unique_id`, so the two join
directly — no path derivation, nothing stored twice, and nothing a later run can
overwrite. The endpoint stays for the live window, where the result does not
exist yet, and as the fallback for a run that never produced one (cancelled or
killed, whose relations the worker settles in the table instead).
`relationOutcome` mirrors the worker's `classify_status` so the colour drawn over
a record agrees with the record: `warn`, `skipped` and `no-op` leave the relation
untouched and stay uncoloured, as do tests and analyses, which match no asset.
Verified in the browser on the run whose model had been stolen: all six
relations green again, both sources correctly uncoloured.
* docs(dbt): record why only dbt-core 1.x has live per-model progress
`emits_node_events()` reads as "the Rust engines produce no node events", which
is false and would close off the option. They produce exactly the same events;
they put them on the console and ignore `--log-format-file json`, which both
accept. Measured on 2.0.0-alpha.5 and fusion 2.0.0-preview.202: 15 node events
each on stdout, 0 in the file log, for a three-model project.
Taking them means owning the job log's presentation to work around a flag that
is documented and simply unimplemented, so the note records the measurement, the
sample event, and that flipping the predicate is the whole change once either
engine honours it.
* fix(dbt): give HighlightCode a dialect-agnostic sql language
`npm run check` had three errors the fast check does not reach: `"sql"` is not a
value `HighlightCode` accepts. Every SQL dialect it knows maps to one grammar,
but a dbt model is compiled by whichever adapter the project targets, so naming
a dialect would be a guess — `sql` is now a value in its own right.
`langOf` was typed `string` and returned `markdown`, `python` and `text`, none
of which the component accepts either, so a dbt project's YAML and Python files
rendered unhighlighted. It now returns the component's own prop type, which is
what caught them, and `undefined` for what has no grammar rather than a name
that silently means the same thing.
Verified in the project panel: SQL 22 tokens, YAML 27, where YAML was plain.
* fix(dbt): stop failing no-op models, drop table triggers client-side, keep cross-selection edges
The sweep that settles a run's stranded relations was marking `no-op`, `warn`
and `skipped` models FAILED on successful runs: reconciliation reports those
nodes without settling their record, so they were indistinguishable from a model
the run never reached. It now excludes every relation the run accounted for, so
only the genuinely abandoned ones are settled.
`table` was removed from the backend's auto-trigger kinds but left in both
client mirrors, so the editor and `pipeline dev`/`docs` kept drawing cascade
arrows the deploy will not create.
`isModuleEntryPoint` scanned for the first `__mod/`, the same bug its sibling
just had: a `legacy__mod/script.ts` nested in a dbt project — dbt owns those
names verbatim — read as that script's entry point. Both now anchor on the
outermost boundary.
A script selecting a model whose parent another script builds dropped the parent
entirely, so no `dbt_edge` could reach it and the two relations sat on the graph
unconnected. The parent is now kept as an endpoint and recorded as a READ, since
this script does not build it — splitting a project across selections only
composes if the seam still draws.
* fix(dbt): don't double-run after-all tests, count only models a script builds
An `after_all` run whose test phase failed saves a `run_results.json` holding
tests alone. Retrying it reran exactly those tests — and then the test phase ran
the whole suite again, appending a second copy of every result: duplicate ids in
the run table, doubled totals. A retry whose saved results are tests alone IS
the test phase, so the suite is not run after it, and the two phases now merge
by node id rather than concatenating.
Keeping a selection's unselected parents as nodes made them count toward the
`×N` badge, whose tooltip says "materializes N models" — a script selecting one
mart claimed the staging models upstream of it, and the number grew with the
seam. The count now comes from the relations the script writes.
That change also made the cross-selection read block dead, with a comment
asserting the inverse of what now happens; it is removed, and the test that
covered it still passes on the new arm. The test I added landed between a
neighbouring test's comment and its `#[test]`, orphaning the attribute so that
test stopped running.
Two display fixes: the run page no longer shows a relation's SQL when the
provenance belongs to another project that materializes the same relation, and
the editor no longer draws an explicit `# on table://` arrow the deploy refuses.
The starter descriptor no longer promises the removed cascade.
* fix(dbt): clear untouched models, reject unknown descriptor fields
A `no-op` model was left `running` forever on a successful run. The previous
attempt at this stopped the sweep marking such models FAILED but gave them no
terminal state instead, so they simply never settled. Reconciliation now returns
what it settled and what the run reported but did not build, and the two get
opposite treatment: a relation the run left untouched has its row DELETED, which
is what the finished run's own result says about it (`relationOutcome` colours a
`no-op` nothing), so the live and settled views agree; only a relation the run
never reached at all is failed.
The descriptor accepted unknown fields, so `selcet:` was ignored and left an
empty selection — building the whole project — and a misspelled `target` fell
back to the profile's default. It rejects them now. That immediately caught two
of our own test fixtures still passing `repo:`, a field removed with the git
path, which is exactly the class of mistake it exists to stop.
`isDbtModulePath` matched `__dbt/` anywhere in a path, the third site with that
bug: `foo__mod/vendor/x__dbt/a.ts` read as a dbt project file, and the push then
looked for `foo.script.yaml` and could skip the edit.
A verbatim dbt bundle dropped any file named `*.lock` before it reached the
module map, so an authored `uv.lock` never deployed and the unmodified-project
round trip quietly lost it. The exclusion now applies only to `__mod` bundles,
where `.lock` really is the script's own lockfile — in the walker that hashes
modules too, or a change to such a file would not register as one.
Also: `langOf` fell back to `undefined`, which HighlightCode resolves to
TypeScript rather than to no highlighting, so seeds and Markdown were coloured
as code; and four comments still gave the removed cascade as the reason for
sharing an asset node, which is now lineage.
* fix(dbt): retry failed tests too, anchor the last __dbt path check
`retry_failed_nodes` only ran after the model phase, which fails before the
`after_all` test phase exists — so a project whose models built and whose tests
failed got no retry at all, exempting exactly the failure mode that separate
phase produces. The loop is now a function, called after both phases.
`isDbtGeneratedPath` matched `__dbt/` anywhere, the fourth site with that bug:
`foo__mod/vendor/x__dbt/target/a.ts` counted as generated dbt output, so
`ignoreF` excluded an ordinary module file and a module-only edit never deployed
its parent script.
`wmill sync push` still dropped an ADDED or DELETED `.lock` three branches
before the module arm, so the earlier fix only covered a first push: adding a
`uv.lock` to a deployed project was reported as a change forever and never
applied, and deleting one left it deployed. Editing worked, which is what made
the round trip look whole.
Also removes a duplicate `#[test]` that was double-registering a test and
detaching its neighbour's comment, and rewrites seven comments that still gave
the cascade as the reason for behaviour that now serves lineage only.
* fix(dbt): bound the excluded-file read, keep the retry budget job-wide
`isBundledModuleFile` read a file in full before deciding it was too big or
binary, so a project sitting next to a multi-gigabyte parquet seed loaded the
whole thing only to reject it. It now takes the size from `stat` and reads at
most the 8 KB the NUL check needs: a 191 MB file is rejected in 0.0ms at 82 MB
RSS.
Calling the retry helper after both phases gave each its own `attempts` budget,
so a job could spend double what the descriptor asked for — the bound exists
because every attempt is a real dbt invocation holding a worker slot. The budget
is now the job's, spent across whichever phases fail, and the field says so.
Extracting that helper had also placed it between `#[allow(clippy::
too_many_arguments)]` and `run_dbt`, taking the attribute off the 12-argument
function it was written for.
* fix(dbt): actually spend the retry budget
`retry_failed_nodes` looped on `while *remaining > 0` and never decremented it,
so a failing job reissued `dbt retry` — logging "attempt 1 of 3" each time —
until the job's deadline instead of `attempts` times. The decrement existed
briefly and was lost when the function was re-extracted by hand.
Claiming and counting are now one operation, `claim_attempt`, because keeping
them apart is exactly how the bound goes missing: the loop cannot iterate
without spending the budget.
Its test is bounded by its own `for` rather than by the function under test. An
earlier version collected `std::iter::from_fn(|| claim_attempt(..))`, which
against a non-spending `claim_attempt` is an infinite iterator — it allocated
until the machine died. A test for a loop bound must fail an assertion when the
bound regresses, not consume the host: it now reports `[1, 1, 1, …]` against
`[1, 2, 3]` in 0.00s.
* fix(dbt): ask before reading, not after
Bounding `isBundledModuleFile` did nothing for the bundle builder, which read
the whole file into memory and only then asked whether to keep it — so a
multi-gigabyte seed beside a project was still loaded in full just to be
skipped. The predicate is now consulted first, and the read happens only for
files the bundle actually carries.
* feat(dbt): animate the ref() edges feeding the model being built
The nodes moved during a run but the edges did not, so the graph showed where
dbt had got to without showing it flowing there.
Reuses the canvas's existing rule rather than adding a second one: an edge
animates when it touches what is happening. For a pipeline that is the running
script; for dbt the unit of work is the model, so a `ref()` edge animates while
its target builds. Same `animated` field, same visual language, no new styling.
Verified mid-run on a 7-model project: of six `ref()` edges only the two feeding
the model then building were animated, and none once the job finished.
* feat(dbt): show what each model wrote, and say when its SQL is another project's
Three things a reader wanted from the run graph and could not get.
Row counts: the worker already records one per relation and `run_progress`
already returned it, but the graph used only `status` and dropped the number. A
model that built green having emitted zero rows is the failure that looks like a
success, so the count is on the node.
The relation's fully-qualified name, copyable: there is no table browser to open,
so the next best affordance is the exact identifier to paste into a SQL client.
It is parsed with `splitRelation`, which honours quoting the way the worker's
`split_relation` does — splitting on every period renders
`"wh"."analytics.v2"."orders"` as a relation `orders` in a schema `v2`, which
does not exist.
And when two projects materialize one relation, the graph keeps a single
provenance winner, so the losing project's node carries the other's model. The
SQL was already suppressed there — correctly, it is not this run's code — but
silently, which reads as a dead click. It now says so.
* fix(dbt): a finished run's graph is the models it built, not today's project
`/assets/graph` is the current deploy, so an old run's graph drifted with the
project: a model added after it appeared as though the run had built it, and the
older the run the wronger the picture. A finished run's node set now comes from
its own result, which named exactly what it touched.
Sources survive the filter regardless — dbt never lists them in
`run_results.json` because it does not build them, but they are the upstream the
run read, and dropping them would leave the models hanging.
The graph is still the current deploy's, so a model renamed or deleted since
cannot be drawn at all. Rather than a silently shorter graph, the count is
stated above it.
Verified by adding a model after a run: the old run renders 7 models without it,
a fresh run renders 8 with it.
* feat(dbt): preview a model's rows with `dbt show`
There was no way to see the data behind a node — only its SQL and its row count.
`dbt show` selects from a model and returns rows, and every engine ships it, so
the preview needs no adapter code of ours: no connection path, no dialect-correct
quoting, no type coercion for ten warehouses. It runs against the profile the
run already renders.
It is a `dbt_command` rather than a new endpoint, so it inherits the whole job
path — authorization, isolation, cancellation, logs, engine provisioning — and
`limit` joins the run form beside it. That the allowlist can admit it at all is a
consequence of dropping the cascade: while a successful job dispatched its
deploy-time write set, a command that wrote nothing woke every consumer for
relations nothing had touched.
Read-only, and treated as such: no graph republish, no materialization records,
no retry state, no test phase. Captured rather than streamed, like `dbt ls` —
these rows are the result, not commentary, and the job-log writer is what
`NO_LOGS_AT_ALL` discards.
Verified: `{"dbt_command":"show","select":["stg_customers"],"limit":3}` returns
three rows; a preview leaves `materialized_partition` untouched (62 → 62, 0 rows
for the job); `clean` is still refused by the allowlist.
* feat(dbt): preview a model's rows from the graph, and keep our locks out of dbt projects
The run page could show a model's SQL and how many rows it wrote, but not the
data. Selecting a model now offers "Preview rows", which runs the script with
`dbt_command: show` and renders the result as a table.
Explicit rather than on-select: a preview is a job, so it costs a worker slot
and the engine's start-up, and previewing on every click would spend both on
mere navigation. Sources are excluded — dbt shows what a model SELECTs, and a
source is not one.
Also: `updateModuleLocks` was the one module helper that never learned about
verbatim bundles, so it walked a dbt project writing `foo.lock` beside `foo.sql`.
None of those files is a Windmill script needing a lockfile, and the bundle
promises to round-trip the project byte-for-byte — our artifacts have no business
in it.
Verified in the browser: selecting `stg_customers` and previewing returns the
columns `id`/`src` and five rows from the warehouse.
* fix(dbt): keep a run's models when another project owns their provenance
Scoping a finished run's graph to the ids it named dropped relations whose
provenance winner belongs to a different project — so a run of a project sharing
a schema showed 3 of the 6 models it had built. An id that was never this run's
package cannot be judged against its result, so it is kept: the relation IS one
the run wrote, and hiding it understates the run. The same rule applies to the
"no longer in the project" count, which otherwise reported deletions that were
only provenance collisions.
Previews are now cached per model and survive the selection moving. One was
thrown away whenever the reader clicked elsewhere, which for a job costing a
worker slot and an engine start-up meant re-running it to see it again — and the
run continues in the background, so leaving and returning finds the rows there.
The spinner also never span: `startIcon` takes the icon and its classes
separately, so the animation has to be passed alongside.
How long it took is shown with the rows. A preview is a job, and its cost should
not be something the reader has to guess at.
* fix(dbt): resolve argument references, clamp the show limit, flag renamed relations
`handle_dbt_job` cloned `job.args` where every other executor calls
`build_args_map`, so a `$var:` / `$res:` / `$encrypted:` argument reached dbt as
the literal string. A placeholder holding a schema or an `enabled` flag would
then build a different slice of the project than the caller asked for.
`--limit` took any positive i64, and the worker buffers the whole of dbt's
stdout to read the rows out of it — so a caller with only run permission could
make it hold an unbounded allocation. It is clamped to a ceiling now, extracted
as `show_limit` so the bound is pinned by a test rather than inline in an async
function nothing can reach.
And a model keeps its id when its alias or schema changes, so an old run's node
showed today's relation while the run wrote another — the page asserting it had
materialized a table that did not exist yet. The run's result carries the
relation each node actually wrote, so the drift is detectable without a graph
snapshot, and the count is stated above the graph. Rendering the run's own
lineage still needs a per-job snapshot; this stops the page claiming otherwise.
* fix(dbt): stop persisting resolved secrets, bound the preview by bytes
Resolving `$var:` / `$res:` / `$encrypted:` for dbt — added in the previous
commit — meant `save_run_state` wrote the resolved PLAINTEXT into
`dbt_run_state.args` and the worker's `state.json`. The row outlives the job, so
a secret stayed in the database and a later `dbt_command: retry` replayed it
after the grant was revoked or the value rotated. The invocation now carries the
args as submitted alongside the resolved ones, run state persists those, and the
restore path resolves them again under whoever is retrying.
Clamping `--limit` bounded the row COUNT, not the size: one column can hold a
megabyte, so a thousand rows is a thousand megabytes, and `run_capturing`
buffers all of it. The captured output has a byte ceiling now.
`limit` became a built-in argument without joining `RESERVED_ARG_NAMES`, so a
descriptor writing `{{ limit }}` was silently handed the preview control's
default instead of being told the name is taken.
Two display fixes: the relation-drift banner compared a canonicalized (lower
case) asset path against the warehouse's own spelling, so it fired on every
model of every finished Snowflake run; and caching a preview's failure left
`Preview rows` dead for that model until reload.
* feat(dbt): key the graph by script version so a run renders its own project
The dbt graph was keyed by path alone, so a deploy overwrote the only copy and a
run page could only ever show today's project — an older run rendered today's
models, SQL and `ref()` lineage no matter what it had run. My previous attempt
filtered that view to the ids the run named, which stopped it lying but could not
show what was gone: the data no longer existed.
`dbt_node` / `dbt_edge` now carry `script_hash` in their primary key, so each
deployed version keeps its own graph, and the run page passes the version its job
recorded. Per DEPLOY, not per run — ten thousand runs of one version share one
graph — and a composite FK to `script (workspace_id, hash)` with ON DELETE
CASCADE means a version's graph dies with it. Nothing pruned these before,
because there was one copy per path; they would otherwise have accumulated with
no sweep.
Two deploys of one path now write disjoint rows, so the graph can no longer be
lost to a race. `claim_graph_publication` remains only for what is still
path-keyed — the `asset` usage rows — and an older deploy finishing late records
its own graph before declining to touch those, where before it published nothing
at all.
A pinned request is scoped by the version's own nodes rather than by `asset`:
that table describes the current deploy, so scoping through it would filter a
model out of the very run that built it.
Verified end to end: deployed v1 (8 models), ran it, deployed v2 with four models
removed and one rewritten. The old run renders 8 models, 6 ref() edges and v1's
SQL; a new run renders 4 and the v2 rewrite.
* fix(dbt): scope graph cleanup to one version, bound the preview capture
Archive and delete both act on a single `hash`, but the graph cleanup they
called deleted every row for the path. Now that the graph is keyed per
version, archiving an old version erased the live one's models, SQL and
lineage, and nothing repaired it. Both callers have the path in hand, so the
by-hash wrapper is gone and they use the version-scoped clear directly.
`dbt show` checked its 8 MB ceiling after `wait_with_output` had already
buffered everything, so the ceiling could not bound what the worker held.
`run_capturing` now reads both pipes incrementally against a caller-supplied
limit and kills the child on overflow. The read buffers are heap-allocated:
as arrays they were baked into the future, which the job poller boxes several
layers deep, and that overflowed the worker thread's stack — a `dbt show` run
aborted the whole worker process.
A retry's `dbt parse` ran on the arguments as submitted while the build ran on
resolved ones, so a `$var:` shaping the graph parsed verbatim. The parse moves
to the caller, after resolution.
A run that names its own `select`/`exclude` now drops the descriptor's
`selector`: dbt resolves `--selector` instead of `--select`, so passing both
made a preview of one model return another's rows.
Also: log instead of silently swallowing a `modules` column that fails to
deserialize (pre-existing, but for dbt it means running with no project at
all); keep the model SQL reachable once a preview has landed; render which
node the rows came from; stringify object-valued cells; document
`dbt_script_hash` in the OpenAPI spec.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: record the per-worktree dev environment and the backend-run check
Three mistakes this guidance would have prevented, each of which cost a cycle:
A worktree has its own database and ports, but AGENTS.md stated the
single-checkout defaults as facts. Pointing `DATABASE_URL` at another
worktree's database makes `cargo sqlx prepare` fail on every query touching a
table your migrations added — and it deletes `.sqlx/` before it fails, so the
cache is gutted rather than merely stale. Starting a backend on the wrong port
leaves the UI up with every call 502ing, which reads as an application bug.
Both values are now discoverable with commands that work as written.
`prepare` is also documented as the wrong tool for a removal-only change: the
cache is already complete for CI, and the only residue is orphaned entries that
can be found by text-matching against the sources without a database.
Nothing told a reader that `cargo check` does not exercise a worker path. A
read buffer declared as an array inside an async block is baked into the
future, and once boxed by the job poller it overflows the worker thread's
stack — compiling and unit-testing clean while aborting the whole worker
process at runtime.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): scope the remaining path-wide reads and clears to one version
Four places still spoke for a whole path after the graph became per-version:
The relation-root drift check read `dbt_node` by path with an unordered
`LIMIT 1`, so with v1 at root A and v2 at root B it could answer with v1's
row, suppress the refresh v2 needed, and leave v2's graph naming relations the
run does not build. It now reads this job's version.
`dbt_dep`'s no-resource branch cleared the path, so a descriptor edited to
bring its own `profiles.yml` emptied every earlier version's graph and with it
every finished run's page. The ownership being given up is the path-keyed
`asset` usages cleared beside it; the graph clear is now this version's.
The graph was inserted before the publication claim checked the version was
still live. Archive and delete only soft-update `script`, so the foreign key
still accepted an in-flight dependency job's rows and the failed claim
committed them — and because pinned queries deliberately serve archived
versions, deleted model SQL became readable again. The write is now gated on a
`FOR UPDATE` liveness check.
`clear_dbt_run_state_by_script_hash` resolved a hash to a path and deleted the
path's saved run. `dbt_run_state` is keyed by path by design — one saved run
per script — so archiving one version discarded the live version's resumable
failure. It clears only once no live version of the path is left; `identity`
already refuses a resume whose project, warehouse or engine moved.
"Preview rows" ran `runScriptByPath` while the SQL beside it was pinned to a
hash, so an old run showed its own SQL over today's rows. Verified end to end:
with v3 deployed, the v2 run's preview runs v2's hash and returns v2's rows.
Also: keep the TAIL of a captured stderr, since dbt prints its summary last;
one `$derived` for the parsed result rather than five; collapse three
near-identical argument accessors onto one generic; fold the single-use
`copy_dir_command` into its caller; and give `parseDbtRun.ts` one status
classifier instead of spelling dbt's failure vocabulary twice.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(dbt): one JobCtx down the executor, one table per adapter
Two changes aimed at the operations this code will keep having: adding a
phase, and adding a warehouse.
`JobCtx` already bundled the five values every phase needs, and nine functions
took it — but the top of the executor threaded the fields apart and rebuilt the
struct at each call, so the same literal appeared eight times and each new
phase meant five more parameters. It is now built once per entry point and
reborrowed. `prepare_project` goes from 21 parameters to 17, `retry_failed_nodes`
from 15 to 11, and `run_dbt` drops below the lint threshold. The two remaining
constructions are the worker boundary, where the pieces genuinely arrive apart.
`DbtAdapter` answered five questions with five parallel matches over the same
eleven variants, plus a sixth list of adapters kept by hand in a test. The
facts now live in one `AdapterSpec` per adapter, reached through one exhaustive
match, so adding a warehouse states its name, driver, package, port, database
key and licensing together and the compiler demands the arm. Each arm spreads
from a Postgres base, which makes the inheritance visible per adapter instead
of hidden in the `_ =>` defaults `default_port` and `database_key` used to
carry. `DbtAdapter::ALL` replaces the list the test kept separately.
Verified by dumping all seven facts for all eleven adapters before and after:
byte-identical.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): job-keyed run progress, and stop path-wide reads and clears
Six findings from the last round, in the order they bite.
The publication liveness gate refused on `archived`, but `create_script`
archives the parent on every redeploy — so deploying v2 while v1's dependency
job was still parsing left v1 without a graph, permanently, which is the exact
case the unconditional write existed to serve. It gates on `deleted` alone now;
an explicit archive is still covered by the `FOR UPDATE` ordering.
A project-owned `profiles.yml` trusted `profile.type` instead of reading the
file. The Rust engines carry every adapter, so a CE script could declare
`postgres` over a target that is `sqlserver` and have dbt connect with the
enterprise adapter. The file is read whichever way, and a descriptor that
disagrees with it is refused.
Renaming a dbt script, or editing one so its newest version is no longer dbt,
cleared the graph for the whole path — every older version's models, SQL and
lineage, which their own finished runs still render. Neither needs it: graph
queries join on `(path, hash)` through a `language = 'dbt'` CTE, so an old
version's rows cannot attach to whatever lives at that path next.
Live progress read `materialized_partition`, whose key is the relation and
whose `job_id` is only the last writer. Two runs of one project took rows from
each other. Progress now has its own job-keyed table; the relation table is
untouched, because one row per relation is right for the pipeline canvas and
fork defer. Verified with two overlapping builds: both keep 6 rows in the new
table, while the old one attributes 6 to one run and 0 to the other.
`Scratch::drop` removed a half-installed virtualenv synchronously from inside
the job future, blocking a runtime thread; it goes to `spawn_blocking`, with a
direct call when there is no runtime to hand it to.
The E2E list asked for a `# on table://` subscription the deploy now refuses.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(dbt): snapshot a dynamic descriptor's graph per run
A `{{ }}` placeholder in `vars` can enable a different set of models per run, so
those runs re-ingest the graph. Keyed by version alone, each re-ingest
overwrote the last: reopening an older run showed the newer run's project, and
a model only the older run built was gone entirely — no SQL, no lineage, and
nothing the saved result could colour, since it can only tint nodes that are
there.
`dbt_node` / `dbt_edge` gain `job_id`. A run of a dynamic descriptor writes its
own snapshot under its job id and its page reads it back; a static descriptor
writes the version's graph once, under a zero-UUID sentinel, and every run of it
reads that. The sentinel is a value rather than NULL because `job_id` is part of
the primary key and Postgres does not treat two NULLs as the same key, so each
re-ingest would add a row set instead of replacing one.
`/assets/graph` takes `dbt_job_id` and prefers a snapshot when one exists,
falling back to the version's graph otherwise — so a run page passes it
unconditionally and static descriptors are unaffected. Snapshots age out after
30 days, pruned by the runs that write them, so no background sweep has to learn
about these tables.
Verified end to end: one deploy, two runs of it with `extra=yes` and `extra=no`
gating a model's `enabled`. The version's graph holds 6 models, run 1's snapshot
7 including `opt_extra`, run 2's 6 without it; the endpoint returns each run's
own and falls back to the version's when the parameter is omitted.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* perf(dbt): only snapshot a run whose graph differs, and prune from every run
Two costs the per-run snapshot carried, both found by measuring it rather than
by reading it.
A snapshot was written for every run of a dynamic descriptor, but marking one
dynamic is conservative: `graph_is_per_run` is true whenever `vars` holds a
`{{ }}` placeholder or `env` holds a `$var:`, which says the arguments reach dbt
and not that they change which models exist. The usual case is a date var, whose
graph is identical run after run, so the table filled with copies of an
unchanging picture — around 1 KB per model per run, which is a gigabyte or so a
month for a 200-model project on an hourly schedule. A row set now carries a
digest of its nodes, edges and relation root, and a run whose digest matches the
version's writes nothing; the read already falls back to the version's graph, so
those pages are unchanged. Only a run whose model set really differs pays.
The prune was hung off the progress reporter, which exists only for engines that
emit node events — so a Fusion or dbt-core-2x instance accumulated snapshots and
never deleted any. Retention that stops working because of an engine choice is
not retention; it runs detached from every dbt run instead.
Verified against a descriptor with a var-gated model: the version's graph holds
8 rows, a run that resolves to that same graph stores none at all, and a run
that enables the extra model stores its own 9. Both pages still render their own
project — 6 assets without the extra model, 7 with it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): scope every dbt_node join to the chosen snapshot
`job_id` joined the key, but only the scoping CTE and `dbt_edge` were taught to
filter on it. The outer node SELECT and the parent/child joins in the edge query
were not, so each model came back once per retained snapshot plus once for the
version's graph, and each edge matched every combination of the two — the model
count multiplied and the edge join fanned out quadratically. Measured against
one stored snapshot: 17 node rows where 8 are wanted, and 28 edge pairs where 7
are. The response dedup hid the edge blow-up from the payload, not from the
plan, and the run page refetches the graph every two seconds.
The progress table gained writers it was missing. `terminalize_running_relations`
settled only the relation-keyed table, so a cancelled or killed run — the case
that function exists for, since it leaves no `run_results.json` — showed every
in-flight model still spinning on the run page for as long as the row lived. An
agent worker cannot write the new table at all, having no database of its own,
so the read falls back to the relation-keyed one when a job has no rows there.
Also: a wrapped string literal missing its backslash put eighteen spaces in the
middle of the profile-disagreement error; a comment still described concurrent
runs of one dynamic version overwriting each other's graph, which is what
keying by job removed; and the `materialized_partition` index justified itself
by a run-page poll that has since moved to another table, though the closing
sweep still earns it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): give a graph snapshot a marker row, and scope what reads it
Five findings, four of which are the same mistake in different places: a
snapshot's identity was inferred from its contents.
Existence was inferred from a `dbt_node` row, so a dynamic run that disabled
every model — a legitimately empty graph — read as "no snapshot" and its page
showed the deployed models instead. The digest was a column repeated on every
node and read back with a `LIMIT 1` carrying no `job_id`, so a run could compare
itself against another run's digest and suppress a snapshot it needed. The
relation-root drift check read the same rows unscoped, so after a drift it could
find a previous run's root and conclude nothing had moved.
`dbt_graph_snapshot` holds one row per stored graph — path, version, job,
digest, timestamp. Existence is that row, the digest lives there once, the drift
check reads the deployed row explicitly, and the retention sweep deletes markers
first and then the rows no marker stands for. The digest is SHA-256 rather than
`DefaultHasher`, whose output is documented as unstable across Rust releases:
this value outlives the process that computed it, so a toolchain bump would have
silently stopped every comparison matching and quietly reinstated the duplicate
snapshots the digest exists to prevent.
`/run_progress` ignored the view token, so a share-link viewer got the graph and
was refused the progress that colours it.
A preview sent only its own three arguments, so a descriptor with a required
`{{ }}` var could not be previewed at all and an overridden one previewed a
different relation than the page was showing. The run's arguments go first now,
with the preview's three overriding.
Verified on a project whose only model is var-gated: the deploy stores a marker
with zero nodes, a run with the var set stores a marker with one, and the
endpoint answers 0 and 1 respectively rather than showing the deployed models
for both.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* refactor(dbt): squash the runtime's migrations into one
Ten migrations reshaping the same three tables is a history no installation
ever had. `dbt_node` gained `script_hash`, then `job_id`, then `ingested_at`,
with its primary key rebuilt twice; `graph_digest` was added by one migration
and dropped by the next after the digest moved to its own table. On a fresh
database all of that replays to arrive at a shape the schema can simply state,
and this feature has never shipped, so there is no upgrade path to preserve.
One migration now creates `dbt_node`, `dbt_edge`, `dbt_graph_snapshot`,
`dbt_run_state` and `dbt_run_progress` in their final shape, carrying forward
the rationale each of the replaced migrations recorded. The enum additions stay
in `add_dbt_lang`, since a value cannot be added and used in one transaction,
and the `materialized_partition` index stays separate because it belongs to a
table this feature did not introduce.
Verified by rebuilding: dropped the five tables, replayed from the single
migration, and confirmed the result is identical — same primary keys, the same
two composite `script` foreign keys, the same seven indexes. Every `sqlx::query!`
in the workspace then compiled against it, which checks each column's name, type
and nullability, and a deploy plus run on the rebuilt schema produced 8 nodes,
7 edges, a snapshot marker and 6 progress rows.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): authorize snapshot reads, bound the prune, give the marker a lifecycle
`dbt_job_id` is caller-supplied and selected straight from `dbt_graph_snapshot`,
which carries no RLS — so a caller who could see the script could read any run's
model set and relation paths, which a dynamic alias or schema can encode. Both
graph queries now require the job itself to be visible, in the authed
transaction, the same gate `raw_code` already applies to the script that
produced it.
The drift check compared against the deployed graph alone, which misses the way
back: a run at root B republishes the path-keyed `asset` usages at B, and
returning the profile to A then matches the deploy and skips the refresh,
leaving those usages at B while dbt builds A. It reads the most recent ingest
for the version instead — the one that last wrote them — ordered rather than an
arbitrary `LIMIT 1`.
The prune anti-joined every non-deployed node and edge with no age predicate, so
each run scanned the whole retained sidecar and concurrent runs duplicated it.
All three deletes share one age bound again, with the sentinel spelled as a
literal so the partial indexes apply — a bound parameter cannot be proven to
match the index predicate.
`dbt_graph_snapshot` was the one dbt table nothing in the script lifecycle
deleted: no `script` foreign key and absent from both `clear_dbt_manifest*`
sites. A marker outliving its rows is read as a snapshot with no nodes, and its
digest still answers the suppression check, so an identical run would write
nothing and then render an empty graph. It cascades like the rows now and both
clears take it.
Also: the preview cleared `exclude` rather than inheriting it, since previewing
a model the run excluded reached dbt as `--select m --exclude m`; and
`terminalize_running_relations` no longer claims to cover a killed worker, which
never reaches it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): quote profile names, record where usages were published, stop polling the graph
A profile name comes from the project's own `dbt_project.yml` and a target from
the descriptor, and both were interpolated into `profiles.yml` as bare YAML —
including as mapping keys. A name like `prod # hidden` truncates the mapping and
a newline opens a sibling key of the author's choosing. Both are rendered as
quoted scalars now, as are the BigQuery keyfile's keys, with a test that asserts
the document still parses to exactly the keys we wrote.
The drift check read the most recent ingest, which latches: a run that returns
to the deployed root re-ingests but stores no snapshot (its digest matches the
version's), so the moved run's rows stay newest and every later run pays an
extra parse and ingest. The publisher now records the root it published the
path-keyed usages at, which is the only thing that answers "where do the current
usages point" — the deploy's own root goes stale as soon as a run republishes.
The run page polled `/assets/graph` every two seconds alongside progress, so it
re-sent every node's SQL for the length of a run — hundreds of KB a tick on a
real project, for a graph that a dynamic descriptor re-ingests exactly once
before the build. It fetches once more shortly after mount and then polls
progress alone.
Node results carry `outcome` beside `status`. `status` stays dbt's own word, but
dbt owns that vocabulary — 1.x and 2.x differ on casing and `no-op` arrived in a
minor release — so publishing only it would force a break or a lie the first
time it moves. `outcome` is the stable half a downstream script branches on.
Also: the worker's dbt entry points are `pub(crate)`, since nothing outside the
crate calls them and they resolve secrets and launch processes; and the snapshot
gate records that it is RLS-only where `/jobs/run_progress` also honours a
share-link token, which is a gap in what a shared page shows rather than in what
it protects.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* test(dbt): pin the graph storage invariants against a real database
Every defect review found in this area was DB-shaped — which row set a read
resolves to, which rows a clear takes, whether a snapshot exists at all — and
none of it is reachable from a unit test on a pure function. Four rounds
established these answers and nothing guarded them, which is why each round kept
finding another.
Six cases, on the harness the repo already uses for schema-shaped behaviour:
an identical run stores no snapshot and leaves no marker; a differing run keeps
its own while the version's is untouched; an empty run graph is still a snapshot
rather than an absent one; clearing one version leaves the others whole; the
path-wide clear takes the markers with it; and the sweep ages out run snapshots
while never touching a version's own graph.
`IngestedNode` gains `Default` so a test can state the two fields a case is
about rather than the eighteen it is not.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* perf(dbt): bound a script's stored graphs by deploy count
Run snapshots expire on a clock, but a VERSION's graph could not: its reader is
every finished run of that version, and a run page is as old as its job. So
nothing reclaimed them — a deploy graph went only when its `script` row was hard
deleted, which Windmill does not routinely do. A CI deploying on every commit
added a full model set with SQL bodies per commit, forever: roughly 200 KB a
deploy for a 200-model project, which is gigabytes a year across an instance.
Bounded by COUNT instead of age, since age is the thing that cannot be right
here. The newest 50 deploys per path keep their graph and older ones are
reclaimed, making growth `versions x models` rather than unbounded in time.
Generous on purpose: reaching the bound empties that version's run pages, so it
exists to stop unbounded growth rather than to be hit in normal use. Ordered by
the script's own `created_at`, so a late-finishing job re-ingesting an old
version cannot promote it.
Pinned by a test that deploys past the bound and asserts both halves: the count
holds, and the newest version is always among the survivors.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): let the run page know when its snapshot has landed
The one-shot graph refetch was wrong: a dynamic descriptor's ingest happens
before the build but after cloning, dependency install and parse, so a fixed
delay either fires too early — and the run page then shows the deployed models
for the whole run, never that run's own — or keeps re-sending the whole graph
for the length of it. Neither is a timing problem to tune; the page had no way
to tell "the snapshot is not written yet" from "this run has none".
`/assets/graph` answers that directly: `dbt_snapshot_job` is the job the dbt half
resolved from, when one was asked for and found. The page polls the graph until
that is its own job, and stops. A static descriptor never snapshots, so an
attempt cap ends it there rather than polling for the run's duration.
`dbt_node.relation_root` is gone. The drift check moved to the marker's
`published_relation_root`, which left the column written on every node and read
by nothing.
`outcome` was published as the stable half of the result contract, but the
in-tree consumer still ranked and coloured from dbt's own word — so the field
existed and nothing used it. `statusRank` takes it, `DbtRunResult` passes it, and
`classifyStatus` is documented as the fallback for results that predate it and
for the live event stream, which carries dbt's word alone.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(dbt): record what a share-link viewer actually sees
The comment at the snapshot gate said the graph "falls back to the deployed
set", which is only the rarer half of it. A share link is an extra grant for a
logged-in user who lacks access to the job, so the usual case is no read on the
script either — and then the `live` CTE matches nothing and the whole dbt half
comes back empty. A blank Models panel over working progress rows, not a
fallback.
`docs/dbt-runtime.md` now carries the analysis a follow-up needs: that relaxing
this leaks nothing, because `v2_job_completed.result` already gives that viewer
every node's `unique_id` and `relation_name` — the graph's only incremental
exposure is `raw_code`, which is gated separately on seeing the script. And the
shape of the fix: `OptViewToken` and `validate_view_token` are self-contained
enough to move into `windmill-api-auth`, which `windmill-api-assets` already
depends on, after which the gate can honour a token for that job's snapshot
alone while `raw_code` stays where it is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): keep model SQL behind the scripts:read scope, and unbreak CI
`/assets/graph` is authorized as `assets:read`, and RLS decides whether the
caller can see the script that produced a node — but RLS is not a scoped
token's grants. A token deliberately narrowed to `assets:read` could therefore
read model source and repository paths for scripts outside its `scripts:read`
paths. The same `build_scope_path_predicate` the macro endpoint already applies
now gates `raw_code` and `original_file_path`; the relation's shape is
unaffected, only its body is withheld.
`DbtAdapter::ALL` exists for the tests that must cover every adapter, so it is
dead in a release build and `-D warnings` failed all four backend checks on it.
It is `#[cfg(test)]` now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): stop the graph poll at the ingest, and type `limit` in the schema
The poll's stop condition was a snapshot appearing, with a 40-attempt cap
behind it — so a STATIC descriptor, which never snapshots, took the cap every
time and re-fetched the whole graph forty times. That is most of what removing
the poll was meant to save, and static is the common case.
The ingest runs BEFORE the build, so the first model to report progress proves
it has already happened: a snapshot absent by then is one this run never
writes. Progress arriving is now the second exit, and the cap is only a
backstop for a run that reports none at all.
`limit` is declared `Typ::Int` but `dbt_arg_schema` had no integer arm, so the
run form and the generated clients saw an untyped default and offered no
numeric control for a value the worker clamps. Covered by the schema test.
`relationOutcome` still re-derived from dbt's word while `statusRank` had moved
to `outcome`; both read it now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): make a retry prove its arguments still resolve the same
The saved arguments are the ones SUBMITTED, so a `$var:` in them is re-resolved
on retry. The identity did not cover the resolved values, so a variable that
changed between the failed run and the retry was accepted — and which graph the
retry then used depended on WHERE it landed: a worker holding the local
snapshot replays the saved manifest, while a database restore reparses with the
new value. Placement decided whether the resumed failures described the
relations being built.
The identity gains a digest of the resolved arguments, and is compared in two
halves because resolution happens between them. Project, warehouse, engine and
env are checkable up front; the arguments are not, because a retry request
carries only `dbt_command` and the ones to compare are the SAVED arguments after
this caller has re-resolved them. Comparing the whole string up front would have
refused every retry — which is what the obvious version of this fix does.
A row written before the digest existed has no last segment, and still restores
rather than being refused.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): keep pre-upgrade retries working, and stop losing a late snapshot
Splitting the identity on its last `|` read a pre-upgrade row's env digest as an
arguments digest and left only `<run_identity>` as the prefix, so every saved
failure on an upgraded instance became unretryable — a regression the previous
commit's own test missed by using an identity with no `|` in it at all, which is
not what an old one looks like. The digest is tagged (`|args=`) rather than
positional, and the test now uses a real pre-upgrade identity.
The graph poll gave up after a bounded number of tries, but provisioning and
`dbt deps` precede the ingest and can outlast that on a cold worker — and the
engines that emit no node events never produce the progress that ends it early.
A finished run now reloads the graph unconditionally, and the poll's own exit
issues one last load: progress proves the ingest happened, not that the previous
tick saw it, and dbt's compile window is wider than one tick.
A `dbt retry` restores the failed run's arguments inside the worker and they are
never written back to the retry job, whose own args are just
`{"dbt_command": "retry"}` — so previewing a row on a retry's page ran without
the vars the run used. The result now carries the invocation's arguments as
SUBMITTED, so a `$var:` stays a reference and no resolved value is published.
The deploy-count sweep ran instance-wide on every dbt run: `FROM script WHERE
language = 'dbt'` has no index to stand on, and both orphan deletes are the
complement of every partial index here. It is scoped to the running script's
`(workspace_id, path)` — which `index_script_on_path_created_at` serves — and
the orphan deletes only run when a marker actually went.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): hide dbt from module-less pickers, and stabilise the retry digests
`processLangs` feeds every language picker, including flow steps and app inline
scripts. Those are raw bodies with nowhere to carry a module bundle, and a dbt
script IS its bundle — so choosing dbt there produced a job that could only fail
once the worker looked for `dbt_project.yml`. Those two surfaces use
`processInlineLangs`, which drops the languages that need modules; a flow still
reaches dbt the way it reaches any script, by path to a deployed one.
`graph_digest` moved to SHA-256 because it is persisted and compared by a later
worker, and `DefaultHasher` is documented as unstable across Rust releases — but
the retry identity's own digests were left on it, and they are persisted in
`dbt_run_state.identity` for exactly the same comparison. A toolchain bump would
have refused every saved failure as a different project. All three go through
one `stable_digest`, length-prefixed so no split of the same bytes collides.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): enforce the tag scope on snapshot reads, reset state between runs
A tag scope is an orthogonal hard restriction: a token limited to some tags must
not read a job outside them however else it is authorized. The snapshot lookup
went through `v2_job` RLS alone, which knows nothing about tags, so such a token
could still retrieve a run's model set and its dynamic relation paths. The same
predicate `require_job_read_access` applies for the progress half of the page is
applied here — `get_scope_tags` is already public in `windmill-api-auth`, and it
is `None` for an unscoped caller, so a normal session pays nothing.
SvelteKit reuses the run graph between run ids, and `graphTries`, `polled` and
`raw` all describe the previous job: a spent retry count stopped the next run's
snapshot poll before it began, and stale progress coloured its models with
another run's statuses. All three reset when the graph key changes.
Also a wrapped string literal missing its backslashes, which put two ~22-space
runs in the middle of the retry-refusal message.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs(dbt): put each digest helper's rationale on its own function
Inserting `stable_digest` above `split_identity` split that function's doc, so
five lines describing where the identity divides ended up introducing the
hasher. Each is back on the function it describes, stated once.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(dbt): read the run-pinned graph through the job, not the asset graph
Pinning the asset graph to one run is job-scoped data, but `dbt_job_id` sat on
`/assets/graph`, authorized as `assets:read`. The job-read contract —
`require_job_read_access` — is five parts that pull in opposite directions (tag
scope restrictive, `created_by` permissive, app-embed restrictive-overriding,
view token permissive, RLS underneath), so plain RLS is neither a stricter nor a
looser approximation of it. Restating the parts near the graph query kept leaving
one out: first the job check entirely, then the share-link asymmetry, then the
tag scope, and the app-embed restriction was still missing and failing open.
The helper cannot be called from `windmill-api-assets`, because `windmill-api`
depends on that crate. So move the read instead of the check: the run-pinned
graph is now `GET /w/{w_id}/jobs/dbt_graph/{id}` in `windmill-api`, on the same
gate as the `run_progress` it colours, and `/assets/graph` has no `dbt_job_id`
parameter at all.
- `asset_graph_for` takes the job as an argument from an already-authorized
caller; the route handler passes `None`.
- Extract the graph response into a `AssetGraph` component schema, now that two
paths return it.
- The run page fetches the job route when it has a job id.
* fix(dbt): charge assets:read on the run-graph route, trust the job gate in SQL
Round 13 findings on the route moved last commit.
The scope domain comes from the URL segment, so putting the read under `/jobs`
asked a scoped token for `jobs:read` alone while returning asset-graph data that
`/assets/graph` charges `assets:read` for. A token narrowed to polling run status
could read workspace topology, and the missing-job fallback made it cheaper still
— any random UUID skipped the job gate. Both scopes are now required: the job
gate reaches this run, `assets:read` reaches asset data at all.
The `chosen` CTE re-decided job visibility under plain RLS after the caller had
already passed `require_job_read_access`. It could only disagree, and did so
silently by falling back to the deployed graph — a share-link viewer entitled to
the run was shown a different run's model set. Dropped; the contract is that a
job reaching `asset_graph_for` is already authorized.
Also: the flow editor's `+` insert menu still offered dbt (the third
`processLangs` caller, missed when the other two moved to `processInlineLangs`),
the docs still described the deleted `dbt_job_id` parameter, and the new handler
had again been inserted between `get_run_progress`'s doc comment and its
function.
* fix(dbt): resolve a pinned run's version from the job row, not script RLS
Local codex review of the branch.
A share-link viewer is entitled to the run and usually has no grant on the
project — that is what the link works around. The graph's `live` CTE resolved the
version by selecting `script` inside the viewer's RLS transaction, so it answered
for their access to the project rather than for the run they were given: the
Models panel came back blank beneath working progress rows.
A pinned run now takes its path and hash from the job row the handler already
read after authorizing the job, so `live` does not consult `script` at all.
`raw_code` keeps its own `EXISTS` against `script`, so the model bodies stay
behind access to the project. Verified under RLS as an unprivileged role: the
shape query goes 0 rows -> 1, the `raw_code` gate stays 0.
Taking the version from the job also means a caller can no longer pin one
project's version while naming another's run, since `dbt_script_hash` is ignored
when a job is given.
The run page's graph fetch is a raw `fetch`, which bypasses the interceptor that
adds `X-View-Token` to generated-client calls, so a shared page was refused
before any of this mattered; it goes through `appendViewToken` now.
Also trims three comments to the AGENTS.md limit, dropping drafting-history
rationale that belongs in docs/dbt-runtime.md.
* fix(dbt): snapshot vars-overridden runs, carry the pinned version everywhere
Second local codex pass.
A `vars` run argument overrides the descriptor's, and vars drive `enabled`,
alias, schema, database and materialization — so such a run builds relations the
deployed graph does not describe. It now snapshots under its own job id, which
per-job keying makes safe: the version's graph stays for runs that did not
override. The old comment claimed gating on it would strand the override's graph
for the next default run, which was true only when the write went to the
deployed slot.
Two sites still read the caller's `dbt_script_hash` instead of the version
resolved from the job, so `/jobs/dbt_graph/{id}` without that redundant
parameter dropped models the run's version had and a later deploy removed.
The `dbt_snapshot_job` marker re-checked `v2_job` under RLS — the recheck the
graph query itself drops. A share-link viewer got the right graph and a null
marker, so the run page refetched it 40 times before giving up.
`wmill script preview` read the bundle with the generic `__mod` suffix and
script-module parsing, so previewing a `.dbt.yaml` omitted the project and failed
on the missing `dbt_project.yml`. It uses the same suffix and verbatim read as
deploy.
* fix(dbt): key retry state by principal, not by script path alone
`dbt_run_state` held one row per (workspace, script path), and a retry replaces
the caller's arguments with the saved ones. Anyone able to run the script could
therefore retry whoever ran it last, replaying that run's literal `select` and
`vars` against the warehouse and publishing them as their own job's
`invocation_args`. Running the script was already theirs to do; seeing another
principal's arguments was not.
`permissioned_as` joins the key, so a retry resumes only state written under the
same authority. Two runs sharing an authority can already act for each other, so
this is the boundary that matches the rest of the job model.
* test(dbt): pin what a caller without access to the project sees of its run
The share-link case had no regression guard, and every fix in this area touched
one of its two halves: the graph's SHAPE has to survive a caller who cannot read
the script, and the model SQL must not.
Two cases against a real database, calling `asset_graph_for` as a member with no
grant on the project's folder: pinned to a run, the models render and `raw_code`
is withheld; unpinned, the same caller sees nothing of it, so making the first
work did not relax the second.
Both assertions were checked by mutation — reverting the `live` bypass empties
the graph, and dropping the `raw_code` script gate leaks `select 1` — so neither
passes on the code it is meant to catch.
* fix(dbt): key the worker-local retry cache by principal too
Keying `dbt_run_state` by `permissioned_as` left its worker-local twin keyed by
workspace and script path alone, so the boundary held only where the database row
was consulted. An agent worker never reads that table — `Connection::Http` leaves
`latest_job` as `None` — so there the local cache was the whole boundary and it
had none: the next principal to retry the script on that worker restored the
previous one's `select` and `vars`.
Also records the sqlx `--all-targets` trap in the update-sqlx skill: it is needed
for queries inside tests, and in a CE checkout it aborts on `tests/otel.rs`
(EE-only `otel_ee`) after having already emptied the cache.
* fix(dbt): log a dropped retry-state save, correct the run-progress contract
Saving retry state is best-effort — losing it costs a retry, not the run that
just finished — but `.ok()` dropped the reason too. The only symptom was `dbt
retry` reporting nothing to resume, which reads as a bug in retry rather than a
failed write. Found by running a real failing build against a worker whose
binary predated the `permissioned_as` column: the insert violated NOT NULL and
said nothing.
The run-progress endpoint's OpenAPI description promised an empty list for a
caller who cannot see the job. It is refused instead; an empty list means the job
recorded nothing yet or is unknown here.
* fix(dbt): return the retry-state write failure the warning was added to report
`save_run_state` discarded the insert result, so the caller's warning could never
fire and a lost retry row stayed silent — the symptom being `dbt retry` finding
nothing on another worker.
The error is held rather than returned at once: the worker-local copy is what an
agent worker resumes from, so a failed insert must not cost that too. Every exit
after it surfaces it, including the ones that give up on the local save.
* fix(dbt): decide a retry's graph from its restored args, keep local state in step
Three from the seventh local review.
A retry submits only `dbt_command`, so the vars-override check ran against an
empty argument set and left `graph_is_per_run` false. The failed run's arguments
are restored afterwards, and those are what the retry builds with — an overridden
one wrote no snapshot for its own job and its page fell back to the deployed
graph, showing the wrong enabled models, aliases and schemas. The decision is
re-asked once the restore has happened.
A failed durable write no longer publishes the worker-local generation either.
`restore` accepts a local generation only when the database row names it, so
publishing one the database never recorded made this worker reject its own newest
state and resume the previous run's — its selection and vars, or "nothing to
retry" if that one had succeeded. An agent worker attempts no durable write, so
it keeps its local copy as before.
A rename that also converts away from dbt moved the old path's retry state onto
the new one, reinstating what the conversion had just cleared and leaving one
user's arguments and results under a path no dbt script occupies. It moves only
while the destination stays dbt, and clears the source otherwise.
* fix(dbt): drop retry state when a run produced none, let module pushes fail loudly
Three from the eighth local review.
A run that never wrote `run_results.json` — cancelled, timed out, or dead before
dbt got there — left the PREVIOUS run's state authoritative in both the database
and the local pointer, so a later `dbt retry` resumed that older invocation's
failed nodes. Producing nothing resumable now clears both copies, so neither can
answer for the other.
`wmill sync push` wrapped the descriptor lookup and its deployment in one
try/catch meant for a missing parent. Any API failure or invalid descriptor was
reported as "no parent found" and swallowed, so a module-only push exited zero
with the remote project unchanged. Only the lookup is tolerated now.
Also condenses a comment that narrated how earlier status comparisons behaved.
* fix(dbt): forget retry state on pre-build exits too, drop cascade claims
A dynamic run whose pre-build `dbt parse` or graph ingest fails returns before
the save that clears stale state, so the previous run stayed authoritative in
both the database and the local pointer and `dbt retry` resumed ITS failed nodes
— writing relations the run that just failed never touched. Both exits now
invalidate, through one helper shared with the no-artifact case.
Two frontend comments described dbt producer rows as driving cascade dispatch.
The executor returns before dispatch for every dbt job and deployment rejects
`table://` subscriptions, so they promised behaviour that cannot occur; they
describe the lineage and ownership that is actually retained.
* fix(dbt): let the database decide retry state where it is reachable
A SQL worker treated "no `dbt_run_state` row" as no opinion and accepted any
worker-local generation. But no row is the authoritative answer that the last
invocation left nothing resumable, so a local pointer that outlived it — an
unlink that failed, a process killed between the delete and the removal, a stale
cache — resurrected a replaced run and let `dbt retry` write relations it never
touched. An agent worker keeps accepting its local copy: it has no authority to
consult.
Invalidation failures are logged rather than dropped, since a silent one is
exactly what leaves the pointer behind.
* docs(dbt): record how to run an agent worker locally, keep archived graphs
Every step of standing one up fails as something else: a normal build cannot
start one at all, the server's routes need a separate feature, and all three
token mistakes surface as a bare 401 on the agent with the reason only in the
server log. Written down with the error each produces.
Also keeps a dbt script's graph when it is ARCHIVED rather than deleted. The
pinned read resolves versions through a CTE that already skips archived rows, so
clearing bought nothing and emptied the Models panel of every completed run of
the project. Deletion still clears it.
* feat(dbt): let an agent worker publish its graph, through one endpoint
An agent worker was refused any dbt script whose profile comes from a Windmill
resource — the common case — because it could neither read the stored relation
root to check for drift nor re-ingest a corrected one.
Those look like two needs but collapse into one: verification exists only to
decide whether the stored graph still describes reality, so a worker that can
PUBLISH never has to ask. It stores what it just parsed.
`POST /api/agent_workers/dbt_graph/{workspace_id}` is the whole addition. It
wraps the same `replace_dbt_manifest` the SQL path calls, so digest suppression,
the marker write and retention cannot drift between the two transports, and it
refuses a job the token's tags do not cover. `IngestedManifest`/`IngestedNode`
gain Deserialize to cross the wire.
Two guards go, both now false: the pre-build refusal, and the `Connection::Sql`
gate added earlier to stop a `vars` override grounding an agent run.
Live progress stays SQL-only — that is a per-model event stream, and routing it
through the API would mean a round trip per node.
* docs(dbt): warn that a differing cargo feature set swaps the shared binary
* docs(dbt): record the verified agent-worker behaviour and the tmpfs quota trap
An agent worker now runs a dbt job end to end, retries, and publishes its graph
— confirmed with a dynamic descriptor whose per-run snapshot came back through
the new endpoint. The doc said it was refused; that was true before the endpoint
existed.
Also `WINDMILL_DIR`: on a dev box the job dies with `Disk quota exceeded (os
error 122)` writing the project's files while `df` shows free space AND free
inodes, because /tmp is a tmpfs carrying a per-USER quota. Point the worker at a
real disk rather than trying to clean up beneath it.
* chore(dbt): pin the EE revision carrying the agent graph endpoint
* fix(dbt): bind the published graph to the job, break the completed-page poll loop
Five from the thirteenth local review.
The EE endpoint took `script_path` and `script_hash` from the payload and checked
only that the supplied job carried one of the agent's tags, so an agent holding
any matching-tag job could name another script and replace its graph. Both are
read from the verified queue row now and the request carries only the job id. A
raw preview has no version, so it no-ops rather than 422ing before dbt runs.
`IngestedManifest`/`IngestedNode` take `#[serde(default)]`: they were
serialize-only, and a field the serializer skips made the whole manifest
unparseable on the receiving side.
A completed run page fetched the graph forever — `load()` assigns `raw`, which
recomputes `settled`, which re-entered the same effect. The final fetch is keyed
to the job by a plain (non-reactive) variable, and `settled` is read untracked.
The pin now names a revision that compiles: the previous one still called
`authed.tags()`, a method that does not exist, because both that fix and the JSON
response landed after it was committed.
* fix(dbt): keep a run snapshot out of the script's deployed ownership
Everything `persist_ingest` writes after the manifest is keyed by PATH — one row
set per script, describing what is deployed there. A run snapshot was still
reaching it, so a one-off `vars` override republished that invocation's relations
as the script's ownership and the workspace graph stayed on the override's
schemas and aliases: an ordinary run of a static descriptor never ingests again
to correct it, so only a redeploy would. A snapshot now stops after recording its
own rows.
The row preview also selected a bare model name, which dbt resolves across every
installed package while `show` takes a single node — a project model sharing its
name with a package's was previewed wrongly or refused. It selects the
package-qualified FQN.
* fix(dbt): forget stale retry state when preparation itself fails
`prepare_project` runs before every path that could clear it, and it fails for
reasons unrelated to the saved run — a profile that stopped resolving, a
provision cancelled, packages that will not install. The invocation still left
nothing resumable, so the previous one must not stay authoritative: a repaired
project would otherwise let `dbt retry` rebuild an older run's selection and
write relations the latest invocation never reached. A retry is exempt, since it
is trying to use that state and failing to prepare says nothing about it.
Also corrects the runtime doc, which still described agent workers as unable to
run dynamic descriptors or Windmill-resolved profiles. They publish their graph
through the API now; what they do not get is live progress and a durable retry
row, and the doc says so.
* fix(dbt): spell the whole FQN for preview, bound retained retry generations
The FQN selector added last commit was `<package>.<name>`, but a dbt FQN is the
resource's path within its package and the matcher must consume the selector and
end on equal lengths — so it matched nothing for a model under `models/marts/`,
which is the layout most projects use and the one this repo's own complex fixture
has. The middle segments come from `original_file_path`, whose first element is
the resource root the FQN excludes. Without a path it falls back to the bare
name: ambiguous across packages, but a selector dbt resolves rather than rejects.
Tested on a nested model, which is the input that separates the three spellings.
Superseded retry generations were removed only when a later run published one,
and never inside the hour-long grace period — so a burst left a manifest and a
results copy per run with nothing afterwards to collect them. At most four now
sit in the grace window, oldest evicted first.
* fix(dbt): scope preview state to the run, seed the project on a language switch
Previews are keyed by `unique_id`, which is the same string for the same model in
every run, and the run-change effect reset only the graph and progress. Opening a
second run of one project therefore showed the previous run's rows immediately,
and `runPreview` treated them as cached and refused to fetch. A generation
counter also drops a preview that resolves after navigation, which the reset
alone cannot catch.
The dbt project was seeded only by the empty-script bootstrap, but dbt is in the
ordinary language picker: reaching it by switching a draft produced a script with
no `dbt_project.yml`, which the runtime refuses to deploy or run. Both entry
points seed now, and neither touches modules that already exist.
* chore(dbt): cache the agent graph endpoint's query for the EE offline build
* fix(dbt): publish the graph a moved profile built
A run snapshot stopped before everything `persist_ingest` keys by PATH, which is
right for a one-off `vars` override and wrong for the other two reasons a run
re-ingests. `graph_is_per_run` was one bool for all of them, and the profile
drift check both sets it and reads back what the publisher recorded: a profile
moved A->B was detected by every run forever, each paying a `dbt parse` for a
snapshot nobody reads while the asset rows went on naming schema A.
The reason is carried now (`GraphRefresh`), and it decides both writes. Drift is
the version's own move, so it rewrites the VERSION's graph and republishes the
ownership that ends the drift; a dynamic descriptor snapshots under its job id
and still publishes; anything the CALLER scoped — an overridden `vars`, a
narrowed `select` — snapshots and publishes nothing, so one invocation's subset
can neither stand as what the script owns nor drop the models it left out from
the version's graph. Where they meet the caller wins, and the next ordinary run
settles the drift.
A restore also rebuilt `run_results.json` by copying the generation directory a
second time, so a burst of saves pruning it mid-restore left `dbt retry` with
nothing to resume and a job that reported success. It is written from the bytes
the restore already read; a manifest that went the same way falls back to the
parse a database restore pays anyway, and a generation that vanished before
either read falls back to the database's row for that same run instead of
reporting there is nothing to retry.
* fix(dbt): select a row preview by package, not by file path
The preview built dbt's FQN by dropping one segment of `original_file_path`,
which assumes the model root is `models/`. A project setting
`model-paths: ["src/models"]` turned `src/models/marts/orders.sql` into
`pkg.models.marts.orders`, and dbt's matcher — equal lengths, compared from the
front — resolves that to nothing: the preview came back empty for every model in
the project.
It selects `<name>,package:<pkg>` instead. The comma is dbt's intersection
operator, so this names the node by its own name and the package it belongs to,
which is what the FQN was reaching for and needs no knowledge of the resource
root. Verified on dbt-core 1.12, dbt-core 2.0.0-alpha.5 and fusion
2.0.0-preview.202, including a package shipping a model whose name the root
project also uses.
* fix(dbt): refuse a lockfile version that is not one, keep a named selector
Two things a preview reaches that a deploy does not vouch for.
A raw preview submits its own `lock`, so `engine_version` arrives from the
caller and was interpolated straight into the engine cache path — `../..` in it
made the download, extraction and rename land anywhere the worker can write,
and provisioning runs on the host rather than inside the dbt jail. Both it and
`adapter_version` (a pip requirement) are now accepted only as a plain version
token.
`effective_selector` also read any submitted `select`/`exclude` as an override
of the descriptor's named selector. The generated run form posts a default back
for every field the caller left untouched, and a selector descriptor's `select`
default is `[]` — so pressing Test, saving a schedule or firing a webhook built
the WHOLE project instead of `--selector nightly`. An override is now one that
DIFFERS from the descriptor's own value; a run that wants the whole project
despite the selector asks with `["*"]`.
* fix(dbt): let a moved profile settle, from the runs that actually happen
Two ways the drift check could never come to rest, both verified against a real
run of a real project on a normal worker.
`add_caller_args` read any submitted `select`/`exclude` as a caller's narrowing.
The generated run form posts a default back for every field left untouched, so
every run from the UI, a schedule, a webhook or a flow step carried them and was
marked caller-scoped: with the profile moved A->B, each one stored its models
under its own job id and left the workspace graph — and the root the check reads
back — at A. Since no UI run omits the field, the "an ordinary run settles it"
escape hatch was unreachable. Both this and `effective_selector` now ask one
question, `selection_is_overridden`: DIFFERENT from the descriptor's, not merely
submitted.
The root was also recorded beside the path-keyed publication rather than beside
the graph it describes, so a version that cannot claim the path — an older one
run by hash, a deploy overtaken by a newer one — rewrote its graph at the moved
root and recorded nothing. Its next run then compared against a root that was
absent or two moves stale and skipped the refresh its own run page needed. It is
written wherever the deployed row's graph is.
Verified end to end: same UI-shaped arguments before and after, the moved
profile now republishes (asset rows and version graph both move to the new
schema), a second run detects nothing and re-parses nothing, and moving the
profile back settles it again.
`prune_dbt_run_graphs` also ran from runs alone, while a deploy writes a whole
node set of its own, `raw_code` per model included: a project redeployed on
every push by CI and run nightly kept one full graph per push until the next
run, and one deployed but never run kept them for good.
* fix(dbt): drop a self-dependent effect in the run graph
`previewGen` was `$state` written by the effect that also reads it, three lines
under a `finalLoadFor` that is a plain `let` for exactly that reason. Nothing
reactive reads it — the only reads are inside `runPreview`, a plain async
function — so it becomes a plain `let` too.
* fix(dbt): pin a retry to the engine versions it resolved
`run_identity` carried the engine KIND but not the version it resolved, nor the
dbt-core 1.x adapter's. Redeploy an unchanged project after a release and it
locks a newer dbt or adapter while the saved `run_results.json` still passes the
check, so `dbt retry` feeds one version's artifacts to another — the exact
reproducibility the lockfile exists to hold. Both resolved versions are in the
identity now; a real failure and retry still resumes.
Also drops three comments that outlived what they describe: two said `[]`
clears a descriptor's selector, which `selection_is_overridden` reversed, and
one pointed at an agent-worker guard that no longer exists — the agent path
reaches the ingest deliberately and publishes through the API.
* fix(dbt): discard a run graph the page has already navigated away from
The component is reused across runs, so a slow `/jobs/dbt_graph` or progress
response could land after the reset and put the previous run's models, statuses
and failure state on the current run's page, where nothing would fetch again to
correct it. Every response is now checked against the generation it was
requested under — the counter the preview path already used, renamed for what
it means.
The graph poll also backs off. Neither of its stops is reachable for a whole
class of runs — `dbt_snapshot_job` never matches a static descriptor, and
`polled` stays empty for the engines that emit no node events — so an ordinary
run walked to the cap, re-sending every model's SQL 40 times in two minutes.
* fix(dbt): forget the previous run when the durable save fails, keep quoting
`save_run_state` returns the database error when its upsert fails, which leaves
run N-1's row and local generation in place: same project, same arguments, so a
`dbt retry` matches them and resumes an older attempt's failed nodes against
this checkout — the outcome the no-results branch twelve lines above calls
`invalidate_run_state` to prevent, reached by another door. It now goes through
the same call. Best effort, since the delete goes to the database that just
refused a write, but the local pointer is what a retry landing back here reads.
The run page also rejoined a relation's parts after `splitRelation` stripped
their quotes, so the one name the button exists to paste —
`"wh"."analytics.v2"."Order Items"` — was copied as something no client
resolves. It copies `relation_name` verbatim.
And a source on a finished run was called another project's: the check that
guards against two projects claiming one relation asks whether this run executed
the node, and a run executes no sources — they appear in no `run_results.json`.
Nothing materializes a source, so that warning could never be true of one.
Docs: the `vars`-override paragraph still said such a run does not refresh the
graph, which the table above it contradicts — it refreshes under its job id and
publishes nothing.
* fix(dbt): keep the asset rows and the version's models describing one graph
The workspace graph takes an asset's relations from the path-keyed `asset` rows
and its models, SQL, tests and lineage from the version's `dbt_node`/`dbt_edge`.
A dynamic descriptor published the former while storing the latter under its own
job id, so a placeholder that moved an alias or a schema left the current graph
with assets no model stands behind — nothing dbt contributes to them survives.
Ownership is published exactly when the VERSION's graph was written now, which
is the only state in which the two agree. Two cases are settled elsewhere by
design: an override's relations are a one-off, and a dynamic descriptor at a
moved profile keeps the deploy's ownership until a redeploy — its runs each show
their own models and it re-parses regardless, so the undetected drift costs it
nothing it was not already paying.
An agent worker has no durable row, so its local `current` pointer is the whole
of what a retry reads — and every local publication failure returned success
with the PREVIOUS run's pointer still in place. Where a row exists that is
harmless (`restore` takes a local generation only when the row names it), so the
abandonment is scoped to the agent case.
A failed `dbt show` also cleared the retry state: the preparation-failure exempts
`retry` but not a read-only command, and the run page's row preview is exactly
that, run as the principal the state is keyed by — so a preview that could not
provision took the retry away from the run being looked at.
Frontend: the run-change reset left `loading` and `failed` behind, so the gap
before the next run's answer rendered "no models in the asset graph" — a claim
about the descriptor — over a project that is fine. And three derivations argued
from "the graph is the current deploy", which the pinned endpoint made untrue;
each is still needed, for the version-graph rewrite and retention reasons now
written down.
* fix(cli): let --skip-scripts cover a script's module files
The module shortcut in `elementsToMap` maps the file and `continue`s before
every skip filter, and a module is deployed as part of its parent script — so
`wmill sync push --skip-scripts` still pushed the script whenever one of its
modules changed, and pull still overwrote them locally. Harmless while a module
was a rare helper file; every file of a dbt project is one of these now.
* docs(dbt): a dynamic descriptor's ownership stays the deploy's
* fix(dbt): read the run out of a failure whose message has braces of its own
`parseDbtRun` anchored on the FIRST `{` in the error message and parsed
everything after it. The worker appends the structured result after the error
text, and dbt's errors carry braces — a Jinja template, the compiled SQL, an
adapter's own JSON — so the failures most worth reading were the ones whose
summary and per-node outcomes the run page dropped. Every brace is tried now,
bounded, and the first that parses as a run wins.
Pins the EE revision that gives the agent publish endpoint the deleted-version
guard the SQL path takes: deletion is soft, the foreign key still accepts graph
rows, and the pinned graph query serves non-live versions, so an agent finishing
during a delete put a deleted project's model SQL back on screen. The query is
byte-identical to `persist_ingest`'s, so the offline cache already covers it —
verified with a full-EE `SQLX_OFFLINE=true` check.
* fix(dbt): seed a project when a modular draft switches to dbt
`seedDbtProject` returned whenever the draft carried any module at all, so a
modular script holding a `helper.ts` reached dbt with none of what dbt needs:
the project view is read-only, and the worker refuses a bundle without
`dbt_project.yml`, so that draft could neither run nor deploy. Keyed on the
project file now, and the seed goes in under whatever is already there — the
previous language's helpers are inert to dbt and the user's to remove.
Also records this runtime's schema in `backend/summarized_schema.txt`: the
`table` asset kind, the `dbt` script language, the five dbt tables and the
`materialization_status` enum the progress table uses.
* docs(dbt): move the pipeline-membership rationale out of the deploy path
* fix(dbt): gate a pinned run's model SQL on the version it belongs to
The `EXISTS` against `script` is the only thing standing between a share-link
viewer and the project's source, and it matched the workspace and path alone.
`extra_perms` is a grant on a ROW: archive a version that granted someone
access, recreate the path with narrower permissions, and that stale grant
satisfied the probe while the query returned the NEW version's `raw_code`. Both
probes name the hash now. The regression test drives exactly that shape and
fails without it, returning `select 2` to a caller granted only on the archived
version.
* fix(dbt): keep a delimiter an identifier escaped by doubling
Every dialect these relations come from escapes its own delimiter by doubling
it, and both split functions closed the quoted section on the first half and
reopened on the second: `"schema"."a""b"` came out as `a.b`. The manifest keeps
the real spelling, so the run wrote its per-model status and row counts under an
asset path no graph node has — the node simply never moves, which is the failure
mode this splitter exists to prevent.
Fixed in the worker and in its frontend mirror, which have to agree, with a case
per delimiter on both sides.
* fix(dbt): key retry state by the caller, not only by the principal it runs as
An `on_behalf_of` script executes every caller's job as its owner, so
`permissioned_as` names one principal for all of them and the retry state — the
durable row and the worker-local generation both — collapsed onto a single
entry. After one caller's run failed, the next could submit `dbt_command: retry`
and resume it: their arguments replayed against the warehouse, and handed back
through `invocation_args`. Nothing else separated them, and on an agent worker
the local directory is the whole boundary.
`created_by` joins the key in both places. For an ordinary script it changes
nothing — `permissioned_as` is already that caller — and a run that was itself
superseded was never resumable anyway.
Includes the offline cache for the four changed queries and the three the
pinned-graph regression test added last commit, which had none: `prepare`
without `--all-targets` does not compile test targets, so CI's
`SQLX_OFFLINE=true ... --all-targets` would have failed on them.
* fix(dbt): compare the schema too when reporting a relation that moved
`relationDrift` compared the leaf name alone, and the move it exists to report —
a profile repointed at another schema, which a later run then writes into the
version's graph — leaves every model's name exactly where it was. So the one
case that reliably produces a graph naming relations this run did not write was
the one case the notice stayed silent for.
The schema segment joins the comparison, qualified against qualified: an
unqualified one means the target's own database, which the relation names
anyway, so comparing that would report a move on every node.
* fix(dbt): bound the retry state now that it is keyed per caller
Keying by `created_by` fixed one caller resuming another's run and created a
growth problem doing it: a shared `on_behalf_of` script kept one row and one
worker directory for everyone who had ever run it, and the generation prune only
bounds files INSIDE a directory.
Three bounds, none of them new machinery. A run with nothing failed or skipped
saves nothing — `dbt retry` builds from those nodes alone, so that state could
only ever be refused — while still clearing what the previous run left, since
its failures are no longer what last happened here. The rows expire on the same
30-day clock as a run snapshot, swept per path by the prune every dbt job
already spawns. And the worker-local directories are swept there too, by the age
of the pointer a save rewrites, because their digest names neither the script
nor the caller.
* docs(dbt): the retry state is worker-affine only on an agent worker
* fix(auth): only the server may set a token label that names a user
`create_token_internal` wrote `NewToken.label` verbatim, and the auth layer reads
some labels as an IDENTITY: `username_override_from_label` maps
`ephemeral-script-end-user-<name>` to exactly `<name>`, which then becomes
`created_by` on every job that token pushes. The label is free-form request
input, so any member could mint a token that speaks as somebody else — the shape
`require_job_read_access` already works around when it refuses to trust
`username_override` and falls back to an RLS probe, and the one that made dbt's
retry-state key (`created_by`) forgeable for an `on_behalf_of` script.
The labels are refused where request input enters: the member-facing
`tokens/create`, and `impersonate`, which names its subject in
`impersonate_email` and has no business renaming the caller too. The legitimate
producers are unaffected — a job's own token comes from `create_token_for_owner`
in the worker, and native triggers and app-embed tokens build their labels
themselves rather than accepting one.
`Ephemeral lsp token` stays allowed: its override is the fixed sentinel `lsp`,
not a name the caller chose, and the editor mints exactly that label through this
endpoint for its language server. The test pins the two lists together, so an arm
added to `username_override_from_label` that lets a label choose a name fails
until it is reserved too.
* Revert "fix(auth): only the server may set a token label that names a user"
This reverts commit
|
||
|
|
a0798a3d82 |
refactor(flows): make the flow-value round-trip preserve display-only fields in one place (#10382)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a544dfde9a |
fix(frontend): preserve top-level flow settings in AI flow tools (#10369)
* fix(frontend): preserve top-level flow settings in AI flow tools Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DaNkfh3YH8VoTNebkRuunB * fix(frontend): treat degenerate agent transforms as unconfigured in chat mode toggle Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DaNkfh3YH8VoTNebkRuunB * fix(frontend): treat persisted static-null agent transforms as unconfigured Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DaNkfh3YH8VoTNebkRuunB --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3b95a2d096 |
feat: reusable AI agent steps with rigid linking and edit/fork (#9825)
* feat: reusable AI agent steps with hybrid linking and evals Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: make linked AI agents rigid (read-only) with unlink-to-fork Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: show inherited agent config read-only on linked step Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: edit/update a saved agent in place via upsert Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: bind linked AI agent tool inputs to host flow context Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: rebind linked AI agent tool inputs via graph tool nodes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: linked AI agent tool nodes, step test, and read-only card Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: remove ai_agent resource type migration, sync from hub instead Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: remove AI agent eval suite and run endpoint, defer to later Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: unwire eval routes, types and UI (completes eval removal) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: update reusable AI agents guide for eval removal and tool rebinding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: regenerate system prompts for AIAgent agent/tool_inputs schema Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: strip brain transforms on link, avoid dirtying flow on tool open Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: flow-local test form and linked-agent marker in read-only graph Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: store linked tool overrides as diff from resource base Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: resolve linked agent tools in read-only viewer with fallback Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: use operating workspace, block non-static provider, warn on unbound tool inputs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: resolve linked parent's tools from resource for nested agent tool lookup Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope linked-agent tools by flow path, thread workspace to path check and embedded viewer Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: strip flow-context tool inputs on agent save, drop unbound-inputs warning Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: persist agent edit mode across tool selection, show linked tool code read-only Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat: show linked agent resource path in node definition panel Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: edit linked tool inputs in step panel, make tool nodes display-only Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: wire step-panel tool bindings (completes display-only pivot) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: single scroll for linked card, agent path as node label, drop fill-inputs in tool cards Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style: align linked-agent UI with design tokens and components Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: separate linked tool select target from module id to unbreak agent clicks Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@aanthropic.com> * fix: save agent tool inputs verbatim, host flows override via tool_inputs Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope agent edit state by flow path, require linked-tools scope at init Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: block saving an agent whose static provider is incomplete Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: type errors in agent tool bindings and save drawer input Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: key agent edit state by workspace, resync tool bindings on external changes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: include workspace in linked-tools scope and tool schema fingerprint Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: remove unused workspace prop from FlowModuleSchemaMap Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: drop linked-agent placeholder tool node, path label suffices Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: workspace-qualified resource links, guard stale tool schema loads Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: keep flow tool overrides out of the agent on edit, fold only on unlink Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: fold preserved tool overrides into the step on edit cancel Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: refuse overwriting non-agent resources on save, show memory kind on linked card Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: consume picker value, invalidate edit state on undo/reinit, cap nested agent tools Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: guard in-flight edit fork against restores, migrate edit state on rename Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor: validate agent edit state by fork identity instead of path keys Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: key agent edit entries by fork marker alone, immune to editor nesting Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: keep agent edit state across structural graph edits and flow renames Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: centralize agent edit reanchor, guard in-flight saves, seed rename scope from flow path Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: ancestry-keyed edit reanchor and doc-scope sweep for republished linked tools Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: guard stale linked-tool fetches and resolve while-loop nested linked agents Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * fix: drop empty tool override entries on revert and correct stale viewer comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B6kq9PYqNdc5q7ubidBYAs * docs: drop stale eval mention from the linked-agent comment Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: deploy linked agent resource, guard viewer fetches, align tools schema Address review findings on the reusable-agent branch: - Cross-workspace deploy never collected a linked step's `agent` resource, so the deployed flow failed at runtime unless the agent already existed there. - The read-only viewer published resolved tools without the generation guard flowState uses, letting a superseded link's tools win a race. Share one guarded publisher (`publishLinkedAgentTools`) between both call sites. - `tools` was still required in the OpenFlow AiAgent schema while the deserializer defaults it, rejecting hand-authored linked steps; make it optional and narrow the call sites. - Overlay `tool_inputs` in the non-linked branch too, so a flow persisted while a step sits in "Editing" mode still binds tools to this flow. - Cap the linked-tools store's scope map; nothing evicted it before. - Drop the orphaned `.sqlx` entry left by the eval removal, regenerate the copilot OpenFlow schema, and fix the generator's nested-`z.record` arity. - Move `refreshFlowStateStore` out of `agentEditStore` into its own module. - Document that linked agents' tool scripts are outside the lock pipeline. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: regenerate system prompts for optional AIAgent tools Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow saved-agent deps on deploy, accept the linked shape in the schema Round-18 review findings: - Deploying a linked flow queued only the outer ai_agent resource. Follow `$res:` refs inside a resource value (every UI-saved agent has a provider resource) and the agent's own tools, which reference scripts, flows, MCP resources and nested linked agents by bare path. - The AiAgent input_transforms schema still required provider/output_type, so it rejected the very shape linking persists (brain transforms stripped, flow-local inputs kept). Only user_message is always present. - dfs traversed `value.tools` unconditionally through a cast, which throws on a linked module that omits it now that the field is optional. - Trim the flow-refresh invariant comment to the 4-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: recurse into inline nested agent tools on deploy, require provider when unlinked Round-19 review findings: - The deploy walk only inspected a saved agent's top-level tools, so an inline nested agent tool's own scripts, flows and MCP resources were skipped. Recurse into it; a linked one is still queued as a resource instead. - Normalize a `$res:`-prefixed MCP tool resource_path like other refs. - Dropping provider/output_type from the schema's required list also let a standalone providerless agent validate, which deploys clean and then fails on every run. The constraint can't go in the schema: an `anyOf` makes AiAgent a union, which breaks the FlowModuleValue discriminated union it belongs to (verified: zod throws "Invalid discriminated union option"). Enforce it in validateFlowModules instead, next to the other cross-module checks, via a shared collectProviderlessAgentIds. - Correct the deploy paragraph in the docs: provider resources are traversed now. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: follow linked tool_inputs overrides on deploy, untrack vitest artifact Round-20 review findings: - A linked step's `tool_inputs` override replaces the resource tool's default at runtime, so a static `$res:`/`$var:` override is the dependency the flow actually uses. The deploy walk queued only the saved agent, leaving runs in an empty target workspace to fail on the missing override target. It also never scanned an aiagent module's own input_transforms, since the scan was gated to script/rawscript/flow. - Extract the pure walkers to deployDependencies.ts and cover them: three rounds have each found a further gap in this one function. - Untrack a vitest cache artifact committed by accident, and ignore a repo-root node_modules/ (only per-package paths were listed). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: collect inline agent provider and tool deps, correct tool_inputs docs Round-21 review findings: - An inline agent's provider credential sits inside an object-valued static transform, so the top-level string check missed it and such a flow deployed without its provider. Walk transform values instead of string-matching them. - An inline agent's own tools were only partly reachable: getAllModules drops MCP and websearch tools, so their resources were never queued. A standalone agent module now recurses through agentResourceDependencies, and the module's own input_transforms are scanned inside aiAgentModuleDependencies so one function owns the whole step rather than splitting it with the caller. - `tool_inputs` was documented as empty/absent for non-linked steps, which contradicts the runtime applying it when `agent` is unset so a flow persisted mid-Edit keeps its bindings. Describe that case in both the Rust doc and the OpenFlow description. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep linked steps brain-free on load, gate stale agent fetches, log linked tools Round-22 review findings: - loadSchemaFromModule filled every AI agent schema key with a placeholder transform, re-adding provider/memory to a linked step that deliberately carries none — persisted on the next save and rejected by the generated Copilot schema. Fill only the flow-local keys when the step is linked. - The linked-resource fetch was neither aborted nor tagged, so switching a step from agent A to B could publish A's tools under B and show A's brain next to B's link. Tag each result with the (workspace, path) it was fetched for and drop the ones that no longer match. - "Test this step" passed no tools for a linked agent, and the log viewer drops tool_call entries it cannot resolve to a definition, so the agent's invocations vanished from the log. Pass the resolved resource tools. - Correct the cancel-edit comment: the runtime does apply tool_inputs on an unlinked step, and folding is what leaves nothing for it to overlay. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: pin the edit session across saves, resolve linked tools in the run viewer Round-23 review findings: - Cancel stays enabled while a save awaits its requests, and it keeps the `tools` array identity, so the old guard passed and the completing save relinked the step and cleared the edits Cancel had just kept. It also accepted any replacement edit marker. Pin the path being saved and require the marker to still hold it, which still tolerates a content-preserving refresh re-anchoring the marker onto a clone. - Resolve linked agents' tools in the run/status viewer too: it reads module.value.tools straight from raw_flow, which is empty for a linked step, so AIAgentLogViewer dropped every tool_call it could not match and the graph drew the agent with no tool nodes. Same gap the previous commit closed for "Test this step" only. - Drop the overlay call-site comment: it claimed resource defaults are discarded and unmatched keys ignored, while overlay_tool_inputs preserves defaults and inserts new keys, as its own test asserts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: scope linked tools without the trigger-node path, keep the standalone save guard Round-24 review findings, both regressions from the previous commit: - Passing `path` to the run viewer's graph also switched on its Trigger node (`triggerNode ? path : undefined`), which reads a TriggerContext that /run/[...run] does not provide — the page threw "Cannot read properties of undefined (reading 'triggersCount')". Give the graph a separate `linkedToolsPath` for the tools bucket so the two stay independent. - The rewritten save guard tracked only the edit path, so a plain "Save as agent" no longer noticed the step being replaced mid-request (undo, session sync): the replacement has no edit path either, so the stale completion relinked it and stripped its brain. Keep the array-identity check when there is no edit session, and use path re-anchoring only when there is one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep recorded tool calls in run history, send tool_inputs from step previews Round-25 review findings: - The agent log viewer dropped any recorded tool_call whose definition it could not find among the supplied tools, so renaming or removing a tool — or losing read access to a linked agent's resource — erased calls that had actually run. Render the recorded call labelled by its function name; its args, logs and result come from the child job, not the definition. - "Test this step" sent tool_inputs only for a linked step, but a step forked for editing has no `agent` while still carrying the flow's bindings, which the runtime overlays. The preview ran resource-authored defaults instead of the bindings under test. Send them from both branches. - Polling a running flow replaces `job` every tick, so the run viewer re-read every linked agent's resource each time. Key the fetch on the set of linked steps instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: never discard edits made during a save, isolate the run viewer tools bucket Round-26 review findings: - The agent editor stays live while a save is in flight, so edits made after the snapshot were not in the resource yet linking stripped them from the step too, losing them outright. Compare the config against the snapshot on completion and, if it moved, leave the step alone and tell the user to save again. - The run viewer published into the editor's `${ws}:${flow path}` bucket, so opening an older run in the preview pane could flip the edited flow's tool nodes to that run's agent. Key it by job instead. - Drop the now-unreachable undefined filter in the agent log viewer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: claim the linked-tools generation on direct publishes and clears Round-27 review findings: - The step editor wrote resolved tools (and cleared them on unlink) straight into the store, leaving the fetch generation untouched. An older in-flight load for the previous agent then still passed its own check and overwrote them, so the graph and binding editor could show agent A while the step links to B. Claim the generation before those writes. - Correct two comments that still described unmatched tool calls as dropped; they are kept and labelled by their recorded name. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: retain the loaded linked agent, rebuild run logs when tools resolve Round-28 review findings: - Rejecting a superseded resource response left the card with nothing: a late reply for a previous agent replaces `linkedResource.current` and no refetch follows, so the linked step lost its brain, tools and provider warning until remount. Retain the last response that matched the current link instead. - The agent log viewer built its module list on mount only, so a linked agent's asynchronously resolved tools never replaced the placeholders, and switching between completed runs reused the first snapshot. Rebuild on a value key — callers rebuild the agentJob object each render, so tracking its identity would reload in a loop. - Refresh a linked-tools scope's recency when it is read, not only when it is published: a run viewer opens one bucket per nested job, which could otherwise evict the bucket a still-displayed run is using. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: supersede stale log reloads and stale tools on a link change Round-29 review findings, both on the reloads added last round: - Every prop change starts another loadToolCalls, and it awaits child-job requests before writing the shared view, so a slower reload for a previous run could restore its logs and tool states over the run now selected — or replace newly resolved definitions with an earlier empty-tools snapshot. Build the states locally and let only the newest load publish, including the parent's index-keyed job cache. - While a newly linked agent resolves, the previous agent's tools stayed in the store, so its bindings were editable against a step already linked elsewhere, and a failed load left them indefinitely. Clear them once the link moves away from what this component published; tools resolved at flow load are untouched, so selecting a step still doesn't flicker. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve a run's linked agents in the run's own workspace Round-30 review finding: the run viewer fetched linked agent resources with the navigation workspace, but session and fork previews render it with `workspaceId` pointing elsewhere. Those runs resolved nothing — or an unrelated resource sharing the path — losing tool nodes and log definitions. Prefer the explicit override, then the job's own workspace. The store scope stays keyed on `workspace` so it still matches what FlowGraphV2 reads; the job id in the key already makes the bucket unique. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refetch a run viewer's linked tools if its scope is evicted Round-31 review nit: the viewer publishes one scope per mounted nested job, hidden ones included, so a loop with many loaded iterations can push a displayed scope past the store's cap. Nothing refetched it afterwards — the set of linked steps had not changed — leaving the run without tool nodes or log definitions. Track the store and republish when the bucket is gone; publishing always writes a key, so this settles instead of looping. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: retain in-use linked-tool scopes instead of refetching evicted ones Round-32 review findings. Republishing an evicted scope settles for one scope but not against the cap: with more than 32 mounted nested jobs holding linked agents, restoring one necessarily evicts another, and that mutation reran every viewer's effect — an endless round of resource requests. Hold a scope for as long as a viewer is mounted and skip retained scopes when evicting, so buckets in use are never dropped and nothing has to refetch. The cap yields to correctness when everything mounted is in use. Dropping the publish key also restores refetching when the fetch workspace changes for an otherwise unchanged job and link. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: guard non-static brain edits during save, retain every displayed scope Round-33 review findings: - The in-flight edit guard compared the saved config, which holds only static brain values. A computed system prompt, memory or temperature changed while the save was awaiting the API therefore compared equal, and linking stripped it with no warning. Compare what linking actually discards — every brain transform and the tools — leaving the flow-local inputs free to change. - Retaining run-viewer scopes made them fill the cap, and eviction then picked any unretained scope, including the editor bucket a user is looking at, with nothing to refetch it. Retain the scope each graph draws from for as long as it is mounted, so every displayed bucket is protected. - A failed agent job has no parseable action list; the loader returned early and left the previously selected step's tool tree under the new header. Clear the view instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: resolve only flow modules in viewer scans, prune scopes on release Round-34 review findings: - Both viewer scans used the default dfs, which descends into agent tools, and published each linked agent under its bare id. Tool ids imported from a resource are not flow-global, so a nested linked agent sharing an id with a top-level step superseded that step's fetch and showed its tools instead. Scan flow modules only — the graph resolves the store per module node. - Scopes skipped while retained were never reconsidered, so closing views left the store over its cap for the tab's life. Prune on release too. - Correct two comments that still argued the premises the retain mechanism and the read-recency policy replaced. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't report success when a save left the step unlinked Round-35 review nits: - persist warns that changes made during the save are not in the resource and leaves the step alone, but both callers then toasted success unconditionally, burying the only actionable message. Report whether the step was linked. - Condense the tool_inputs invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: seed the published link at mount, keep run history for toolless agents Round-36 review findings: - `publishedFor` started unset, but initFlowState has already published for the step's link by then. A link change landing before this component's own request therefore skipped the clear, leaving the previous agent's tools under the new link — indefinitely if the new one fails. Seed it from the link at mount. - A standalone agent that omits `tools` kept `undefined` here, and the gate downstream then hid the AI message and tool-call history behind the generic result view. Default to an empty list like the other consumers. - A save that lands after the step was replaced writes the resource but leaves the step alone; say so instead of closing the drawer with no outcome. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: qualify nested agent tool store keys, keep an empty tools identity stable Round-37 review findings: - The step editor keyed the linked-tools store by the bare module id for nested agent tools too. Those ids come from a resource and are not flow-global, so a nested linked agent sharing an id with a top-level step read that step's tools — then overwrote them once its own fetch landed. Qualify the key by the parent agent, as the edit store already does; flow modules keep the bare id the graph looks up. - The `tools` binding handed the editor a fresh [] on every read when the module omits the field — a shape this PR made valid — so the save guard's identity check never matched and such a step could never link. Read through one shared empty array instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: accept the first tool on an agent module that omits tools Round-38 review nit: the graph's tool insert required an existing `tools` array, so a module authored without the field — valid since `tools` became optional — swallowed the insert while still pushing history and dispatching a change. Create the array on first use. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't evict a scope on the write that created it, and cover the store Round-39 review findings: - A rename removed the retained old key from the order but the new one is not retained until readers re-run, so eviction deleted the fresh bucket immediately. Reorder without evicting; the next publish or release enforces the cap, by which point the new key is held. - Writing the test for that surfaced the same shape in touchScope: it evicts right after appending, so once every older scope is retained the scope just published was the only eligible victim and was dropped at once. Exclude the scope being written. Add the store's first test: retention, eviction past the cap, pruning on release, and the rename handoff — four rounds landed fixes here with nothing pinning the behaviour. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-resolve linked agents when a wholesale edit changes the links Round-40 review findings: - Undo/redo, YAML apply, AI apply and session restore swap a step's `agent` without re-running initFlowState, and the step editor only watches the step it is mounted on — so an unselected step kept showing, and binding against, the previous agent's tools. Re-resolve from the editor whenever the set of links changes. - Document that linked resolution is live rather than pinned: an edit landing mid-run affects steps that have not started, and a nested agent tool looks its definition up by id when its own job starts, so it can run a changed definition. Pinning would mean carrying the resolved definition into the child job instead of its id; inline agents are unaffected because their tools are snapshotted with the flow value. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: per-module empty tools identity, invalidate tools when a link is replaced Both findings are over-corrections in the two preceding commits: - The shared empty-tools array made identity stable, but stable everywhere: a wholesale edit that keeps the module id reuses the component, so when both the old and the replacement module omit tools the save guard saw no change and could link and clear the replacement. Hand out one empty array per module value, which a replacement always renews. - The editor's link watcher resolved the replacement agent without dropping the previous one's tools first, so a step selected before the fetch landed still showed agent A under link B — and the freshly mounted editor seeds itself from B, so it could not tell. Clear the entry when the link for a module changes, seeding the map from the graph so the first run doesn't refetch what initFlowState just resolved. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reserve graph space for linked tools, re-resolve only changed links Round-41 review nits: - The layout reservation read the module's own `tools`, which is empty for a linked agent, so its display-only tool nodes were drawn over the node above in read-only viewers. Count the resolved tools for a linked step. - The editor's link watcher refetched every linked agent on each run. Resolve only modules whose link actually changed, and skip the pass entirely on a rename, where the scope sweep has already carried the buckets over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: protect a renamed scope until it is retained, drop the phantom tool row Round-42 review nits: - Readers release the old scope before retaining the new one, so a migrated bucket is unretained in between and, over the cap with everything else held, was the only thing eviction could take. Protect a just-migrated scope until a reader retains it, and cover that release/retain order in the store test. - The layout reserved an add-tool row for linked agents, which have no add-tool node, leaving dead vertical space. Match computeAIToolNodes. - Re-resolving links no longer short-circuits on a rename: comparing each module still costs nothing when only the path changed, and a restore that renames and relinks in one tick now gets both. - Hoist the duplicated linked-tools lookup in the graph's store update. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: kill a scope's in-flight fetches before migrating it Round-43 review finding: fetch generations are keyed by (scope, module), so a resolution still running against the pre-rename scope keeps a valid generation there. It publishes into the old bucket after the rename, and the doc-scope sweep — which gives the source precedence — carries it forward over a link resolved since under the new scope, leaving the graph and binding editor on the previous agent's tool ids with nothing to refetch them. Invalidate the source scope's fetches before each migration, and pin the behaviour: the new test fails without the invalidation. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-resolve links a scope sweep cancelled, and only sweep a real bucket Round-44 review findings, both on the previous commit: - Invalidating the source scope killed fetches that were perfectly current — a link still loading when the rename landed — and nothing restarted them, because the watcher already records that link. Resolve again, in the destination, every link the migration left without tools. - The doc-scope sweep ran on every store version bump, so during a draft refresh the first completed fetch cancelled the others mid-flight. Skip the sweep entirely when the source scope holds nothing. - Condense a six-line invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: split rename from doc sweep, hide brain fields of nested linked agents Round-45 review findings: - Two reviewers disagreed about invalidating a scope whose bucket is empty, because the two callers differ. A rename is a cut-off: every fetch still running against the old scope is stale whether or not anything resolved there, so it always invalidates. The doc-scope sweep has no cut-off — those fetches belong to the refresh in progress — so it still waits until that scope holds something. - Recording the swept links as published undid the rename+relink fix: a restore that renames and swaps a link in one tick would keep the previous agent's tools with nothing to refetch them. Leave that comparison to the watcher, which compares links rather than presence. - A nested agent that is itself linked was offered the whole agent schema in the tool bindings, but the runtime overlays only its flow-local inputs, so the rest were collected and dropped. Show what actually applies. - Condense the hybrid-linking comment to the constraint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: don't resolve a shared agent's tool defaults when loading it Round-46 review finding: the whole agent resource was interpolated before tool_inputs was overlaid, so each tool's default `$res:`/`$var:` resolved first. A host flow overriding a default that points at the author's resource still had to resolve that resource, and an unused tool whose default is unreadable in the consumer's permission context failed the agent outright — defeating the point of sharing an agent across contexts. Read the resource raw, overlay the host's overrides, and interpolate only the brain; each tool resolves its effective inputs when it executes. The nested tool lookup reads raw too, since it only needs definitions. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: interpolate the brain before overlaying caller inputs Round-47 review findings, all on the previous commit: - user_message and user_attachments were inserted before interpolation, so they went through it a second time: a user message of `$WM_TOKEN` expanded to the job token and was sent to the model provider. Interpolate the resource first, then overlay the already-resolved flow-local inputs. - The relink watcher skips tool nodes, so a linked agent nested as a tool kept the previous agent's entry through undo, YAML/AI apply or a session restore, and the step editor seeds itself from the new link and cannot tell. Emit the ancestry-qualified key for those too. - Correct the guide, which still named the interpolation path this branch replaced, and condense two invariants to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: deploy $jsonvar deps, key run logs by tool identity, seed only top links Round-48 review findings: - The deploy walkers recognised `$res:` and `$var:` but not `$jsonvar:`, which the worker resolves too, so a secret referenced that way by an agent brain, a saved tool default or a host override never reached the target workspace. - The run log rebuilt only when a tool's name or the tool count changed, so a refreshed resource that altered a tool's path, code or id behind the same name kept showing the old definition. Key on the array identity instead: the store swaps it exactly when the contents differ. - Nested linked agents were seeded as already published, but initFlowState resolves only top-level links, so their tools never loaded until their editor was opened. Seed what initFlowState actually publishes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: let the watcher's fetch survive the step editor's stale-clear Round-49 review nits: - On a relink the step editor claimed the fetch generation before clearing the previous agent's tools, which discarded the watcher's already-running fetch for the new link. The tool nodes then only appeared if the step stayed selected until the editor's own refetch landed. Clear without claiming: the watcher superseded the old fetch when the link changed, so nothing stale can return. Unlink still claims, since no watcher fetch covers it. - Condense the store's opening invariant to the four-line limit. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: condense the stale-clear invariant Round-50 review nit. Also records why the branch deliberately doesn't claim a fetch generation: a reviewer asked for the opposite this round, but writing `agent` re-runs the editor's watcher, which supersedes the old fetch and starts one for the new link — claiming here would discard it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: guard Edit/Unlink by step identity, not just the link path Round-51 review finding: forkFromResource compared only the agent path after its fetch, so a module replaced mid-request while keeping the same link passed the check — the stale continuation then wrote the fetched brain and tools into the replacement and unlinked it. Compare the step's own `tools` array too, which is one instance per module value and so identifies the step. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: report an Edit or Unlink abandoned because the step changed Round-52 non-blocking note: forkFromResource returns undefined when the step was replaced mid-request, and both callers treated that as do-nothing, so the click looked ignored. Say what happened, as the save path already does. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: hugocasa <hugo@casademont.ch> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@aanthropic.com> |
||
|
|
0f62891d43 |
fix(ai): stop teaching nonexistent while-loop iter.value state-carrying (#10345)
* fix(ai): stop teaching nonexistent while-loop iter.value state-carrying Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): scope while-loop results guidance to cross-iteration reads only Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): drop unverified wmill state-helper fallback from while-loop guidance Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): document supported cross-iteration results state in while loops Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(ai): rescope while-loop fast-path rule and add results-carrying example Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ecb1a92070 |
fix(copilot): stop write_flow forcing rawscript code into nested JSON (#10260)
* fix(copilot): stop write_flow forcing rawscript code into nested JSON The global-chat write_flow tool made the model embed rawscript bodies inside the modules JSON string, so code had to survive three levels of escaping (tool arguments -> modules string -> content string). Models routinely mangled the quotes/newlines and flow creation failed on the first tries. Bring write_flow to parity with flow mode's set_module_code escape hatch: detect rawscript modules left empty or as inline_script placeholders and tell the model to fill them via set_flow_module_code, add a code-escaping hint to the JSON parse error, and update the guidance to keep code out of the modules structure. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(copilot): only warn on saved write_flow; add regression tests Address review: writeFlowDraft reports conflicts/persistence errors as {success:false} rather than throwing, so the empty-body warning must be folded into the JSON result only on a successful save — otherwise the model is told to set_flow_module_code on a flow that was never saved (stale or nonexistent draft). Add core.test.ts coverage for the empty-body warning (top-level, nested, preprocessor, failure; populated suppressed), the no-warning-on-failed-save path, and the malformed-JSON escaping hint. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(copilot): warn on patch_flow_json inline_script placeholders Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test(ai-evals): add global case for quote-heavy inline flow code Exercises write_flow creating a rawscript whose body is multi-line and quote-heavy (the scenario the write_flow fix targets), so the global-mode A/B can measure that code lands out-of-band via set_flow_module_code rather than being escaped into the modules JSON string. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(copilot): resolve inline_script placeholders in global flow writes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor(copilot): condense patch_flow_json warning comment Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(copilot): soften write_flow guidance to inline by default Benchmarks showed the aggressive "empty content + set_flow_module_code for any multi-line/quoted body" guidance pushed even capable models onto the multi-round-trip fill path, inflating per-iteration overhead with no reliability gain when inline escaping would have succeeded. Default to inlining and reserve the empty+fill escape hatch for bodies that are genuinely hard to escape or when a write_flow call returns a JSON parse error — the case that actually benefits escaping-prone models. The warning, parse-error hint, and set_flow_module_code recovery path are unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(ai-evals): drop global quote-heavy inline-code case The manual global-mode A/B (Sonnet, Gemini 3 flash/pro, GPT-4o) showed no pass-rate delta: the GPT-5 inline-escaping failure this change targets does not reproduce on any available model, so the case guards nothing measurable. Keep the unit tests as the regression guard instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: Guilhem Lemouel <guilhemlemouel@gmail.com> |
||
|
|
af3e3fe667 |
fix(ai-chat): size AI-created flow notes to fit their text (#10091)
* fix(ai-chat): size AI-created flow notes to fit their text Free notes created via the flow AI chat omit `size` (the tool prompt tells the model to let the editor size them). validateFlowNotes seeded a fixed 275x60 box, but free notes never grow to fit content, so multi-line markdown overflowed the box. Estimate height from the text instead. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ai-chat): stack auto-placed flow notes by height to avoid overlap Auto-placed free notes were staggered by a fixed index*84px step, but notes can now be up to 600px tall, so consecutive generated notes overlapped. Track a running y-cursor and advance it by each note's real height. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ai-chat): advance note stack cursor past preserved column notes A round-tripped note keeps its existing auto-column geometry ({-375, y}); the stack cursor ignored it, so a newly added geometry-less note landed on top. Preserved notes overlapping the auto-stack column now advance the cursor. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(ai-chat): trim estimateFreeNoteSize comment per AGENTS.md Keep only the non-obvious fixed-height renderer constraint; drop the implementation narration. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
7ebfad382a |
feat(ai-agent): give tools a real description instead of the tool name (#10083)
* feat(ai-agent): use a real tool description instead of the tool name Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ai-agent): render tool-name error full width and hoist it above the description Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ai-agent): make tool description field hug its content so a single line is vertically centered Add an optional minHeight param to the autosize action (default unchanged at 30px) and pass minHeight 0 for the tool description so an empty/one-line field no longer reserves the 30px floor and leaves dead space below the text. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(ai-agent): regenerate OpenFlow-derived prompts, CLI guidance, and copilot zod schema for tool description Fixes the check-freshness CI failure (system_prompts + skills.gen.ts) and makes the flow copilot's openFlow.json / openFlowZod.gen.ts aware of the new AgentTool.description field so AI-authored tools can set it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
c5060a1e9a |
fix: scope AI-session flow/script editors to the session workspace (#10025)
* fix: scope flow script-edit drawer to session workspace and fix scroll Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope flow schema inference to session workspace Thread an optional workspace through loadSchemaFromPath/loadSchemaFlow/ loadSchemaFromModule/loadFlowModuleState/initFlowState/pickScript/pickFlow and pass the op (session) workspace at fork-context call sites, so a flow opened in an AI session resolves path-referenced scripts/subflows against the session workspace instead of the nav workspace. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: scope script editor log panel and git-repo pickers to op workspace LogPanel and the ansible git-repo viewer/picker read the nav workspace directly; pass the script editor's op workspace so past-test results/logs and git-repo resource/file lookups target the session workspace in a fork. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct session pipeline trigger-editor workspace comment The comment claimed session activation syncs $workspaceStore; SessionPicker intentionally does not, so trigger create/edit/delete from a fork session's pipeline canvas writes to the nav workspace. Document the known limitation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: forward op workspace to git-repo S3 file browser GitRepoViewer scoped its own calls to the op workspace but rendered the nested S3FilePickerInner without workspace={ws}, so the file list/preview/ metadata still queried the nav workspace with a session-workspace prefix. Addresses Codex review on #10025. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5a460dbec6 |
fix: accept bunnative language in AI chat flow step validation (#10030)
* fix: accept bunnative language in AI chat flow step validation Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: regenerate copilot flow schema from openflow spec Run gen_openflow_schema.sh + minifiedOpenflowJson.sh instead of hand-patching. Also syncs three fields the checked-in generated files had drifted from since the last regen (reasoning_effort, reasoning_token_delta streaming event, aiagent tag). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore: regenerate system prompts for bunnative openflow schema Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
9036ac789f |
fix(frontend): name the draft in AI chat test-run confirmation (#10024)
* fix(frontend): name the draft in AI chat test-run confirmation
The confirmation card shown before an AI-chat test run displayed a
static, generic header ("Run script test"). Make it name the target
and clarify it runs the user's draft.
- Tool.confirmationMessage now accepts a function of the parsed args;
shared.ts resolves it before setting the tool status.
- test_run_script/flow/step (global chat) build a dynamic header
naming the script/flow/step, e.g. "Run a test of your draft of X".
- In-editor script/flow test-run tools say "Run a test of your draft".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(frontend): fall back to tool name in YOLO tooltip for function messages
The auto-accept ("bypassed in current mode") tooltip rendered
confirmationMessage directly. Now that it can be a function of the call
args, render the tool name instead of the function source there.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(frontend): use neutral wording in test-run confirmations
The global test-run tools fall back to deployed content when no draft
exists, so "your draft" could contradict what actually runs. Drop the
draft claim and just name the target: "Run a test of X".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
d9b080f57f |
feat(ai): add Azure AI Foundry as a native AI provider (#9879)
* feat(ai): add Azure AI Foundry as a native AI provider Adds `azure_foundry` as a new AIProvider variant wired through the AI chat (copilot) and AI agent flow steps. Foundry's chat completions API is OpenAI-compatible and uses Azure conventions (api-key header, Azure URL building), so it reuses the existing OpenAI-compatible query builder and proxy path via the shared `is_azure` helper (renamed from `is_azure_openai`). Backend (windmill-ai): - New `AzureFoundry` enum variant (serde `azure_foundry`) - `get_base_url` requires a resource base URL (like Azure OpenAI / Custom) - `is_azure()` covers Azure OpenAI + Foundry (api-key auth, Azure URL) - Added to OpenAI-compatible proxy support and HttpForward proxy mode - New proxy URL unit test Frontend (copilot): - New provider entry, completion config, model-token handling, streamed usage tracking, and reasoning registry (all model-id-gated, so a no-op for Foundry's non-OpenAI catalog) - Treated as a chat-completions provider, not the OpenAI Responses API OpenAPI: - `azure_foundry` added to AIProvider (openapi.yaml) and AIProviderKind (openflow.openapi.yaml); regenerated CLI guidance Note: the `azure_foundry` resource type (base_url + optional api_key) is hub-managed and must be published to the Windmill Hub separately. Fixes WIN-2122 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ai): add azure_foundry to copilot flow Zod provider enum The tracked copilot flow schema (openFlowZod.gen.ts and its openFlow.json source) still carried the old AIProvider enum, so validateFlowModules / validateSpecialFlowModule rejected AI-generated flow edits that create or update an aiagent module with provider kind "azure_foundry" before they could be saved. Add the value to both (preserving the generated single-line format) and a regression test over the flow-module validation path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(ai): lead provider list with OpenAI, Anthropic, Google AI Reorder AI_PROVIDERS so the three primary direct providers come first. The AIProviderPicker renders the first three entries as quick-access buttons, so these become the defaults (previously OpenAI, Azure OpenAI, Azure Foundry); Azure OpenAI / Azure Foundry stay adjacent right after. No logic depends on provider order (only per-provider defaultModels[0] is read). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
d5cb944cf9 |
fix(frontend): ensure type:object in test_run_flow tool schema for Anthropic (#9721)
Flows with no defined inputs can produce a sparse schema (e.g. { order: [] })
that lacks the "type": "object" field. buildSchemaForTool spread this schema
into the tool parameters as-is, so the Anthropic API rejected the tool
definition with `400 invalid_request_error:
tools.N.custom.input_schema.type: Field required`. The existing fallback in
anthropic.ts only triggers when parameters is falsy, but the sparse schema is
truthy.
Default type:object before spreading the schema in buildSchemaForTool, and
backfill type/properties/required in FlowAIChat's getFlowInputsSchema as a
defense-in-depth measure.
Fixes WIN-2087
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
f2f0812a04 |
feat(flows): opt-in to include the stopping step's result in early-stop errors (#9446)
* feat(flows): early stop can include the stopping step's result in the raised error
When a step uses Early Stop with "Raise an error message if stopped", the
flow result was entirely replaced with a static error object
({"error": {"name": "EarlyStopError", "message": "..."}}), discarding the
stopping step's own output. This made it impossible to stop+fail a flow
while preserving the data the step produced (e.g. an API that returns
HTTP 200 with a userErrors payload).
Add an opt-in `error_include_result` flag on StopAfterIf. When enabled on
the raise-error path, the raised payload becomes
{"error": {...}, "result": <step result>} instead of dropping the result.
Default is false, so existing behavior is unchanged. The option is threaded
through the worker's stop-after-if handling (including stop_after_all_iters_if
for loops/branchall) and exposed in the flow editor's Early Stop panel.
Fixes WIN-2012
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(flows): cover early-stop error_include_result payload shaping
Add a regression test asserting that a step using Early Stop with a raised
error message and error_include_result=true fails the flow while preserving
the step output as {"error": {..}, "result": <step result>}, and that with
the flag off the result is the bare {"error": {..}} object.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(flows): nest early-stop step result inside the error object
Embed the stopping step's result under `error.result` rather than as a
top-level sibling of `error`. This keeps the flow result shape as
`{ "error": { .. } }` — identical to a normal error — so consumers that
key off the top-level shape (single `error` key) keep working, while the
data is still preserved for those that look inside the error object.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(flows): always include the stopping step's result in early-stop errors
Drop the opt-in `error_include_result` gate. Since the step result is nested
inside the error object (`error.result`), the top-level result shape stays
`{ "error": .. }` — identical to a normal error — so consumers that detect or
parse failures by the top-level shape are unaffected. Gating it added schema
surface, plumbing, and a UI toggle for no real compatibility benefit.
Now, whenever a step early-stops with a raised error message, the flow fails
and the raised error embeds the stopping step's own result under
`error.result` (aggregated iteration results for loops/branchall). This
reverts the `StopAfterIf.error_include_result` field, its threading, the
OpenAPI/generated-client surface, and the editor toggle; the "Raise an error
message" tooltip now notes that the step result is included.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* feat(flows): gate early-stop result inclusion behind opt-in flag
Re-introduce the per-step `error_include_result` flag (default off) instead
of always embedding the step result. Although nesting the result under
`error.result` keeps the result *shape* backward-compatible, it does not
address data exposure: a failed flow's result is propagated to synchronous
webhook callers, the flow's failure module, and the workspace/global error
handler (commonly a Slack/email/outbound-webhook notifier). Always including
the step output would surface previously-redacted intermediate data to all of
those sinks for every existing error-stop flow.
Gating keeps the existing behavior (bare `{ "error": .. }`) as the default and
only embeds `error.result` when the flow author explicitly opts in, matching
the original issue's intent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(flows): omit error_include_result when false; refresh generated prompts
- Add `skip_serializing_if = "is_false"` to `StopAfterIf.error_include_result`
so serialized flows are byte-identical when the flag is off. Fixes the
`flowmodule_serde` round-trip test (cargo_test) and avoids churn on existing
flows.
- Regenerate `system_prompts/auto-generated/` and `cli/src/guidance/skills.gen.ts`
for the new OpenFlow `error_include_result` property. Fixes check-freshness.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(flows): cover error_include_result for the loop "stop after all iters" path
Add a regression test for the stop_after_all_iters_if branch, where `nresult`
already holds the aggregated iteration results — confirming `error.result`
carries each iteration's output (distinct from the per-step fallback path).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
||
|
|
e4e0984e55 |
feat: let flow AI chat create and edit sticky notes (#9412)
* feat: let flow AI chat create and edit sticky notes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: strengthen flow AI guidance to prefer groups for organizing flows Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: harden flow note validation (validate position/size, document color default and group acceptance) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix: make AI-created free notes draggable by seeding default position and size Free notes need explicit geometry to be draggable/resizable in the editor; UI-created notes always set position+size but agent-created notes omitted both, so they couldn't be moved until resized. Seed defaults in validateFlowNotes. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> |
||
|
|
5c20d6b4f7 |
feat: add global ai chat test tools (#9391)
* feat: add global ai chat test tools
* fix: avoid session id in flow test preview
* test: cover global flow preview ids
* test: require script and flow test tools
* fix: harden global flow test fallback
* Revert "fix: harden global flow test fallback"
This reverts commit
|
||
|
|
e4213c1ab8 |
feat(flow-ai): constrain flow-group colors to the NoteColor palette (#9343)
The flow AI chat's set_flow_json tool lets the model set a `color` on each semantic flow group, but nothing told it which colors are valid, so it would sometimes emit hex codes / arbitrary CSS color names. Those render with default styling at best and break the group color picker at worst. - core.ts: the set_flow_json schema `.describe()` and the `groups` system-prompt bullet now spell out that `color` MUST be one of the palette names (yellow, blue, green, purple, pink, orange, red, cyan, lime, gray) — no hex, no CSS colors — and that omitting it lets the editor auto-assign one. - helperUtils.ts: validateFlowGroups now rejects any color outside that palette, sourced from the NoteColor enum so the two can't drift. - helperUtils.test.ts: tests for reject-unknown / accept-known / accept-omitted. Split out of the sessions branch (gl/layout-ai), where it had been bundled into the large feature commit; it's an independent flow-AI improvement. Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com> |
||
|
|
eadeac248b |
feat: sessions page with isolated AI chat + flow editor (#9034)
* feat(sessions): chat + editor side-by-side with multi-session state
Introduces the Sessions feature: a workspace where the AI chat and an
editor (flow / script / app / raw-app) sit side-by-side, with each session
having its own AIChatManager instance, history, and target item. Sessions
are persisted across reloads and can be staged into forks for review.
Key pieces:
- sessions/ — SessionWrapper (the split-pane shell), SessionPicker
(sidebar list), SessionForkBar, SessionWorkspaceBar, FlowEditorView /
ScriptEditorView / AppEditorView / RawAppEditorView, ForkDiffDrawer,
sessionRuntime (per-session AIChatManager + draft state),
sessionState (in-memory + persisted index), sessionUnread, sessionScope,
appDraftCodec / flowDraftCodec, forkEditUrl, /sessions route.
- WorkspaceItemDrillPicker refactor — extracts WorkspaceItemRow + adds
surfaceAI drafts, stale-while-revalidate. workspacePicker.ts drops
explicit invalidate() in favor of always re-fetching in the background.
- ForkDiffDrawer + WorkspaceItemDiffViewer — per-kind diff bodies
reusable from the compare page. FlowGraphDiffViewer / FlowGraphV2 gain
inlineDiff forwarding + onHeight callback for equal-height layout.
- Global AI chat sessions plumbing — AIChatManager exports the class +
adds disabledModes, beforeSend hook, scoped instance context. AIChat /
AIChatDisplay accept session-only props (wideLayout, emptyHint,
inputPreface, hideHeader, hideModeSelector, forceDisabled). Chat
preserved across /flows/add → /flows/edit, /scripts/add → /scripts/edit.
- Draft-first loaders — sessions open drafts when present, otherwise
seed a draft from the last deployed value via globalDraftStore.
RawAppEditor / AppEditor / AppEditorHeaderDeploy get newApp prop +
fixes so draft-only apps can deploy.
- Compare page (/forks/compare) — bigger overhaul to plug into the new
drawer.
- Sidebar — Sessions entry + unread badge + status dot in
SidebarContent / MenuButton / SideBarNotification.
- Misc fixes — chat group color palette constraint, deploy_workspace_item
confirmation dropped, open_preview tool, picker drafts surfacing,
fork archive/delete buttons on compare page.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(sessions): bypass UserDraft inside session panes + sessionUnread crash
After merging main's UserDraft PR (#9121) into the sessions branch, two
integration issues surfaced:
1. AppEditor.svelte calls `UserDraft.use<App>('app', path)` at the
component level — keyed by ($workspaceStore, 'app', path). Sessions
that haven't materialized a fork yet stay at the user's main
workspace, so a session targeting an app at the same path as a
regular /apps/edit tab shared the same LS key. The session would
read the regular tab's autosave and write its fork-edits back over
it.
Gate UserDraft.use on `!getContext('aiChatManager')` — sessions
inject the manager via setContext, so inside a session pane the
handle is `undefined`, stateApp falls through to the `app` prop
the session loaded, and the auto-save $effect bails. Same gate on
the four UserDraft.remove call sites in AppEditorHeader and
RawAppEditorHeader so save/deploy from a session pane doesn't wipe
the LS draft of a non-session tab at the same path.
2. sessionUnread.svelte.ts called useLocalStorageValue at module
scope. Main's PR added a deep-mutation $effect inside that helper,
which now requires component-initialization context — every page
crashed at import time with `Svelte error: effect_orphan`.
Replaced with a plain module-level $state + manual localStorage
persist; same reactivity contract for callers.
3. ScriptEditorView.svelte was passing a `replaceStateFn` prop that
ScriptBuilder dropped on main. Removed.
Verified end-to-end with Playwright:
- /flows/edit/{path} regression: UserDraft handle still created, no
console errors
- /sessions loads, sessionUnread doesn't crash
- Session targeting non-raw app `u/admin/userdraft_collision_test`
displays the fork content (FORK_ONLY_MARKER) even with an LS
poison at `userdraft/w/local/app/{path}` containing a
POISONED_BY_REGULAR_TAB_AUTOSAVE marker; poison remains untouched
after the session loads and renders
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(sessions): stop fork-create retry loop on first user message
Removed the SessionWrapper $effect that retroactively committed the
session's workspace from the in-memory chat history. When opening a
session whose previous commit attempt had failed (or whose response was
lost) the effect ran in a tight retry loop, flooding the user with
`workspace_pkey` violations from `create_workspace_fork`.
The send path already commits through `AIChatManager.beforeSend` →
`commitSessionWorkspace`, which is the deterministic moment-of-action.
The $effect was a redundant reactive bridge that turned every backend
failure into an infinite retry.
Also hardens `materializeFork`/`commitSessionWorkspace` so the most
common cause of the duplicate-key error self-heals:
- `materializeFork` short-circuits when `fork.id` is already in
`$userWorkspaces` (the previous create actually succeeded, we just
lost the response). On a `workspace_pkey` catch, refresh the workspace
list and adopt the existing row instead of toasting an error.
- On a real `materializeFork` failure, `commitSessionWorkspace` now
drops `pending_fork` so the session falls through to the
workspace-pick fallback instead of looping on the same broken intent.
* feat(sessions): show EditorHeader breadcrumb in the not-found state
When a session's target item has been deleted or moved, the editor pane
used to render a bare "Script not found at path X" line — leaving the
user with no way to navigate to a different target without backing out
of the session.
Each editor view now renders a `SessionItemNotFound` shell instead: a
real `EditorHeader` (read-only summary, no pen popover) with a
breadcrumb keyed to the missing kind+path, plus the "not found" copy
below. Clicking any breadcrumb segment opens the workspace picker
scoped to that level — pick a replacement and the session swaps target
via the existing `onNavigate` callback.
`SessionItemNotFound` maps `raw_app` to `EditorHeader`'s `kind: 'app'
+ raw_app: true` so the picker routes through `/apps_raw/...`; the
local label still says "Raw app not found" (not "App not found") so
the user knows which surface is missing.
* fix(picker): stop self-feeding fetch effect that OOM'd the tab
The drill picker's $effect watched `scope` and called `ensureLoaded`
on every change. `ensureLoaded` reads `loaded[kind]` synchronously
(to decide whether to show a spinner), so the effect ended up
subscribed to the very signal it fills. Each fetch result wrote
`loaded[kind] = items`; Svelte 5's $state proxy notifies on every
property set even when the reference is unchanged from cache, which
refired the effect, which called `ensureLoaded` again, which awaited
the cached fetch, which wrote `loaded[kind]` again... runaway loop.
In `/scripts/edit/...` the picker's lifecycle stabilised quickly
enough to mask the loop, but in a session pane (multiple warm
sessions, picker kept alive by the surrounding state) the cycle
spun freely — 29.8 million iterations in <100 ms during testing,
enough to OOM Firefox / kill the Chromium tab.
Two changes:
- Replace the scope-watching $effect with an explicit `setScope()`
helper called from `drill()`, `goUp()`, and `onMount`. Fetch is
now a callback reaction to user navigation, never a reactive
consequence of one. No closed feedback cycle is possible.
- Untrack the `loaded[kind]` read inside `ensureLoaded`. The search
$effect (which loads every kind on first keystroke) is still a
reactive caller; the untrack stops it from subscribing to the
signal `ensureLoaded` fills, so the same loop can't form there.
* feat(script-editor): wire initialTestPanelCollapsed through ScriptBuilder
The `initialTestPanelCollapsed` prop was already declared on
`ScriptBuilderProps` (used by the session preview to start the editor
with the run/test pane closed) but never destructured in
`ScriptBuilder.svelte`, so the value silently dropped on the floor
and the test pane always opened.
- `ScriptBuilder.svelte` — destructure the prop and forward it to
`<ScriptEditor>`.
- `ScriptEditor.svelte` — accept the prop and seed `rawTestPanelSize`
to 0 when true, while keeping `storedTestPanelSize` at the default
30 so the user's first toggle expands the pane to a sensible width
rather than 0.
Regular `/scripts/edit/...` doesn't pass the prop → default `false`
→ panel still opens by default.
* fix(sessions): resolve aiChatManager via context in AskUserQuestionDisplay
Inside a session the chat uses a per-pane AIChatManager injected via context. AskUserQuestionDisplay imported the global singleton, so answers clicked in a session dispatched to the singleton's callback map and the AI loop stalled. Resolve via getContext with singleton fallback, matching ChatMode / ToolExecutionDisplay.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(raw_apps): let preview start in single-view on the preview tab
Add a defaultSplitWithPreview prop (default true). When false (session preview), the editor boots in single view with the preview tab selected: gate the onMount default-file activation, the setActiveDocument auto-activation, and iframeShouldMount so the UI Builder bundler iframe still mounts when preview is the active tab. RawAppEditorView passes defaultSplitWithPreview={false}.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(copilot): add get_preview_status tool and make open_preview idempotent
So the assistant can tell whether the session preview already shows the item it just edited, instead of re-opening or re-offering it. Mirrors the open_preview handler plumbing (setGetPreviewStatusHandler) and the session runtime registers it alongside open_preview. open_preview now returns 'already open' when the requested target matches the active session's current target. The system prompt steers the AI to check status before offering. Unit tests cover the no-arg schema, the session-only error, and handler dispatch.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(sessions): make script preview reactive to AI draft writes
ScriptEditorView read the draft via static UserDraft.get inside an effect, which only subscribes to UserDraft's reactive cell when a live entry exists. None did for the preview path, so the chat's writes (UserDraft.save) only touched localStorage and the open preview never updated. Hold a live handle via UserDraft.useMany (reactive getter so it re-acquires when open_preview swaps the path without remounting) and read inbound through handle.draft, materializing the shared $state cell that bridges the chat's writes to the editor.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(sessions): make raw-app preview reactive to AI draft writes
Mirror of the script-preview fix. RawAppEditorView read the draft via static UserDraft.get inside an effect, which only subscribes to UserDraft's reactive cell when a live entry exists. None did for the preview path, so the chat's raw-app writes (UserDraft.save / setDraftAndMeta, from write_app_file / patch_app_file / write_app_runnable) only touched localStorage and the open preview never updated. Hold a live handle via UserDraft.useMany (reactive getter so it re-acquires when open_preview swaps the path without remounting) and read inbound through handle.draft. Verified in-browser: an external UserDraft.save live-updates the bound summary in the open preview.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* fix(sessions): make flow preview reactive to AI draft writes
Mirror of the script/raw-app preview fixes, completing two-way binding for all three session editor kinds. FlowEditorView read the draft via static UserDraft.get inside an effect, which only subscribes to UserDraft's reactive cell when a live entry exists — none did, so the chat's writes (write_flow / patch_flow_json / set_flow_module_code) only touched localStorage and the open preview never updated. Hold a live handle via UserDraft.useMany (reactive getter so it re-acquires when open_preview swaps the path without remounting) and read inbound through handle.draft. Verified in-browser both directions: an external UserDraft.save live-updates the flow header summary and rebuilds the module graph; a preview edit propagates through the debounced save to both UserDraft.get and the chat's getGlobalDraft adapter.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* feat(sessions): surface local-storage drafts in fork diff & compare page
Augments the backend fork-vs-parent comparison with browser-local (UserDraft) drafts so a session's uncommitted AI/user changes are visible in the Fork Diff Viewer and the /forks/compare page. Adds forkDraftDiff.ts (augmentForkComparisonWithLocalDrafts + getForkItemValue), a 'local changes detected' / new-draft warning surface (checkbox-slot warning icon, no-op-baseline filtering, dedup), a 'Local draft <> fork' tab in DiffDrawer, and selectTooltip/nonSelectableTooltip plumbing in Row/WorkspaceDeployLayout.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
* Revert "feat(sessions): surface local-storage drafts in fork diff & compare page"
This reverts commit
|
||
|
|
ac26aa4e4c |
feat: add yolo mode for ai chat tools (#9258)
* feat: add yolo mode for ai chat tools * nit * fix: align chat footer controls * feat: add ai chat autonomy modes * feat: add autonomy mode dropdown * fix: highlight yolo autonomy icon * fix: auto accept flow edits * fix: hide unsupported autonomy modes * fix: handle auto-accept flow editor races |
||
|
|
413404a788 | fix: collapse successful ai tool details (#9265) | ||
|
|
110384580e |
refactor: add global ai chat mode with workspace-item draft tools (#9056)
* docs: add global ai mode plan
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add global ai draft mode
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: scope global ai mode to scripts and flows
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* refactor: simplify global ai workspace item shape
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* refactor: split global ai write tool into per-type tools
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add global ai schedule and trigger workspace item tools
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add dev-only /global_drafts route to inspect ai draft store
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add edit_script and patch_flow_json global ai tools
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add deploy_workspace_item global ai tool with confirmation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: emit open-resource action card after deploy_workspace_item
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add delete_workspace_item global ai tool with confirmation
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* chore(system_prompts): emit RESOURCES_BASE and resource/variable zod schemas
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* feat: add global ai resource and variable workspace item tools
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* fix: search_resource_types uses listResourceType to avoid embedding feature dep
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>
* Revert "fix: search_resource_types uses listResourceType to avoid embedding feature dep"
This reverts commit
|
||
|
|
b883f9a9d2 |
feat: add ai chat schedule and trigger tools (#8961)
* feat: add ai chat schedule and trigger tools * refactor: use zod for ai chat workspace tools * refactor: let ai provide runnable target fields * refactor: generate ai chat workspace tool schemas * fix: add object type to composed tool schemas * fix: avoid top-level trigger schema unions Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: block undeployed workspace ai tools Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: inject ai workspace tool target Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: add ai evals for workspace tools Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: make workspace tool eval prompts realistic Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: surface workspace tool errors * fix: show workspace tool success details * fix: describe workspace tool path format * fix: clarify workspace path examples * fix: tighten workspace tool validation * fix: align workspace tool prompts * chore: mark generated chat schemas * chore: mark generated cli skills --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
932d183311 |
fix: persist flow groups from AI chat tool calls (#8906)
* fix: persist flow groups from AI chat tool calls Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: validate group ids and coerce empty groups to undefined Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com> |
||
|
|
a5363ea4ed |
refactor: unify flow chat tree operations (#8862)
* refactor: make flow chat code edits explicit * refactor: centralize flow tree lookups Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: simplify flow chat tree mutations Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: reuse flow tree lookup in schema map Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: remove flow lookup alias Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: reuse flow tree in previous results Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: reuse canonical flow module lookup * fix: align rebased flow helpers Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * docs: remove flow chat cleanup plan Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: remove flow chat helper wrappers Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: preserve non-flowmodule AI agent tools in skeleton and previous_result Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: consolidate flow module ID collectors into flowTree Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: search full flow tree in test_run_step to find special modules Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: recurse into aiagent tools in collectAllFlowModuleIds Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
51b09ace45 |
feat: add empty inline script warnings to flow chat (#8853)
* fix: seed empty inline flow scripts Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: cap frontend eval chat turns * fix: roll back failed inline script seeding * refactor: simplify inline flow script warnings * refactor: share flow module traversal * refactor: make flow chat code edits explicit * fix: resolve ai tool review actions * refactor: remove dead flow rawscript helper --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
49844eb240 |
fix: encourage subflow reuse in AI chat flow builder prompt (#8839)
* docs: encourage subflow reuse in AI chat flow builder prompt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test: add workspace flow reuse benchmark Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: centdix <farhadg110@gmail.com> |
||
|
|
b39671d933 |
feat: add compact json patch tool to flow chat (#8840)
* fix: use compact json for flow patches Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: improve flow eval harness Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: record flow benchmark history Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: preserve schema in set flow json Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * style: clean set flow json schema guard Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: clean flow patch review followups Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
d3cb0c6220 |
fix: improve flow chat and benchmark coverage (#8825)
* fix: support special flow modules in evals Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: extract shared flow helper logic Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: make special flow tools openai-compatible Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: improve flow eval prompts and validation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: relax flow benchmark overfits Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * test: record updated flow benchmark history Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: address flow review findings Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: source flow chat special module prompt Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: narrow rawscript helper return type Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * refactor: dedupe flow chat prompt guidance Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * fix: relax flow test10 validation Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
cdcc56461b | feat: add black-box ai eval benchmarks (#8618) | ||
|
|
e44504c6e9 |
feat: add PDF input support to AI agent (#8525)
* feat: add PDF input support to AI agent with user_attachments field Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test: add integration tests for PDF input and backward compat Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add ContentPart::File variant for PDF support across all providers Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: address review feedback on PDF support - Extract parse_data_url_bytes and mime_to_document_format helpers in Bedrock - Add is_document_mime helper in ai_types for centralized MIME routing - Extract s3_object_to_content_part helper to deduplicate image_handler/openai - Rename AnthropicImageSource to AnthropicBase64Source - Derive Bedrock DocumentFormat from MIME type instead of hardcoding Pdf Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: merge user message and attachments into single message for Bedrock Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> |
||
|
|
5d79f33590 |
Final Svelte 5 migration (#8211)
* Remove $$props.field usage * Rename slots to ensure no hyphen * _props * _trigger * OnSelectedIteration type correct capitalization * rename _content * Remove afterUpdate * Migrate everything to svelte 5 * array bind * Fix popover * type never * nit fixes * Fixed many trivial errors * onClick * Fix errors * use let: * nit typing * fix: wrap state_referenced_locally vars with untrack() Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * Add untrack import * Fix all syntax errors due to untrack migration * Fix undefined errors * Fix more undefined errors * untrack(() => initialOpen) * svelte-ignore * Fix state_descriptors_fixed error in Chart.svelte Use $state.snapshot() to pass plain copies of data/options to Chart.js instead of $state proxies. Chart.js's listenArrayEvents tries to define property descriptors on data arrays, which Svelte 5 proxies reject. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * nit typing * Merge issue * Fix "path is not set" error in resource picker / editor * Fix InputTransformForm error when rerunning some flows * fix npm run check --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
a7e269f9f3 |
feat: add workspace search and runnable details tools to AI chat modes (#7874)
* feat: add workspace search and runnable details tools to navigator mode Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: correct uFuzzy search result indexing in workspace search uFuzzy.search() returns [idxs, info, order] where order contains indices into idxs, not into the original haystack. The code was using order values directly as array indices, returning wrong results. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: mutualize search_workspace and get_runnable_details tools - Move search_workspace tool def + implementation into shared.ts as createSearchWorkspaceTool() factory, used by navigator and flow modes - Move get_runnable_details tool into shared.ts as createGetRunnableDetailsTool() factory, used by navigator, flow, and script modes - Replace flow mode's scripts-only search_scripts with search_workspace that searches both scripts and flows - Add search_workspace and get_runnable_details to script mode - Remove duplicated WorkspaceScriptsSearch class from flow/core.ts Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add get_runnable_details to flow mode system prompt Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add hard limit on runnable content passed to AI context Truncate script content and flow value at 20k chars in get_runnable_details to avoid flooding the context window. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: make search_workspace type param required for strict schema OpenAI strict mode requires all properties in required array. Make type a required enum ('all', 'scripts', 'flows') instead of optional. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * cleaning * nit * cleaning * refactor: use shared createSearchWorkspaceTool in app mode Replace app mode's local list_workspace_runnables tool with the shared createSearchWorkspaceTool() factory, consistent with navigator, flow, and script modes. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * search by keyword * cleaning * fix: document search_workspace and get_runnable_details in script mode system prompt Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add get_runnable_details tool to app mode Without it, the AI can find scripts/flows but can't inspect their schema/content when configuring backend runnables with correct inputs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: race condition in WorkspaceRunnablesSearch workspace caching Track scriptsWorkspace and flowsWorkspace separately instead of a single shared workspace field. Previously, initScripts could update the shared workspace field, causing initFlows to skip re-fetching when the workspace changed (it saw the workspace already matched), returning stale data. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> |
||
|
|
a7ce5484b8 |
feat(local-dev): create Claude skills when doing wmill init (#7699)
* use skills * add prompts * update system prompts * generate skills on init * add prompts in cli * better for raw apps * nit * test pipeline draft * better * yaml for triggers and schedules * cleaning * better * add descriptions to ai agent fileds * adjust * better openapi * better * nit * feat: add typed provider and memory schemas for ai agent in openapi Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * feat: improve zod validation errors with dynamic schema extraction Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * regen * fix * cleaning * refactor: deduplicate skill descriptions in generate_skills_ts_export Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * cleaning --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
d4a1b4abed |
add aiagent module support to inline script extraction/replacement (#7773)
* dual build for utils-internal * bump version * feat(cli): add aiagent module support to inline script extraction/replacement Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com> * add missing field in openapi * bump yaml validator version * cleaning * cleaning * cleaning * nit * cleaning * cleaning --------- Co-authored-by: Claude Opus 4.5 <noreply@anthropic.com> |
||
|
|
62cb147847 | fix tool validation (#7482) | ||
|
|
15a4b26d44 |
feat(aichat): add get_lint_errors tool for script and flow mode (#7431)
* feat(aichat): add get_lint_errors tool for script and flow mode This adds a new `get_lint_errors` tool to the AI chat for script and flow modes, similar to what exists for app mode. For script mode: - Added `getLintErrors` function to Editor.svelte that returns lint errors from Monaco - Added `ScriptLintResult` and `ScriptLintError` interfaces - Added `get_lint_errors` tool definition and implementation - Updated system prompt to instruct AI to use the tool after code changes For flow mode: - Added `FlowLintResult` interface for flow-level lint results - Added `get_lint_errors` tool that gets lint errors from the currently selected module - Updated system prompt to include linting in the tool selection guide The AI is now instructed to always use `get_lint_errors` after making code changes and fix any errors before proceeding with testing. Closes #7430 🤖 Generated with [Claude Code](https://claude.com/claude-code) Co-Authored-By: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> * fix for script * fix for flow * cleaning * fix DatatableCreationPolicy --------- Co-authored-by: claude[bot] <41898282+claude[bot]@users.noreply.github.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: centdix <farhadg110@gmail.com> |
||
|
|
3e2565f710 | fix flow not sent (#7417) | ||
|
|
61a3c81d5d |
chore(appchat): improve prompt and tools (#7376)
* nit flow * better prompt * remove files from user message * truncated files * nit * f |
||
|
|
d229d469a1 |
chore(appchat): add tests pipeline (#7374)
* draft test app * gitignore * add app test pipeline * add lot of tests * add variant * remove unrelated changes * fix * fix |