mirror of
https://github.com/windmill-labs/windmill.git
synced 2026-09-06 00:02:13 +00:00
dfc06bf821f7846eacf8fafbf2aed978992ea263
1916
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
88d6e11ff0 |
fix(datatables): a permissioned data table cannot be pointed at another database
Its roles live in the database it points at: the logins were created there and every grant they hold is recorded there. Carried onto another database they authenticate against a cluster that never heard of those grants. Opting out first is what drops them from the database they belong to. Also: the ATTACH test now calls the parser the executor runs, rather than a byte-identical copy of its regex that no regression could reach. |
||
|
|
709cf8aecb |
Merge remote-tracking branch 'origin/main' into datatable-perms-3
# Conflicts: # backend/ee-repo-ref.txt |
||
|
|
b72ccc3593 |
fix: key build artifact caches on a runnable's inline modules (#10819)
* fix: key build artifact caches on a runnable's inline modules Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: seal the cache-key base and skip prebundling multi-file bun scripts Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: tighten cache-key invariant comments and name the retained-artifact residual Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: version the build artifact keyspace so pre-fix artifacts are abandoned Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: namespace the artifact cache by keyspace version instead of the hash preimage Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: namespace module-bearing artifacts instead of versioning the whole keyspace Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: pin the cache-name base seal and name the retained-artifact residual Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: bump ee ref for agent-worker module resolution fix Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: align agent-worker module resolution with the worker for previews by hash Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: drop calculate_hash imports left unused by artifact_cache_name Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: update ee-repo-ref to 2d6c66b32f20d9605c6a677727473ab66fcc8a87 This commit updates the EE repository reference after PR #743 was merged in windmill-ee-private. Previous ee-repo-ref: efce983cae3d53175bbb286a10205a2a360c2a9e New ee-repo-ref: 2d6c66b32f20d9605c6a677727473ab66fcc8a87 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
8f349c032a |
fix: nested template literals in step inputs, and unresolvable $args tags (#10856)
* fix(frontend): keep nested template literals intact in template inputs Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: fail a flow step with an unresolvable $args tag instead of hanging Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): surface input expression errors when running a step test Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): treat an escaped \${ as literal text when escaping backticks Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: accept the string "null" as a tag component, reject only JSON null Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: leave a same_worker step's inert tag alone, log an unresolved flow tag Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): escape every backtick when the template walk desynchronizes Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: leave a dedicated runnable's inert step tag alone Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: reroute a step only when its own tag is what failed to resolve Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: name the inert-tag guard step_is_pulled_by_tag Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: reject a tag only when it interpolates to nothing at all Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): validate the template walk instead of trusting a balanced stack Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: describe what an unresolvable tag actually interpolates to Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor(frontend): decide template escaping with a real parser, not a hand-rolled scan Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: name is_flow_step on push now that it is load-bearing Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): heal an expression escaped before nested templates were handled Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: reroute a step whose tag reads args that failed to evaluate Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: never hand a job that failed before running to a dedicated runner Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: reroute only a step whose args failed, leave other tags untouched Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: drop the post-preprocessor tag fallback, leaving tag resolution untouched Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * refactor: leave interpolate_args exactly as it was Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: use a generic example in the template literal tests Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): show an expression escaped by the old rule as it was authored Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): surface input expression errors from every step-run entry point Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: state what is_dedicated_worker actually reads Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): heal only text whose backticks were all escaped by the old rule Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): match the old rule textually so an authored backslash still heals Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(frontend): heal only expressions the old rule broke, never ones that parse Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8b80b09f33 |
fix: restrict filesystem workspace storage to debug builds (#10864)
* fix: restrict filesystem workspace storage to debug builds Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q7p2VbtYqaXHGaAskgwVk5 * chore: update ee-repo-ref to b58ad414b098d3d7787001a352bfbb13e43a335f This commit updates the EE repository reference after PR #747 was merged in windmill-ee-private. Previous ee-repo-ref: 1b4dada77a8fe2224579c643550c63b1ac2616de New ee-repo-ref: b58ad414b098d3d7787001a352bfbb13e43a335f Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
ae162b6d3e | Merge branch 'main' into datatable-perms-3 | ||
|
|
9c557859c5 |
feat: AI agent evals: datasets, scored runs and comparison (#10633)
* feat: eval datasets and standalone runs for reusable AI agents Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: agent eval drawer with case editor, runs and capture entry points Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: document AI agent eval datasets and standalone runs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: say how many eval cases the list is not showing Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review findings on eval datasets - keep an edited case's conversation and tool inputs: serde(flatten) silently drops Box<RawValue> fields, so the update payload is spelled out - remount the case editor per case so one case's turns cannot leak into another - require jobs:read / flow_conversations:read on the capture endpoints, which UserDB does not gate by token scope - take the dataset lock in create and update so a delete cannot be undone by a concurrent metadata write, and delete cases before metadata - load more cases beyond the first page, and stop capping the agent picker - record that the version stamp is taken at enqueue, not at resolution Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-2 review findings on eval datasets - block operators from dataset and case writes - pass the editor's operating workspace through the drawer and the capture request, instead of assuming the navigation workspace - discard superseded case-list responses so switching datasets cannot land the previous dataset's cases - reject a dataset without a case_id (or vice versa) rather than running an inline case under a dangling association - run unsaved edits inline instead of silently running the stored case - surface the API error body on a failed run - fetch dataset metadata concurrently when listing - $bindable() without a default on the optional open prop - correct the permission and enqueue-time-version wording in the docs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: run an untouched saved case by reference again The editor writes back keys the stored case omits, so comparing the raw objects reported every unedited case as edited: the run went inline and lost the dataset/case stamp its history depends on. Compare a normalized form, and pin it with a test. Also scope the history query to the drawer's workspace and drop superseded responses. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: show a dataset's cases as a table, and fix round-4 review findings The case list showed one case at a time with no overview. It is now a table with the case, where it was captured from, and its last run — the last-run column is a single jobs query on the path stamp rather than a request per row. Review fixes in the same file: - keep the edit baseline on the selected case rather than looking it up in the loaded page, so a case beyond page 1 is not treated as unedited and run stale - release the loading state when a superseded case load returns early - reload every loaded page after a write instead of collapsing to page 1 - last remaining 'resolved to' wording in the version tooltip Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: run a dataset as an experiment, with scorers as runnables An experiment runs every case of a dataset against one subject and records the exact case set it executed, so a result set stays reproducible while the dataset keeps changing. Each case runs as its own small flow — the agent, then a step per scorer — so a case keeps the run stamp, history query and trajectory view a single run already has, and scorers need no orchestration of their own. Results are read back per step by node id rather than by walking a nested loop's status. A scorer is any runnable taking (input, output, expected): a script, a flow, or a reusable agent used as a judge. A judge is prompted with the case and the answer as one JSON message; a script or flow receives them as named arguments. Scores accept a bare number, a boolean or {score}. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: results table for an experiment, with scorer columns One row per case: status, the agent's answer, and a column per scorer, with the mean per scorer above the table and a link into each case's run for its trajectory. Averages skip cases a scorer produced no number for — counting a missing score as zero would read as a regression. The drawer's left pane becomes Cases / Results, and Results carries the scorer picker and Run dataset. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: compare an experiment against a baseline Per-scorer deltas on each row and on the mean, and a filter down to the rows that regressed. Rows join by case id, so a case added after the baseline ran has no delta instead of counting as a change. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-5 review findings on experiments - match scorers by label when diffing two experiments; joining by array position subtracted one scorer from another whenever the scorer sets differed - report a row's status from the case job, not the agent step, so a case whose scorer failed no longer reads as a success - delete a dataset's experiments with it: they hold copies of its cases, and a recreated dataset of the same path would have exposed them - select the experiment that Run dataset just started instead of leaving the table on the previous one - expected is scored now, so stop describing it as having no consumer Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-6 review findings on experiments - hold the dataset lock across an experiment launch, so a delete landing between reading the cases and writing the experiment cannot recreate the deleted dataset's inputs - match scorers between experiments on kind and path, not on label: labels default to a path's last segment, so f/a/quality and f/b/quality compared against each other - average mean deltas over the cases both runs scored; comparing each run's own average reported a regression from a case the baseline never ran, with no regressed row to point at - openapi: the row status is the job's, which is also canceled/skipped; runEval takes scorers; the update-case body no longer advertises source, which the handler deliberately ignores - record why the experiment prefix cannot reach a sibling dataset Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-7 review findings on experiments - release the dataset lock for the push loop and retake it for the write, re-checking the dataset still exists: holding it across the whole launch made every capture and case edit on that dataset 409 until the last job queued - assemble experiment results with bounded concurrency; a 100-case, 3-scorer experiment was 400 sequential lookups, each itself several queries - clear the baseline when it becomes the selected experiment, which was comparing a run against itself and reporting zero deltas - take the header mean over the same cases as its delta while comparing, so the two numbers beside each other describe the same set - a canceled or skipped case is no longer the same grey dot as a running one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-8 review findings on experiments - verify the dataset's identity, not just its existence, before recording an experiment: the path can be deleted and recreated during the push loop, and the experiment holds copies of the old dataset's cases - give the recording lock a longer budget than a case edit, since its jobs are already queued and giving up strands them, and say so when it fails - keep score lookups sequential within a case: nesting two bounded streams multiplied into 32 in-flight queries against a 50-connection pool - clear a baseline that no longer belongs to the loaded experiments, so switching datasets does not leave comparison mode on with nothing to compare - keep a scorer's own mean when the baseline never ran it, instead of blanking a column full of numbers - EvalCaseDraft.expected no longer claims nothing scores it Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: do not trust an experiment's job ids, and require write to record one Experiment objects live in workspace object storage, which a script can write directly, and results are read on the unrestricted pool — so a forged experiment naming another flow job returned output the jobs API would have refused. Only jobs this server stamped with that experiment's id are read now. Also from round 9: - recording an experiment requires write on the dataset, not read: it persists into the dataset's namespace and its shared list - clear the results table when the selection changes and surface a failed load, instead of labelling the previous experiment's numbers as the new one's - a storage fault is no longer reported as a deleted dataset - the lock-timeout message at the recording site no longer says to retry, which would run the whole dataset again on top of the jobs already queued - ExperimentRow.status documents canceled and skipped Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bind the experiment trust check to the requested dataset The previous check matched jobs on the experiment id alone, which the stored object supplies — so copying another dataset's experiment JSON under a readable key carried its jobs' output along with it. A job is now only read if it was stamped for this experiment *and* for the dataset the caller's read access was checked against, and an experiment that names a different dataset is not served from this key at all. Also from round 10: - add the .sqlx entry for that query; without it every SQLX_OFFLINE build failed - serve results over GET: as POST the route-scope middleware classified a read as ai_evals:write, locking read-only tokens out of their own results - clear the selected and baseline experiments synchronously when the dataset changes, so the previous dataset's id is not requested under the new one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-11 review findings on experiments and scorers - give scorers the whole case input, not just the message: an answer that came from attachments or a replayed conversation could not be judged on it - accept a judge's boolean and structured {score} answers, including stringified ones, and pin every documented scorer shape with a test - record an experiment for the cases that did launch when a later push fails, instead of leaving those jobs running with nothing to attribute them to - do not capture a preview parent's synthetic runnable_path as a host flow; the saved case could not be rerun - clear the case table before loading a dataset and surface a failed load, so a failure cannot leave the previous dataset's cases under the new name - keep the results table through a refresh of the same experiment - exclude flow-step jobs from the per-case last-run lookup - drop case sets from the experiment list, which is only used to pick a run - report a database failure at the recording lock as itself, not as contention Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-12 review findings on capture and run history - load flow_node.flow for flownode parents: an agent inside a deployed branch or loop captured without its agent, host flow or tool bindings - decide host_flow_path by whether the path resolves to a flow, not by job kind: excluding previews wholesale also dropped the flow editor's step test, whose path is real - page the per-case last-run lookup by created_before until the loaded cases are covered; one page of 200 reported older cases as never run - do not record an experiment when nothing launched - only attach the case input to a job when a scorer will read it - keep the case table through a save; only a different dataset clears it - drop the superseded duplicate comment on the score parser Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop refetching run history on every case write Reading the case list before the first await made the whole job-history query a dependency of it, so every save, delete and Load more refetched up to 1000 job rows and blanked the column. Read untracked instead. - an empty Last run cell now distinguishes never-ran from not-found-within the page bound, which the comment already claimed and the cell did not - reloading a dataset no longer replaces a populated table with a skeleton - keep the score-parser comment that describes every shape it handles Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: keep eval datasets in Postgres instead of object storage Datasets, cases and experiments become rows (`eval_dataset`, `eval_case`, `eval_experiment`, `eval_experiment_case`) rather than objects under a `wmill_eval_datasets/` prefix. What a run produced is still the job's: only case inputs and an experiment's case snapshot are stored. This removes the machinery the object store needed: - The advisory lock and the read-modify-write of a per-dataset JSONL. A case is a row, so there is nothing to serialize. - The launch-time identity check on the dataset. The foreign key makes a concurrent delete fail the transaction instead. - The trust guard on an experiment's job ids, which existed because a script can write workspace object storage directly and could forge an experiment naming somebody else's job. An experiment now chooses every job id and records itself before pushing anything, so a launch that dies partway leaves a recorded case whose job is missing rather than a running job nothing accounts for; cases that never reached the queue are removed again. Row-level security on `eval_dataset` is the authority on who may read or write a dataset, so `extra_perms` grants work and the rule is not mirrored in Rust. Cases and experiments carry a read policy derived from their dataset and no write policy: they are written on the unrestricted pool after the dataset row itself has been asked, with `SELECT ... FOR UPDATE`, whether the caller may write it. Cases are capped at 256 KiB each and 10 000 per dataset, refused rather than truncated. Attachments are S3 references, not inline bytes, so a case that approaches either cap is a mistake rather than a use case. Evals no longer need the `parquet` feature or a configured workspace object storage. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * style: align the eval drawer with the design system - Scorer chips are `Badge`s rather than a hand-rolled bordered span, and the section header is a `Label` with its tooltip, as are the case editor's fields (which also gets the label colour right). - The results table showed status as a coloured bullet, which says nothing to a colour-blind reader. It now carries the same icons the runs table uses, with the status as its accessible name. - Feedback colours move to the `-500` shades the brand guidelines name. - The conversation JSON error uses `TextInput`'s `error` prop for the border and the caption style for the message, as elsewhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: author an expected answer, tags and attachments on a case Every scorer is handed `(input, output, expected)`, but nothing could produce an `expected` except a conversation capture: the case editor had no field for it and a captured run left it empty. So: - The editor gains Expected, Tags and a read-only list of the attachments a captured case carries. Expected is plain text, or JSON when the answer has structure. - Capturing from an AI agent run keeps what that run answered, which is the only moment a reference answer exists for free. The results table also laid itself out by content, so a long answer pushed the scores — the numbers the table exists for — off the edge of the pane. It is fixed-layout now, with the text columns bounded. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: expected is captured from a run and can be authored Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: link a saved agent when inserting an ai agent step "AI Agent" in the step picker was a leaf that always created a blank step, so reusing a saved agent meant inserting a blank one, opening its step input and linking it there. It is a category now, like Flow and AI Sandbox, listing the workspace's `ai_agent` resources next to a blank option, filtered by the picker's own search. A picked agent produces a step that is already linked rather than one linked afterwards: `agent` set, no tools, and only the flow-local `user_message`/`user_attachments` transforms. Seeding the brain keys there would leave transforms a linked step never reads and that `AgentResourceBar` strips on its next link change. Each `on:new` forwarder rebuilds the insert detail field by field instead of spreading it, so a new field is dropped unless the forwarder names it. `agentPath` is typed on both `GraphEventHandlers.insert` and `FlowGraphV2`'s `onInsert` so the next one to forget it fails the check. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: restore the link on cancel and simplify the agent bar Cancel on an agent edit forked the step into a standalone copy, which is the opposite of what the word means and needed a paragraph under the card to explain. It discards the edits and re-links the step now, leaving the agent untouched; diverging from an agent is Unlink's job, on the linked card. This flow's `tool_inputs` survive the round trip as overrides, so Cancel no longer folds them into the tools the way Unlink does. Linking a step to a saved agent happens in the step picker at insert time, so the bar's own resource picker is gone and "Save as agent" is the one action left. Its `+` button was a trap besides: it opened the generic resource form, where an agent would have to be written as raw JSON. The card itself was `surface-secondary`, the sections token, so in dark mode it was darker than the pane and read as a sunken well rather than an elevated card. It uses `surface-tertiary` as the brand table prescribes, its tool chips are `Badge`s, and the editing card no longer overflows the pane and clips its own buttons. The remaining tooltip follows the inline `Label` convention rather than sitting in a flex row whose gap stacked on the trigger's own margin. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: rework the AI agent evals surface into one table Evals become a single pane: a dataset of cases, one column per scorer, one row per case, with the run being looked at chosen from the toolbar. Runs are permanent. Running the whole dataset opens one; running a single case records nothing at all — it is a job, and looking at what it did is not a claim that it belongs in the history. Its result and its scores sit over the row until they are saved as a run, which carries the cases that were not rerun and the scoring jobs themselves, so the number that is saved is the number that was looked at. A scorer is a runnable: a judge agent or a script, created in one click and edited in place. Scores carry a reason and per-assertion checks, shown on hover with a rescore button. What ran is always named. A run records the agent version, or — for a configuration that is not deployed — a hash of it, so a table can say that its numbers describe an agent that no longer exists: those rows dim and the table offers to rerun. An agent's draft can be run directly instead of the deployed value, and once those edits are deployed the runs that made them are recognised as that version. A step with no agent of its own is evaluable too, and saving it as an agent moves its history onto it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: keep an agent's in-progress edits on the agent Editing a linked agent forks it into the step, which is what makes the edits runnable there — but the agent is what is being edited, so that is where the unsaved state belongs. The edit is mirrored into the agent's own resource draft as it is made. It then survives leaving the flow, shows the agent as drafted wherever it appears, and is what evals run when asked to run the draft rather than what is deployed. Deploying or cancelling clears it; opening Edit without changing anything does not mark the agent as drafted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: shape the evals surface around a saved agent Evals hang off an `ai_agent` resource, so the surface is now only ever about one: the `draft` subject kind, the standalone-step subject and the move that carried a step's history onto a newly saved agent are gone. - A run is permanent and numbered per agent. Running a single case is a trial: it answers in the panel and never touches the table. - "Run scorers only" opens a run of its own that reuses the answers of the run you are looking at, so a scorer added later measures what already ran without calling the agent again. - A draft run whose configuration is later deployed is stamped, once, to the version it became, so its label stops reading `v23 + edits` forever. - A scorer can carry a pass threshold, read off the scores already recorded. - The table is the case, its answer and one number per scorer; datasets are created and edited in a drawer; a run that executed an earlier state of the current draft says so above the table, in one line. - Which agent a step is, whether it is being edited, and which version it is on is a strip above the step's tabs, because it is true of every tab. - Capturing a case from a step test or a conversation is dropped, and with it the `memory` override on a linked step that nothing set. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: run past versions of an agent, and number versions per resource The evals home becomes one table of every run of the agent, whichever dataset each is of, with one badge per scorer. A list spanning datasets cannot hold every dataset's scorers to look a name up, so a score carries its name and kind with its number, and thresholds are joined in per run and column. Run now asks what to run: the latest agent, resolved when the run executes as a flow step does, any past version, or the unsaved edits. Pinning is a subject kind of its own, since a linked step resolves the resource live and inlining is the only way to run a version that is no longer current. Scorers move into the edit-dataset drawer. The column header over a run reports and nothing else: a run is permanent, and a control there that changed the columns would edit the past from the one place that must not. Adding one offers four ways rather than two, writing and reusing being different jobs, and both new kinds open with a summary filled in. Versions are numbered per resource. `resource_version.id` is one identity sequence for the whole table, so an agent saved nine times read v4 ... v24, and the gaps counted writes in workspaces the reader cannot see. The id stays how a version is addressed; the new number is what it is called, in the resource history drawer as well as here. It is assigned on write rather than counted on read because trimming past the cap and clearing a history both take the oldest rows, and counting the survivors would renumber a version a run already names. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: read the dataset a remembered selection names Reopening the evals modal restored the last dataset from storage as a bare path, without reading the row it names. Every "is this already the one?" test compared against that selection, so all of them short-circuited and the dataset was never loaded: editing it opened a drawer with no summary, no scorers and no cases. The remembered path is now brought into context the same way any other choice is, and the tests compare against the dataset that is loaded rather than the one that is selected, so a selection can no longer stand for a read that did not happen. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: give dialogs a trail in their header A dialog deep enough to navigate had nowhere to say where you were: the header held a fixed title, and the way back was a control each body placed for itself, somewhere in a toolbar that moves with everything else the toolbar holds. The header is the one part of the surface that does not move, which is where the trail belongs. `Modal` takes an optional `trail` of levels below its title, rendered as a breadcrumb whose ancestors are the way back. Declarative on purpose: callers of this depth already hold the state that says where they are, so the dialog reads it rather than owning a stack they would have to push and pop in step with it. Escape follows the trail. Leaving a level is what someone deep in a dialog means by it, and closing the whole surface throws away the navigating they did to get there; at the root it closes as before. That only works if a dialog can tell it is the surface being addressed, so `Disposable` now answers `isTopmost()` and the dialog asks before acting: it keeps Escape for itself, so nothing else was arbitrating between it and a drawer opened from inside it, and both were acting on one key press. Evals is the first caller: its runs list is the root, a run is a level in it, and the back button that used to sit above the table is gone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: portal dialogs out of wherever they were opened from A dialog rendered in place inherits whatever the calling component happens to sit inside. One `transform`, `filter` or `overflow` anywhere above it makes its `fixed` positioning resolve against that ancestor instead of the viewport, and a surface meant to cover the app is then confined to a box it never asked for: the nav rail paints over it and its own edges are clipped. Drawers have always portalled for this reason. Dialogs only did so when an enclosing pane claimed them, and rendered in place otherwise, so the same screen could show a drawer over everything and a dialog trapped behind the nav. They now portal the same way: to the pane when one claims it, to `body` otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: make the dialog's title the first step of its trail The trail listed levels below the title, so a dialog one level deep read "Evals > All runs > Run 20 · v6": three steps for two places, the first two of them the same place under different names. The title is the root, so it is the root's own segment, and the trail a dialog is given is now the whole path with that segment at its head. Its height stopped moving too. A heading carries a line-height of its own, so a header holding only an h3 stood six pixels shorter than one holding segments as well, and the dialog's whole top edge stepped as you navigated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: sharpen the evals controls around where you are standing Each screen now offers what belongs to it. The list starts runs; a run is a record, so it offers only the one thing that acts on the record itself, which is measuring the answers it already stored. Starting a fresh run from inside one asked which agent and which dataset from the screen least about either, and scoring an existing run was offered from the list, where there is no run to score. Which run and what it is read against are one question asked twice, so they sit together rather than at opposite ends of a row. Choosing what to run is now a toggle over the two states worth naming, the draft and the saved agent, with every earlier version one click further: running an old version is deliberate, and a list made all three look alike. The draft is read when the dialog opens rather than taken from the caller's polled copy, which could be seconds behind an agent edited a moment ago and would leave the option out exactly when it is the reason for opening the dialog. The dataset field carries its path under it and its edit button on hover, as a resource picker does, so the closed field says what the open list said. Edits waiting on an agent are a "draft" here as everywhere else in Windmill, rather than "+ edits". The dialog runs an evaluation rather than "the agent", which is what it was already called everywhere it is recorded. An agent being edited keeps its evals button on a line of its own, clear of the decision to save or discard. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: settle the evals controls on the patterns Windmill already has The version choice uses ToggleButtonMore, as the AI provider picker does: the two states worth naming stay in the group, the rest are behind the overflow menu, and the one you pick joins the group rather than appearing in a second control below it. The deployed one says which version it resolves to. A run offers nothing to start. Scoring an existing run again was the last thing left there, and it was one button explaining a distinction that the run and the dataset already make between them. The warning that a run executed an earlier draft is about the run on screen, so it goes when the run does rather than following you back to the list, and it sits against the table instead of inside a frame of its own. A dataset just created stays open for its scorers and cases: those are what a dataset is, they can only be added to one that exists, and closing on create sent you to find it again to add them. Scorer settings are a cog rather than a word, now that the row holds three actions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: close the gap in the version toggle and say what naming a dataset does The overflow trigger is not a pill, so the room it reserves showed as a gap between it and the button before it; it is pulled in by that much. The dataset field gets its clear button, which is also the slot the edit button is positioned against, so the two now sit where a resource picker puts them. Naming a new dataset said nothing about what happens next, and the drawer looked like it was missing the rest of itself. It says so instead: a scorer and a case both belong to a dataset, so there is nothing to attach either to until this one exists, and creating it leaves the drawer open on them. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: choose a dataset's scorers while naming it A scorer is a reference to a runnable, not a child of the dataset, so it needs the dataset's name but not its row. The list is collected in the drawer while the dataset is being named and sent with the create, which already accepts one, so a dataset arrives holding the columns that were chosen for it rather than being made empty and then edited to hold them. Cases stay where they were: a case *is* a row of the dataset, so there is nothing for it to be a row of until one exists. The drawer says which of the two is which instead of leaving the screen looking like it is missing the rest of itself. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: level the version toggle and name the dataset in its own field The overflow trigger stands a row taller than a toggle button, so the group grew to its height and left the sunken background showing under every pill beside it. Every child of the group is the same height now, which is why the AI provider picker never had the band: it sizes them all alike. The dataset field says the summary with the path after it rather than carrying the path on a line below. The list stacks the two, which a one-line field cannot do, so it says both the other way round. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: tidy the evals forms and the run's own controls Picking a scorer that exists chooses between two sources rather than showing both: the ones already measuring something, and everything else in the workspace. The first list says what each is called with its path under it and what it already measures on the right, instead of three columns that were the same path truncated three ways whenever a scorer had no name of its own. A dataset's drawer says what it is for on the page rather than under an icon, and its summary is sized like the field beneath it. The run's own row lines up with the table under it, the warning above that table is spaced off the rule rather than sitting on it, and adding a case is gone from a run: a run is a record of cases that were answered, so curating them from it is editing what it measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: create a dataset holding the cases written for it Creating a dataset takes the cases to create it with, so one can be assembled in a single act instead of made empty and then filled in. The drawer holds them while the dataset is being named, gives them ids of its own to be edited by, and sends them with the create. Every case is checked before the dataset is written. `eval_case` grants users no write, so the rows cannot be inserted in the transaction that creates the dataset under the caller's own policies; validating first is what keeps "created holding these cases" from becoming "created, holding some of them", and the rows that do follow go in one transaction of their own. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: name the button for what it opens, and say what each version is Starting an evaluation asks which state of the agent and which dataset, and both cost a provider bill, so a button that read as spending one on the way past was lying about the click. It opens something, and says so. Running one case from the panel keeps its own name and its play icon, because that one does run on click. The version options say what they are rather than what they are not: what a flow step would or would not run is a fact about somewhere else, and someone choosing what to evaluate is not standing in it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: give the editing card two rows and mark evals as beta At the width of a step panel the card's one row wrapped: the line naming the agent, the line saying what saving does, and the two buttons deciding the edits' fate all fought for it. Deciding gets a row of its own, and evals sits against the line it is about, since evals of an agent being edited run the edits. Evals is named wherever it is offered. It read as a word in one state of the card and as an icon in the other, which is two things to recognise for one door. The dialog carries a beta badge against its own name, before any level below it: every way in lands there, so it is said once and stays put as you navigate. The version toggle spells out which is which. Both are the agent at v2 and the difference between them is the whole choice, so it is worth the width. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: name a new dataset, and lay the scorer's settings out like a step's inputs A new dataset arrives called "Dataset 1", which the path follows as it follows any summary: a dataset with none was one every table could only call by its path, and the two seeds are what the summary rule already produces. Scorer settings put each field's description between its label and its input, where a step's inputs put theirs, and its inputs are the size the rest of the drawer uses. The runnable behind the column is a link to it with its kind's icon, since it is a resource of its own and the one thing about it these fields cannot change. The line explaining that a pass line re-reads recorded scores went: the threshold is a number to set, and how it is applied is not a decision being made here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: curate a dataset in the drawer and save it in one act The drawer holds the cases while they are edited and writes them when it is saved: added, changed and dropped, whichever it is. Typing no longer writes, so a set is never half saved while someone is still deciding what is in it, and Save means the same thing whether the dataset exists yet or not. A case panel offers reading rather than acting. Running one case now and editing one from a run were the last two ways to change a record from the screen showing it, and the machinery behind the first went with it. The answer is rendered as the prose it is, under what it is: the case's result, whichever run is selected above it. The rest is what the run's table was doing to its own edges: a column name is clipped to its column rather than running into the next, the table squares off against an open panel, and that panel closes with the run it belonged to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: one border above a table, and a link to the run's job The row above the table drew a bottom border and the table draws its own top edge, so every table sat under two lines. The row keeps its spacing and the table keeps its edge. A column header no longer spins while its scores arrive: the cells under it are where the numbers are missing, and they say so themselves. The beta badge is the height of the word beside it rather than of the line it sits on. A run is one flow and therefore one job, so the run says where that job is: what it is doing, what it cost and what it logged are all there rather than reconstructed from the table. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: stream scores as each scorer finishes, and show them per case A scorer runs after the agent inside the case's own iteration, so its verdict can be read as soon as its step is done. Waiting for the iteration to end held every column of a case back until the last of them finished, which is why answers arrived one at a time and scores all at once. Reading a job that is still running needs one guard: a module with nothing in it is a step that has not run, not one that produced nothing, and recording the second makes a failure that never goes away. The panel beside the table shows what each column made of the case and why. The reason a judge gave was stored and never shown, which is the half of a score that says anything. It stops repeating the question the header already asks, and a case still running reads as waiting rather than as an answer that says "Running". A run is a number beside a dataset, so the list puts the two together. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: score a case with every scorer at once The scorers of a case read the answer and never each other, so they ran one after another for no reason: measuring a case now takes as long as its slowest column rather than as long as all of them. Each is a branch of its own, kept from failing the others, so a judge that errors costs its own column and no more. An iteration is three steps again — answer, payload, scores — rather than one per scorer, and each branch is named for the column it produces, so the graph of a run says which scorer did what instead of spelling out an id. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: read a judge's score out of the JSON it nearly wrote A judge quoting the agent inside its own reason writes those quotes unescaped, which is invalid JSON and also the most ordinary sentence for it to produce. The whole verdict was being thrown away over it, so a column that had a number reported having none. The number and the reason are now read straight out of such text. Deliberately not a second JSON parser: it finds the two keys and takes what follows, which is what survives a quote in the middle of a sentence. A case still running says so with a spinner rather than with the word "Running" sitting where its answer goes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: ask a judge for a shape instead of trusting it to write one A new judge carries an output schema, so the provider holds it to `{score, reason}` rather than the prompt asking it to. Windmill already delivers a schema whichever way the model takes it, a tool for Claude and Bedrock and the native parameter elsewhere, so there is no list of models to keep here. An agent with no runs offers its first one where the first row would be, rather than from a toolbar above a table that has nothing in it. Starting a run no longer picks a dataset for you. It fell back to whichever came first, which on an agent that has never run means offering another agent's set as though it were the obvious one; and with no dataset at all it says so and offers the one move there is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: report a column that failed throughout, and hold the run dialog The runs overview dropped any column that produced no number, so a judge that failed on every case of a run vanished from the row and read as a column nobody had asked for. The aggregate now reports every column that has cells, with the count of the ones it failed on, and the badge says "failed" where there is nothing to average. A column with no cells at all is still left out: that one was added after the run and has nothing to say about it. Creating a dataset closes the drawer rather than turning it into an edit of what it just made: scorers and cases already ship with the create, so there is nothing left to stay open for. Reached from the run dialog, it gives the screen back with the new dataset selected, and the dialog keeps the version you had already chosen. Also: - the case panel's job link moves to the panel's own header, where its scope is: the job is the whole iteration, not the answer it sat over - one action in the scorer drawer's header, as its neighbours have. The reuse list picks rather than adds, and says which dataset each column already measures - adding a case is the last row of the list it lands in - the pane shows what it has read rather than an empty state it has not earned yet, and its rows say they open - the linked agent card loses a border it had inside another one * fix: keep the linked agent card's outline The card is a thing inside the step's inputs rather than a section of them, and the outline is what says so. Only the rule inside it goes: the detail it separates is already set apart by being detail. * refactor: fit the eval surface to the shipped design * feat: give a nested dialog a back control and the runs list its own moves Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: put a dialog's description under its title Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: fold a dialog's back control into the crumb it returns to Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a dataset's cases as a table rather than a list beside a form Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a dataset's cases in the grid the data tables are edited in Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: edit a grid cell of prose in place, and cap a dataset at one page Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: keep the cell editor's styles beside it, not in the vendored theme Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: keep an empty cell empty and cap the editor's growth Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: name the step that assembles a run for the scorers Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: run the payload step natively, and say so when nothing serves that tag Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: report an answer as answered while its scorers are still running Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * feat: let a scorer say a case is not one it measures Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: score the answer, and leave a case with no expected answer unmeasured Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: split the evals backend into modules Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: record what a run produced so it outlives its jobs * fix: read only the agent step's own tool jobs into the payload * fix: pin a run's configuration and give the judge the attachments * feat: write a dataset's cases in one transaction * chore: refresh the sqlx cache for the eval queries * fix: drop results a newer selection has superseded * fix: keep a draft the agent editor never opened on * feat: let a run record what it produced instead of waiting to be read * fix: serialize the replacements of a dataset's cases * fix: stop the poller from superseding a read slower than its interval * chore: refresh the sqlx cache * fix: keep a failed read from settling a cell as a case with no answer * fix: hold the case grid while its save is in flight * fix: keep a failed collect step from failing the run it recorded * chore: refresh the sqlx cache * fix: commit an open cell into the save that reads it * refactor: size the eval buttons with unifiedSize * docs: describe a run as the one flow it is * fix: show a run's recorded rows when part of it cannot be collected * refactor: size the remaining PR-added buttons with unifiedSize * fix: save the dataset name that was submitted, not the one typed after * fix: force an open cell into the save that was pressed for it * fix: refuse to score a run whose evidence could not be read * fix: hold one lock over a dataset's case count and its writes * fix: keep one unreadable run from costing the whole runs list * refactor: drop the banned bindable-default from the eval props * fix: hold the scorer controls while the dataset is written * fix: read only the caller's own draft of an agent * docs: say in the contract that a run pins its configuration * fix: say a scorer did not run rather than blaming a missing answer * feat: resume the agent draft you already had when you press Edit * refactor: build the trail and dataset controls from Button * fix: clear the open-cell flag when the drawer reopens * chore: refresh the sqlx cache * fix: read a run's configuration and its version from one snapshot * fix: refuse a dataset path or summary the column cannot hold * refactor: handle the agent draft the way the resource editor does * fix: run only a configuration the launch actually read * docs: bound dataset path and summary where they are submitted * fix: surface a stalled agent draft instead of claiming it is kept * fix: stop claiming a draft holds edits a failed write never sent * fix: word a missing score only once the run says whether the case answered * fix: let a breadcrumb crumb shrink so its truncation applies * docs: describe where an agent's unsaved edits live and what drops them * fix: keep harvesting scores when the run cannot yet word a missing one * fix: report a refused draft write the card was reading as a save * fix: drop the refused draft write when the server copy is taken instead * refactor: build the scorer and dataset pickers from the design system * fix: say what removing a scorer column actually does * fix: drop a refused draft write wherever the server copy is read * fix: let a picker row be as tall as the two lines it holds * docs: record what removing a scorer column does to recorded runs * fix: send a queued draft write before reopening, and drop only what it refuses * refactor: write the agent draft at commit points instead of mirroring keystrokes Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * refactor: run an agent's edits from the step instead of keeping them as a draft Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: make the diff badge keyboard operable and refuse an edits run without its edits Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop the dataset icon from the scorer picker rows Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: size the evals buttons like the rest of windmill and call a run of edits edits Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: count a brain expression as an edit of the linked agent * fix: cap scorers per dataset and report a launched run as launched * fix: harvest scores in one read, refuse duplicate case ids, allow group paths * fix: mint scorer ids server-side, save a dataset edit in one request, check attachments * fix: write a dataset edit and its cases in one transaction * fix: atomic dataset create/edit, reset eval pane per agent, stable pending scorer ids * refactor: govern eval_case writes by RLS so a dataset edit is one transaction * fix: pin launch snapshot, order case locks, cap dataset size, guard stale load * fix: cap dataset bytes on single-case writes, reset run-dialog flag on load failure * feat: migrate eval datasets on username change, settle unspawned cases, drop unused case endpoints * fix: resolve scorer scripts as the caller and pin their hash; migrate scorer paths on rename * fix: bound a failed tool call's error to the payload truncation cap * fix: pin scorer hash as a hex string, reject missing judges, migrate eval authorship * fix: record an out-of-range scorer result as an error, not a score * fix: resolve judges in one caller-scoped read, pin deployed scripts, bound pass_if * fix: settle unspawned cases only when the run completes, and their score cells too * feat: reassign eval datasets and their path references when offboarding a user * fix: use the regex backreference in offboarding eval path rewrites * fix: register eval datasets in offboarding registries, keep resource-version param name * refactor: name the resource-version path param id, since it is the row id not the version * fix: validate dataset paths canonically, clone eval data on fork, surface eval load and launch failures * docs: note MCP tool results are not yet surfaced to eval scorers * fix: show the eval error state on any load failure, not only an empty dataset list * fix: preserve eval case order across a batched save Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * docs: scope the eval launch delete-safety guarantee to the assembly window Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: only offer deployed scripts as eval scorers, drop unbuilt rescore claim Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: enforce 0-1 scorer threshold in the settings drawer and clear stale eval load errors Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: scope subject version/hash reads to the caller and keep a 0 pass threshold Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: select the saved dataset when creating or renaming from the Run dialog Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: gate eval dataset rename on path ownership, not just write access Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * fix: tolerate a malformed agent config when resolving the deployed label Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LVjewUvXFEjNLLw7kxz41h * refactor: trim eval code and comments, fix shared select and modal paths Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: drop the rename warning when editing an eval dataset path Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: add eval dataset delete, keep summary on partial edits, settle resultless scorer cells Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * test: cover parseThreshold and subjectLabel Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix: hold dataset Save during a scorer write, derive draft_hash only from the carried draft Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
5974f5491d |
Merge remote-tracking branch 'origin/main' into datatable-perms-3
# Conflicts: # backend/ee-repo-ref.txt # backend/parsers/windmill-parser/src/asset_parser.rs # backend/windmill-common/src/workspaces.rs |
||
|
|
01fc4f1568 |
fix: name the requested storage when a workspace storage lookup finds nothing (#10803)
* fix: name the requested storage when a workspace storage lookup finds nothing * chore: point ee-repo-ref at the merged ee commit |
||
|
|
4b406e37c0 |
fix: size the ephemeral job token to the job timeout it must serve (#10804)
* fix: size job token to the premium cloud job timeout Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: give the job token setup headroom and drop dead MAX_TIMEOUT_DURATION Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: cap job token setup slack so self-hosted tokens stay at 7d Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
01891cd732 |
fix: keep workflow-as-code scripts off dedicated workers (#10805)
* fix: keep workflow-as-code scripts off dedicated workers A dedicated subprocess calls the script's `main`. A workflow-as-code v2 entrypoint exports none, so a WAC script configured as a dedicated worker failed every run with `entry.module.main is not a function`, and its checkpoint/dispatch round-trip never ran at all. Leave such a script unregistered in the dedicated worker map instead. The worker still holds the script's dedicated tag, so the job falls through to the regular executor on the same worker and runs correctly; rejecting it at push time would strand it, since nothing else pulls that tag. `is_wac_v2` covers only the languages whose executor actually routes a workflow through the WAC runner: Deno runs a WAC-shaped script as a plain `main`, so claiming it is WAC would deny it a path it uses correctly today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VMNiDoBuUY9jzqcuFLT2wU * chore: update ee-repo-ref to ac02c4696ea0be6a8b8ae154ddd7521bc1c3bbc0 This commit updates the EE repository reference after PR #740 was merged in windmill-ee-private. Previous ee-repo-ref: bf742f6ea4d435bd47c9ee0ac5ad800925d79672 New ee-repo-ref: ac02c4696ea0be6a8b8ae154ddd7521bc1c3bbc0 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
5088e13705 |
fix(ci): unbreak the windows test jobs and the discord comment relay (#10799)
* test: assert the unpacked repo symlink without following it `unpack_keeps_a_link_that_stays_in_the_repo` read through the link it had just unpacked. Windows stores a symlink's target verbatim and its object manager rejects the `/` in a POSIX one, so `read_to_string` came back with `ERROR_INVALID_NAME` and the release's `cargo_test_windows` job was red. Pin what the function is responsible for on every platform — the link is kept and materialized — and read through it only where a POSIX relative target resolves. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P3WRtxdKNGdomWX9vaAYGx * test: key the cli sync-map fixtures with the platform separator A sync map is keyed with the platform separator on both sides — `FSFSElement` walks the tree with `path.join`, and the remote `ZipFSElement` starts at `"." + SEP` and joins from there — while an `!inline` reference is always forward-slash. `lock_dedup.ts` follows that convention; the fixtures did not, so on Windows they built a map shape the CLI never produces and 12 of them failed. `getTypeStrFromPath` is the same story: it matches `"dependencies" + SEP`, and the test handed it a forward-slashed path. Build the fixture keys through the separator, leaving the `!inline` references and the `present` map forward-slash, as `sync.ts` hands them over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P3WRtxdKNGdomWX9vaAYGx * ci: skip the discord comment relay when the thread lookup returns none A rate-limited or unauthorized Discord response carries no thread list, and under `bash -e` that aborted the step — jq cannot iterate null, nor parse the HTML error page Cloudflare answers a 429 with — before it reached the "thread not found, skipping" branch right below. Three comment relays failed that way on the 1.794.0 head. Keep the step green for both, but tell them apart: a response with no thread list is a delivery that was dropped for a reason worth seeing, so it warns with the body it got, while a PR that genuinely has no thread stays quiet. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01P3WRtxdKNGdomWX9vaAYGx --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
3c8e4b43fd |
fix: resolve a script path to its new version as soon as the lock lands (#10794)
* fix: resolve a script path to its new version as soon as the lock lands Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011H5ygpzQHkPeYsjiP9GzBy * fix: tell MCP script deploy callers to stop polling on a lock error Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011H5ygpzQHkPeYsjiP9GzBy --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
d85050f505 |
feat: upgrade bun to 1.4.0 and demote deno in the language picker (#10784)
* chore: upgrade bun to 1.4.0 in dockerfiles and CI pins * chore: move deno last in the language picker and relabel it Deno * chore: move deno last in the pipeline language picker too * chore: pin debugger image to bun 1.4.0 and trim the deno picker comment * chore: state the deno picker constraint without referencing the old order * fix: stamp bun lockfiles back to v1 while the fleet predates bun 1.4 * fix: ask bun for a v1 lockfile instead of rewriting one, and refuse an escalated lock * chore: warn instead of silently storing a lockfile with no readable version |
||
|
|
5099f405d4 |
feat: make the Git Repo Viewer work with GitHub App repositories (#10765)
* fix: resolve the head commit of GitHub App repos in the git repo viewer `get_git_commit_hash` ran `git ls-remote` against the raw resource URL. A GitHub-App-backed repository stores a tokenless URL, so the probe failed with "could not read Username" and the viewer never got past its first step. Resolve the head over the GitHub REST API with a server-side installation token instead, reusing the lookup the auto-pull poller already uses for app repos. Non-app repositories keep the ls-remote path. Also picks up the EE-side allowlist fix that lets the clone hub script request an installation token. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 63c67e2a2db198af26a0334f5be14af7d9987eb1 This commit updates the EE repository reference after PR #732 was merged in windmill-ee-private. Previous ee-repo-ref: 2a260961fa0a9bb5631c17e2f718cb8efb4f9aa2 New ee-repo-ref: 63c67e2a2db198af26a0334f5be14af7d9987eb1 Automated by sync-ee-ref workflow. * fix: honour the app-repo head lookup's not-app-backed result `get_app_repo_head_for_autopull` documents `Ok(None)` as "this repo is not app-backed, use the ls-remote path", which is what the other two callers do. Fall through to `ls-remote` on `None` instead of turning it into a 500, and drop the handler's own `is_github_app` read now that the callee's answer is honoured. Also bumps ee-repo-ref to pick up route-safe ref handling in that lookup. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: serve GitHub App repositories as an archive instead of a token The viewer's clone script asked the server for an installation token and put it in the clone URL. That token is installation-wide and carries the App's full permissions, so minting one requires a workspace admin, and the viewer was therefore admin-only for app-backed repositories. The server now streams a tarball of the commit instead, authorized by read access to the git_repository resource, so no GitHub credential reaches the job. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: run delegate_to_git_repo playbooks from GitHub App repositories An Ansible job's runnable_path is the user's own script, which no entry in the git-sync script allowlist can match, so `delegate_to_git_repo` could never obtain a token for an app-backed repo. It also gave up entirely on agent workers, whose connection has no database to mint one from. A playbook run only reads a working tree: the clone is followed by one rev-parse for a log line, and nothing after that touches git. So take the same archive route the viewer uses, extracting the commit's tarball into the job's repository directory. No GitHub credential reaches the worker, and agent workers work because the route is HTTP. Archive entries are joined onto the target by hand so a crafted archive cannot write outside the job directory. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: drop the now-immutable secret_url binding Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * feat: point the repo viewer at the archive-based clone script hub/28905 reads app-backed repositories through the server's archive route instead of minting an installation token, which the backend in this release no longer grants it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: stream repository archives to disk rather than into memory The archive download went through `AuthedClient::get`, whose client caps a request at 20 seconds and whose response was then buffered whole. A repository is arbitrarily large, so that cut off slow downloads and put every job on the worker at risk of running the process out of memory. Add `get_streaming`, the read counterpart to the streaming upload path, and write the response out chunk by chunk. Extraction now creates each entry's parent directory: a tar carries directory entries only by convention, and the traversal guard now has tests, one of which caught the missing parent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: require admin to read an app-backed repository A `git_repository` resource names the repository rather than holding a credential for it, so read access to one authorizes nothing: anyone who can write a resource path can point one at any repository the GitHub App installation reaches, then read their own resource. The head lookup now requires admin for app-backed repos, matching the archive route and the repository picker, which already limits itself to workspaces where the caller is an admin. Repos that aren't app-backed are untouched and stay open to any reader. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: describe the repo viewer's hub script as it stands The file read as a patch waiting to be applied, against a hub version two releases stale. Describe what the published script does, including the archive route app-backed repositories now take. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: run the archive fetch under the job poller, off the job directory Three defects in the delegate path's fetch: The download and extraction ran outside the job poller that the git clone paths go through, so a cancelled or timed-out run kept streaming and extracting an arbitrarily large repository while holding the worker. There is no wall-clock bound on the download itself, by design, which is exactly why it needs the poller. The archive was written to a fixed name inside the job directory, where `create_file_resources` has already laid down the run's own files at paths the playbook chooses. A run naming a file `repo_archive.tar.gz` had it truncated and then deleted. It goes to a per-job temp path now. Link entries were unpacked with their target unchecked. `Entry::unpack` writes the link verbatim, so a link out of the tree plus a later entry descending through it writes wherever it points. Targets now face the same containment check as entry paths. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: keep repo symlinks, refuse only writes that go through them The link check rejected any target containing `..`, which is ordinary in a repository — `docs/x -> ../README.md` resolves inside the tree, and a git checkout keeps it. Rejecting it failed the whole extraction for repositories the clone path handles, and app-backed repos have no clone path to fall back to. Targets are preserved as git preserves them. What would let one escape is a later entry written at or underneath the link, so that is what is refused. Extraction also polls an abort flag now: a `spawn_blocking` task outlives the join handle its caller drops, so a cancelled job left it unpacking in the background. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: refuse hard links in a repository archive Leaving link targets verbatim is right for symlinks — git checks them out that way, and an escape needs a second entry descending through the link, which is refused. A hard link is not like that: unpacking one creates it against a target resolved there and then, so an escaping target is useful on its own. No git tree can express a hard link, so an archive carrying one did not come from a repository. Refuse it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: update ee-repo-ref to 21f79bbbd39ae89665d1a89738630978616aa309 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * chore: update ee-repo-ref to 37695a769b25d16b34107eedc1076793a8b388c8 This commit updates the EE repository reference after PR #737 was merged in windmill-ee-private. Previous ee-repo-ref: 21f79bbbd39ae89665d1a89738630978616aa309 New ee-repo-ref: 37695a769b25d16b34107eedc1076793a8b388c8 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
c2deea13b7 |
fix(security): a WM_TOKEN job token can never be a global superadmin (GHSA-hfh4-cx4h-3fcr) (#10124)
* fix(security): a WM_TOKEN job token can never be a global superadmin (GHSA-hfh4-cx4h-3fcr)
Privilege escalation: an app/flow/schedule/trigger execution policy's `on_behalf_of`
(which a `wm_deployers` member can set) could point at a superadmin email. The
resulting job `WM_TOKEN` then passed the email-based superadmin checks, granting
instance superadmin. `forbid_superadmin_job_token` only guarded ~15 of ~75 routes.
Fix at the token layer: a WM_TOKEN must never satisfy a superadmin gate,
regardless of whose email it runs as (sentinel OR a real superadmin).
- `ApiAuthed` gains a `job_id` field, stamped once in `AuthCache::get_opt_job_authed`
from the resolved token's job_id (correct even on cache hits).
- `require_super_admin(db, email)` -> `require_super_admin(db, &ApiAuthed)`, rejects
`authed.job_id.is_some()`. `require_super_admin_email` kept for the few internal
callers without an ApiAuthed.
- `is_super_admin_authed(db, &ApiAuthed)` for the boolean `is_super_admin_email`
authorization branches on request handlers (workspace deletion, fork drops,
dev-workspace attach/archive, object-storage SSRF exemption, custom dbname, EE GHES
+ connected repositories, ...). Migrate ~75 sites (OSS + EE).
- CUSTOM_INSTANCE_DB reads the *authenticated* job_id, not the caller-supplied
`?job_id` query param. Worker-tag check takes a precomputed job-aware `is_super_admin`
on the request path.
Execution-time on-behalf checks (scheduled/flow worker-tag, Cloud enqueue quota,
is_devops_email) are hardened in a follow-up — see
docs/followup-onbehalf-execution-privilege-hardening.md.
Regression tests: a superadmin-email WM_TOKEN is rejected on `require_super_admin`
routes, on `DELETE /workspaces/delete/{w}` (403, workspace preserved), and on the
CUSTOM_INSTANCE_DB lookup with no `?job_id` (401); real superadmin tokens still succeed.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix: cap devops role at workspace admin and reject reserved on_behalf_of identities
Extends the job-token cap with three pieces:
- `require_devops_role` takes `&ApiAuthed` and rejects job tokens.
`is_devops_email` is true for superadmin emails, so every worker-management,
instance-config and service-log route was reachable by the same superadmin
`WM_TOKEN` that `require_super_admin` already rejects.
- A `job_id` claim that does not parse as a uuid rejects the token rather than
resolving to `None`, which would clear the job provenance and uncap it. Applies
to the internal JWT and the external `jwt_ext_` path.
- Defense in depth at store time: `validate_on_behalf_of` refuses the reserved
internal sentinels as an `on_behalf_of` on apps/flows/scripts/schedules/triggers,
and app execution refuses a policy carrying one — covering already-persisted and
forked-app rows that predate the cap. Deploying on behalf of a real user,
including a real superadmin, stays allowed; the cap handles that at execution.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(mcp): preserve job-token provenance when minting the proxy JWT
The MCP endpoint-tool proxy re-mints a JWT from the caller's ApiAuthed to
forward the proxied request, but passed job_id: None. A job's WM_TOKEN is
capped at workspace admin (GHSA-hfh4-cx4h-3fcr); dropping the job_id here
re-minted an uncapped token that satisfies require_super_admin /
require_devops_role on the proxied route (e.g. listWorkers exposing worker
IPs, job/workspace IDs, and sensitive tags).
Carry api_authed.job_id into create_jwt_token. Adds an in-module regression
that decodes the forwarded JWT and asserts the job_id is preserved for a job
caller and absent for a non-job caller.
Reported by Codex CI review (P1) on #10124.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: cap the admin-or-devops gate at workspace admin for job tokens
require_admin_or_devops (the EE critical-alerts endpoints) grants when the
caller is a workspace admin OR an instance devops. is_devops_email is true
for superadmins, so a WM_TOKEN running on-behalf of a superadmin who is not a
member of the target workspace could clear the devops branch and read/ack that
workspace's critical alerts (GHSA-hfh4-cx4h-3fcr). This gate takes a bare
email, not an ApiAuthed, so the token-layer cap could not see it.
Thread the caller's job-token provenance and reject the devops branch for job
tokens, matching require_devops_role. The workspace-admin branch stays allowed
— that is the cap ceiling. Adds an enterprise-gated regression proving the
bypass is closed and a real superadmin token still clears the gate.
Found while auditing the PR for bare-email gates the choke-point cap misses.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: cap instance-global is_admin gates at workspace admin for job tokens
Three instance-global routes gate on the caller's own `is_admin` claim, which
`ApiAuthed.is_admin` carries into a WM_TOKEN (it is a workspace-admin claim,
true for superadmins too). A job token is capped at workspace admin
(GHSA-hfh4-cx4h-3fcr), so its is_admin claim must not authorize instance
actions on a route with no workspace binding:
- `unarchive_workspace` — unarchive an arbitrary workspace by id
- `prune_concurrency_group` — delete a global concurrency group
- `list_worker_groups` — return unobfuscated `env_vars_static` (may hold secrets)
Add job-token-aware `is_instance_admin` / `require_instance_admin` helpers (the
same shape as `require_super_admin` / `require_devops_role`) and use them at
these three sites. Workspace-scoped `require_admin(authed.is_admin, ...)` gates
are intentionally left unchanged — a workspace-admin job token is within the
cap there. Regression added covering all three; verified it lets a WM_TOKEN
unarchive/leak without the fix and is blocked with it.
Reported by Codex CI review (P1) on #10124.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(mcp): drop orphaned path_field_renames from EndpointTool test helper
The merge with main adopted main's mcp path-substitution refactor (#10162),
which removed the `path_field_renames` field from `EndpointTool` and its
consumer (`substitute_path_params` no longer takes per-field path renames).
main's `runner.rs` `ep` test helper still constructed the struct with
`path_field_renames: None`, so the workspace test build (cargo test --all,
which compiles windmill-mcp's own #[cfg(test)] module under the `server`
feature) failed with E0560. A plain `cargo check` does not compile that test
module, so it only surfaced in CI's cargo_test.
Remove the orphaned field to match the struct.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* test: describe the sentinel-rejection policy the forged-identity test asserts
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: complete ApiAuthed initializers in feature-gated tests after merge
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: stop job tokens minting credentials that shed their provenance
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: cap the MCP OAuth approval mint at the same elevated-job-token gate
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: cap the self-service password reset at the elevated-job-token gate
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: cap app embed/SDK mints and scope widening at the elevated-job-token gate
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: keep job tokens from destroying the account they run on behalf of
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: deny job tokens a foreign-workspace admin claim and workspace ejection
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* docs: keep the follow-up inventory in the PR instead of the repo
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* fix: make the session workspace status gate job-token aware
session_workspace_status derived its superadmin branch from a bare email
check, so a job token carrying a superadmin identity resolved the existence
of workspaces it has no relationship with rather than seeing them as
deleted. Switch to is_super_admin_authed, matching every other instance
gate reached from a request ApiAuthed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* revert: leave the global concurrency-group listing on the plain admin gate
The listing exposes concurrency keys across workspaces, which is metadata
rather than a capability, and it 401s rather than degrading. Keep the guard
on the prune route next to it, which is the destructive one.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix: keep the instance-admin gate on the global concurrency listing
The listing spans every workspace's concurrency keys, and the gate rejects
only job tokens: the !is_admin branch is the pre-existing check, so
workspaced tokens and interactive admins are unaffected.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* chore: update ee-repo-ref to d30af67d38954f9012f7bad08da23e347344b4c6
This commit updates the EE repository reference after PR #664 was merged in windmill-ee-private.
Previous ee-repo-ref: 7870573dbc3360f99bada143f094c67dce0d9e9c
New ee-repo-ref: d30af67d38954f9012f7bad08da23e347344b4c6
Automated by sync-ee-ref workflow.
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: hugocasa <hugo@casademont.ch>
Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com>
|
||
|
|
fa7fbd348d |
fix(security): validate ansible git repository URLs before invoking git (#10759)
The Ansible executor passed the user-controlled git repository `url` (from playbook YAML or a `git_repository` resource) straight into `git clone`, `git ls-remote` and `git remote add` on the worker host. A URL that git parses as an option — e.g. `--upload-pack=<cmd>` — turns `git ls-remote <url> HEAD` into arbitrary command execution on the host, outside any job sandbox. Non-http transports (`ext::`, `file://`, local paths) similarly run programs or read host files. Add `validate_git_repo_url` in windmill-common: reject a leading `-`, reject remote-helper `::` syntax, and allow only the `http(s)`, `ssh`, `git` and scp-like `[user@]host:path` transports. Also reject a `branch`/`commit` that starts with `-`. Validation runs at every ansible entry point that spawns git, covering both the inline-YAML and resource-provided URL paths. CWE-88 (argument injection) / CWE-78. Reported by Nitin Gavhane. Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
ab3c0206d7 |
fix: support @typechecked decorator in Python relative imports (#8495)
WindmillFinder's ModuleSpec lacked origin, so __file__ was never set on loaded modules. inspect.getfile() then raised "is a built-in module", breaking typeguard's @typechecked and anything else that introspects module source. Use spec_from_file_location() which sets origin correctly. Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: hugocasa <hugo@casademont.ch> |
||
|
|
633d7bcb2e |
feat: add trigger_history table with source tracking (#10696)
* feat: add trigger_history table with source tracking Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate trigger history reads on scopes and harden its writers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: filter trigger history scopes in SQL and match the cleared-handler diff Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: record a trigger restore from the trashbin in its history Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: record bulk http trigger creates and document the recording boundary Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: lock the trigger row when capturing its history preimage Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: only record an auto-disable that actually flipped the schedule Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: state the auto-disable invariant once instead of at four call sites Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: render trigger history changes as a structured field diff Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: make a server-initiated disable atomic with its history row Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: note that the auto-disable savepoint takes no pool connection Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: note the flow fallback is the last chance to disable Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: never leave a trigger enabled because its history row failed Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: retry the disable history row instead of dropping it on first failure Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: use the design-system Button for the change-value expander Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: hold the trigger row lock across its disable history row Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the history-loss alert out of the listener cancellation race Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: read the history workspace through the trigger-workspace seam Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
878b8ef4c4 |
perf: cache resolved python interpreter path across worker restarts (#10701)
* perf: cache resolved python interpreter path across worker restarts Every worker process start spawned two `uv python find` subprocesses to re-discover an interpreter path that had not changed, and every python job spawned one more. The resolved paths are now memoized in a small JSON file next to PY_INSTALL_DIR, which outlives the process, so a restarted worker (notably under EXIT_AFTER_N_JOBS) reuses what the previous one resolved. An entry is only served when the uv binary is the same one that produced it and the interpreter is still on disk; otherwise it falls through to a real `uv python find`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review findings on the python path cache - resolve uv through PATH on windows, where `metadata("uv")` looked in the worker's current directory and silently disabled the cache - stat uv with tokio::fs instead of blocking the runtime, and compute the identity once per resolution instead of once per read and twice per write - store one file per version instead of a shared map, so workers resolving different versions concurrently cannot drop each other's entry Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the windows uv PATH probe off the async runtime The lazy static resolving uv through PATH stats candidate entries synchronously, so its first use is moved onto a blocking thread. Also records why an entry keyed on a minor-only version does not pin a patch: uv answers such a request with its minor-version link and re-points it on a patch install, so the memoized path follows the upgrade. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
578d5e9a7d |
perf: back off the interactive worker shell under EXIT_AFTER_N_JOBS (#10700)
* perf: back off the interactive worker shell under EXIT_AFTER_N_JOBS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: address review nits on the shell backoff docs and periodic warning Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: only give the worker shell its sub-second cadence during a live session Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
22eadab67d |
perf: resolve the worker external IP in the background (#10697)
* perf: resolve the worker external IP in the background `run_workers` awaited `external_ip::get_ip()` — an HTTPS GET to hub.windmill.dev — before spawning any worker, so every worker process paid that round trip before its first job pull. Measured on a CE debug build it was 120-450 ms of a ~200-500 ms startup, and behind a firewall the call does not fail fast: it burns its whole 5 s connect timeout, on every process start. That cost is per-job under EXIT_AFTER_N_JOBS. The value is informational (it is only written to `worker_ping.ip`, which the workers list displays so users can whitelist the address), so nothing needs to wait on it. It now resolves into a process-wide cache off the startup path, and `WORKER_EXTERNAL_IP` supplies it explicitly for deployments that know their egress address or have no egress at all. Until it resolves the ping carries no IP, which `insert_ping_query` now COALESCEs so a reclaimed row keeps the address the previous process wrote instead of being blanked. The main loop reports the IP as soon as it lands rather than on the next periodic tick, so a short-lived process still records it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep unknown worker IPs out of the whitelist alert Review follow-ups: - `WhitelistIp` filtered only the `'unretrievable IP'` sentinel, so the `'NO IP'` one a pending or failed lookup now leaves in the row would be offered as an address to whitelist. It filters both. - Register `WORKER_EXTERNAL_IP` in `ENV_SETTINGS` so operators can confirm from the instance settings view that it took effect. - The worker tracked whether it had reported the IP by re-reading the cache after each ping rather than remembering what the ping carried, so a lookup landing mid-ping marked it reported without it reaching the row. The value is read once and threaded through `insert_ping` / `update_worker_ping_full`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: report a sentinel IP once the lookup has definitively failed Keeping the previous process's address on a reclaimed `worker_ping` row is right while the lookup is still in flight, but not once it has failed: the row would advertise an address nothing has confirmed, and the whitelist alert would offer it. A failed lookup now reports `UNKNOWN_IP`, leaving NULL to mean "in flight". `WORKER_EXTERNAL_IP` is rejected when longer than the `varchar(50)` column rather than panicking the worker on its initial ping, which is a hard failure. Adds the regression guard for the `ON CONFLICT` semantics: reverting to `ip = EXCLUDED.ip` would compile and blank every reclaimed row. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the agent initial ping acceptable to older servers An agent worker routinely runs against a server of a different version, and one predating the background lookup rejects an initial ping carrying no IP — which `run_worker` turns into a panic, so a newly upgraded agent would crash-loop against it. The not-resolved-yet case goes over the wire as the sentinel instead, and the server maps it back so a reclaimed row still keeps its address while resolution is pending. Also documents `ip` as the one conditional exception to `insert_ping_query`'s "only `started_at` and `jobs_executed` survive a restart", and adds `WORKER_EXTERNAL_IP` to the README env-var table. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: deliver the resolved IP to servers that only take it at registration A server predating the background lookup applies `ip` from the initial ping only, and ignores it on the periodic ones. An agent registering before its lookup resolves would therefore keep the sentinel forever on such a server, where it used to report its real address. It registers a second time once the address is known, skipping that when the address is still unknown, when the server is reached over SQL and needs no second registration, or once a job has run, since registering clears the row's current job. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: re-register the resolved IP even after a job has run Gating the second registration on "this process has not run a job yet" meant an agent that pulled queued work before its lookup resolved never delivered the address to a server that only takes one at registration. No job of the worker is in flight where that runs, so the gate bought nothing beyond the last job's id, which the next job refills. Documents the two cases where WORKER_EXTERNAL_IP stops being an optimisation and becomes the only way to report an address: an agent against such a server, and a process shorter-lived than the lookup. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * revert: drop the WORKER_EXTERNAL_IP escape hatch Supplying the address by hand skips the hub lookup, which is not something to make easy. Resolving it in the background is what keeps it off the startup path; opting out of it is a separate decision this does not need to take. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: distinguish an IP never established from one that could not be retrieved `NO IP` was doing double duty: the column default for a row whose lookup has not resolved, and the marker for one that failed. An operator reading the workers list could not tell "not resolved yet" from "this instance cannot reach the hub", and the latter is the actionable one. A failed lookup now reports `unretrievable IP`, which is also what it reported before the lookup moved off the startup path. That leaves `NO IP` meaning only "no address established", which is what an agent sends while its lookup is in flight and what the server maps back to "unresolved" — so the wire sentinel no longer collides with the failure marker, and an agent delivers the failure to a server that only reads an IP at registration. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
93b811fd8d |
fix: git sync missed metadata-only deploys, deploy check missed job link (#10662)
* fix: git sync missed metadata-only deploys, deploy check missed job link Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: skip the deploy hook when the mute toggle matched no row Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: pin ee ref forward of main so the bump only adds this change Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to a65162b22b127b54c0686095ee1b16b04e3111f7 This commit updates the EE repository reference after PR #724 was merged in windmill-ee-private. Previous ee-repo-ref: ac5f646c3ace7e5841200c6b83b34fb4371340d9 New ee-repo-ref: a65162b22b127b54c0686095ee1b16b04e3111f7 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
71b9989daa |
feat: auto-build binaries to object storage on deployment (#10673)
* feat: auto-build binaries to object storage on deployment Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: queue the auto-build from pre-locked deploys and off the lock slot Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: materialize companion modules before a deploy-time build Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep a build job from stamping lock_error_logs on a healthy script Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: de-flake test_flow_lock_all and surface the lock error it hides Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: trim drafting history from the flow-lock fixture comments Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: stop a binary build from restarting dedicated workers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the build-job marker off the agent wire and out of user args Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2fcce4526a |
feat: add EXIT_AFTER_N_JOBS worker mode for environment cleanup (#10671)
* feat: add EXIT_AFTER_N_JOBS worker mode for environment cleanup Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review findings on the EXIT_AFTER_N_JOBS worker mode Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-2 review findings on EXIT_AFTER_N_JOBS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address round-3 review findings on EXIT_AFTER_N_JOBS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bound WORKER_SUFFIX length and document the same-worker drain Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: validate the assembled worker name length Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
4cb51cf7bc |
feat: add memory limits to the go build subprocess (#10666)
* feat: bound go compilation memory with GOMEMLIMIT Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bound the whole go build tree, not each toolchain process Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the go build memlimit and parallelism atomic Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: log the go limits actually installed and stop serializing small workers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: make go build parallelism authoritative over persisted GOFLAGS Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: canonicalize the go build -p value and floor the module-step budget Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: parse GOMAXPROCS for -p the way the go runtime does Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: read GOMAXPROCS with go's own grammar and report limits neutrally Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: derive go build parallelism from the cgroup quota over its own period Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep go's minimum build parallelism under sub-CPU quotas Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the windows 1CU cap out of go's two-compiler floor Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: record that a worker runs one job at a time Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: scope the one-job-at-a-time rule away from native workers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
dad4c10c8b |
fix: stream ansible playbook logs in real time (#10669)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
84f3b0094d |
fix: harden custom env var name handling in the nativets/bun prologue (#10634)
* fix: escape and validate custom env var names in the nativets prologue
Custom workspace environment variable names were spliced verbatim into the
generated NativeTS/Bun JS prologue (both the `const {name}` binding and the
`process.env['{name}']` assignment), while only the value was escaped. A
non-identifier name could therefore alter the generated program.
- Add `escape_js_single_quoted` / `is_valid_js_identifier` helpers.
- worker.rs and bun_executor.rs: escape the name as a string literal, and only
emit the `const {name}` binding for valid identifiers.
- set_environment_variable: reject non-identifier names on write (deletion stays
unrestricted so existing rows remain removable).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: address review — reserved-word const gate, grandfathered-name editability
- Gate the `const {name}` prologue binding on `can_bind_as_prologue_const`, which
additionally excludes JS reserved words and the prologue's own bindings
(`process`, `BASE_URL`, `BASE_INTERNAL_URL`); such names would otherwise emit a
SyntaxError that breaks every NativeTS run. They are still exposed via
`process.env['{name}']`.
- set_environment_variable: only enforce the identifier check for names that don't
already exist, so editing the value of a pre-existing non-identifier name (the
edit UI resubmits the name) isn't rejected with no in-product fix.
- Document the name constraint on the endpoint in openapi.yaml.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: exclude eval/arguments from const gate; skip existence query on valid names
- Strict-mode ES modules forbid `eval` and `arguments` as binding names, so add
them to the non-bindable set — otherwise an env var named `eval`/`arguments`
emits `const eval = ...`, a SyntaxError that breaks every NativeTS run.
- set_environment_variable: run the existence check only when the name isn't a
valid identifier, so the common (valid-name) path skips the extra query; trim
the rationale comment.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix: allow `async` as a prologue const binding; note reserved-bindings coupling
`async` is a contextual keyword, not a reserved word — `const async = ...` is
valid, so it needn't be excluded from the const binding. Also cross-reference the
prologue head from PROLOGUE_RESERVED_BINDINGS so the two stay in sync.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
|
||
|
|
201d7c4eb2 |
fix: bound postgres result collection so an oversized result cannot OOM the worker (#10644)
* fix: bound postgres result collection so it cannot OOM the worker Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: render the sql result limit exactly so the error can be set verbatim Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: point the fraction rationale at the renderer that still emits them Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * perf: stop re-parsing every collected row to rebuild it as a RawValue Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * style: drop a dangling doc line and an unrelated rustfmt reflow Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
00822a7435 |
fix: bound duckdb result collection so an oversized result cannot OOM the worker (#10641)
* fix: bound duckdb result collection so an oversized result cannot OOM the worker Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep the duckdb cap a worker-survival limit rather than a cloud product one Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: refuse an oversized blob before it expands to one json value per byte Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: share one expansion budget across a row's values, nested ones included Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bound the row's own serialization so escaping cannot outgrow the budget Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: charge a json column before parsing it into a value tree Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: trim the json budget rationale and name what the budget does not cover Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: gate the sql result size limit on the duckdb feature Its only consumer is the duckdb executor, so the minimal build compiled it as dead code and failed under -D warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a6157ba104 |
feat: bound how much disk a single duckdb job can spill (#10645)
* feat: bound how much disk a single duckdb job can spill * fix: name the env var and correct duckdb's unreachable spill-cap advice * fix: do not blame an unset env var for duckdb's default spill cap * style: keep the duckdb spill-cap invariant comments within four lines * docs: size the duckdb spill cap against the disk cloud pods actually use |
||
|
|
18ae0bdfbf |
feat: keep duckdb spilling behind the local-filesystem fence (#10607)
* fix: explain duckdb failures caused by job isolation Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: apply the isolation policy to the schema-sync pre-pass Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: bump ee ref for the out-of-memory hint wording Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: bump the bundled DuckDB engine to 1.5.5 The 1.5.5 duckdb crate no longer hands back a 96-bit `rust_decimal`, so a DECIMAL wider than that renders instead of panicking inside an `extern "C"` frame — which, being unable to unwind, aborted the whole worker process and left the job running as a zombie. `SELECT '1234567890123456789012345678.9012345678'::DECIMAL(38, 10)` was enough. Adapting to the crate's API: `Value` is now `#[non_exhaustive]` and gained `UHugeInt` and `Geometry`, and `rust_decimal` became an optional feature that the `decimal`/`numeric` argument path still needs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review findings on the duckdb bump Run the FFI crate's own tests in CI: it is excluded from the workspace, so the `cargo test --all` in backend-test never reached them and the new guard against the worker-aborting DECIMAL would not have run. build_dev.sh now honors a caller-pinned CARGO_TARGET_DIR so the test build reuses that compile instead of building the bundled engine a second time. Also pin UHUGEINT rendering, and correct the rust_decimal rationale — `Decimal::new` is public without the feature, so the reason is that the feature reproduces the exact binding the crate used to derive, not that nothing else can. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: address review nits on the duckdb bump Name the unsupported DuckDB type rather than dumping the value, which may be arbitrarily large or hold data that does not belong in an error message, and say which column it came from. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: pin the ee ref to the narrowed duckdb extension allowlist Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: keep duckdb spilling behind the local-filesystem fence * chore: repin the duckdb fork after adding the reset-test exclusion * docs: stop claiming the duckdb patch has been filed upstream * docs: point the backend duckdb bullet at the fork's rationale * fix: place lock_temp_directory so no existing struct member moves * fix: skip the extension-load guard when the repo is unreachable * refactor: trim the fork comments and fail the extension guard in CI * chore: repin the duckdb fork onto upstream duckdb-rs main * fix: keep the engine patch applying on a CRLF checkout * docs: link the upstream issue tracking the underlying problem * chore: repin the duckdb fork onto the patch as filed upstream Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to 88568d11162ffa11723e7955e613224bab4f0568 This commit updates the EE repository reference after PR #720 was merged in windmill-ee-private. Previous ee-repo-ref: 22f075c1164d9dd5a3ba92d682905aabd071d273 New ee-repo-ref: 88568d11162ffa11723e7955e613224bab4f0568 Automated by sync-ee-ref workflow. * chore: repin the duckdb fork onto the cmake/fmt build fix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test: pin the immutability half of lock_temp_directory The spill test proves the exemption works; nothing proved the lock that makes it sound. A rebase could drop the refusals and leave every other tripwire green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
8c6211c277 |
feat: offer more dev workspace environment labels (#10570)
* feat: allow custom dev workspace environment labels Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reject dev labels that shadow a tracked branch's namespace Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: guard dev labels against a repo's assumed default branch Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: state the badge-cap rationale once and drop unenforceable openapi constraints Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: offer a fixed list of environment labels instead of free text Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: match the accepted label set to the openapi enum exactly Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: stop describing the label set as dev/staging only Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c61404a0f4 |
feat: preview merge result in git-sync PR diff check (#10542)
* feat: preview PR merge result in git-sync diff check * chore: update ee-repo-ref * fix: match pr diff sentinels as structured field, tighten comments * chore: update ee-repo-ref * fix: neutral verdict for unfetchable pr head, testable sentinel parse * chore: bump git-sync pull script pin to hub/28889 * chore: bump git-sync pull script pin to hub/28890 * fix: cover failed history deepening in unavailable-head check text * chore: update ee-repo-ref to 181fa0c206d7f84a289b4396a7f7764bc815d284 This commit updates the EE repository reference after PR #712 was merged in windmill-ee-private. Previous ee-repo-ref: 36f5c0e9d147f9eed63ebc316f2aef9f86b500af New ee-repo-ref: 181fa0c206d7f84a289b4396a7f7764bc815d284 Automated by sync-ee-ref workflow. --------- Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> Co-authored-by: Ruben Fiszel <ruben@windmill.dev> |
||
|
|
956210ea06 |
feat: improve duckdb isolation (#10565)
* [ee] feat: improve duckdb isolation Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * chore: update ee-repo-ref to c4e6cbc1a1efeca5b71920c7903db2347d0eeda0 This commit updates the EE repository reference after PR #713 was merged in windmill-ee-private. Previous ee-repo-ref: f630f7e73cb863e312430738d81d802a3971f7cd New ee-repo-ref: c4e6cbc1a1efeca5b71920c7903db2347d0eeda0 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
c59b60c729 |
fix: keep the same_worker pin when a suspend ends without approval (#10552)
* fix: keep the same_worker pin when a suspend ends without approval A disapproved or timed-out approval gate hands the flow back through the UpdateFlow channel with unrecoverable = true. That flag means "the previous step's worker died", and it is read by six sites. Five of them happen to want what it does here, but continue_on_same_worker and continue_with_runners do not: the worker that ran the approval step is alive, so unpinning the error handler and routing it by tag breaks the ./shared contract of a same_worker flow and can land it on a worker group that cannot run it — the same defect #10551 fixed for the three producers that hand back a live flow. Replace the boolean with StepFailureKind so the suspend producer can say "worker alive, but this failure is not the module's to handle" instead of overstating a worker death. The failed module's error policy is deliberately still bypassed: the failure is recorded against the step the gate was holding back, which never ran, so its retry would re-open the gate and its continue_on_error would skip it outright (verified: the gated step is marked Failure with a nil job id and the flow jumps past it). suspend. continue_on_disapprove_timeout remains the way to continue past a gate. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(flow-editor): flag that continue on error does not cover the approval gate A resolved approval is recorded against the step the gate holds back, not the step carrying the suspend, so continue_on_error never sees it: the flow still stops on a disapproval or timeout. Point users at suspend.continue_on_disapprove_timeout, which is what actually continues past a gate, whenever both settings are on and that one is not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
154f8f461e |
feat(debugger): install debug session deps from the instance registry settings (#10550)
* feat(debugger): install debug session deps from the instance registry settings Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(debugger): keep install-time registry credentials out of the session-visible tree Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: drop em dashes from the debugger registry docs and comments Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(debugger): stop installing for a session that went away during the settings fetch Also serves nativets sessions the npm settings their installer reads. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
1aee22296e |
fix: keep the same_worker pin across a flow module that spawns no job (#10551)
When a module completes without spawning a job — an empty branch, an empty for-loop, or a module already marked Success — the flow hands itself back through the UpdateFlow channel, and the result processor resumed it with unrecoverable = true regardless of what sent it. That flag means "the previous step's worker died", which holds for none of the three producers except a suspend that ended without approval. The stale argument was inert until continue_on_same_worker and continue_with_runners started reading it, since when the step after such a module is pushed as an ordinary queued job. It is then routed by tag and can land on any worker in the pool, breaking both the ./shared directory contract and the guarantee that a same_worker flow stays on a worker able to run it — a step whose tag resolves to a worker group that cannot execute its language fails instantly, taking the flow with it. Carry the flag on the UpdateFlow message so each producer states its own case, rather than having the shared receiver assume the worst. The three that hand back a live flow forward whatever their caller reported, so a genuinely unrecoverable failure still crosses the hop unchanged. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
74c418570b |
fix: forward TLS trust roots to debug sessions and honor INIT_SCRIPT on windmill_extra (#10532)
* fix: forward proxy and TLS settings to debugger subprocesses Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reach uv and the bun debugger with the forwarded network settings Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: map every CA variable spelling onto the one uv reads Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: keep package-index credentials out of debugged user code Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: install debugger dependencies outside the interpreter running user code Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: sandbox and bound the debugger dependency installer Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: correct the installer timeout rationale Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * docs: scope the uv --cert note to the commands prepare-deps runs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: build the debug venv against the interpreter that runs the script Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: do not start the debuggee for a session that already went away Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: remove the debug script when the session is gone before it starts Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
7d153d5750 |
fix(debugger): pass python index settings to prepare-deps and report failures (#10533)
* fix: honor python index settings in prepare-deps and report install failures * fix: forward python registry env to the debugger's prepare-deps * fix: scope registry credentials to the prepare-deps subprocess * fix: install python debug dependencies from the service, not the session * fix: bound the debugger dependency install and keep the proxy bypass default * docs: name the nsjail config that isolates debug sessions |
||
|
|
340d3cd565 |
feat(dbt): reach any dbt adapter through a dbt_profile resource, and constrain the warehouse picker (#10525)
* feat(dbt): reach any dbt adapter through a dbt_profile resource, and constrain the warehouse picker
The workspace dbt warehouse picker listed every resource in the workspace, so a
slack or github resource was an offerable answer to a field that can only be a
warehouse. Constraining it exposed that the set of resource types that actually
work is both smaller than the docs claim and too small to be useful:
- `render_profile` translates only six adapters from a Windmill resource; the
rest (clickhouse, duckdb, salesforce, mssql, oracle) refused one outright.
- `redshift` and `duckdb` name no resource type anywhere, so two of the
adapters the quickstart advertises were unreachable.
- the `databricks` resource carries `workspace_url`, while the renderer demanded
`host`, so that warehouse could never render at all.
So the picker gets a constraint and dbt gets an escape hatch wide enough to make
it honest. `dbt_profile` is a resource whose value IS a `profiles.yml` target —
`{ type, target }` — passed to dbt unchanged, so any adapter and any key it
documents works.
`DbtAdapter` is now open: it carries dbt's own `type:` spelling plus an optional
`KnownAdapter` (the eleven Windmill has facts about — a field mapping, a pip
package, the license gate). Anything else is carried by name and installed as
`dbt-<name>`, the convention every adapter on PyPI follows, so "whatever dbt
supports" no longer means "whatever this enum lists". The license gate is
unaffected: `sqlserver`/`oracle` still resolve to their `KnownAdapter` and are
still gated. The name is confined to `[a-z0-9_-]` starting alphanumeric because
it reaches a pip requirement and a venv path on the host.
Two adjacent fixes fall out: the project's own `profiles.yml` and the
descriptor's `profile.type` now accept any adapter instead of the closed list,
and a databricks resource renders its `host` from `workspace_url`.
The picker is constrained to `dbt_profile` plus the translated types, so nothing
it offers can fail for want of a mapping.
Fixes WIN-2320
* fix: drop the unused DbtAdapter::from_resource_type wrapper
Nothing calls it: a Windmill resource type maps through
KnownAdapter::from_resource_type, and the executor resolves an adapter from
the resource's own dbt spelling or by inference. CI builds with -D warnings,
so the dead wrapper failed every backend check.
* fix(dbt): make dbt_profile the block itself, and address the review findings
**A `dbt_profile`'s value IS a `profiles.yml` output block**, `type` included.
It was `{ type, output }`, which asked the user to restructure their block
before pasting it — a translation step, in the one type that exists to avoid
translation. The schema now declares no properties, so the resource form renders
a single JSON editor over the value.
That means the value's shape can no longer say what it is: a `dbt_profile` and
Windmill's bigquery resource are both objects with a `type` (the latter says
`type: service_account`). So the warehouse carries its resource's type
(`DbtWarehouseConnection.resource_type`), and detection is exact. It also makes
decision 9's "the resource type name is the authority" true at runtime for the
translated path, which until now resolved its adapter by sniffing fields.
Review findings, all three reviewers:
- **[P0] an author-chosen adapter became an unsandboxed PyPI install.** `dbt-` is
not a reserved prefix, and `provision_core_1x` installs through `run_tool`,
outside the nsjail ordinary dependency installation uses — so `dbt-<name>` from
a script author's `type` could run a PEP 517 build backend as the worker. Now
gated on a list of published adapters plus `DBT_EXTRA_ADAPTERS`, so trust stays
the admin's call. The open set survives: the engines that ship their adapters
install nothing and take any type.
- **[P1] `type: fabric` rendered as `sqlserver`.** dbt's `type:` was resolved
through the resource-type table, where `fabric` is a Windmill alias for SQL
Server — so a Fabric profile installed dbt-sqlserver, was enterprise-gated, and
failed on an ODBC driver without ever naming Fabric. dbt types now have their
own table.
- **[P1] two spellings of one adapter compared unequal.** `PartialEq` covers the
carried name, so `postgres` != `postgresql` even resolving to one adapter, and
the descriptor/resource check rejected valid configs with a message naming the
same adapter twice. The name is normalised to the adapter's dbt spelling.
- **[P2] identity keys.** `database_key` is what a Windmill resource spells it,
and only translated adapters have one; the rest read dbt's `database`.
- **[P2] duplicate `sslrootcert`** when a block carried both a PEM and a path.
Verified with three real dbt builds: a flat `dbt_profile` postgres block, the
same with `type: postgresql` under a `profile.type: postgres` descriptor (the
alias case, which failed before), and trino for the unknown-adapter path.
* docs(dbt): say that installing an adapter is gated, not just using one
The open-adapter text promised every future adapter is installed as dbt-<name>,
which ensure_adapter_installable refuses outside PUBLISHED_ADAPTERS and
DBT_EXTRA_ADAPTERS. Separates the two: rendering, licensing and identity are open
to any adapter, and only the dbt-core 1.x PyPI install is gated, because that is
the step that runs outside the sandbox.
* fix(dbt): keep a dbt_profile's own sslrootcert when Windmill writes none
The previous round skipped the block's sslrootcert unconditionally to avoid
emitting the key twice, which drops a path-only CA reference — a certificate
baked into the image or mounted on the worker, which is the block's own trust
source. Skipped now only when a root_certificate_pem is present, which is when
Windmill writes a replacement.
* fix(frontend): let a resource type declare no properties
A schema without `properties` is a JSON-edited resource type, not a broken one -
`dbt_profile` is a profiles.yml block whose keys belong to its adapter, so there
is nothing for Windmill to declare. Both editors assumed properties exist:
- ResourceEditor threw on Object.keys(undefined) while deriving the field order,
which left the drawer on its loading skeleton forever, so the resource could
not be viewed or edited at all.
- ApiConnectForm caught the same throw and reported the type as missing from the
workspace, offering to sync a type it already had.
Both now fall back to the raw JSON editor, which is what usesRawEditor already
intended for a schema with no properties.
* chore: cut the new comments to AGENTS.md's four-line cap
Each still states its constraint once; the long-form rationale belongs in
docs/dbt-runtime.md and the PR, not beside the code.
* fix(dbt): keep a dbt_profile's empty and nested collections intact
A block with no children reads back as null, so `extensions: []` reached the
adapter as a missing value rather than the empty list dbt was handed, and a
nested array went through the scalar path and arrived as a quoted JSON string.
Both are keys dbt passes to the adapter as it finds them, so the type has to
survive: empty collections are emitted inline, and the value half of an entry
recurses instead of bottoming out at a scalar.
The test parses the rendered YAML back rather than string-matching it, since
what matters is what a YAML reader sees.
Also cuts DbtWarehouseConnection.resource_type's comment to the four-line cap.
|
||
|
|
58df573b15 | feat(datatables): pick the role the database manager connects as | ||
|
|
c054018c9b |
require an explicit user for azure workload identity on postgres (#10521)
* fix: name the pg login in the job log for token auth modes Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * refactor: route remaining pg login defaults through login_name Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: require an explicit user for azure workload identity on postgres Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
23cd7737b6 | feat(datatables): opt-in role-based permissions | ||
|
|
9f3d15583a |
fix: log the db auth mode used and hint at the ms_entraid sentinel (#10508)
* fix: surface which auth mode a sql connection used and hint at ms_entraid Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: scope the ms_entraid hint to azure hosts and pin the sentinel trim Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8364dd96ef |
feat: send prompt_cache_key on the openai responses api (#10507)
* feat: send prompt_cache_key on the openai responses api * fix: bound prompt_cache_key to the provider limit and scope it to retryable paths * fix: keep a digest suffix when bounding long frontend cache keys * docs: attach the cache-key doc block to the function it describes |
||
|
|
56d53d3f32 |
fix: locate coursier artifacts by coordinate so private maven registries work (#10501)
* fix: locate coursier artifacts by coordinate so private maven registries work Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: match maven coordinates by path component, longest first Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: skip the empty directory a 404 leaves at a maven coordinate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: ignore coursier's dot-prefixed bookkeeping when claiming a coordinate Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
f365929eaa |
feat(fork): merge a fork deletion on evidence, not on the counters (#10484)
* feat(fork): merge a fork deletion on evidence, not on the counters `workspace_diff.ahead`/`.behind` record that a write happened on a side, not what it was or who made it. That leaves one row shape undecidable: an item the parent has and the fork does not can mean the parent added it, the fork deleted it, or a git-sync pull reverted a deploy that had just brought it in. #10467 kept every such row out of the merge direction, which killed the phantom but also dropped the only way to propagate a fork-side deletion and left a rename's old path behind in the parent. Record the evidence instead: - `workspace_diff` gains, per side, the last event's kind (`write` / `delete` / `rename_from`) and origin (`authored` / `sync`). Rows written before the migration have neither and keep #10467's behavior. - The kind is probed from whether the path still holds an item once the write has committed; an item kind the probe doesn't map records no evidence rather than a deletion. Create and update are not split — nothing at that point tells them apart for every kind, and the comparison already recomputes existence per side. - The origin comes from an `X-Windmill-Deploy-Origin` header the API scopes into a task-local for the request. It is the load-bearing half: recording `delete` alone would read a git-sync revert as a fork deletion and reproduce the original bug. Two clients set it — `wmill sync push` (which the git-sync auto-pull runs inside a job) and the compare page's parent→fork "Update fork". Merging the other way stays authored so a deletion keeps propagating up a fork chain. - The merge direction admits a parent-only row only when the fork's last event was an authored delete or rename-away. Such a row stays opt-in, never bulk-selected, and reads "Removes in <parent>"; the update direction keeps offering it back as "New". A fork deletion and a rename now merge into the parent, a rename leaves no duplicate behind, and a fork the parent also edited surfaces in both directions instead of the parent silently winning. Fixes WIN-2289 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): address review — detached tallies, enum wire values, doc duplication Codex P1: a dependency job tallies its deploy whenever it happens to finish, and the event kind is probed from the state at that moment. If anything removed the path in between (a git-sync revert), the stale tally read that deletion as its own and filed it as authored — handing the merge exactly the removal this is meant to withhold. `tally_deployed_object_changes` now takes `Option<DeployOrigin>`; `None` bumps the counter and leaves the evidence columns as the last vouching tally left them, and the worker path passes it. Covered by extending the removal-origin test: a detached tally after the sync archive must not disturb `(delete, sync)`. Also from review: - `fork_removed_it` compares through `DeployOrigin::as_str()` / `DeployEventKind::as_str()` rather than repeating their wire values, so a renamed variant can't silently make the predicate always false. - `deploy_origin`'s module doc no longer claims `sync` is inert: it cannot make the merge propose a removal, but it does drop a row out of both sides of the `all_ahead_items_visible` comparison. - `WorkspaceDiffRow` says why only the fork half of the evidence is consumed. - The delete-vs-revert rationale is stated once (the migration) instead of restated in eight files. - `PATH_KEYED_TABLES` is swept by a test: its query is built at runtime, so a wrong table name is not a compile error and would only surface as a failed tally for that trigger kind in a fork. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): let only a request task vouch for a deploy event Round 2 found the first fix incomplete. Detaching only the failed/cancelled dependency path left the common route untouched: a dependency job that succeeds calls `handle_deployment_metadata` from the worker, where `deploy_origin::current()` read as `Authored`. A sync archiving the script while its lock generation was pending then had its deletion probed on completion and refiled as authored — the same fabricated removal, on the path most deploys actually take. `current()` now returns `Option`, `Some` only inside the request scope the API always enters. Having no scope means "not the task that served this write", which is true of every worker-side call and needs no marking at the call site. The integration test drives the real `handle_deployment_metadata` off a request task instead of the tally directly, and fails without this. Two more from the same round: - The script dependency handler passed no `renamed_from`, unlike the flow and app handlers next to it. A lock-generating create has no earlier tally, so that was the only chance for the path a rename vacated to be recorded at all — renames of Python/TS scripts left the old path in the parent, which the bash-only manual check missed. - The tally now drops a `renamed_from` equal to the path itself. Callers pass the previous path whether or not the deploy moved the item, so an unfiltered one both counted the path twice and stamped it `rename_from` when nothing was renamed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): carry a deploy's origin into the dependency job it queues Round 3 caught the previous fix cutting too deep. Refusing a detached tally any claim also refused its rename evidence, and a lock-generating deploy has no other tally — so the `renamed_from` added alongside it was inert, and a renamed flow, app or Python script still left its old path in the parent with nothing to merge. Flows and apps always generate, so renames worked essentially nowhere. The two capabilities are now separate. `TallyEvidence` says whether the tallying task served the write (`Served`, may probe what the path holds now) or is reporting one that committed earlier (`Deferred`, may not), and each column is written only from a source that answers for it. The origin itself is a fact of the deploy either way, so the request stamps it into the dependency job's args and the worker re-enters the scope with it — the last place that knows it handing it to the only tally that will run. Also from round 3: `WorkspaceDiffRow`'s event fields skip serializing `None` rather than emitting `null`, matching what the schema declares (OpenAPI 3.0.3 ignores a `description` sibling of `$ref`, so those moved onto the shared schemas). Verified against a live worker: renaming a flow in a fork records `(rename_from, authored)` on the vacated path and the merge offers its removal, while the deployed path claims nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): mark the CLI's parent-to-fork merge as sync `wmill workspace merge --direction to-fork` is the CLI's "Update fork" and deletes items in the fork, but without the marker the compare page sets. Its deletions were recorded as authored fork decisions, so once the parent recreated such a path the merge would offer deleting it there. Also from review: an unrecognized deploy-origin arg now reads as no evidence rather than as authored — strict where a request header is lenient, since an unmarked request really is authored but an unreadable stored value is skew. Reading the arg moved next to `stamp_origin_arg`, the half that writes it, so the round trip a lock-generating deploy depends on is covered by one test. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: drop the imports the shared arg reader made unused CI compiles with `-D warnings`, so this was four red Backend jobs rather than a lint. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): stop a stale deferred rename from restating a removed path Nothing orders these events. A tally that served the write made its claim inside its own commit, but a deferred one reports a write that landed at an unknown remove. So a lock-generating rename whose dependency job finished after a sync had removed the vacated path could overwrite `(delete, sync)` with `(rename_from, authored)` — the path is gone either way, so the merge would then offer removing it from the parent on the strength of the older event. A deferred claim now only writes where the side has none, which is the case it exists for: a vacated path that nothing else has spoken for. The regression asserts the ordering directly, and fails without the guard. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): record a rename's vacated path from the request that made it The deferred mechanism could not be made correct, as round 7 showed: its guard protected an existing row, but that row is deleted as soon as the two workspaces agree on the path — so a rename job finishing after the reconciliation inserted fresh, and the stale claim reappeared against whatever the parent later recreated there. Ordering cannot be recovered outside the row, because the row is disposable. So the vacated path is now recorded by the request, which is inside its own commit and whose row shares the counter's lifetime. A deploy that hands its metadata to a dependency job — every flow and app, and any script needing a lock — calls `tally_rename_vacated_path` once its transaction has committed; scripts reach it through the post-commit hook they already had, which grew a second variant rather than new plumbing. That lets the whole deferred apparatus go: `TallyEvidence`, the origin job arg and its round trip. `deploy_origin::current` is `Some` only inside a request scope again, and `handle_deployment_metadata` hands `renamed_from` to the tally only when it can answer for it — git-sync still gets it either way, so the rename keeps naming itself in the commit message. The vacated path's kind now reads `delete` rather than `rename_from` for these deploys, since it is probed rather than declared. The merge treats the two alike; only the row's tooltip is less specific. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(fork): cover raw-app renames, and stop firing CI before the lock exists Two things the vacated-path call broke or missed: - `create_script` reads its third return value as "no lock generation needed" to decide whether the script is runnable now, and the new `VacatedPath` variant made that true for renames that do generate. Those fired dependent CI tests from the API against a version with no lockfile, and again from the dependency job. The variant now decides it explicitly. - Raw apps rename through `update_app_raw`, a separate route into `update_app_internal`, which the new call had not been attached to. Both routes now go through one helper. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(fork): assert the kind only an inline rename can record `rename_from` is what a deploy says when it knows it moved the item, which only the path that reports both halves from its own request can. Nothing pinned it, and that is the side the vacated-path change touched. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore: update ee-repo-ref to a45bec03922d305aad5893ed354dc029c7f97bb4 This commit updates the EE repository reference after PR #709 was merged in windmill-ee-private. Previous ee-repo-ref: 62f494b2a51de0dfc0cfa0c3530ff19a1d32667c New ee-repo-ref: a45bec03922d305aad5893ed354dc029c7f97bb4 Automated by sync-ee-ref workflow. --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-authored-by: windmill-internal-app[bot] <windmill-internal-app[bot]@users.noreply.github.com> |
||
|
|
beef6e295c |
feat: azure workload identity auth for mssql and postgres resources (#10470)
* feat: azure workload identity auth for mssql and postgres resources * refactor: keep mssql config lines untouched by the auth-mode change * fix: single-flight token refresh, cache eviction and identity-aware pg cache key * fix: back off after a failed entra id refresh and normalize blank pg identity fields * fix: re-check the fallback token lifetime after a failed refresh * refactor: select workload identity with a sentinel password instead of resource fields * fix: log the workload identity mode on the postgres path too |
||
|
|
d5095515ed |
fix: return result.json and stdout results from sandboxed containers (#10460)
* fix: return result.json and stdout results from sandboxed containers Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: capture unmasked stdout-only last line, validate container result.json Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: reject image WorkingDir that escapes the container root Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: verify the whole result mount destination against the extracted rootfs Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix: name the right skip reason and gate the symlink test to unix Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> |